节点文献

基于变分自编码器的音乐生成技术研究

Research on Music Generation Technology Based on Variational Auto-Encoders

【作者】 杨超;

【导师】 许洪光;

【作者基本信息】 哈尔滨工业大学 , 信息与通信工程, 2024, 硕士

【摘要】 随着深度学习在生成领域的快速发展,音乐的自动生成逐渐成为了融合艺术与科技的引人注目的研究热点。利用神经网络深度学习音乐数据的特征与风格,可以以快速高效的方式生成新颖的音乐作品。这些生成的音乐不仅能够作为短视频的背景音乐,还可以成为影视剧情的感情注脚,或者为游戏场景打造出氛围独特的音效。然而,在现有音乐模型中,生成效果并不尽如人意,常常出现节奏感缺失和旋律质量较差的问题。这些问题主要源于两个方面。首先,变分自编码器在训练过程中存在一个难以克服的问题,即隐藏空间中的后验分布快速逼近先验分布,也就是KL散度减小,导致模型无法有效学习输入数据的特征,从而生成的音乐多样性较差。其次,由于使用了循环神经网络作为编码器和解码器,梯度消失问题导致模型难以学习到输入数据的长期特征,进而影响生成音乐的连续性。为解决这些问题,本文提出了两个研究方案。首先,引入离散码本到变分自编码器中,以解决KL散度快速减小导致模型无法学习数据多样性特征的问题。针对离散码本可能出现的坍塌问题,引入码本重启机制以避免此类情况,也就是当码本向量的使用次数超过一定阈值的时候,就将其进行重新初始化,从而使得不同的码本向量可以得到充分的训练。在解码器方面,采用多层解码器结构,通过传导层将量化向量分解为不同的子序列,再由分别通过解码器进行解码,从而缓解梯度消失问题。实验结果对比表明,改进后的模型生成的音乐表现出更丰富的音高变化,旋律更为出色。其次,考虑到文本翻译领域中注意力机制的良好表现,本文尝试将其引入变分自编码器的编解码器中。与循环神经网络不同,注意力机制在模型训练过程中能够全局关注信息,更好地关联前后内容。并且在训练的过程中使用了随机掩码,模型可以看到部分未来的数据,从而使得最终生成的音乐具有更好的前后相关性,也就是前后的连续性更好一些。经过实验评估,改进后的模型生成的音乐在空拍率和合格音符比例方面均有所提升,表现出更好的连续性和节奏感,从而提高了生成音乐的质量。

【Abstract】 With the rapid development of deep learning in the field of generation,the automatic generation of music has gradually become a captivating research focus that integrates art and technology.Utilizing the features and styles of music data through neural network deep learning enables the rapid and efficient generation of novel musical compositions.These generated musical pieces not only serve as background music for short videos but also act as emotional accents in film and television narratives or contribute to creating unique atmospheric sounds for gaming scenarios.However,in existing music models,the generated effects are not always satisfactory,often exhibiting issues such as rhythm deficiencies and lower melodic quality.These problems primarily stem from two aspects.Firstly,variational autoencoders(VAEs)face a challenging issue during training,where the posterior distribution in the latent space quickly approaches the prior distribution,leading to a reduction in KL divergence.This makes it difficult for the model to effectively learn the features of input data,resulting in poorer diversity in the generated music.Secondly,the use of recurrent neural networks(RNNs)as encoders and decoders introduces the problem of gradient vanishing,making it challenging for the model to learn the long-term features of input data and,consequently,impacting the continuity of the generated music.To address these issues,this thesis proposes two research solutions.Firstly,discrete codebooks are introduced into the variational autoencoder to overcome the rapid reduction in KL divergence,enhancing the model’s ability to learn diverse features of the data.To tackle the potential collapse issue with discrete codebooks,a codebook reset mechanism is introduced to avoid such situations.Specifically,when the usage count of a codebook vector exceeds a certain threshold,it is reinitialized,ensuring different codebook vectors receive sufficient training.In the decoder,a multi-layer decoder structure is employed.Through propagation layers,quantized vectors are decomposed into different subsequences,which are then decoded by the decoder.This approach alleviates the problem of gradient vanishing.Experimental results show that the improved model generates music with richer pitch variations and outstanding melodies.Secondly,considering the successful performance of attention mechanisms in the text translation field,this paper attempts to introduce it into the encoder and decoder of the variational autoencoder.Unlike recurrent neural networks,attention mechanisms can globally focus on information during the model training process,better associating the content before and after.Additionally,random masking is used during training,allowing the model to see parts of future data,enhancing the final music’s continuity.Through experimental evaluation,the improved model shows advancements in the rate of empty beats and the proportion of qualified notes,exhibiting better continuity and rhythm,thereby enhancing the quality of the generated music.

  • 【分类号】TP18;J60-05
节点文献中: