节点文献

基于生成式模型的生态系统声音景观重构的研究

Ecosystem Soundscape Reconstruction Using Generative Models

【作者】 王梅;

【导师】 刘方邻;

【作者基本信息】 中国科学技术大学 , 生态学, 2025, 博士

【摘要】 全球生物多样性持续下降,各国都在加大生态监测与保护力度。生态声学通过记录和研究生物与环境声构成的声音景观,成为揭示生物多样性与评估生态状况的重要途径。随着人工智能的发展,机器学习被广泛用于解析声音景观的物种和结构变化。然而,现实生态系统中的声源多样且相互叠加,现有基于声学指数或判别式模型的方法难以准确识别与分离声源,限制了声景解析的精度和效率。因此,需建立一种能够实现生态声景精准解析的方法框架,以突破现有技术瓶颈,为生态监测与生物多样性评估提供更可靠的支撑。生成式模型是一种区别于判别式模型的数学建模方法,其理论体系相对完善,但在生态声学尚处于探索阶段。声学数据是否具备可被生成模型有效学习的分布结构,以及生成模型能否据此解析和重构声景中的生态信息,成为需要验证的关键问题。为此,本文围绕生成式模型的潜在能力开展探索性研究,评估其在声学数据解析与生态信息重构中的可行性与优势。本研究以鹞落坪国家级自然保护区为例,布设监测站点,采集音频数据,旨在构建一套基于生成模型的声景分析流程。主要内容如下:1. 利用VGGish模型提取音频的高维特征嵌入,采用高斯混合模型(Gaussian Mixture Model,GMM)进行无监督聚类,识别出7类声景成分。在方法层面,验证了高维声学特征表征声景的效果,初步探索了生成式概率模型在声源识别中的可行性。在生态层面,揭示了区域生物声与非生物声的主要组成比例及昼夜、季节和海拔梯度差异。2. 在声景成分初步建模的基础上,构建典型群落和物种水平声谱图的成对数据,采用生成对抗网络(Generative Adversarial Network,GAN),训练了群落和物种生成模型。在像素级重构中,群落水平声源的平均F1分数为0.79,物种水平声源的平均F1分数为0.76;在图像级分类中,GAN与基线分类器性能相当;在降噪与声源分离中,GAN的均方误差表现优于基线方法。在方法层面,表明了GAN模型的声景重构和解析能力。在生态层面,揭示了群落水平的频率分布差别以及不同物种的声学空间划分情况。3. 依据上述核心算法,使用Python语言,开发了生态声学分析工具包Eco Spect。集成特征提取、聚类建模、声源分离等模块及交互界面,实现了声景的可视化与自动化处理。综上所述,本研究构建的生成式分析框架,能够在声源复杂的环境下重现声景的结构与动态变化,为评估生物多样性、监测生态扰动及理解生态系统健康状况提供了新的技术途径和分析思路。

【Abstract】 Global biodiversity continues to decline,and countries around the world are intensifying efforts in ecological monitoring and conservation.By recording and studying soundscapes composed of biological and environmental sounds,ecoacoustics has become an important approach for revealing biodiversity and assessing ecological conditions.With the development of artificial intelligence,machine learning has been widely applied to analyze species composition and structural changes in soundscapes.However,in real ecosystems,sound sources are diverse and often overlap with one another.Existing methods based on acoustic indices or discriminative models have difficulty in accurately identifying and separating sound sources,which limits the precision and efficiency of soundscape analysis.Therefore,there is an urgent need to establish a methodological framework capable of precise ecological soundscape analysis,so as to overcome current technical bottlenecks and provide more reliable support for ecological monitoring and biodiversity assessment.Generative models,as a mathematical modeling approach distinct from discriminative models,possess a relatively well-developed theoretical system but remain in the exploratory stage within ecoacoustics.Against this background,whether acoustic data possess a distributional structure that can be effectively learned by generative models,and whether generative models can thereby analyze and reconstruct ecological information within soundscapes,become key questions to be verified.To this end,this study conducts exploratory research on the potential capability of generative models,evaluating their feasibility and advantages in the analysis and reconstruction of ecological information from ecoacoustic data.Taking Yaoluoping National Nature Reserve as a case study,this research established monitoring stations and collected audio data,with the aim of constructing a soundscape analysis workflow based on generative models.The main contents are as follows:High dimensional feature embeddings of audio data were extracted using the VGGish model,and unsupervised clustering was conducted with a Gaussian Mixture Model,GMM,through which seven categories of soundscape components were identified.At the methodological level,the effectiveness of high dimensional acoustic features in representing soundscapes was verified,and the feasibility of generative probabilistic models in sound source identification was preliminarily explored.At the ecological level,the major composition ratios of biological and abiotic sounds in the study area,as well as their differences across diel periods,seasons,and altitudinal gradients,were revealed.On the basis of the preliminary modeling of soundscape components,paired spectrogram data at the levels of typical communities and species were constructed,and Generative Adversarial Networks,GAN,were used to train community and species generation models.In pixel level reconstruction,the average F1 score of community level sound sources was 0.79,and the average F1 score of species level sound sources was 0.76.In image level classification,the performance of GAN was comparable to that of the baseline classifier.In denoising and sound source separation,the mean squared error of GAN was better than that of the baseline method.At the methodological level,the soundscape reconstruction and analysis capability of the GAN model was demonstrated.At the ecological level,differences in frequency distribution at the community level and the acoustic niche partitioning of different species were revealed.Based on the core algorithms described above,an ecoacoustic analysis toolkit named Eco Spect was developed in Python.It integrates modules and an interactive interface for feature extraction,clustering based modeling,and sound source separation,thereby enabling the visualization and automated processing of soundscapes.In summary,the generative analytical framework constructed in this research is capable of reproducing the structure and dynamic changes of soundscapes under acoustically complex environments,and provides a new technical approach and analytical perspective for assessing biodiversity,monitoring ecological disturbances,and understanding ecosystem health.

  • 【分类号】Q14;TN912.3;TP18
节点文献中: 

本文链接的文献网络图示:

本文的引文网络