节点文献

基于深度学习的增强子鉴定及其活性预测方法研究

Research on Deep Learning-Based Enhancer Identification and Activity Prediction Methods

【作者】 张瑶

【导师】 吴昊;

【作者基本信息】 山东大学 , 软件工程(专业学位), 2025, 硕士

【摘要】 基因表达与调控构成细胞分化和个体发育等生物过程的核心分子机制。增强子作为关键的基因调控结构,与生物发育及疾病发生等多种过程密切相关,识别增强子对于理解基因组调控机制、确定关键元件和研究控制基因表达和疾病相关机制的网络至关重要。尽管已有多种实验技术用于检测增强子,但这些方法往往成本高、耗时长,限制了其广泛应用。近年来,生物信息学家提出了一些计算预测方法,但由于特征编码单一、模型结构简单,预测性能与泛化能力仍旧受限。针对这一问题,本研究聚焦DNA序列数据,探索了不同的序列特征编码方案,并分别构建了增强子识别模型(Enhancer-MDLF)和增强子定量预测模型(EAP-LSTM),以提升增强子的预测准确性和泛化能力。此外,研究还深入分析了增强子中的关键转录因子结合位点(TFBS)基序,为未来研究提供新的视角。主要内容包括:(1)基于卷积神经网络的增强子识别算法现有增强子识别方法在特征提取和模型设计方面仍有局限,且对增强子区域的转录因子基序探索不足。本研究提出了一种全新的多模块输入深度学习框架——Enhancer-MDLF。实验结果表明,Enhancer-MDLF在八种不同的人类细胞系(GM12878、HEK293、HMEC、HSMM、HUVEC、K562、NHEK 和 NHLF)上均优于现有方法Enhancer-IF,并在通用增强子数据集和增强子-启动子数据集上展现出更优的性能,进一步验证了其鲁棒性。此外,研究引入迁移学习,以提供一种有效且潜在的解决方案,来应对增强子特异性预测的挑战。进一步研究还利用模型解释技术识别出可能与增强子区域相关的TFBS基序,这对于研究增强子调控机制具有重要意义。(2)基于Bi-LSTM的增强子活性定量预测算法当前增强子定量预测方法受限于单一特征编码和简单模型架构。本研究提出了一种全新的深度学习框架EAP-LSTM,用于跨物种和不同细胞系的增强子活性定量预测。该模型融合了多种特征模块,在六种细胞系(五种人类细胞系A549、HCT116、HepG2、K562和MCF-7,以及一种果蝇细胞系S2)上的评估结果均优于最先进的模型。实验结果进一步验证了 EAP-LSTM的鲁棒性和高准确性,尤其是在小样本数据集中通过迁移学习有效缓解了性能下降问题。进一步研究还探讨了增强子区域内TFBS的作用,并识别出与增强子活性相关的关键基序,为深入解析增强子功能的分子机制提供了重要见解。

【Abstract】 Gene expression and regulation constitute the core molecular mechanisms underlying biological processes such as cell differentiation and individual development.As key gene regulatory elements,enhancers are closely associated with various processes,including biological development and disease occurrence.Identifying enhancers is crucial for understanding genomic regulatory mechanisms,pinpointing key elements,and investigating networks that control gene expression and diseaserelated mechanisms.Although several experimental techniques exist for detecting enhancers,these methods are often costly and time-consuming,limiting their widespread application.In recent years,bioinformaticians have proposed several computational prediction methods;however,their predictive performance and generalization ability remain constrained due to simplistic feature encoding and model structures.To address this issue,this study focuses on DNA sequence data and explores different sequence feature encoding schemes.Two models are constructed to improve enhancer prediction accuracy and generalization ability:Enhancer-MDLF,an enhancer identification model,and EAP-LSTM,an enhancer activity quantitative prediction model.Furthermore,the study conducts an in-depth analysis of key transcription factor binding site(TFBS)motifs within enhancers,providing new perspectives for future research.The main contributions are as follows:(1)Enhancer Identification Algorithm Based on Convolutional Neural NetworksExisting enhancer identification methods exhibit limitations in feature extraction and model design,with insufficient exploration of transcription factor motifs within enhancer regions.To address these challenges,this study proposes a novel multimodule input deep learning framework,Enhancer-MDLF.Experimental results demonstrate that Enhancer-MDLF outperforms the previous method,Enhancer-IF,across eight distinct human cell lines(GM12878,HEK293,HMEC,HSMM,HUVEC,K562,NHEK,and NHLF).Additionally,it achieves superior performance on generic enhancer datasets and enhancer-promoter datasets,further validating its robustness.Moreover,transfer learning is introduced as an effective and potential solution to address the challenges of enhancer specificity prediction.Furthermore,model interpretation techniques are employed to identify TFBS motifs potentially associated with enhancer regions,offering valuable insights into enhancer regulatory mechanisms.(2)Bi-LSTM-Based Enhancer Activity Quantitative Prediction AlgorithmCurrent enhancer activity prediction methods are limited by single-feature encoding and simple model architectures.To overcome these limitations,this study proposes a novel deep learning framework,EAP-LSTM,for cross-species and multicell-line quantitative prediction of enhancer activity.The model integrates multiple feature modules and achieves superior performance across six cell lines,including five human cell lines(A549,HCT116,HepG2,K562,and MCF-7)and one Drosophila cell line(S2),outperforming state-of-the-art models.Experimental results further confirm the robustness and high accuracy of EAP-LSTM,particularly in mitigating performance degradation in small-sample datasets through transfer learning.Additionally,the study explores the role of TFBSs within enhancer regions and identifies key motifs associated with enhancer activity,providing valuable insights into the molecular mechanisms underlying enhancer function.

  • 【网络出版投稿人】 山东大学
  • 【网络出版年期】2026年 05期
  • 【分类号】TP18;Q811.4
节点文献中: 

本文链接的文献网络图示:

本文的引文网络