节点文献

基于电子病历的临床试验标准分类研究

Research on Clinical Trial Criteria Classification Based on Electronic Medical Records

【作者】 张帆

【导师】 周小兵;

【作者基本信息】 云南大学 , 计算机技术, 2019, 硕士

【摘要】 近年来,深度学习和自然语言处理技术在飞速发展,医疗领域与计算机领域的联系越来越密切,计算机不单用于电子计算通信,还用来进行辅助性的临床研究,运用其强大的计算能力和处理能力帮助产生医疗决策。本文所解决的临床试验队列选择任务即是自然语言处理在临床研究领域的应用。伴随着计算机决策在医学研究方面的准确性不断提高,自然语言处理技术在医疗领域的应用受到更多研究人员的关注。本文引入了一个临床实验队列选择的任务和相关数据集,这个任务旨在回答这样一个问题:NLP模型能否使用叙述性医疗记录来确定哪些患者符合临床试验的选择标准?本文并没有使用传统的语法、语义分析和构建规则的方式,而是采用了深度学习的方法,分别基于词向量和句向量来构建NLP分类模型。在基于词向量的方法中,本文使用经过预训练的词向量模型和通过词向量扩充方法得到的词向量模型,为临床记录文本构造词向量,用词向量来表征出文本的语义信息,然后使用CNN模型、BiLSTM模型、注意力机制、RNNCNN模型等深度学习模型构造了一系列的临床试验标准分类模型,并在此基础上设计出基于词向量的双通道分类模型;在基于句向量的方法中,本文使用经过预训练InferSent模型将文本句子编码为句向量,然后改进本文提出的模型得到基于句向量的双通道分类模型,最后通过不同的对比实验来验证本文模型的分类性能。基于词向量的双通道模型在测试集上实现了0.7810的micro-F1,在使用了词向量扩充方法之后将micro-F1值提升到了0.7905;基于句向量的双通道模型在测试集上实现了0.7961的micro-F1,达到了模型最好效果,证明了本文设计的改进模型在临床试验标准分类任务上的有效性。

【Abstract】 In recent years,the technology of deep learning and natural language processing has been developing rapidly,and the connection between medical field and computer field is getting closer and closer.Computer is not only used for electronics,computing and communication,but also for auxiliary clinical research,and its powerful computing and processing ability can help produce medical decisions.The clinical trial cohort selection task in this thesis is the application of natural language processing in medical research.With the increasing accuracy of computer decision making in clinical research,the application of natural language processing technology in medical field has attracted more researchers’attention.This thesis introduces a cohort selection of clinical trials task and related data set,which aims to answer the question:can NLP models use narrative medical records to identify which patients meet selection criteria for clinical trials?Instead of using the traditional methods of grammar,semantic analysis and rule construction,this thesis adopted the method of deep learning to build the NLP classification model based on the word embedding and sentence embedding respecti’vely.In the method based on word embedding,this thesis use pre-training word embedding models and the word embedding models which is obtained by word embedding expansion methods to construct word vector for clinical record text and use word vector representing the text information.Then,a series of clinical trial criteria classification models were constructed using deep learning models such as Convolutional Neural Network,Bidirectional Long-Short Time Memory model,Attention Mechanism and RNNCNN model.And based on these models,a two-channel classification model based on word embedding was designed;In the method based on sentence embedding,this thesis used the pre-training InferSent model to encode text sentences into sentence vectors.Then the model which was proposed in this thesis was improved to obtain a two-channel classification model based on sentence vectors.Finally,the classification performance of our model is verified by different comparison experiments.The two-channel model based on the word vector realized micro-fl of 0.7810 on the test set.After the word vector expansion method was used,the micro-fl value was increased to 0.7905.The two-channel model based on sentence vector achieved 0.7961 micro-fl on the test set,achieving the best effect of the model,and proving the effectiveness of the improved model on clinical trial criteria classification task.

  • 【网络出版投稿人】 云南大学
  • 【网络出版年期】2020年 03期
  • 【分类号】TP18;TP391.1;R197.323
  • 【被引频次】1
  • 【下载频次】87
节点文献中: 

本文链接的文献网络图示:

本文的引文网络