节点文献
基于前后文词形特征的生物医学文献句子边界识别
Sentence Boundary Detection in Biomedical Texts Using Context Morphological Features
【摘要】 针对生物医学文献的特点及信息抽取的特殊要求,提出了基于前后文词形特征和有教师学习的句子边界识别算法.与针对一般英语书面语设计的句子边界识别算法不同,本文提出的算法不使用特殊的辅助词表和语法层面的特征信息,只使用前后文单词的词形信息作为句子边界识别和消歧的依据.利用这些特征设计了最大信息熵识别器和支持向量机识别器,并在Medline摘要上进行了实验,达到了超过99%的正确率.实验结果表明,最大信息熵法和支持向量机法在句子边界消歧问题上具有相近的性能,同时还表明,对生物医学文献句子边界识别,只使用词法层面的特征,不使用辅助词表和词性等语法层面的信息,仍可达到其它算法在一般英语书面语上利用辅助词表和词性信息所达到的性能.
【Abstract】 A sentence boundary detection algorithm is proposed for information extraction from biomedical texts according to characteristics of the texts and special requirements of information extraction. The algorithm is based on context morphological features and supervised learning technology. In contrast to algorithms developed for sentence boundary detection in common English texts, the algorithm does not use special vocabulary and grammatical level information, and makes decision about sentence boundary just based on morphological information of the context words. A maximum entropy detector and a SVM detector are developed by using these features. Experiments done on Medline abstracts show that the algorithm has achieved accuracy of recognition above 99%, and maximum entropy and SVM methods have the approximate performance for the problem of sentence boundary disambiguation. The experiments also show that just using morphological level information without supplementary vocabulary and grammatical level information remains the approximately same performance as the other algorithms using supplementary vocabulary and grammatical level information for common English texts do.
【Key words】 natural language processing; biomedical information extraction; sentence boundary detection; machine learning;
- 【文献出处】 小型微型计算机系统 ,Journal of Chinese Computer Systems , 编辑部邮箱 ,2006年01期
- 【分类号】TP391.43
- 【被引频次】4
- 【下载频次】187