节点文献

基于词典法和机器学习法相结合的蛋白质名识别

Protein names recognition based on dictionary and machine learning method

  • 推荐 CAJ下载
  • PDF下载
  • 不支持迅雷等下载工具,请取消加速工具后下载。

【作者】 李刚郭崇慧林鸿飞杨志豪唐焕文

【Author】 LI Gang, GUO Chong-hui, LIN Hong-fei, YANG Zhi-hao JANG Huan-wen ( Department of Applied Mathematics, Dalian University of Technology, Dalian, 116024, China; Department of Computer Science and Engineering, Dalian University of Technology, Dalian, 116024, China)

【机构】 大连理工大学应用数学系大连理工大学计算机科学与工程系

【摘要】 生物实体名识别对生物医学文献的信息抽取有重要的意义。本文针对如何识别蛋白质名进行了有益的尝试,主要采用了基于词典的方法,其中运用了近似搭配算法和首词查询的方法进行蛋白质名识别,同时结合机器学习方法训练了一个分类器来过滤候选词以提高识别的准确率。

【Abstract】 Identification of biomedical entities is one of important techniques to extract information from biomedical documents. This paper proposes an effective model based on dictionary to identify protein names. The approximate string searching method and first name searching are used to identify the candidate protein names, and a Naive Bayes classifier filtering the candidates is applied to improve the accuracy.

【关键词】 候选词编辑距离分类器
【Key words】 candidatesedit-distanceclassifier
【基金】 国家自然科学基金资助项目(90103033,60373095)。
  • 【会议录名称】 大连理工大学生物医学工程学术论文集(第2卷)
  • 【会议名称】大连理工大学生物医学工程学会研讨会
  • 【会议时间】2005-12
  • 【会议地点】中国辽宁大连
  • 【分类号】TP18;TP391.4
  • 【主办单位】大连理工大学
节点文献中: 

本文链接的文献网络图示:

本文的引文网络