节点文献

基于词频分布变化统计的术语抽取方法

Terminology Extraction Based on Statistical Word Frequency Distribution Variety

  • 推荐 CAJ下载
  • PDF下载
  • 不支持迅雷等下载工具,请取消加速工具后下载。

【作者】 周浪张亮冯冲黄河燕

【Author】 ZHOU Lang1,2 ZHANG Liang2 FENG Chong2 HUANG He-yan2_(College of Computer Science and Technology,Nanjing University of Science and Technology,Nanjing 210094,China)1(Research Center of Computer & Language Information Engineering,CAS Beijing 100089,China)2(Dept.of Computer Science and Technology,Nanjing University,Nanjing 210093,China)3

【机构】 南京理工大学计算机科学与技术学院计算机语言信息工程研究中心南京大学计算机科学与技术学院

【摘要】 提出了一种规则与统计相结合的术语抽取方法,用于抽取包含多个词语的词组型术语。目前,绝大多数的统计方法都侧重于衡量术语的结构完整性,但这些方法并不能体现术语与专业相关的领域特征。通过对术语在各文档中的分布情况进行观察,提出了一种利用术语在语料中词频分布变化程度的统计信息来检验术语的领域相关性的方法,同时结合机器学习方法获取的语言知识,从计算机领域的语料中抽取领域特征明显的词组型术语。实验证明,该方法对低频术语和高频普通词串有较强的分辨能力。

【Abstract】 A hybrid terminology extraction system combined with linguistic knowledge and statistical information was introduced to extract compound terms which contain more than one word.There have been many statistical strategies used in automatic terminology extraction,most of which emphasize particularly to measure the integrality of the terms,other than domain features.To measure the domain relativity of terms,a mew method utilizing term frequency distribution variety was proposed.Incorporating with linguistic knowledge acquired by machine learning method,an automatic extraction system was implemented to extract multi-word terms from the corporate of computer domain.The results show that this approach is effective especially to distinguish terms with lower frequency and common words with higher frequency.

【基金】 国家863高技术研究发展计划项目(2006AA01Z152);国家自然科学基金项目(60672149)资助
  • 【文献出处】 计算机科学 ,Computer Science , 编辑部邮箱 ,2009年05期
  • 【分类号】TP391.1
  • 【被引频次】62
  • 【下载频次】709
节点文献中: 

本文链接的文献网络图示:

本文的引文网络