节点文献
基于词频分布变化统计的术语抽取方法
Terminology Extraction Based on Statistical Word Frequency Distribution Variety
【摘要】 提出了一种规则与统计相结合的术语抽取方法,用于抽取包含多个词语的词组型术语。目前,绝大多数的统计方法都侧重于衡量术语的结构完整性,但这些方法并不能体现术语与专业相关的领域特征。通过对术语在各文档中的分布情况进行观察,提出了一种利用术语在语料中词频分布变化程度的统计信息来检验术语的领域相关性的方法,同时结合机器学习方法获取的语言知识,从计算机领域的语料中抽取领域特征明显的词组型术语。实验证明,该方法对低频术语和高频普通词串有较强的分辨能力。
【Abstract】 A hybrid terminology extraction system combined with linguistic knowledge and statistical information was introduced to extract compound terms which contain more than one word.There have been many statistical strategies used in automatic terminology extraction,most of which emphasize particularly to measure the integrality of the terms,other than domain features.To measure the domain relativity of terms,a mew method utilizing term frequency distribution variety was proposed.Incorporating with linguistic knowledge acquired by machine learning method,an automatic extraction system was implemented to extract multi-word terms from the corporate of computer domain.The results show that this approach is effective especially to distinguish terms with lower frequency and common words with higher frequency.
【Key words】 Terminology extraction; Machine learning; Distribution variance; Knowledge acquisition; Termhood; Unithood;
- 【文献出处】 计算机科学 ,Computer Science , 编辑部邮箱 ,2009年05期
- 【分类号】TP391.1
- 【被引频次】62
- 【下载频次】709