节点文献
利用统计量和语言学规则提取多字词表达
Extracting Multiword Expressions with Statistics and Linguistic Rules
【摘要】 基于特定领域的语料库,利用统计和语言学规则相结合的方法提取多字词表达(Multiword expressions)。首先利用领域高频词作为种子词提取候选串,进一步利用各种统计量、多字词表达边界过滤规则对候选串进行噪声剔除,得到多字词表达。实验结果表明,该方法对于处理大规模真实文本效率很高,可以有效提高多字词表达的获取,可以更有针对性地在特定领域提取多字词表达。
【Abstract】 Multiword Expressions(MWEs) are one of the bottlenecks for more precise Natural Language Processing(NLP) systems.Particularly,the lack of coverage of MWEs in resources can impact negatively on the performance of tasks and applications.For special domains,a significant portion of the vocabulary is composed of MWEs.This paper puts forwards an automatic method for extracting Chinese MWEs with help of statistics and linguistic rules.Seed words of high frequency in special domain are selected to extract candidate strings.By means of statistical measures and linguistic rules,noises in candidate strings are filtered.After filtering,Chinese MWEs are obtained finally.The result of our experiment shows that the method in this paper is efficient to deal with large-scale real texts.The method can extract Chinese MWEs rapidly.Chinese MWEs extracted in this way can be used in many application fields.
- 【文献出处】 太原理工大学学报 ,Journal of Taiyuan University of Technology , 编辑部邮箱 ,2011年02期
- 【分类号】H087
- 【被引频次】21
- 【下载频次】327