节点文献
计算机领域术语的自动获取与层次构建
Computer Domain Term Automatic Extraction and Hierarchical Structure Building
【摘要】 设计一种能够自动获取计算机领域术语的方案,提出基于规则与统计相结合的抽取方法,使用亚马逊网站的计算机类图书作为语料库,通过分词、去停止词预处理以及词频统计的方法提取出计算机类领域术语,并插入到由ODP构建的树中,形成计算机领域术语的层次结构。实验结果表明,与人工标注结果相比,使用该方法自动获取的术语有很高的准确率与召回率。
【Abstract】 This paper presents a computer domain term automatic extraction method based on rules and statistics.It uses computer book titles from Amazon.com website as corpus,data are preprocessed by words splitting,stop words and special characters filtering.Terms are extracted by a set of rules and frequency statistics and inserted into a word tree from ODP to build the hierarchical structure.Experimental results show high precision and recall of the automatically extracted results compared with manual tagged terms.
【关键词】 计算机领域术语;
术语获取;
层次结构;
ODP项目;
【Key words】 computer domain term; term extraction; hierarchical structure; Open Directory Project(ODP);
【Key words】 computer domain term; term extraction; hierarchical structure; Open Directory Project(ODP);
【基金】 国家“863”计划基金资助项目(2006AA10Z232)
- 【文献出处】 计算机工程 ,Computer Engineering , 编辑部邮箱 ,2011年02期
- 【分类号】TP391.1
- 【被引频次】9
- 【下载频次】214