节点文献

面向Internet的中文新词语检测

Internet-oriented Chinese New Words Detection

  • 推荐 CAJ下载
  • PDF下载
  • 不支持迅雷等下载工具,请取消加速工具后下载。

【作者】 邹纲刘洋刘群孟遥于浩西野文人亢世勇

【Author】 ZOU Gang 1,LIU Yang 1,LIU Qun 1,MENG Yao 2,YU Hao 2,Nishino Fumihito 2,KANG Shi-yong 3 (1 Institute of Computing Technology,Chinese Academy of Sciences, Beijing 100080,China; 2 Fujitsu Research & Development Center Co.,LTD, Beijing 100081,China; 3 Yantai Normal University Chinese Department, Yantai, Shandong 264025,China)

【机构】 中科院计算技术研究所数字化实验室富士通研究开发中心有限公司烟台师范学院中文系 北京100080北京100080北京100081山东烟台264025

【摘要】 随着社会的飞速发展 ,新词语不断地在日常生活中涌现出来。搜集和整理这些新词语 ,是中文信息处理中的一个重要研究课题。本文提出了一种自动检测新词语的方法 ,通过大规模地分析从Internet上采集而来的网页 ,建立巨大的词和字串的集合 ,从中自动检测新词语 ,而后再根据构词规则对自动检测的结果进行进一步的过滤 ,最终抽取出采集语料中存在的新词语。根据该方法实现的系统 ,可以寻找不限长度和不限领域的新词语 ,目前正应用于《现代汉语新词语信息 (电子 )词典》的编纂 ,在实用中大大的减轻了人工查找新词语的负担。

【Abstract】 With the fast development of the society,more and more new words come out in our life. It is one of the important topics in Chinese natural language processing to collect those new words. A method is presented for detecting these new words automaitcally in this paper. Through analysing webpages grabbed from the Internet, a large word and string set is built, which new words are detected from and filtered by rules. At last new words which exist in the webpages grabbed are extracted. The system built in this way can find new words in any length and in any field.Now it is applying to the compilation of Modern Chinese New Word Information Dictionary. It reduced human labor a lot in practise.

  • 【文献出处】 中文信息学报 ,Journal of Chinese Information Processing , 编辑部邮箱 ,2004年06期
  • 【分类号】TP393.09
  • 【被引频次】169
  • 【下载频次】838
节点文献中: 

本文链接的文献网络图示:

本文的引文网络