节点文献

基于Web的语料库建设

A Preliminary Research on the Construction of Web Corpus

  • 推荐 CAJ下载
  • PDF下载
  • 不支持迅雷等下载工具,请取消加速工具后下载。

【作者】 俞倩兰王国新邹永林

【Author】 Yu Qianlan Wang; Guoxin; Zou Yonglin (Department of mechanics and Electronics Changshu College; Changshu 215500)

【机构】 常熟高等专科学校!常熟215500

【摘要】 对网上中文信息语料库搜集技术的实现原理和关键技术进行了讨论和分析,介绍了基于Web网络的通讯及网上自动获取信息的原理,讨论了中文信息处理中的分词技术及其发展,提出了一个网上《人民日报》语料库搜集技术的实现方案.

【Abstract】 With the internet getting increasingly popular in China and the, information in Chinese on WWW becoming ever greater in volume, the importance of automatic data search technique in the Chinese information corpus on the line is more obvious than ever. The development and improvement of the technique is of great significance for bettering the process level of information in Chinese. The present paper, based on a discussion and analysis;of the realization laws and essential technology of data search technique in the Chinese information corpus, attempts to introduce the principles of realizing net communication at the Web and obtaining automatically the information on the line. A scheme of search technique for the corpus of People, s Daily is suggested with the classification and combination technology so far developed in the process of information in Chinese discussed and analyzed.

【关键词】 Web语料库分词
【Key words】 WebCurpusdividing words
  • 【文献出处】 常熟高专学报 ,JOURNAL OF CHANGSHU COLLEGE , 编辑部邮箱 ,2000年02期
  • 【分类号】G356
  • 【被引频次】3
  • 【下载频次】168
节点文献中: 

本文链接的文献网络图示:

本文的引文网络