节点文献

基于大规模语料的中文词聚类研究与实现

An experimental study of Chinese words cluster on large-scale corpus

  • 推荐 CAJ下载
  • PDF下载
  • 不支持迅雷等下载工具,请取消加速工具后下载。

【作者】 蒋宏飞曹海龙杨沐昀

【Author】 hfjiang1 hlcao2 myyang3(Harbin Institute of Technology. Harbin 150001. China)

【机构】 哈尔滨工业大学计算机系

【摘要】 词聚类算法对自然语言处理具有重要意义。Brown 1990年提出了一个经典的词聚类算法,但是由于算法本身的复杂度较高,故难于对大规模语科进行处理(Brown文中提到词数超过5000便是不可行的)。本研究中我们尝试着对上万词数的中文词语料进行了实现。并且,针对算法时间复杂度高,不能应用于更大规模语料库的问题,提出了一个加速改进思想。在近似的情况下,它可以降低原算法一阶复杂度。本实验所用的语料来自人民日报1998年1月份的部分内容。

【Abstract】 Word-cluster algorithm is a key technique in NIP and many other research fields. In 1990, Brown presented a classic word-cluster algorithm. Unfortunately, the complexity of the algorithm is so high and it is impracticable for large-scale corpus (Brown’s paper mentioned that the algorithm is impracticable if the number of words beyond 5000). In this research we attempt to apply the algorithm to ten thousands of Chinese words. Moreover, we bring forward an accelerated idea, which can reduce the complexity of this algorithm for one order. The corpus we used in this research is a part of the contents of the People’s Daily in 1998.1.

【基金】 国家自然科学基金基于双语信息的英汉译文消歧技术研究(60375019)
  • 【会议录名称】 第二届全国学生计算语言学研讨会论文集
  • 【会议名称】第二届全国学生计算语言学研讨会
  • 【会议时间】2004-08
  • 【分类号】H087
  • 【主办单位】中国中文信息学会
节点文献中: 

本文链接的文献网络图示:

本文的引文网络