节点文献

基于《知网》的中文信息结构抽取研究

An Approach Based HowNet for Extracting Chinese Message Structure

  • 推荐 CAJ下载
  • PDF下载
  • 不支持迅雷等下载工具,请取消加速工具后下载。

【作者】 尤昉李涓子王作英

【Author】 You Fang 1 Li Juanzi 2 Wang Zuoying 11 (Dept.of Electronics Engineering,Tsinghua University,Beijing100084) 2 (Dept.of Computer Science Technology,Tsinghua University,Beijing100084)

【机构】 清华大学电子工程系清华大学计算机科学与技术系清华大学电子工程系 北京100084北京100084北京100084

【摘要】 文章提出了一种在真实文本中抽取中文信息结构的方法—利用大规模基于语义依存关系的语料库对《知网》的中文信息结构模式进行训练,用这些带概率的模式作为规则建立部分依存分析器,从而从真实文本中最大限度地抽取符合知网中文信息结构定义的短语。该研究除了对将要建立的基于语义依存关系的语言模型是个有益的补充外,对于文本理解、对话系统甚至语音合成中的重音预测、韵律建模等等方面都有十分广阔的应用前景。

【Abstract】 An approach of extracting Chinese Message Structure from real texts is presented in this paper.The authors used the annotated corpus based on semantic dependent relations for Chinese Message Structure patterns’ training.With those patterns as rules,they built a partial dependency parser,so as to extract CMS from real texts as most as possible.The description of the training algorithm,experimental results and some conclusion are given.

【基金】 国家863高技术研究发展计划项目(编号:863-306-ZD03-02-1);985重大项目“人机自然语言交互技术”(编号:985校-22-攻关-06)资助
  • 【文献出处】 计算机工程与应用 ,Computer Engineering and Applications , 编辑部邮箱 ,2002年18期
  • 【分类号】TP391.1
  • 【被引频次】13
  • 【下载频次】293
节点文献中: 

本文链接的文献网络图示:

本文的引文网络