节点文献

匹配树和决策树方法识别英语句子中的BaseNP

USING MATCHING TREE AND DECISION TREES TO IDENTIFY BaseNP IN ENGLISH SENTENCES

  • 推荐 CAJ下载
  • PDF下载
  • 不支持迅雷等下载工具,请取消加速工具后下载。

【作者】 荀恩东李生赵铁军

【Author】 XUN En Dong, LI Sheng, and ZHAO Tie Jun (Department of Computer Science and Engineering, Harbin Institute of Technology, Harbin 150001)

【机构】 哈尔滨工业大学计算机科学与工程系!哈尔滨150001

【摘要】 提出了语料库和机器学习相结合的方法识别英语句子中的简单的、非递归的名词短语 (Base NP) .在含有词性标注和 Base NP边界标注的训练语料中 ,抽取所有不同类型 Base NP短语对应的词性序列 (Base NP规则 ) ,通过规则排序和语言学知识 ,对其中正确率低且明显不符合语法的规则进行剔除 .在识别时 ,采取规则匹配树的方法进行最大长度匹配 ,通过归纳机器学习 C4.5算法引入上下文信息 ,由 C4.5算法学习出有效 (或无效 )应用 Base NP规则的条件 ,参照上下文条件 ,约束应用 Base NP规则 .实验结果表明 ,提出的方法具有很高的正确率和召回率 .

【Abstract】 A new method, which combines the corpus approach with the machine learning approach, is put forward in this paper to identify simple, non recursion noun phrases (BaseNP). Firstly, all different part of speech (POS) strings (BaseNP rules) which are corresponding to BaseNP are extracted from the training corpus tagged with POS and the boundary of each BaseNP. By means of training and based on linguistics knowledge, some BaseNP rules which have lower precision and have no linguistics sense apparently are deleted. Secondly, the remaining BaseNP rules are employed to identify BaseNP in new sentences. In the process, a heuristic algorithm of longest match, which is combined with the machine learning method of inductive decision trees to consult contexts, is applied. Experiments show that this new method results in higher precision and recall precision.

【关键词】 BaseNP名词短语匹配树决策树
【Key words】 BaseNPnoun phrasematching treedecision tree
【基金】 国家自然科学基金;国家“八六三”高技术研究发展计划基金
  • 【文献出处】 计算机研究与发展 ,JOURNAL OF COMPUTER RESEARCH AND DEVELOPMENT , 编辑部邮箱 ,2000年07期
  • 【分类号】TP391.1
  • 【被引频次】14
  • 【下载频次】208
节点文献中: 

本文链接的文献网络图示:

本文的引文网络