节点文献
基于小规模标注语料的机器学习方法研究
Analysis and Prospect Machine Learning Methods Based on Limited Corpus
【摘要】 文中通过讨论机器学习和自然语言处理之间的关系,论述了语料库语言工程中机器学习的困境,概述分析了应用半监督学习的现状,研究有限样本下结合未标注样本的方法和统计学习理论框架的结合前景。
【Abstract】 Many difficulties exist when applying machine learning techniques to statistical natural language processing. This paper surveys the status of semi-supervised machine learning methods combined with the unlabeled data and gives a prospect of method based on the semi-supervised and statistical learning theory under limited data, which can be break through the bottleneck to some degree.
【关键词】 机器学习;
语料库;
未标注样本;
Cotraining;
主动学习;
统计学习理论;
【Key words】 machine learning; corpora; Unlabled data; Co-training, active learning; statistical learning theory;
【Key words】 machine learning; corpora; Unlabled data; Co-training, active learning; statistical learning theory;
- 【文献出处】 计算机应用 ,Computer Applications , 编辑部邮箱 ,2004年02期
- 【分类号】TP181
- 【被引频次】9
- 【下载频次】334