节点文献
主动学习算法研究进展
Recent advances in active learning algorithms
【摘要】 主动学习的主要目的是在保证分类器精度不降低的前提下尽量降低人工标注的成本.主动学习算法通过迭代方式在原始样例集中挑选可以提升模型性能的样例进行专家标注,并将其补充到已有的训练集中,使被训练的分类器在较低的标注成本下获得较强的泛化能力.首先对主动学习算法中3个关键步骤的研究进展情况进行了分析:1)初始训练样例集的构建方法及其改进;2)样例选择策略及其改进;3)算法终止条件的设定及其改进;然后对传统主动学习算法面临的问题及改进措施进行了深入剖析;最后展望了主动学习需进一步研究的内容.
【Abstract】 Active learning mainly aims at reducing the cost of manual annotation without decreasing the accuracy of the classifier.Active learning algorithm gets high quality training sample set by selecting the informative unlabeled samples which are labeled by domain experts later.The selected sample set is used to train the classifier.This improves the generalization ability of trained classifier while minimizes the cost of the labeling.Firstly,the recent advances in the three key steps in active learning algorithm was summarized,including:1)the method for constructing the initial training sample set and its improvement;2)the sample selection strategy and its improvement;3)the termination condition and its improvement.Then,the problems in active learning were analyzed and the corresponding countermeasures were presented.Finally,the future works in active learning were addressed.
【Key words】 active learning; initial training sample set; sample selection strategy; termination condition;
- 【文献出处】 河北大学学报(自然科学版) ,Journal of Hebei University(Natural Science Edition) , 编辑部邮箱 ,2017年02期
- 【分类号】TP181
- 【被引频次】38
- 【下载频次】854