节点文献
基于Bootstrapping的文本分类模型
Semi-Supervised Text Categorization Using Bootstrapping
【Author】 Chen Wenliang Zhu Muhua Zhu Jingbo Yao Tianshun (Natural Language Processing Lab,Northeastern University,Shenyang,110004)
【机构】 东北大学自然语言处理实验室;
【摘要】 本文提出一种基于Bootstrapping 的文本分类模型,该模型采用最大熵模型作为分类器,从少量的种子集出发,自动学习更多的文本作为新的种子样本,这样不断学习来提高最大熵分类器的文本分类性能。文中提出一个权重因子来调整新的种子样本在分类器训练过程中的权重。实验结果表明,在相同的手工训练语料的条件下,与传统的文本分类模型相比这种基于Bootstrapping 的文本分类模型具有明显优势,仅使用每类100篇种子训练集,分类结果的F1值为70.56%,比传统模型高出4.70%。同时,使用大约一半或者更少规模的标注训练语料作为种子集,就可以达到传统的分类模型的相当结果。该模型通过使用适当的权重因子可以更好改善分类器的训练效果。
【Abstract】 This paper proposes a semi-supervised text categorization using bootstrapping.The System uses theMaximum Entropy Model as the text classifier.It learns more automatic labeled samples as new seed training samples fromunlabeled samples using a small size of seed training samples.In this paper,we use a weighted factor to adjust the weight ofnew seed samples during the following training process.The experimental results show that the proposed system performsbetter than the conventional system with the same labeled documents.And it yields 70.56% F1 using only 100-1abeleddocuments for each category,4.7% over the conventional system does.And it can provide the same performance as theconventional system using 50% or less training samples.The results also show that the weighted factor can improve theperformance.
- 【会议录名称】 NCIRCS2004第一届全国信息检索与内容安全学术会议论文集
- 【会议名称】NCIRCS2004第一届全国信息检索与内容安全学术会议
- 【会议时间】2004-11
- 【会议地点】中国上海
- 【分类号】TP391.1
- 【主办单位】复旦大学计算机科学与工程系、上海市智能信息处理重点实验室