节点文献
检索主题难易度预测
Prediction of Topic Difficulty
【Author】 L(?) Xue-qiang LAI Zhi-guo ZAN Hong-ying XIANG Kun (Institute of Computational Linguistics,Peking University,Beijing 100871)
【机构】 北京大学计算语言学研究所;
【摘要】 TREC2004 Robust 任务有一项新要求,就是要把检索主题按照从易到难的顺序排列.针对新要求,该文提出了基于单词歧义性大小的检索主题难易度模型.根据WordNet 和它附带的Brown 语料库构造了单词义项分布词典,然后把检索主题中的单词按歧义性大小分为七类,通过计算平均单词容易度来度量检索主题的难度.实验结果表明该模型有一定的预测能力.最后预测了TREC2004 Robust 任务的250个检索主题的难易度.
【Abstract】 TREC2004 robust track requires predicting the relative difficulty of the topics.A topic difficulty model basedon word sense ambiguity was proposed in this paper.After constructing a sense distribution dictionary using WordNet andBrown corpus,the words in a topic could be put into seven classes.Average word easiness reflected the topic difficulty.Experimental results showed that the model can predict topic difficulty to some extent.According to the model,the relativedifficulty of 250 topics in TREC2004 robust track was predicted.
【Key words】 Information retrieval; TREC; robust track; topic difficulty; sense distribution;
- 【会议录名称】 NCIRCS2004第一届全国信息检索与内容安全学术会议论文集
- 【会议名称】NCIRCS2004第一届全国信息检索与内容安全学术会议
- 【会议时间】2004-11
- 【会议地点】中国上海
- 【分类号】TP391.3
- 【主办单位】复旦大学计算机科学与工程系、上海市智能信息处理重点实验室