节点文献

检索主题难易度预测

Prediction of Topic Difficulty

  • 推荐 CAJ下载
  • PDF下载
  • 不支持迅雷等下载工具,请取消加速工具后下载。

【作者】 吕学强赖治国昝红英项锟

【Author】 L(?) Xue-qiang LAI Zhi-guo ZAN Hong-ying XIANG Kun (Institute of Computational Linguistics,Peking University,Beijing 100871)

【机构】 北京大学计算语言学研究所

【摘要】 TREC2004 Robust 任务有一项新要求,就是要把检索主题按照从易到难的顺序排列.针对新要求,该文提出了基于单词歧义性大小的检索主题难易度模型.根据WordNet 和它附带的Brown 语料库构造了单词义项分布词典,然后把检索主题中的单词按歧义性大小分为七类,通过计算平均单词容易度来度量检索主题的难度.实验结果表明该模型有一定的预测能力.最后预测了TREC2004 Robust 任务的250个检索主题的难易度.

【Abstract】 TREC2004 robust track requires predicting the relative difficulty of the topics.A topic difficulty model basedon word sense ambiguity was proposed in this paper.After constructing a sense distribution dictionary using WordNet andBrown corpus,the words in a topic could be put into seven classes.Average word easiness reflected the topic difficulty.Experimental results showed that the model can predict topic difficulty to some extent.According to the model,the relativedifficulty of 250 topics in TREC2004 robust track was predicted.

【基金】 国家863项目(2002AA117010-8);国家自然科学基金项目(60203022)
  • 【会议录名称】 NCIRCS2004第一届全国信息检索与内容安全学术会议论文集
  • 【会议名称】NCIRCS2004第一届全国信息检索与内容安全学术会议
  • 【会议时间】2004-11
  • 【会议地点】中国上海
  • 【分类号】TP391.3
  • 【主办单位】复旦大学计算机科学与工程系、上海市智能信息处理重点实验室
节点文献中: