节点文献
基于加权关联模式挖掘与规则后件扩展的跨语言信息检索
Cross-Language Information Retrieval Based on Weighted Association Patterns and Rule Consequent Expansion
【摘要】 【目的】针对自然语言处理中查询主题漂移和词不匹配问题,提出一种基于加权关联模式挖掘和规则后件扩展的跨语言信息检索模型及其算法。【方法】该模型采用新的加权关联模式支持度和基于最大项目权值的项集剪枝策略挖掘频繁项集,利用置信度和相关度评价加权关联规则,根据扩展模型从规则中提取优质扩展词实现规则后件扩展,扩展词与原查询词项组合为新查询再次检索文档得到最终检索结果。【结果】实验结果表明,与单语言检索基准比较,本文检索模型的R-prec和P@10平均增幅分别为42.49%和25.53%;与跨语言检索基准比较,其平均增幅分别为91.87%和64.61%;与现有基于加权关联规则挖掘的跨语言检索方法比较,R-prec和P@10最高平均增幅分别可达93.20%和34.60%。【局限】只进行实验性研究,需要探讨在实际跨语言搜索引擎中的具体应用。【结论】本文检索模型能有效地减少查询主题漂移和词不匹配问题,改善和提高检索性能。
【Abstract】 [Objective] This paper proposes a new Cross-Language Information Retrieval(CLIR) model, aiming to address the issues facing natural language processing, such as query topic drift and word mismatch. [Methods] First, we explored the frequent item-sets with the weighted association patterns and the pruning strategies based on maximum item weight. Then, we used the confidence and relevance degrees to evaluate the weighted association rules, which helped us extract the high quality expansion terms. Finally, we combined the new terms with the original ones to create new queries for the final lists. [Results] Compared with the monolingual retrieval benchmark, the average increases(AIs) of R-prec and P@10 of the proposed model were 42.49% and 25.53%. Our results were 91.87% and 64.61% higher than the cross language retrieval benchmark. Compared to the existing CLIR methods, the maximum AIs of R-prec and P@10 were 93.20% and 34.60%. [Limitations] The proposed model needs to be examined with more cross language search engines. [Conclusions] Our model improves the performance of CLIR.
【Key words】 Information Retrieval; Cross Language Retrieval; Text Mining; Association Rule; Natural Language Processing;
- 【文献出处】 数据分析与知识发现 ,Data Analysis and Knowledge Discovery , 编辑部邮箱 ,2019年09期
- 【分类号】TP391.1
- 【被引频次】4
- 【下载频次】281