节点文献

基于加权关联模式挖掘与规则后件扩展的跨语言信息检索

Cross-Language Information Retrieval Based on Weighted Association Patterns and Rule Consequent Expansion

  • 推荐 CAJ下载
  • PDF下载
  • 不支持迅雷等下载工具,请取消加速工具后下载。

【作者】 黄名选卢守东徐辉

【Author】 Huang Mingxuan;Lu Shoudong;Xu Hui;Guangxi (ASEAN) Financial Research Center, Guangxi University of Finance and Economics;Guangxi Key Laboratory of Cross-border E-commerce Intelligent Information Processing, Guangxi University of Finance and Economics;School of Information and Statistics, Guangxi University of Finance and Economics;

【通讯作者】 黄名选;

【机构】 广西财经学院广西(东盟)财经研究中心广西跨境电商智能信息处理重点实验室(广西财经学院)广西财经学院信息与统计学院

【摘要】 【目的】针对自然语言处理中查询主题漂移和词不匹配问题,提出一种基于加权关联模式挖掘和规则后件扩展的跨语言信息检索模型及其算法。【方法】该模型采用新的加权关联模式支持度和基于最大项目权值的项集剪枝策略挖掘频繁项集,利用置信度和相关度评价加权关联规则,根据扩展模型从规则中提取优质扩展词实现规则后件扩展,扩展词与原查询词项组合为新查询再次检索文档得到最终检索结果。【结果】实验结果表明,与单语言检索基准比较,本文检索模型的R-prec和P@10平均增幅分别为42.49%和25.53%;与跨语言检索基准比较,其平均增幅分别为91.87%和64.61%;与现有基于加权关联规则挖掘的跨语言检索方法比较,R-prec和P@10最高平均增幅分别可达93.20%和34.60%。【局限】只进行实验性研究,需要探讨在实际跨语言搜索引擎中的具体应用。【结论】本文检索模型能有效地减少查询主题漂移和词不匹配问题,改善和提高检索性能。

【Abstract】 [Objective] This paper proposes a new Cross-Language Information Retrieval(CLIR) model, aiming to address the issues facing natural language processing, such as query topic drift and word mismatch. [Methods] First, we explored the frequent item-sets with the weighted association patterns and the pruning strategies based on maximum item weight. Then, we used the confidence and relevance degrees to evaluate the weighted association rules, which helped us extract the high quality expansion terms. Finally, we combined the new terms with the original ones to create new queries for the final lists. [Results] Compared with the monolingual retrieval benchmark, the average increases(AIs) of R-prec and P@10 of the proposed model were 42.49% and 25.53%. Our results were 91.87% and 64.61% higher than the cross language retrieval benchmark. Compared to the existing CLIR methods, the maximum AIs of R-prec and P@10 were 93.20% and 34.60%. [Limitations] The proposed model needs to be examined with more cross language search engines. [Conclusions] Our model improves the performance of CLIR.

【基金】 国家自然科学基金项目“基于深度学习和迁移学习的东盟跨语言查询扩展研究”(项目编号:61762006);广西应用经济学一流学科(培育)开放性课题“中国-东盟贸易商务数据挖掘及应用研究”(项目编号:2018MA07);广西(东盟)财经研究中心开放性课题“东盟财经文本大数据关联模式挖掘及其跨语言检索研究”(项目编号:2018DMCJYB08)的研究成果之一
  • 【文献出处】 数据分析与知识发现 ,Data Analysis and Knowledge Discovery , 编辑部邮箱 ,2019年09期
  • 【分类号】TP391.1
  • 【被引频次】4
  • 【下载频次】281
节点文献中: 

本文链接的文献网络图示:

本文的引文网络