节点文献

基于深度主动学习的柬语单文档抽取式摘要方法

EXTRACTIVE SUMMARY OF KHMER SINGLE DOCUMENT BASED ON DEEP ACTIVE LEARNING

  • 推荐 CAJ下载
  • PDF下载
  • 不支持迅雷等下载工具,请取消加速工具后下载。

【作者】 余兵兵严馨周枫徐广义莫源源

【Author】 Yu Bingbing;Yan Xin;Zhou Feng;Xu Guangyi;Mo Yuanyuan;Faculty of Information Engineering and Automation, Kunming University of Science and Technology;Yunnan Provincial Key Laboratory of Artificial Intelligence, Kunming University of Science and Technology;Yunnan Nantian Electronic Information Industry Co., Ltd.;School of Southeast & South Asia Languages and Culture, Yunnan Minzu University;Institute of Language Studies, Shanghai Normal University;

【机构】 昆明理工大学信息工程与自动化学院昆明理工大学云南省人工智能重点实验室云南南天电子信息产业股份有限公司云南民族大学东南亚南亚语言文化学院上海师范大学语言研究所

【摘要】 深层神经网络在文档摘要方面取得了很好的效果,其优势只有在大数据集下才能显示出来。为了解决在使用深度学习做柬语单文档抽取式摘要时语料标注不足的问题,提出一种将主动学习和深度学习相结合的方法。利用主动学习抽样策略选择出定量的文档,通过专家标注,结合深度学习中编码器解码器模型进行训练模型抽取得到摘要。实验结果表明,在训练语料显著标注不足的情况下,该方法能够有效地提升柬语单文档摘要的质量。

【Abstract】 The deep neural network has made a lot of progress in document summarization, and its advantages can only be displayed under the big dataset. In order to solve the problem that the Khmer uses the deep learning to make the single document extraction abstract corpus insufficient labeling, a method combining active learning and deep learning is proposed. The active learning sampling strategy was used to select the quantitative documents, marked by the experts, then combined with the encoder decoder model in deep learning, and the training model was extracted to obtain a summary. The experimental results show that even if the training corpus is not markedly marked, the result of extracting the abstract can effectively improve the quality of the Khmer single document abstract.

【基金】 国家自然科学基金项目(61562049,61462055)
  • 【文献出处】 计算机应用与软件 ,Computer Applications and Software , 编辑部邮箱 ,2021年04期
  • 【分类号】TP391.1;TP18
  • 【下载频次】109
节点文献中: 

本文链接的文献网络图示:

本文的引文网络