节点文献

基于语言特征增强的汉-缅平行句对抽取方法

Method for Extracting Chinese-Burmese Parallel Sentence Pairs Based on Language Feature Enhancement

  • 推荐 CAJ下载
  • PDF下载
  • 不支持迅雷等下载工具,请取消加速工具后下载。

【作者】 赵子霄王昊申涛江姝婷张思琦赖华黄于欣余正涛

【Author】 ZHAO Zixiao;WANG Hao;SHEN Tao;JIANG Shuting;ZHANG Siqi;LAI Hua;HUANG Yuxin;YU Zhengtao;Faculty of Information Engineering and Automation, Kunming University of Science and Technology;Yunnan Key Laboratory of Artificial Intelligence, Kunming University of Science and Technology;

【通讯作者】 赖华;

【机构】 昆明理工大学信息工程与自动化学院昆明理工大学云南省人工智能重点实验室

【摘要】 针对低资源语言平行句对抽取中标注资源稀缺、模型表征能力不足的问题,本文提出一种基于语言特征增强的汉-缅平行句对抽取方法。该方法从数据增强、模型架构、训练机制3方面进行优化:首先,基于孪生网络构建汉语与缅甸语双编码器以形成跨语言语义表示空间;其次,引入基于词向量L2范数的信息量评估机制,对高信息特征进行替换与样本增强,以缓解低资源下的数据稀疏问题;最后,通过正负样本构造与对比学习的动态建模,优化样本边界,实现更精准的汉-缅语义对齐。实验表明,所提方法在汉-缅平行句对抽取任务上F1值达95.03%,优于基线模型。此外,该文构建了5×105句对规模的高质量汉-缅通用的数据集,为低资源语言相关研究提供数据支撑。

【Abstract】 To address the scarcity of labeled resources and the limited representational capacity of models in extracting parallel sentence pairs in low-resource languages, this paper proposed a language-feature-enhanced method for Chinese-Burmese parallel sentence pair extraction. The method was optimized from three aspects: data augmentation, model architecture, and training mechanism. First, a Chinese-Burmese dual encoder based on a Siamese network was constructed to build a cross-lingual semantic representation space.Second, an information-content evaluation mechanism based on the L2 norm of word vectors was introduced to replace high-information features and perform sample augmentation,thus alleviating the data sparsity problem under low-resource conditions. Finally, positive and negative samples were constructed and dynamically modeled through contrastive learning to optimize sample boundaries and achieve more accurate Chinese-Burmese semantic alignment. Experimental results show that the proposed method achieves an F1 score of95.03% on the Chinese-Burmese parallel sentence pair extraction task, outperforming the baseline model. In addition, this paper constructs a high-quality general-domain ChineseBurmese dataset containing 5 × 105 sentence pairs, providing data support for research on low-resource languages.

【基金】 国家自然科学基金(No.U24A20334,No.62366027,No.62266027);云南省基础研究计划重大项目(No.202401BC070021);云南省重大科技专项(No.202402AG050007,No.202303AP140008,No.202502AD080014);昆明理工大学“双一流”建设联合专项(No.202201BE070001-021)
  • 【文献出处】 应用科学学报 ,Journal of Applied Sciences , 编辑部邮箱 ,2026年03期
  • 【分类号】TP391.1;TP18
  • 【下载频次】9
节点文献中: 

本文链接的文献网络图示:

本文的引文网络