节点文献
基于语言特征增强的汉-缅平行句对抽取方法
Method for Extracting Chinese-Burmese Parallel Sentence Pairs Based on Language Feature Enhancement
【摘要】 针对低资源语言平行句对抽取中标注资源稀缺、模型表征能力不足的问题,本文提出一种基于语言特征增强的汉-缅平行句对抽取方法。该方法从数据增强、模型架构、训练机制3方面进行优化:首先,基于孪生网络构建汉语与缅甸语双编码器以形成跨语言语义表示空间;其次,引入基于词向量L2范数的信息量评估机制,对高信息特征进行替换与样本增强,以缓解低资源下的数据稀疏问题;最后,通过正负样本构造与对比学习的动态建模,优化样本边界,实现更精准的汉-缅语义对齐。实验表明,所提方法在汉-缅平行句对抽取任务上F1值达95.03%,优于基线模型。此外,该文构建了5×105句对规模的高质量汉-缅通用的数据集,为低资源语言相关研究提供数据支撑。
【Abstract】 To address the scarcity of labeled resources and the limited representational capacity of models in extracting parallel sentence pairs in low-resource languages, this paper proposed a language-feature-enhanced method for Chinese-Burmese parallel sentence pair extraction. The method was optimized from three aspects: data augmentation, model architecture, and training mechanism. First, a Chinese-Burmese dual encoder based on a Siamese network was constructed to build a cross-lingual semantic representation space.Second, an information-content evaluation mechanism based on the L2 norm of word vectors was introduced to replace high-information features and perform sample augmentation,thus alleviating the data sparsity problem under low-resource conditions. Finally, positive and negative samples were constructed and dynamically modeled through contrastive learning to optimize sample boundaries and achieve more accurate Chinese-Burmese semantic alignment. Experimental results show that the proposed method achieves an F1 score of95.03% on the Chinese-Burmese parallel sentence pair extraction task, outperforming the baseline model. In addition, this paper constructs a high-quality general-domain ChineseBurmese dataset containing 5 × 105 sentence pairs, providing data support for research on low-resource languages.
【Key words】 parallel sentence pair extraction; information augmentation; contrastive learning; Siamese network;
- 【文献出处】 应用科学学报 ,Journal of Applied Sciences , 编辑部邮箱 ,2026年03期
- 【分类号】TP391.1;TP18
- 【下载频次】9