节点文献
文本嵌入技术的研究与应用进展
Progress in Research and Application of Text Embedding Technology
【摘要】 [目的]本文对国内外已经发表的自然语言处理领域有关文本嵌入的研究进行较深入的分析和对比,详细描述文本嵌入的知识结构和发展脉络,以及针对不同领域、不同数据集的模型改进方法,讨论流行的嵌入模型,比较每个模型在文本嵌入中的优缺点,同时指出文本嵌入所面临的挑战,提出可能的解决方案。[方法]检索Web of Science数据库、CNKI数据库和万方数据,获取国内外文本嵌入研究的相关文献,运用内容分析法对文献做系统梳理分析,对这些文献中利用的文本嵌入技术以及改进方案、建模思想、生成过程等方面进行对比与分析。[结果]经过去重和合并,保留内容最相关的61篇文献。文本嵌入方法可以归纳为三类:基于频率的文本嵌入、基于神经网络的文本嵌入和基于主题建模的文本嵌入。针对语料库的规模大小、多义词嵌入、通用嵌入的域适应等文本嵌入所面临的挑战,从被调查的研究文章中提出了可能的解决方案。
【Abstract】 [Objective] This article conducts an in-depth analysis and comparison of the research on text embedding and describes the basic model of text embedding and the model improvement methods for different fields and different data sets.Popular embedding models are discussed and the advantages and disadvantages of the models are compared.[Methods] The relevant documents of text embedding research at home and abroad are obtained from the Web of Science database、CNKI database and WanFang database and the text embedding technologies,improvement schemes,and modeling ideas are systematically analyzed.[Results] After deduplication and merging,61 documents with the most relevant content are retained.Text embedding methods can be summarized into three categories:text embedding based on frequency,text embedding based on neural network,and text embedding based on topic modeling.Given the challenges faced by text embeddings such as the size of the corpus,polysemous word embedding,and universal embedding domain adaptation,possible solutions are extrated from the research articles under investigation.
【Key words】 text embedding; natural language processing; content analysis;
- 【文献出处】 数据与计算发展前沿 ,Frontiers of Data & Computing , 编辑部邮箱 ,2023年03期
- 【分类号】TP391.1
- 【下载频次】33