节点文献

基于张量空间及张量网络的语言模型

Language Modeling by Tensor Space and Tensor Network

【作者】 张立鹏

【导师】 张鹏; 张作职;

【作者基本信息】 天津大学 , 软件工程, 2018, 硕士

【摘要】 语言模型是自然语言处理领域的一个非常重要且基础的研究课题,应用在很多自然语言处理任务中,如语音识别,机器翻译,对话系统等。现有的语言模型大体可以分为统计语言模型和神经语言模型两大类,他们各有优缺点,统计语言模型建模要求的参数量过大而难以估计,神经语言模型虽然建模效果较好但又理论性不足。本文提出了一个基于张量空间及张量网络构建的语言模型,并命名为张量空间语言模型(Tensor Space Language Model,TSLM)。在张量空间语言模型中,我们用张量积来捕捉词与词之间的交互信息,并且在此基础上模型能构建一个表达能力更强的语义空间。理论上,我们证明了这样的张量表示是n-gram语言模型的一般形式,另一方面,用张量分解推导了语言建模中的递归的条件概率计算过程。综上,我们提出的基于张量空间的语言模型是一个更一般的语言模型,这意味着统计语言模型和神经语言模型都可以作为张量空间模型的特例。换言之,我们构建的张量空间语言模型统一了统计语言模型和神经语言模型。在语言建模数据集Penn Tree Bank(PTB)和WikiText上的实验验证了张量空间语言模型的有效性。

【Abstract】 Language modeling(LM)is a fundamental research topic in the field of natural language processing.It is applied in many natural language processing tasks,such as speech recognition,machine translation,dialogue system and so on.The existing language models can be divided into statistical language models and neural language models,each of which has its own advantages and disadvantages.The parameters required for statistical language models are too large to be estimated.Although the neural language model has good modeling effect,it is insufficient in theory.In this work,we propose a language model named Tensor Space Language Model(TSLM),by utilizing tensor networks and tensor decomposition.In TSLM,we capture the interaction information between words through tensor product,based on which we can build a high-dimensional semantic space with stronger expressive power.Theoretically,we prove that such tensor representation is a generalization of the n-gram language model.We further show that this high-order tensor representation can be decomposed to a recursive calculation of conditional probability for language modeling.In summary,our proposed tensor space language model is a more general language model,which means that both statistical language model and neural language model can be used as special cases of tensor space language model.In other words,the tensor space language model we constructed unifies the statistical language model and the neural language model.The experimental results on Penn Tree Bank(PTB)dataset and WikiText benchmark demonstrate the effectiveness of TSLM.

  • 【网络出版投稿人】 天津大学
  • 【网络出版年期】2020年 06期
节点文献中: