节点文献

移动营销领域的文本相似度计算方法

Text similarity calculation method for mobile marketing

  • 推荐 CAJ下载
  • PDF下载
  • 不支持迅雷等下载工具,请取消加速工具后下载。

【作者】 孙志远王伟马迪毛伟

【Author】 SUN Zhiyuan;WANG Wei;MA Di;MAO Wei;University of Chinese Academy of Sciences;Computer Network Information Center, Chinese Academy of Sciences;KNET Corporation Limited;

【机构】 中国科学院大学中国科学院计算机网络信息中心北龙中网(北京)科技有限责任公司

【摘要】 针对移动营销文本中长度偏短、用词多变、语句残缺等问题,提出了在文本表示过程中采用word2vec进行词项加权语义映射的方法。首先在全语料库中采用word2vec训练词向量,对整体词向量进行聚类操作来汇聚相近语义词语形成语义簇特征空间,在文本向量化过程中,将词语与聚类中心的相似度和词语本身权重结合完成特征权值计算,向量化之后的文本采用欧氏距离计算相似度。将该算法应用于移动营销短文本测试集,通过K近邻(KNN)分类实验表明,该方法在分类性能上比基于词统计特征的方法在各类的F1值有平均6%的提升,能够更有效地衡量移动营销类别短文本的相似度。

【Abstract】 In this paper, the authors proposed a weighted semantic mapping method based on word2 vec in the short text representation process, aiming at the shortness of text length, the variability of words and the incomplete sentences in mobile marketing text. Firstly, word2 vec was used in the whole corpus to train the word vector, and the whole word vector was clustered to form semantic cluster feature space by similar semantic words. In text vectorization process, feature weights were calculated using similarity between the word and the cluster center integrate with weight of the word itself. The similarity of the text after vectorization was calculated by Euclidean distance. The K Nearest Neighbor( KNN) classification experiments show that this method has a 6% improvement on average F1 value compared to word-based statistical method and is more effective in measuring the short text similarity of mobile marketing.

  • 【文献出处】 计算机应用 ,Journal of Computer Applications , 编辑部邮箱 ,2017年S1期
  • 【分类号】TP391.1
  • 【被引频次】10
  • 【下载频次】191
节点文献中: 

本文链接的文献网络图示:

本文的引文网络