节点文献
一种词聚类LDA的商品特征提取算法
An Algorithm Based on Words Clustering LDA for Product Aspects Extraction
【摘要】 商品评论中经常会使用一些词义近似或上下文相关的中低频词来描述商品特征,如何有效辨识这些中低频词是商品特征抽取的一个难点.由于缺乏先验知识,主题模型难以发现并抽取中低频特征词.提出基于词义相似度和上下文相关度相结合的词聚类度量算法,在此基础上构建了一种基于词聚类先验知识的潜在狄利克雷分配的商品主题特征提取模型.首先对词项按词义相似度、上下文相关度进行聚类;然后在商品主题特征抽取中引入词聚类因素作为权重影响因子,使得同一个聚类簇中的词项属于同一主题的概率增加.相关实验结果表明,本文提出的词聚类和特征提取算法具有较好的效果.
【Abstract】 Product reviews often use some low-frequency synonyms or context-dependent words to describe the product aspects,and howto effectively identify these low-frequency words is a difficult problem in aspect extraction. Due to the lack of prior knowledge,it is difficult to find and extract the low-frequency aspect words by topic model directly. This paper proposes a method for word clustering in corpus of product reviews,and it takes semantic similarity and contextual relevance of words into account. Then based on the method we present a topic model by adding word clustering as a priori knowledge into the LDA for aspects extraction,we call it WCLDA. In the process of WC-LDA,word clustering can be implemented according to the distance of each two words calculated by similarity and contextual degree; Secondly,word clustering is introduced as a weighting factor in LDA for aspect extraction,which can increase the probability belonging to the same topic of the words that in the same cluster. Experimental results showthat the word clustering algorithm and WC-LDA model presented in this paper have a better effect.
【Key words】 word clustering; contextual relevance; Latent Dirichlet Allocation(LDA) model; aspect extraction;
- 【文献出处】 小型微型计算机系统 ,Journal of Chinese Computer Systems , 编辑部邮箱 ,2015年07期
- 【分类号】TP391.1
- 【被引频次】23
- 【下载频次】456