节点文献

基于近似文本分析的意见挖掘

Research on Approximate Text Analysis Based Opinion Mining

【作者】 刘健

【导师】 吴耿锋;

【作者基本信息】 上海大学 , 控制理论与控制工程, 2007, 博士

【摘要】 本文对如何将部分解析技术引入意见挖掘,以提高分析的有效性问题进行了研究,其成果概括如下:(1)提出了基于部分解析的超解析方法及其增量式实现近似文本分析(ATA)(见第2章)。超解析通过放宽归约的严格程度(即允许非紧邻的成份进行归约,允许一个语言成份参与多个归约关系),从而最大限度地寻找关于给定文本(或文本片段)的各种可能解释。为了优化归约时对有效语言成份组合穷举的过程,提出了候选者队列算法(CLA)。作为超解析的实现,本文提出近似文本分析,其系统由两部分构成:分析部件与广义归约部件。分析部件以缓冲队列作为核心的数据结构,从而将超解析的问题转化为了广度优先的搜索问题。而广义归约部件是CLA算法的增量式实现,负责语言成份的归约。(2)提出了基于“近似文本分析”的情感分类方法ATA-SC,及其软件实现ATAFilter(见第3章)。ATA-SC方法考虑了实体词汇与情感词汇之间的语义关联,因此对于情感的识别能力要强于基于单对象假设的情感分类方法。而情感分类模块ATAFilter已集成于邮件过滤软件VIHunter中,在技术测试中展示了良好的性能,同时在实际应用中也获得令人满意的效果,取得了较好的社会效益。(3)提出了一种新的意见抽取任务即意见实例抽取(OIE),及其解决方法FC-OIE;提出了基于位置线索的语义关系识别(SARPC)方法,用于在FC-OIE中识别对象与特性之间语义关联(见第4章)。意见实例抽取任务的目标是保持意见表达的数据结构与源文本之间的关联,使得我们可以通过考察意见元组中各构成要素在原文中的地位,来发掘更深层次的信息。为了解决这一新的抽取任务,FC-OIE采取的策略是:通过SARPC方法为每个特性实例寻找语义关联最强的对象实例,构成“对象实例-特性实例”对偶;对于每个对偶,通过ATA-SC对所含的对象实例与特性实例周围的文本进行情感分析,判断语义方向。(4)提出并实现了用于意见实例抽取与检索的意见搜索系统(OSS)(见第5章)。OSS的目的是从网络评论中抽取意见实例,并根据用户的检索兴趣进行反馈。该系统通过网络爬虫从互联网上抓取评论网页,通过文本清洗得到正文;然后以FC-OIE技术从文本中抽取意见实例,构成意见库;最后通过人机交互将意见库中的信息直观地反馈给系统的用户。

【Abstract】 In this paper, in-depth research on promoting the effectiveness of sentiment analysis in opinion mining by means of partial parsng is described. The achievements of the paper are as follows:(1) A partial-parsing-based method and its incremental implementation are proposed. The new parsing method is named Super Parsing, whose purpose is to seek all the possible interpretation for a given text (or text segment) by relaxing the constraints on reducing. It allows non-adjacent constituents to be merged, and allows one constituent to join multiple reduction relations. To optimize the enumerating process of eonsitutent combinations, the Candidate List Algorithm (CLA) is proposed. Approximate Text Analysis (ATA) is the incremental implementation of Super Parsing. It consists of two parts: Analyzing Component and Global Reduction Component. Analyzing Component takes the buffer queue as core data structure, which converts the Super Parsing problem into the Breadth-first Searching problem. While Global Reduction Component is the incremental implementation of CLA for constituent reduction.(2) A novell sentiment classification algorithm and its software implementation are proposed. The new algorithm is named ATA-based Sentiment Classification (ATA-SC). ATA-SC considers the semantic relationship between entity words and sentiment-relevant words, so its recognizing ability is better than traditional methods based on single subject hypothesis. And ATAFilter is the ATA-based sentiment classification module. This module has been intergated into the Mail Filtering Software VIHunter, and achieved resonable effect in both testing and application.(3) A new opinion extraction task and its solution are proposed. The new task is named Opinion Instance Extraction (OIE). It keeps the association between opinion storage and the source text, so that more context information can be utilized in future mining task. To solve the task, the algorithm Featurecentered Opinion Instance Extraction (FC-OIE) is proposed. It takes two steps: (ⅰ) for each feature instance (FI), to find the most semantically related subject instance (SI) with SARPO approach and make a "SI-FI" pair; (ⅱ) to determine the sentiment of each "SI-FI" pair by casting ATA-SC on the text segment near SI and/or FI.(4) A system for opinion instance extraction and retrieving is proposed and implemented, named Opinion Searching Systen (OSS). The system extracts opinion instances with FC-OIE from web reviews, and feeds them back via human-machine interface to users according to their searching interest.

  • 【网络出版投稿人】 上海大学
  • 【网络出版年期】2008年 04期
节点文献中: 

本文链接的文献网络图示:

本文的引文网络