节点文献
中文WEB文本倾向性分类研究
The Research of Chinese Web Text Orientation Classification
【作者】 单大力;
【导师】 刘云;
【作者基本信息】 北京交通大学 , 信息网络与安全, 2008, 硕士
【摘要】 21世纪是信息爆炸的时代,随着Internet的高速发展,越来越多的信息表现为电子文档的形式,而绝大多数的电子文档都是无序的。文档自动分类技术可以解决信息杂乱现象的问题,并有效地组织和管理这些信息,从而快速、准确、全面地从中找到用户所需要的信息。而文本褒贬倾向性分类更是当前研究的一个热点,是网络舆论安全的重要方面。一些常用的文本分类系统是对电子文档进行内容分类,如将文本分为:军事类,医学类,体育类。这种分类系统无法识别文档的倾向性所在。为了维护网络舆论的安全性,我们开发了Web文本倾向性分析系统,该系统的主要任务是通过对文档内容的分析,给出文档的褒贬倾向性分类,从而直接地确定某网络舆论的支持度,对网络舆论的安全维护起到重要作用。本文讨论了多种应用于文本褒贬倾向性分类的核心技术,分析各种方法的优劣所在。建立了合理的语料库,并依据分类实现的准确性和便捷性原则,用C#语言完成了基于四字哈希词典的中文分词模块,分词准确率超过90%。采用褒贬义特征提取技术,向量空间模型构造技术和简单距离向量分类技术完成Web文档褒贬倾向性分析系统的特征提取模块和倾向性分析模块,其分类精度超过80%。
【Abstract】 21th century is the times when information increases explosively. More and more information exists in the way of electro-document along with the quick development of Internet, and most of these documents are unorderly. Document automatically classifying technique can solve the unoderly information problem, organize and manage the information effectively and help the users acquire the information which they need quickly, accurately and comprehensively. Documents’ orientation text classification is a research hotspot and an important aspect of network consensus security.Some conventional documents classification systems classify the documents by their content, such as: military affairs, medicines, sports. These systems can not identify documents’ orientation. In order to maintain the network’s security, we develop the Web documents’ orientation classification system. Its main task is analyzing the content of documents to classify them by their orientation. It can identify the support of some network consensus and has an important effect of maintaining the network consensus security.This thesis discusses some core techniques used in document orientation classification, and analyses the advantages and disadvantages of them. It Provides logical documents. According to the accuracy and convenience principle, use C# to complete the text participle module based on four-word hashtable dictionary and its text participle accuracy is past 90 percent. We use commendatory and derogatory character pick-up technique, vector space model structure technique and SVM to implement the systems’ documents orientation classification function and its classification accuracy is past 80 percent. At last, we analyze the result of text participle and text orientation classification and bring forward the disadvantages of the system and the future research orientation.
【Key words】 Text Orientation; Text Model; Participle Dictionary; Chinese Participle Mechanism; Text Similarity; Character Pick-up; Text Classification;
- 【网络出版投稿人】 北京交通大学 【网络出版年期】2008年 07期
- 【分类号】TP391.1
- 【下载频次】317