节点文献

中文文本体裁的自动分类机制

Automatic Categorization of Text Genre in Chinese

【作者】 方鸷飞

【导师】 林鸿飞;

【作者基本信息】 大连理工大学 , 计算机应用技术, 2005, 硕士

【摘要】 二十世纪末以来,计算语言学很多的文本分类研究者认识到按照形式分类的重要性,并出现了一个重要的理论转向,即由重视内容的分类转而重视内容与形式并重的研究。体裁属于形式的范畴,与写作风格、句法分析联系紧密,对文章的写作有着明显的制约和规范作用。把体裁分类信息附加于信息搜索引擎的方案,可以显著改善其效能。此外,体裁信息用于协助数字图书馆系统的可视化表示。因此,研究体裁自动分类,有着极高的理论价值和深远的现实意义。 然而,如何识别、描述、利用文本体裁是一项复杂而具有挑战性的工作。首先,体裁概念体系很大程度上是人类思维的抽象归纳,研究者认知受限和体裁自身动态演变等因素使得其概括和表述工作相当困难。其次,这个课题交叉于传统的汉语修辞学与计算语言学之间,需要有较深的语言学功底和计算语言学理论基础。因此,在其研究道路上还存在一些必须要克服的障碍。整体来看,体裁分类研究尚处于全面探索阶段的初期,其技术还不够成熟。而且,国内汉语体裁自动分类的研究工作也刚刚起步。 本文参照英语体裁分类机制,提出了一种基于浅层特征的中文体裁自动分类机制。其中,利用样本分类决策选出十三个中文特征项,借鉴模糊隶属度理论接合定性与定量指标,采用支撑向量机技术计算特征值。该分类机制已经在科学体、政论体、诗歌体、公文体、新闻体共五类体裁的典型文本的语料上得到实现,并获得了较好的效果。系统的局限性是特征提取程序缺乏通用性,必须随着体裁分类体系的每一项变化而做大幅度的调整。本课题尽管取得了一些进展,但必竟只是体裁自动分类研究的一个初步尝试,更多后续理论及应用研究尚待完成。

【Abstract】 Since the end of 20 century, genre problem has become one of the hottest research points of computational linguistics and traditional linguistics in the world. Many researchers of text classification in the field of computational linguistics have realized the importance of form classification. Though they have gotten some preliminary accomplishments, genre classification research is still in the initial time of completely exploring stage. Genre is defined as a category assigned on the basis of external criteria, that is, it refers to form of the text. It is connected close with writing style and the analysis of sentence structure. The research of genre automatic classification has very high theoretical value and profound realistic significance.However, how to distinguish, describe and use text genre" is a complex and challenging work. Firstly, its concept system is an abstract summary of human thought. The limited experience and knowledge of researchers and its constant development make it difficult to summarize and describe genre classification roundly, accurately and efficiently. Secondly, genre automatic classification, which intersects between computational linguistics and the Chinese traditional rhetoric, needs to have the deeper theoretical foundation of linguistics theory and computational linguistics. Therefore, there is still some obstacles that must be surmounted on its research road.The major contribution of this paper is to put forward the automatic system of Chinese text genre classification. This system is divided into three parts, those are corpus collection, feature items selection and classification algorithm realization. It has already got achievement on corpus of scientific type, political comment type, poem type, official document type and news type, the typical texts of five kinds of genre. Compared with foreign related research, the number of our chosen feature terms is fewer, but classification precision and results are good enough. Modular design has facilitated systematic adjustment and test greatly. The program of feature choosing being short of flexibility is limitation of the system. And it must be changed along with each change of genre classification system. Though this paper has make some progress, actually it is surely only one preliminarily try of genre automatic classification research.

  • 【分类号】TP391.1
  • 【被引频次】6
  • 【下载频次】273
节点文献中: 

本文链接的文献网络图示:

本文的引文网络