节点文献

现代维吾尔语语料库词频统计实验性研究

The Experimental Study on the Corpus Based Word Frequency Statistics of Modern Uighur Language

【作者】 毕丽克孜

【导师】 哈力克·尼亚孜;

【作者基本信息】 新疆大学 , 少数民族语言文学, 2003, 硕士

【摘要】 近年来语料库语言学的发展较为迅速,它为语言研究打开了一条新的道路。英语、汉语等语言语料库语言学及词频统计研究为少数民族语言语料库的不同层面进行定量研究奠定了可靠的,坚实的基础和借鉴的经验。维吾尔文信息处理技术的发展和维吾尔语研究的成果为开展维吾尔语语料库研究和进行词频统计创造了条件。 本文在介绍语料库语言学理论和方法的基础上,以一篇维吾尔文中篇小说为语料,探索计算机自动识别和处理维吾尔文浯料的途径,阐述与维吾尔语词频统计技术相关的具体步骤与方法并公布实验性研究结果的词类自动标注文本、词频统计表及统计结果分析。本项实验性研究为计算机自动处理维吾尔语语料,为建立大容量的语料库奠定基础。本项研究所选用的语料样品经数据库和语料库的连接自动标注词类,得出的词频统计结果如下:语料样品中使用的词为1823个,总词次为11771。论文包含每一个词的出现频度、频率。本项实验性研究证明,运用语料库语言学的方法,编制反映维吾尔语的特点的语料库自动处理程序,实现计算机自动识别并统计维吾尔语语料是完全可以的。

【Abstract】 The rapid development of corpus linguistics in recent years has opened up a new way of studying languages. The corpus based research achievements of many languages such as English, Chinese in corpus linguistics and word frequency statistics laid a reliable, solid foundation and the experience for reference to minority languages to go on quantitative analysis of different fields of corpus. The development of Uighur word processing technology and language research achievements has created good condition to promote the research work on the corpus and word frequency statistics of Uighur language. Based on the theories and skills of Corpus Linguistics, this paper tries to find out the ways and methods of automatic recognition and computer processing of Uighur Corpus by using one of the medium-length novel of modern uighur language and states the related procedures and particular techniques and announced preliminary research achievements such as automatic word-classtagged text, word frequency table and the statistic analyse of them-This experiment will help to the automation of processing of Uighur corpus and set the foundation of making large corpus. Through connecting corpus with database, and automatic word-class tagging text, the author draw the conclusion that the whole corpus consisted of total 11771 words and 1823 stems, and calculated the word frequency and word occurrence of each word. This experiment prove that it is ML possible to bring about automatic corpus recognition and statistics of uighur language by combining the methods of corpus with the programs of reflecting the characteristics of uighur language .

  • 【网络出版投稿人】 新疆大学
  • 【网络出版年期】2003年 04期
  • 【分类号】H215
  • 【被引频次】23
  • 【下载频次】716
节点文献中: 

本文链接的文献网络图示:

本文的引文网络