节点文献

蒙古文编码转换研究

A Study on Coding Conversion of the Mongolian

【作者】 巩政

【导师】 高光来; 杨玉明;

【作者基本信息】 内蒙古大学 , 计算机技术, 2008, 硕士

【摘要】 蒙古文信息处理开端于上个世纪70年代末期,随着计算机技术在蒙古文信息处理中的应用,许多研究单位在蒙古文文字处理方面都取得了一些重要的进展。进行蒙古文信息处理的研究工作,首先要解决蒙古文编码、蒙古文输入、蒙古文字库等关键基础技术问题。就蒙古文编码而言,由于研究工作相对独立,且国家没有及时制定蒙古文编码统一标准,各研究单位一般都采用了自定义的基于字形的蒙古文编码系统。在1993年,ISO/IEC10646国际编码标准中才定义了蒙古文国际标准编码。蒙古文信息处理的研究工作最先是在文字排版方面展开的,由于文字排版系统对文字而言比较关注的是文字的“形”,一个单词只要能够出现正确的形状即可。因此基于形码的蒙古文编码方案也应运而生。蒙古文显现字符中普遍存在“一形多音”现象,并且有些组成一个字符的部分结构,在其它多个字符中都可能重复出现,不同的研究单位在制定各自的形码方案时有的采用一个字符,只定义一个编码,但可以表示多个不同发音的字母;有的采用一个字符定义多个编码,相同字形编码不同,可表示不同发音的字母;有的采用将多个字母中都会出现的部分结构,重新定义为一个“字符”或从文字书写的习惯和美观角度出发,将字母中的部分笔画进行了重组,并为每一个“字符”定义一个编码。随着蒙古文信息化的不断深入,人们开始逐渐意识到蒙古文编码差异造成的问题。由于蒙古文编码系统的互不兼容,经常导致技术上的重复开发,在不同编码系统上开发的信息资源无法共享,造成人力、物力和财力的极大浪费。本文主要讨论蒙科立蒙古文编码、智能蒙古文编码、赛音蒙古文编码和蒙古文国际标准编码的转换问题。这里提到的蒙古文特指传统蒙古文,而不包括托忒文、锡伯文、满文和阿礼嘎礼字符。蒙科立蒙古文编码、智能蒙古文编码、赛音蒙古文编码采用的是基于UNICODE的形码编码方案,转换后的蒙古文标准编码拟采用正在报批过程中的蒙古文国家标准编码。整个编码转换工作分三个步骤进行。第一步:分析编码特征,制定编码转换规则,由计算机程序实现编码初步转换。第二步:建立蒙古文词典库,用来校对转换单词的准确性。第三步:建立平行语料库,补充词典库词汇量不足问题,进一步校准不确定的编码转换。

【Abstract】 The processing of information in Mongolian started in the late 1970s. With the application of the computer technology to the information processing in Mongolian, some significant improvements have been made by quite a lot of research units in terms of the Mongolian word processing. If the work of the information processing in Mongolian is to be carried out, such fundamental technological problems have to be solved as the Mongolian coding, the Mongolian inputting, and the Mongolian word bank. As for the Mongolian coding, due to the relative independence of the research work, all the research units have respectively adopted the Mongolian coding systems based on the word shapes, for the state hasn’t set any unitary standards timely. Only in 1993 was the international standardized coding defined in the international coding standards ISO/IEC 10646.The research work of the information processing in Mongolian was initially started in the aspect of the word composition. Because what the word composition system concerns more is the ’shape’ of the words, it serves its purpose when only a word can appear in the right shape. Therefore, the Mongolian coding scheme was created based on the shape code. Such a phenomenon is quite common in the Mongolian characters that one shape often has multiple pronunciations. Besides, some partial shapes which compose some characters can appear repeatedly in many other characters. Thus, when stipulating their own shape coding, different research units have adopted various schemes. For instance, some adopt the scheme of one coding for one character with letters liable to have different pronunciations, while others favor the scheme of one character for more than one coding with the same character but different coding. Still some others take the scheme of using partial compositions in several letters to redefine a character, or recombine some strokes in letters for the sake of writing habits or esthetic sense and define one coding for each character.With the further development of the information in the Mongolian language, people gradually realize the problems brought about by the differences in the Mongolian coding. Since different Mongolian coding systems are not containing mutually and many information resources of the different coding systems cannot be shared, a lot of human resources, material and financial resources have been wasted for they are often repeated technological exploitations.This paper mainly discusses the conversion between Menkeli Mongolian coding, Oyuta Mongolian coding, Saiyin Mongolian coding and Mongolian international standardized coding. The Mongolian in this paper especially refers to the traditional Mongolian, excluding others like TODO,SIBE,MANCHU and Ali Gali. Menkeli Mongolian coding, Oyuta Mongolian coding, Saiyin Mongolian coding have adopted the shape coding scheme based on Unicode while the Mongolian standardized coding after conversion intends to adopt the Mongolian state-standardized coding which is being reported for official approval. The whole coding conversion can be conducted in three steps. First, analyze coding characteristics, stipulate coding conversion rules, and then achieve the preliminary coding conversion by the computer procedure. Second, establish the Mongolian dictionary bank to check the accuracy in converting words. Third, a parallel language data base is established to replenish the vocabulary of the dictionary bank and to adjust uncertain coding conversion.

【关键词】 蒙古文字符编码编码转换
【Key words】 Mongoliancharacter codingcoding conversion
  • 【网络出版投稿人】 内蒙古大学
  • 【网络出版年期】2009年 04期
  • 【分类号】TP391.1
  • 【被引频次】10
  • 【下载频次】576
节点文献中: 

本文链接的文献网络图示:

本文的引文网络