节点文献
高频最大交集型歧义字段问题研究
The Research of High Frequent Maximal Overlapping Ambiguity String
【Author】 Li Bin Chen Xiaohe Fang Fang Xu Yanhua School of Chinese Language and Literature, Nanjing Normal University, Nanjing 210097
【机构】 南京师范大学文学院;
【摘要】 交集型歧义是中文分词的一大难题,建立大规模高频最大交集型歧义字段(MOAS)的数据库,对于掌握其分布状况和自动消歧都具有重要意义。本文采用全切分方法,在4亿字人民日报语料上采集严格定义的高频MOAS14906条,随机抽取了相应的1354270条带有上下文信息的MOAS实例进行人工判定。数据分析表明,大多数真歧义MOAS存在着强势切分现象,词表词字段也应纳入MOAS的探测范围。
【Abstract】 Overlapping ambiguty is still an open issue in Chinese word segmentation. This paper performs a thorough investigation on Maximal Overlapping Ambiguity String (MOAS) which is almost context free. By word omni-segmentation method., we collect 14906 most frequent MOASs from People’s Daily corpus which contains about 400M characters. For the extracted MOASs, 1354270 sentences that contain them are randomly selected and manually labeled. The results show that most MOASs with true ambiguity have a bias towards one segmentation and that lexicon string should be considered when detecting MOASs.
【Key words】 maximal overlapping ambiguity string; lexicon string; word omni-segmentation; biased segmentation;
- 【会议录名称】 全国第八届计算语言学联合学术会议(JSCL-2005)论文集
- 【会议名称】全国第八届计算语言学联合学术会议(JSCL-2005)
- 【会议时间】2005-08
- 【会议地点】中国南京
- 【分类号】TP391.1
- 【主办单位】南京师范大学、清华大学智能技术与系统国家重点实验室