节点文献
面向文本语料库的数据模型及其查询问题
On Modeling and Querying of Text Corpora
【摘要】 语料库为语言研究和自然语言处理提供基础数据服务.传统语料库数据缺乏规范的数据模型,导致无法科学的评价查询结果,大大降低了数据可用性.针对该问题,提出一种面向语料库的数据模型,并讨论了其上的查询问题.首先,给出语料库数据的形式化定义,其次,在关系模型的基础上提出一种面向文本语料库的数据模型,并证明了模型的完备性,在此基础上,扩展传统语料库以KWIC(Key Word In Context)输出为中心的查询语义,定义了语料库数据的查询问题KWIC-EXTENTION.最后,证明这些查询问题的数据复杂度,其中,正匹配查询、负匹配查询、析取匹配查询、n-临近匹配查询的数据复杂度是AC0的,临近正匹配查询的数据复杂度是PTIME(Polynomial Time)的,临近负匹配查询问题的数据复杂度是PSPACE(Polynomial Space)的.这些结论为语料库数据模型和查询方法的研究奠定了理论基础.
【Abstract】 The representation of a text corpus without an efficient data model causes the intractability of the query evaluation over most of the existing text corpora and thus leads to the reduction of data availability. This article proposes a data model for the text corpora data and also discusses the issues on corpus query. First,a formalized definition of the text corpus data is presented. Second,a data model towards the corpus data is proposed in terms of the relational model,which is also proved to be complete. We extend the query semantics of the traditional corpus query that generates KWIC( Key Word In Context) concordances and firstly defines the query problems. Finally we investigate the data-complexity of these querying problems.
【Key words】 corpus; data model; query; data complexity; relational model;
- 【文献出处】 小型微型计算机系统 ,Journal of Chinese Computer Systems , 编辑部邮箱 ,2015年08期
- 【分类号】TP311.13;TP391.1
- 【被引频次】2
- 【下载频次】230