节点文献

依存语法视域下并列结构的理论与计量研究

Theoretical and Quantitative Investigations of Coordination in Dependency Grammar

【作者】 叶子;

【导师】 刘海涛;

【作者基本信息】 浙江大学 , 外国语言文学, 2023, 博士

【摘要】 并列是语言研究中最常见也是最难处理的现象之一,目前仍存在较多尚无定论的争议。在理论研究方面,一方面现有文献中多种句法分析方式并存,且持不同观点的学者互相之间难以说服对方,而另一方面不采用形式化的句法框架会导致无法清晰地对句法模型的表征能力进行判断。在实证研究方面,大量传统研究采用内省方法;现有的几项基于语料库、针对汉语并列现象的定量研究较为零星和初步,描写不够透彻,方法不够严谨,涉及语料处理部分的交代不明,对范畴判别、计数和统计的方法往往没有详细说明。鉴于此,本文在依存语法的框架下,从理论和定量两个角度考察了四类与并列有关的问题,分别是并列的句法结构、并列词与其他词类的模糊边界、语体对并列词频数的影响,以及同/异类并列。针对第一个问题,本文提出了一套整体性句法分析的方法,分为罗列需要表征的全部句法现象、设定所使用的形式化框架以及制定句法分析黄金标准三步。随后,我们参照该框架,汇总了主要的六种并列句法现象(共有中心词、根节点并列、无/多标记并列、共有从属词、嵌套并列、中心词空缺),并罗列了依存语法框架下所有可用于表达并列句法的形式化手段,提出了一种以尽可能简单的方式表征全部上述现象的分析方案。对于后三个研究问题,本研究使用兰卡斯特汉语普通话语料库(LCMC)进行实证考察。LCMC是一个百万词规模、由4大类15小类语体构成的语料库。我们以通用依存(UD)作为标注方案,将其重新句法分析为Co NLL-U这种富标注格式的依存树库后定量考察了上述并列现象。结果发现:在并列词和介词的模糊边界问题上,“和”“与”“跟”“同”这四个虚词中,并列词用法占比从高到低为:“和”>“与”>“跟”>“同”。其中“和”主要偏向并列词用法,“与”的两种用法各占一半,而“跟”与“同”则几乎大部分情况都不用作并列词。此外,对各词形而言,均仅有不到30%的例子存在潜在歧解语境(PAC),且这些歧解语境大部分可通过消歧手段消除,最后只存在5%左右的真实歧解。结果表明真实文本中确实存在无法通过任何测试判断的歧解,但数量不多,该发现支持了一种弱式的歧解观。从允准潜在并介歧解结构的三个条件来看,连续性是最强的要求;就句法位置而言,潜在歧解结构只出现在主语位置、名词修饰语位置、把字宾语位置以及致使结构的名词性论元位置;就谓词类型而言,只出现在交互性谓词、交互性名词以及派生交互性谓词中。在消歧原因中,潜在并列支本身的因素占到了主要地位,而前后句子中的语篇因素则是次要的。在对各个语体中各类并列词频数的考察中,各具体并列词形中,“和”的数量远高于其他并列词,而剩下的并列词的排序在不同语料库中存在差异。在全库中,三大语义类并列词的频数排序为合取类>析取类>转折类。合取类词的数量远高于后两者,但析取类与转折类的差别较小。通过英汉对比发现,汉语合取类并列词在新闻、学术两类书面语体中,数量大于英语;至于小说语体,若以不同前人研究作为参照,该数量关系的比较结果存在差异。从与本文可比性更强的一项研究结果来看,汉语在大部分小说类细分语体中也大于英语。此外,绝大部分并列词频数受语体影响较大,无论在语体大类间还是语体大类内的语体小类间,其频数的变异范围非常大。该影响不仅停留于定量层面,甚至可能颠覆一些定性结论。本研究的发现对所有基于语料库词频测量的研究提出了质疑与挑战,我们也在此基础上提出了一些关于如何汇报频数信息、如何进行标准化测量以及跨研究对比的建议。在对同/异类并列现象的考察中,本研究将句法范畴看成一个三维变量,由词性、句法功能和句法层级构成,并考察了各维度上的频数分布,同时探究语体与并列词类型是否影响这些分布。研究发现异类并列所占总数很少,大多在10%以下,且这一结论与语体和并列词类型无关。有一些维度受两者影响,有一些只受并列词类型影响,也有的规律几乎不受两者的影响。并列支句法范畴之间不相似程度越大,数量越少,且大部分并列结构中并列支完全相似。就我们所考察的全部项目中,不存在句法范畴的三个方面完全不同的并列支。在三个维度上,词性不相同的比例最高,其次为句法层级,而所有并列结构中并列支的句法功能均相同。就句法功能大类的总体分布而言,顺序从高到低依次为论元>谓述>修饰语。不过该分布与语体无关,但与并列词的类型有关。就微观句法功能而言,宾语位置的数量占比最高,其次为定语和主语。在同类并列和异类并列中,两者的顺序有所不同。在异类并列中,词性组合的顺序从高到低分别为动形组合>名动组合>名形组合。该分布与语体和并列词类型的相关性都不大。而在并列支句法层级的组合上,所有考察项均为“短语+小句”的组合。我们通过贝哈格尔定律(即重度原则)来解释该现象。而在同类并列中,并列支词性的分布从高到低分别为名词性>动词性>形容词性。该分布与语体关联不大,但与并列词类型存在强相关。“和”类虚词强烈倾向于名词性并列支,“并”类虚词倾向于动词性并列支,且也有一部分为形容词性并列支,而“而”类虚词强烈倾向于使用形容词性并列支。这也符合我们对三类并列词的一般认识,即汉语中大致可以区分所谓的名词性并列词、动词性并列词以及形容词性并列词,或者可以说存在类型学上所谓的专用标记。就句法层级而言,绝大部分均为短语,极少部分为小句。最后,就同类并列中句法功能和词性的组合而言,以名词性论元和动词性谓语为最高频以及次高频的项目。不过该组合也与语体无关,与并列词类型有关。有趣的是,形容词性并列支并非充当修饰语最多,而是充当谓语最多,这与Croft的经典二维词性模型的预测相悖。上述关于并列的难题背后都涉及语言学中几类具有代表性的争议问题。为此,不同于先前语言研究中通常采用的解释角度,我们提出通过几个与测量相关的概念(如测量精度、标准化测量方式、单位等)来解释上述现象,并探讨了测量在语言研究中的重要性。本文具有若干项创新与贡献:在语言本体研究上,使用基于依存树库的方法进行了精确的信息提取与定量考察,弥补了一些关于并列现象在并列词频数测量、并列的句法结构以及并列支的特点等方面原本所存在的空缺,加深了对并列现象的理解。在研究方法上,在理论研究方面,本文提出了一种整体性句法分析方法,分为三步,旨在为理论句法分析提出一种高效而全新的思路。在实证研究方面,验证了基于树库的语言研究方法比传统基于语料库的方法更高效和精确。从研究的技术手段上而言,本文在语料预处理部分,基于统计方法的句法标注器与基于具有语言学意义的规则的后处理相结合以达到更高的标注准确率。二是研究过程中开发了“通用依存Excel平台”(UDEP)。其涵盖文本预处理、快速查找、条件查找、自动和手动标注、查改错、抽样、统计、计量指标的计算等。该平台旨在为未来的语言研究提供一个强大、高效、精确的工具。在对语言现象的解释上,我们提出从测量的角度来解释一些争议现象,并尝试构建计量语言学的新分支——语言测量学。测量几乎是所有经验科学的第一步,重视语言中的测量问题对于推动整个学科发展具有基础性作用。文中为测量中的若干基础概念(如测量尺度、操作化、单位等)在语言学中寻找对应的类比,提出句法分析也是一种测量等乍一看反直觉的观念,并进行了论证。

【Abstract】 Coordination is one of the most important phenomena and toughest problems in language studies.There are still many controversies that remain unsettled.In terms of theoretical research,we point out that the previous research only lists part of the evidence of the phenomenon at a time,which leads to the existence of multiple syntactic analyses,and neither side is able to convince the other.In addition,not using a formalization leads to the inability to clearly judge the representational power of the syntactic model.In terms of empirical research,most traditional studies adopt introspective methods,while the existing corpus-based,quantitative studies on Chinese coordination are sporadic and preliminary.The descriptions are mostly shallow,and the methods not rigorous enough.Moreover,the preprocessing procedure of most previous investigations are not specified,with the category annotation,counting and statistical methods often not offered in detail.In view of this,this dissertation examined four types of coordination-related phenomena,namely,the syntactic structure of coordination,the fuzzy boundaries between coordination and other phenomena,the measurement of the frequency of coordinators,and the phenomenon of(un)like coordination from the perspective of dependency grammar.For the first question,this dissertation proposed the method of Holistic Syntactic Analysis,which is divided into three steps: listing all the syntactic phenomena that need to be characterized,setting the formal framework used,and formulating the gold standard for syntactic analysis.Subsequently,following the new method,we summarized six main syntactic phenomena related to coordination,listed all the formal strategies for representation within the framework of dependency grammar,and proposed a solution to represent all the above-mentioned phenomena in the simplest way.For the latter three research questions,we used the Lancaster Corpus of Mandarin Chinese(LCMC)for empirical quantitative investigation.The corpus is a million-word corpus consisting of 4 genres and 15 sub-genres.We then adopted Universal Dependencies(UD)as the annotation scheme,and the Co NLL-U format dependency treebank as our material,which is richly annotated,to study the above-mentioned coordination phenomenon quantitatively.Results show that:In terms of the fuzzy boundary coordinators and adpositions,for the four hé-type word forms(hé,yǔ,gēn and tóng),the proportion of coordinator usage from high to low is: hé > yǔ >gēn > tóng.Among these,the use of hé is biased towards coordinators,both uses of yǔ take about half of proportion,and gēn and tóng are barely used as coordinators.In addition,the study found that 7%-30% of the cases have potential ambiguity contexts,and most of these potential ambiguity contexts could be eliminated by disambiguation,and finally there is only about 5% of the cases where true ambiguity exists.The results indicate that there are indeed ambiguities in real texts that cannot be judged by any test,but the number is small,which supports a weak view of ambiguity.With respect to the three conditions for licensing ambiguity constructions,continuity is the strongest requirement.As far as syntactic positions are concerned,potential ambiguity contexts(PAC)only appear in the subject position,noun modifier position,the object position of ba-construction,and the nominal argument position of the causative construction.As for predicate type,PAC only exist in interactive predicates,interactive nouns,and derived interactive predicates.Among the reasons for disambiguation,the coordinand itself plays a major role,while the discourse factor in the preceding or following sentences is secondary.In term of the investigation of the frequency of various types of coordinators in different genres,the number of hé is much higher than other coordinators,while the order of the remaining coordinators is different in each corpora.In the whole corpus,the frequency order of the three major semantic types is conjunctive > disjunctive > adversative.The number of the conjunctive type is much higher than the latter two,but the difference between the disjunctive and adversative type is small.Through the comparison between English and Chinese,it is found that the frequency of Chinese conjunctive coordinators in news and academic written styles is higher than that in English;as for the fiction style,taking different previous studies as the reference point leads to different quantitative relationship.Based on the results of a study that is more comparable to this study,Chinese is also greater than English in most of the subgenres of fiction.In addition,the frequency of most coordinators is greatly affected by the genre,and it varies to a large extent whether between the genres or between subgenres within each genre.The impact is not only at the quantitative level,but may even challenge some qualitative conclusions.The findings of this study question all studies based on corpus-based word frequency measurements.On this basis,we also put forward some suggestions on how to report frequency information,how to standardize measurements,and how to do cross-study comparisons.In terms of the(un)like coordination,the present study regards the syntactic category as a three-dimensional variable composed of part of speech,syntactic function and syntactic level,and examines the frequency distribution on each dimension.Meanwhile,we explore whether the genre and the type of coordinators affect these distributions.The study found that the total number of unlike coordination is small,mostly less than 10%,and this conclusion has nothing to do with genre and coordinator type.Some dimensions are affected by both,others only affected by the type of coordinators,and still others hardly affected by either.We also found that the greater the degree of dissimilarity between the syntactic categories of coordinands,the lower the frequency.In other words,most coordinands are completely similar.In terms of all items we have examined,there is no case where all three aspects of the syntactic category are totally different.Among the three dimensions,the proportion of different parts of speech is the highest,followed by the syntactic level.The syntactic functions of the coordinands in all coordinate constructions are the same.As far as the overall distribution of syntactic functions is concerned,the order from high to low is argument > predicate > modifier.However,the distribution has nothing to do with genre,but with the type of coordinators.In terms of micro-syntactic functions,the number of object positions takes the highest proportion,followed by attributives and subjects.The orders are different between the cases of like and unlike coordination.In unlike coordination,the order of part-of-speech combinations from high to low is N-A combination > N-V combination > N-A combination.The distribution has nonsignificant correlation with genre and coordinator types.On the combination of syntactic levels of coordinands,all show the pattern of "phrases + clauses".We explain this phenomenon by the Behagel’s law,i.e.,the principle of heaviness.In like coordination,the distribution of the parts of speech of the coordinators from high to low is nominal > verbal > adjectival.This distribution also has little correlation with genre,but shows a strong correlation with coordinator types.The hé type coordinators tend to be in conjunction with nominal coordinands.The bìng type coordinators favors verbal coordinands,while some of them are used with adjectival ones.Finally,the ér type is strongly inclined to the adjectival coordinands.The finding is in line with our general understanding of the three types of coordinators.That is to say,Chinese distinguishes between so-called nominal coordinators,verbal coordinators,and adjectival coordinators.To put it differently,there are so-called dedicated markers in linguistic typology.As far as the syntactic level is concerned,most of them are phrases,and very few of them are clauses.Finally,we examined the combination of syntactic function and part of speech in like coordination,nominal arguments and verbal predicates are the most and the second most frequent combinations.However,this combination has no correlation with genre,but with coordinator types.Interestingly,the adjectival coordinands do not function as modifiers but as predicates in most cases,which is contrary to the prediction of Croft’s classical two-dimensional model of parts of speech.All problems mentioned above involve some common controversies in linguistics.To this end,we point out that these deficiencies may be due to a failure in clear measurement,the first step of all empirical scientific research.Since the concept of measurement has long been neglected in linguistics,this dissertation proposes a glottometric(linguistic measurement)theory about language with reference to the measurement theories from other disciplines,and explains the previous phenomenon from this new perspective.The present dissertation has several innovations and contributions:In terms of the language studies per se,through accurate quantitative investigation,we have filled some gaps in the field of coordination in terms of coordinators frequency measurement,syntactic structure of coordination,and the characteristics of coordinands.In doing so,our findings have deepened our understanding of this crucial phenomenon.Theoretically,the present dissertation proposes the method of Holistic Syntactic Analysis,which is divided into three steps,aiming to propose an efficient and new idea for theoretical syntactic analysis.Conceptually,we try to construct a new branch of quantitative linguistics-glottometrics.Measurement is the first step in almost all empirical sciences,and it makes a fundamental contribution to advancing the entire discipline.We find corresponding analogies in linguistics for some basic concepts in measurement(such as scale of measurement,operationalization,units,etc.),propose that syntactic analysis is also a kind of measurement,which at first glance seem counterintuitive,and offer arguments for it.Technically,we have also made several contributions.First,on the basis of the syntactic parsers based on statistical methods,rules with linguistic significance are used for post-processing,and higher labeling accuracy is achieved by combining statistics and rule methods.We try to use this method for the corpus preprocessing part.Second,for the purpose of conducting research in this dissertation,we develop the Universal Dependency Excel Platform(UDEP).Its functions cover text preprocessing,quick query,conditional query,automatic and manual annotation,error detection,sampling,statistics,calculating quantitative indicators,etc.The platform aims to provide a powerful,efficient,and precise tool for future language research.

  • 【网络出版投稿人】 浙江大学
  • 【网络出版年期】2025年 08期
  • 【分类号】H0-0
节点文献中: 

本文链接的文献网络图示:

本文的引文网络