节点文献

三种阅读理解测试题型测量能力的对比研究

A Comparative Study on the Measuring Capacity of Three Reading Item Types

【作者】 张晓

【导师】 唐跃勤;

【作者基本信息】 西南交通大学 , 外国语言文学, 2017, 硕士

【摘要】 人们通常把语言学习分为听、说、读写四种技能。其中,阅读是学生获取语言知识,培养语言能力最直接、最有效的方法。特别是学习一种外语,国内听说环境较差,阅读是主要的信息来源。因此,阅读能力成为衡量一个人语言能力高低的一种非常重要的指标。在国内外各种语言测试中,阅读理解都占有相当大的比重。然而,国内的阅读测试,大多采用单一的测试题型,即多项选择题型,而对其他阅读理解题型的研究和使用都相对较少。因此,导致很多阅读理解测试工作者很难科学地选择恰当的阅读理解题型来取得更为准确有效的测试效果。该研究分别对阅读理解的多项选择题型、句子简化题型以及问答式题型的综合测量能力进行了对比研究。作者运用了格兰特·海宁(Grant Henning)在其著作《语言测试指南:发展,评估与研究》中提出的语言测试统计手段,测量了效度、效度、区分度、难度等决定阅读理解测量能力的因素。其中,结构效度的研究是基于莱尔·巴克曼(Lyle F.Bachman)对阅读能力的划分,然后运用了麦克唐纳(McDonough)的反省理论得出的结果。该论文着力于解决三个问题:1)受试者在完成三组分别含有不同类型的阅读理解测试后,结果是否明显不同?2)多项选择阅读题型、句子简化阅读题型和问答式阅读题型的信度、效度、难度、区分度是否各有不同?如果存在不同点,有哪些不同点?3)哪一种阅读理解题型在测量受试者的综合阅读能力上相对有效可信?基于以上三个问题,本研究对阅读理解多项选择题型、句子简化题型和问答式题型的综合测量能力进行了对比研究。受试者是西南交通大学的60位大学三年级非英语专业的学生,要求受试者在三个不同的时期分别完成三组分别包含三种不同题型的阅读理解测试题。论文作者将测试分数录入电脑,用社会学数据统计软件包SPSS 19.0对所得数据进行了分析。随后,作者又对60名受试者做了一个问卷调查,以获得各个阅读理解题型的表面效度数据。最后,要求受试者根据他们对答题过程的回忆完成一份口头自陈报告,并由论文作者进行录制,以此取得阅读理解测试的结构效度数据。在对各项数据进行分析和比较后,论文作者得出了以下结论:首先,受试者所取得的三种不同题型的阅读理解成绩有显著的不同,多项选择题型式阅读理解最为简单,问答式阅读理解最为困难。其次,基于这对三种题型的阅读理解的信度、效度、区分度和难度等测量能力的对比分析研究,作者得出了这样的结论:问答式阅读理解和句子简化式阅读理解都有较好的区分度和令人满意的信度和效度。然而,一些受试者认为,问答式阅读理解太难,使他们在测试过程中产生了一些焦虑感;但使用最广泛的多项选择式题型在这几个测量能力要素上的数值却都令人不太满意。第三,通过研究,作者得出结论:相比之下,问答式阅读理解题型比另外两种题型能够较为有效地测试出受试者的综合阅读能力。基于以上结论,本文作者建议:阅读理解试题设计者应根据不同的测试目的和用途来选择适合的阅读理解题型;多项选择式阅读理解不适合再作为大型水平测试或选拔性测试的唯一阅读测试形式,而是应该多结合其它如问答式阅读理解这样具有良好测量能力的阅读理解题型。由于实验条件及时间等限制,该研究使用了较小的样本,对不同题型的阅读理解测试的反拨效应也没有涉及,只是对阅读理解的不同题型这一单一变量进行的分析研究,这些都是后期此方面的研究需要改善和提升的地方。

【Abstract】 Language learning is usually divided into four skills,consisting of listening,speaking,reading and writing.Reading is the most direct and effective way to acquire language knowledge and cultivate language ability.Especially for learning a foreign language,because the domestic hearing and speaking environment are poor,reading is the main source of obtaining information.Therefore,reading ability becomes a very important criterion to measure a person’s language competence.Reading comprehension has a large proportion in a variety of language tests.However,most of the domestic reading tests employ a single item type of reading comprehension which is multiple choice question.And the research and use of other item types of reading comprehension are relatively limited.Therefore,it is difficult for many reading comprehension test workers to choose the appropriate item type of reading comprehension scientifically to obtain more accurate and effective test results.This study compared the general capacities of three item types of reading comprehension—MCQ,SSQ,and SAQ respectively.The author employed the statistical method of language testing proposed by Grant Henning in his book A guide to language testing:development,evaluation and research to measure the factors such as difficulty,discriminability,reliability and validity,which determine measuring capacity of reading comprehension test.The study of construct validity is based on Lyle Bachman’s division of reading ability,and using McDonough’s theory of introspection to get the results.The paper focuses on the three questions:1.Are subjects’ performances different w-hen different reading item types are utilized?2.Are there any variances in difficulty,discriminability,validity,and reliability of the three different reading item types?3.Which kind of reading item types can yield a better effective and reliable measurement of subjects’ overall reading competence?In order to answer the above three questions,the author has done a comparative study on the measuring capacities of MCQ,SSQ and SAQ.The subjects are 60 third-grade undergraduates from non-English majors at Southwest Jiaotong University.The subjects were requested to complete three sets of reading comprehension test with identical reading material and different item types in three different time periods.The test scores were input to computer,and the data were analyzed by SPSS 19.0,a statistical software of sociology.Subsequently,the author did a questionnaire survey among the 60 subjects to obtain the surface validity data of each item type.Finally,the subjects were asked to complete oral self-reports based on their memories of the answering process.The author recorded these reports to obtain the construct validity data.After analyzing and comparing the data,the author draws the following conclusions:First,the reading comprehension scores of the three item types are significantly different.And the MCQ is the easiest type,while SAQ is the most difficult type.Second,based on the comparison and analysis of the reliability,validity,discriminability and difficulty of the three item types of reading comprehension,the author draws the following conclusions:SAQ and SSQ both have a good discriminating ability and satisfactory reliability and validity.However,some subjects expressed they had experienced anxiety because SAQ was too difficult for them.However,the data of the most widely used item type MCQ on these measuring factors is not very good.Third,based on the study of data,the author draws the conclusion that SAQ can effectively test the general reading ability of the subjects,comparing with the other two types of reading comprehension test.Based on the above conclusions,the author suggests that reading comprehension test designers should select appropriate item type according to different testing purposes.And MCQ is not suitable for large-scale proficiency test or selective test as the only item type in reading comprehension test.In other words,it should be combined with other item types with good measuring capacities such as SAQ.Due to the limitation of experimental conditions and time,the study employed a small number of samples,and the washback effect of different item types of reading comprehension is not involved.And the author only confined one variate to research which is the item type.These are the parts need to be improved in the future research.

【关键词】 信度效度难度区分度测量能力
【Key words】 ReliabilityValidityDifficultyDiscriminabilityMeasuring Capacity
节点文献中: 

本文链接的文献网络图示:

本文的引文网络