节点文献
A Study on Quality of CET Multiple-choice Items-Item Response Theory Approach
【作者】 陈韵;
【导师】 张文鹏;
【作者基本信息】 电子科技大学 , 外国语言学与应用语言学, 2005, 硕士
【摘要】 语言测试伴随语言教学而出现,两者有着密不可分的关系。英语语言测试是用于评价学生英语水平和教学效果的重要手段。其评价是否准确、客观、符合实际取决于测试的质量。目前常用的测试质量评估手段经典测验理论具有明显的缺陷,其计算结果会随受试样本的变化而改变。起源于二十世纪三十年代的项目反应理论这种测验理论具有受试能力参数与项目参数相互独立的优越性,但却因其数学计算的复杂性而令很多应用语言学研究者对它的使用只能望洋兴叹。客观性多项选择题是大学英语测试的主要题型,其质量直接影响测试质量,因此有必要进行全面质量评估。本研究采用项目反应理论,以一份大学第二学期期末英语测试的客观性多项选择题为研究对象,对其质量进行了综合评估并深入分析了导致低质量项目的因素。研究者在受试学生中随机抽取两个样本,样本一包含207 名受试,用于综合分析,样本二包含文理专业受试各两百名,用于检验项目偏差。质量评估的主要方法是对项目特征曲线、项目信息函数、测验信息函数、项目预测效度的分析和对项目偏差的检测。其中,项目信息函数和测验信息函数分别是单个项目和整体测试的信度指标。研究结果表明该测试部分客观性多项选择题质量不高,测试信度较低。虽然测试具有较好的表面和内容效度,但其预测效度很差。测试中没有偏差项目,因此受试的专业背景不会影响其对试题的反应情况。导致低质量项目的因素有两类:选择题设计违反命题的基本原则;造成听力和阅读理解试题低质的特定因素。前者包括题项有不止一个答案,选项难度、结构不一致,包含无用选项等。后者包括对话中对一个观点提供的语境过多或不充分,阅读理解中设计干扰项时对读者的心理反应做了不恰当的预测,在阅读理解中考常识等。本研究是将项目反应理论应用于语言测试质量评估的一次尝试,充分利用了项目反应理论的优越性。研究表明由于经典测验理论的计算简单性和项目反应理论的估计精确性,在测试质量评估中可以将两者结合使用。研究也给予英语语言测试客观性多项选择题的编写一定启示。
【Abstract】 Language testing emerges with the development of language teaching and the two are closely interrelated with each other —English language testing is an important approach used for evaluating the student’s English ability and the teaching effects. Whether the evaluation is objective, accurate and appropriate for the students’ability level relies on the test quality. Classical test theory (CTT), the most used approach for test quality evaluation nowadays, has an obvious defect that the computation result fluctuates a lot with the change of the samples tested. Item response theory (IRT), the testing theory originated from the 1930s, has the superiority that the examinee’s ability parameter is independent with the item parameters. But because of its complicated mathematical computations, many researchers of applied linguistics can only bemoan their inadequacy to apply this method. The characteristic of objective multiple-choice items directly influences the quality of a test, for it is a main test format in college English test, and thus it is necessary to make an all-round quality evaluation on them. Taking the multiple-choice items in a college English test paper for second semester final examination as the object, this study adopts IRT to make a comprehensive quality evaluation on the items and conducts an in-depth analysis of the factors leading to poor quality items. Two samples are drawn randomly from the population. The first sample consists of 207 examinees, used for comprehensive analysis; the second sample contains 200 examinees majoring in science and 200 in arts, used for checking biased items. The methods applied in data analysis are the analysis of item characteristic curve, item information function, test information function and item validity, as well as the checking of biased items. Item information function and test information function are the reliability indexes for individual items and the whole test respectively. The results of data analysis indicate that some items are of poor quality and the test reliability is not very high. Though the test has good face validity and content validity, its predictive validity is very low. The test does not contain biased items, which means the background knowledge concerning an examinee’s major will not influence his performance. The factors leading to poor quality items fall into two categories, violation of the basic multiple-choice-item writing principles and the unique factors leading to poor
【Key words】 language testing; item response theory; multiple-choice items; quality evaluation;
- 【网络出版投稿人】 电子科技大学 【网络出版年期】2005年 07期
- 【分类号】H319
- 【下载频次】370