节点文献
大语言模型赋能中文二语阅读测试的研发与应用研究
The development and application of CSL reading tests empowered by large language models
【摘要】 本研究面向中文二语阅读测试研发的真实场景,通过多环节多视角的实证研究,系统验证了大语言模型赋能中文二语测试研发的可行性,初步构建出一条人机协同模式下的测试研发路径,涵盖“模型自动生成-质量系统评估-人机协作修改-学习者实测应用”四个核心环节。基于主客观指标的测评结果发现,模型生成试题与HSK真题在语篇输入、预期回答和整体质量上具有可比性。面向240名学习者的实测分析表明,生成试题具备良好的信效度和难度区分能力,能够较为准确地评估中文二语学习者的阅读能力。最后,我们总结了GPT-4应用于中文二语阅读测试生成的优势与局限,为国际中文教育考试现代化和智慧考试提供参考建议。
【Abstract】 This research focuses on real-world scenarios for the development of Chinese as a Second Language(CSL) reading tests. Through empirical studies from multiple stages and perspectives, it is the first to validate the feasibility of adopting large language models(LLMs) to empower CSL test development. A preliminary human-LLM collaborative model for test development was constructed, covering four core stages: LLM-generated tests, quality assessment, human-LLM cooperative modification, and actual measurement of CSL learners. Evaluation results based on both subjective and objective metrics indicate that LLM-generated tests are equivalent to HSK reading tests. This equivalence is observed in terms of the characteristics of the input, the characteristics of the expected response, and the overall quality. Actual measurement analysis based on 240 CSL learners indicates that the modified tests exhibit good reliability, validity, and item discrimination. Furthermore, measurement results demonstrate an accurate capacity to assess the reading ability of CSL learners. Finally, we summarized the advantages and limitations of applying GPT-4 in generating CSL reading tests, providing reference suggestions for the modernization and intelligent assessment of International Chinese Language Education exams.
【Key words】 reading test; large language models; automated item generation; intelligent assessment;
- 【文献出处】 语言教学与研究 ,Language Teaching and Linguistic Studies , 编辑部邮箱 ,2025年04期
- 【分类号】H195.3
- 【下载频次】381