节点文献
基于多特征融合的大语言模型中文生成摘要检测方法
Multi-feature Fusion Approach for Detecting Chinese Summaries Generated by Large Language Models
【摘要】 ChatGPT等大语言模型以其强大的文本生成能力给学术届和工业届的文本真实性和合法性带来了极大的挑战。为探索有效的大语言模型生成文本检测方法,以确保论文和专利的原创性和真实性,通过直接提示、学术润色、检索模仿三种提示方式构建论文和专利摘要的检测数据集,分析人类撰写与大语言模型生成摘要在长度、句法、词汇、困惑度等特征上的差异性,评估多种检测方法在本文所建数据集上的性能,并提出一种融合独特字词数、统计词频、深层语义、困惑度与标记分布特征的多特征融合的中文生成摘要检测MF-CSGD(Multi-Feature Fusion Chinese Summary Generation Detection)方法。相较于人类撰写的摘要,大语言模型生成的摘要表现出结构简洁性、用词多样性,语言困惑度低等特点,所提的MF-CSGD模型在两个数据集上与支持向量机(Support Vector Machine,SVM)、TextCNN(Text Convolutional Neural Network)、RoBERTa(Robustly optimized BERT pretraining approach)等主流机器学习和深度学习模型相比,MF-CSGD的准确率、F1值与AUROC值均为最高。MF-CSGD能够有效识别大语言模型生成的摘要文本。
【Abstract】 Large language models such as ChatGPT, with their formidable text generation capabilities, have posed significant challenges to the authenticity and legitimacy of texts in academia and industry. To explore effective methods for detecting texts generated by large language models, ensuring the originality and authenticity of papers and patents, this study constructs a detection dataset for summaries of papers and patents using three types of prompts: direct prompting, academic polishing, and retrieval imitation. It analyzes the differences in length, syntax, vocabulary, perplexity, and other features between humanwritten and model-generated summaries, evaluates the performance of various detection methods on the constructed dataset, and proposes a multi-feature fusion Chinese summary generation detection method, MF-CSGD(Multi-Feature Fusion Chinese Summary Generation Detection). Compared to human-written summaries, summaries generated by large language models exhibit characteristics such as structural simplicity, lexical diversity, and low language perplexity. The proposed MFCSGD model achieves the highest accuracy, F1 score, and AUROC value on two datasets, outperforming mainstream machine learning and deep learning models such as Support Vector Machine(SVM), Text Convolutional Neural Network(TextCNN), and Robustly optimized BERT pretraining approach(RoBERTa), effectively identifying summaries generated by large language models.
【Key words】 multi-feature; text classification; large language model; thesis patent;
- 【文献出处】 电脑与信息技术 ,Computer and Information Technology , 编辑部邮箱 ,2025年03期
- 【分类号】TP391.1;TP18
- 【下载频次】13