节点文献
基于双重预训练的商品属性分类方法
Commodity Attribute Classification Method Based on Dual Pre-training
【摘要】 商品属性分类任务是指对一段商品的描述文字进行属性分析并进而对多个属性进行分类的过程,其有助于人们从多个角度了解商品,为市场营销、产品管理等提供帮助。当前大语言模型的使用也愈加广泛,但在商品属性分类问题上,通用大模型由于缺乏领域知识和属性关联等信息,性能不够理想。为此,提出了一个基于双重预训练的商品属性分类方法,旨在通过使用特定的预训练方式提高大语言模型在商品属性分类任务中的性能。在T5模型的基础上,引入了领域内文本预训练和基于属性间关联性的预训练两种方法。在Clothing Fit Data数据集上的实验结果显示,使用了双重预训练的T5模型较未经过预训练的模型以及其他基准模型,在各个属性上的分类效果都取得了一定提升。实验结果证明了所提方法的有效性。
【Abstract】 The commodity attribute classification task refers to the process of analyzing the attributes of a piece of merchandise based on its descriptive text and subsequently categorizing multiple attributes.This process aids in providing insights into merchandise from various perspectives, thereby assisting in marketing and product management.While the utilization of large language models is increasingly prevalent, their performance in commodity attribute classification tasks remains suboptimal due to the lack of domain knowledge and attribute correlations.To address this issue, this paper proposes a dual pre-training-based method for commodity attribute classification, aiming to enhance the performance of large language models in such tasks by employing specific pre-training techniques.Building upon the T5 model, this paper introduces two methods: domain-specific text pre-training and attribute correlation-based pre-training.These methods enhance the model’s understanding of the specific task from both input and output text perspectives, facilitating the classification of multiple attributes of merchandise.Experimental results on the Clothing Fit Data dataset demonstrate that the dual pre-trained T5 model outperforms both non-pre-trained models and other baseline models in attribute classification, validating the effectiveness of the proposed approach.
【Key words】 Dual pretraining; Multi-attribute classification; Large language model; T5; Commodity attribute classification;
- 【文献出处】 计算机科学 ,Computer Science , 编辑部邮箱 ,2025年S1期
- 【分类号】TP391.1
- 【下载频次】26