节点文献
面向出版社富媒体知识的文本分类研究
Research on the Processing of Rich Media Knowledge for Publishers
【摘要】 大数据环境下,出版行业面临着富媒体数据带来的跨媒体数据组织和海量历史数据的挑战。为了形成有效的知识组织,针对富媒体出版社的文本数据具有数据量巨大、标签分层级的特点,本论文使用截断奇异值分解进行降维,应用线性分类核支持向量机模型,并且设计了多层级分类方法,对富媒体文本进行文本分类。实验表明,在富媒体出版社的文本数据下,本文方法取得了较好的文本分类结果。在150维的文本特征下,区域分类的第二级分类效果最好,其中准确率达到0.98,召回率达到0.76,F1指标达到0.87。
【Abstract】 The publishing industry faces the challenge of cross-media data organization and massive historical data brought by rich media data in big data area. The text data for rich media publishing houses has the characteristics of huge data and hierarchical labels. In order to form an effective knowledge organization, this paper uses TSVD to reduce dimensionality, applies LinearSVM model, and designs Multi-level classification method for text classification of rich media texts. Experiments show that under the texts of rich media, our method has achieved good results. Under the 150-dimensional text feature, the second-level effect of regional classification is the best, with the accuracy rate reaching 0.98, the recall rate reaching 0.76, and the F1 index reaching 0.87.
- 【文献出处】 情报工程 ,Technology Intelligence Engineering , 编辑部邮箱 ,2019年02期
- 【分类号】G254.1;G230.7
- 【被引频次】3
- 【下载频次】93