节点文献

面向出版社富媒体知识的文本分类研究

Research on the Processing of Rich Media Knowledge for Publishers

  • 推荐 CAJ下载
  • PDF下载
  • 不支持迅雷等下载工具,请取消加速工具后下载。

【作者】 刘琼昕宋祥王鹏

【Author】 LIU Qiongxin;SONG Xiang;WANG Peng;Beijing Engineering Applications Research Center on High Volume Language Information Processing and Cloud Computing;School of Computer Science and Technology, Beijing Institute of Technology;The Key Laboratory of Rich-media Knowledge Organization and Service of Digital Publishing Content Institute of Scientific & Technical Information of China;

【机构】 北京市海量语言信息处理与云计算应用工程技术研究中心北京理工大学计算机学院中国科学技术信息研究所富媒体数字出版内容组织与知识服务重点实验室

【摘要】 大数据环境下,出版行业面临着富媒体数据带来的跨媒体数据组织和海量历史数据的挑战。为了形成有效的知识组织,针对富媒体出版社的文本数据具有数据量巨大、标签分层级的特点,本论文使用截断奇异值分解进行降维,应用线性分类核支持向量机模型,并且设计了多层级分类方法,对富媒体文本进行文本分类。实验表明,在富媒体出版社的文本数据下,本文方法取得了较好的文本分类结果。在150维的文本特征下,区域分类的第二级分类效果最好,其中准确率达到0.98,召回率达到0.76,F1指标达到0.87。

【Abstract】 The publishing industry faces the challenge of cross-media data organization and massive historical data brought by rich media data in big data area. The text data for rich media publishing houses has the characteristics of huge data and hierarchical labels. In order to form an effective knowledge organization, this paper uses TSVD to reduce dimensionality, applies LinearSVM model, and designs Multi-level classification method for text classification of rich media texts. Experiments show that under the texts of rich media, our method has achieved good results. Under the 150-dimensional text feature, the second-level effect of regional classification is the best, with the accuracy rate reaching 0.98, the recall rate reaching 0.76, and the F1 index reaching 0.87.

【关键词】 富媒体文本分类支持向量机降准
【Key words】 Rich mediatext classificationSVMreduce dimension
【基金】 富媒体数字出版内容组织与知识服务重点实验室开放基金项目(ZD2018-07/02):“富媒体数字出版内容的知识挖掘及发现技术研究”
  • 【文献出处】 情报工程 ,Technology Intelligence Engineering , 编辑部邮箱 ,2019年02期
  • 【分类号】G254.1;G230.7
  • 【被引频次】3
  • 【下载频次】93
节点文献中: 

本文链接的文献网络图示:

本文的引文网络