节点文献
一种面向专利文献数据的文本自动分类方法
Automatic text categorization for patent data
【摘要】 中文专利文献自动分类目前尚无成熟适用的方法。分析了文本自动分类的关键技术,并结合专利数据的特点对无词典分词和权重计算进行了改进,提出了一种适用于专利数据分类的层次分类方法,给出了面向专利文献数据的文本自动分类系统的框架模型。实验表明,该系统具有较好的分类精度与效率。
【Abstract】 At present, there are no practical and mature automatic text categorization methods for patent data. Therefore, this paper made a research on several key techniques about text categorization, improved the non-dictionary segment and weight calculation, and then proposed a hierarchical categorization method and an automatic text categorization framework for patent data. The experiment testifies that the system has a good classification accuracy and efficiency.
【关键词】 文本分类;
专利文献;
国际专利分类码;
K-近邻;
【Key words】 text categorization; patent; International Patent Classification(IPC); K-Nearest Neighbor(KNN);
【Key words】 text categorization; patent; International Patent Classification(IPC); K-Nearest Neighbor(KNN);
【基金】 江苏省自然科学基金资助项目(BK2006095);教育部高等学校博士学科点科研基金资助项目(20040286009)
- 【文献出处】 计算机应用 ,Journal of Computer Applications , 编辑部邮箱 ,2008年01期
- 【分类号】TP391.1
- 【被引频次】27
- 【下载频次】476