节点文献
药物专利的数据挖掘技术研究
Study on Data Mining Technique of Pharmaceutical Patents
【作者】 梁静;
【导师】 程文堂;
【作者基本信息】 大连理工大学 , 物理化学, 2007, 硕士
【摘要】 目前,英、美、法等发达国家已经建成了世界权威的专利数据库,对药物化学专利文献处理方面的技术比较成熟,我国近几年也十分重视药物化学信息资源的建设和计算机处理水平的发展并取得了一定的成果。事实证明对专利文献深度挖掘和高技术处理能够明显提高数据库的查全率和查准率,本文以此为出发点,使用目前被广泛应用于各个领域的数据挖掘技术全面处理了药物专利中包含的化学结构图形和文本信息。本论文运用面向对象编程技术,使用C++编程语言完善了本课题组开发的化学结构图形输入输出软件StruDraw,实现了文字向结构图形的翻译功能。用户只需输入要查找的化合物名称便可在图形输出界面得到所需的化学结构图形,免去了费时费力查找资料的过程。本文的重点是药物专利文本信息的处理。保证查全率和查准率的关键在于专利文献的分类准确度,数据挖掘类型之一便是文本的自动分类,机器学习算法是实现数据挖掘技术的手段。本文为实现药物专利分类的机器处理,结合药物专利本身特点,使用机器学习算法实现了专利文本自动分类。首先对2000余份药物专利按照治疗功能分类,抽取其中五类作为训练样本,对每一类提取特征文本,使用向量空间模型将非结构化的文本进行数字化表示,分别使用支持向量机(Support Vector Machine,SVM),朴素贝叶斯(Na(?)ve Bayes,NB),径向基神经网络(Radical Basis Function Network,RBFNetwork)对专利样本进行分类测试,并通过各种分类模型评估指标对这三种分类算法进行了分类性能评估,证明SVM算法在药物专利自动文本分类方面的优越性。使用机器学习算法对药物化学专利分类,取代了以往人工分类的方法,为专利信息检索奠定了基础。
【Abstract】 Pharmaceutical patents have become one of the most importance information widely usedin many fileds, especially in innovative drug design. However, our techniques of storage andretrieval of patent information by computers are far behind those developed countries. Manycommecial pharmaceutical patent databases have been built up in several countries, e.g.,British, U.S.A and French. And we have attanched importance to it in recent years. A copy ofpharmaceutical patent is different from other kinds of patents due to its contents consistingboth generic structures and corresponding descriptive texts. In this paper, the advanced datamining techniques are applied to handle the text information in order to facilitate the retrievalof patent information.I first improve StruDraw, one of chemical software designed specifically for genericstructure input and output in our group. The function of translating text into chemicalstructure may be helpful to those front-end users who have little chemical background toindex chemical structures directly and easily. It is worthy to mention that the software waswritten in C++ and its component-based architecture makes it easy to add new functions witha little modification.As text categorization, the first step of storing a chemical patent by computer is to classifythe patent to which kind it belongs to. Data mining, or machine learning algorithms are morecompetitive to those traditional manual methods. The applications of several machine learningmethods to the categorization of pharmaceutical patents are presented in this paper. About2000 pieces of pharmaceutical patents are categorized into five classes according to theircurative effects and are selected as training instances. Features in text form are first extractedfrom each class and then are expressed in numerical vector form. Three machine learningalgorithms, i.e., Support Vector Machines, Na(?)ve Bayes and RBF Neutral Network are testedby 5 or 10 folds corss validation methods. Their performaces are compared by a series ofexperiments. And results show SVM algorithms outperforms than the other two algorithms.Methods proposed in this paper maybe helpful to the pharmaceutical patent categorization.
【Key words】 Data Mining; Machine learning; Pharmaceutical Patent; Text Categarizating; Translation from Character to Structure;
- 【网络出版投稿人】 大连理工大学 【网络出版年期】2008年 02期
- 【分类号】G306.3
- 【下载频次】478