节点文献

一种基于文本的概率因果图推理方法的研究及其应用分析

Research on a Method of Text-based Probabilistic Causal Graph Reasoning and Its Application and Analysis

  • 推荐 CAJ下载
  • PDF下载
  • 不支持迅雷等下载工具,请取消加速工具后下载。

【作者】 史雪洋程兵

【Author】 SHI Xueyang;CHENG Bing;Academy of Mathematics and Systems Science, Chinese Academy of Sciences;School of Mathematical Sciences, University of Chinese Academy of Sciences;

【机构】 中国科学院数学与系统科学研究院中国科学院大学数学科学学院

【摘要】 因果是人类由弱人工智能时代迈向类人智能时代的关键,加之大数据时代下,数据的多样性及其信息的丰富性促使挖掘文本数据中的因果知识成为新的研究热点.目前的因果推断方法更多地应用在低维的、结构化观测数据,对文本数据的利用并不充分.为了实现对非结构化文本数据的因果分析,本文首先结合现代汉语的句型系统,提出了一种基于规则的事件抽取方法.之后,提出了一种基于文本的概率因果图推理方法.具体来说,针对已抽取出的事件,该方法采用聚类算法抽象并泛化语义相似事件的公共语义特征,以定义文本数据中的变量及观测的概念,并基于语义依存关系抽取因果关系来指导文本中因果事件链条的抽取,以进一步发现文本蕴含的因果网络,进而采用因果图模型完成了对文本数据中因果效应的推断.最后,本文分别选取司法文书及金融研报作为语料进行实验,具体展示了针对文本数据的概率因果推理过程.

【Abstract】 Causality is the key for human beings to move from the era of weak artificial intelligence to the era of human-like artificial intelligence. In addition, under the background of big data, the diversity of data and the richness of information make mining causal knowledge in text data become a new research focus. Recently,causal inference methods are all mostly applied to low-dimensional and structured observation data, which makes no full use of text data. In order to realize the causal analyses on unstructured text data, this paper presents a rule-based event extraction method based on the sentence structure system of modern Chinese. And then, a method of text-based probabilistic causal graph reasoning is presented. Specifically,for the extracted events, it adopts a clustering algorithm to abstract and generalize the common semantic features of them with semantic similarity, so as to define the concepts of variables and observations in the text data. At the same time, it extracts the causal relations based on semantic dependency relations to guide the extraction of causal event chains in text, so as to further discover the causal network contained in the text. Then the causal diagram model is applied on the network to infer the causal effect in text data. In the end, this paper respectively selects the judicial documents and financial research reports as corpus for trails to demonstrate the probabilistic causal reasoning process on text data.

【基金】 中国科学院随机复杂结构与数据科学重点实验室(2008DP173182);科技创新2030——“新一代人工智能”重大项目(2021ZD0111204)~~
  • 【文献出处】 计量经济学报 ,China Journal of Econometrics , 编辑部邮箱 ,2023年02期
  • 【分类号】TP391.1
  • 【下载频次】29
节点文献中: 

本文链接的文献网络图示:

本文的引文网络