节点文献

贝叶斯学习的先验分布的研究

Studying on Prior Distribution in Bayesian Learning

【作者】 胡振宇

【导师】 林士敏;

【作者基本信息】 广西师范大学 , 计算机软件与理论, 2001, 硕士

【摘要】 贝叶斯方法起源于著名的贝叶斯公式(又称为贝叶斯定理)。其后,发展成为一种系统的统计推断和决策的方法.90年代进一步研究可学习的贝叶斯网络,用于机器学习.由于概率统计与数据采掘的天然联系。数学采据兴起后贝叶斯网络日益受到重视,再次成为引人注目的热点.与非贝叶扬方法相比,贝叶斯方法的特出特点是其学习机制可以综合先验信息和后验信息,既可避免只使用先验信息可能带来的主观偏见,和缺乏样本信息时的大量盲目搜索与计算,也可避免只使用样本信息带来的噪音的影响.只要合理地确定先验,就可以进行有效的学习。因此,适用于具有概率统计特征的数据采掘和机器学习(或发现)问题,尤其是样本难得的问题.贝叶斯方法遇到的一个重大的问题是先验分布的确定依据的只是一些准则,没有可操作的完整的理论.在许多情况下先验分布的合理性和准确性难以评价.本文主要对先验分布的迭取进行研究. 1.本论文讨论贝叶斯学习中的一个基本问题—相容性问题.本文主要讨论相容性的定义和性质,证明了在相客性前提下贝叶斯学习的先验无关性:通过一些反例说明不相容性的贝叶斯学习是存在的:指出贝叶斯学习及其先验分布的选取的前提—相容性原则.也给出了一个相容性的相对简化的充分条件。并证明了在此条件下贝叶斯学习的渐近正态性。主要有以下结论: 定义 1.l如果对于θ0的任意一个邻域V,几乎必然有 limμn(V)=l,即任给η>0巾和任意θ0的一个邻域 V存在 n≥no使得几乎必然则称在处是相容的· 引理1.1设θ0是Θ的一个点,π1。、π2是两个在θ0正值连续先验分布,如果在相容, 则几乎必然 定理1.l如果μn和vn均为θ的后验分布.且是相容的,则对于任意的有界连续 !函数/,有卜(川+]松\)。 n—— 假设孔。回,称以下条件为正则条件: (CI)先验密度。(8)在民大于 0且连续。 —刀。。—u丛一h一、。一巳。口讹,。刀.aIOg f(X旧)、十o C)logf(卜)在90的邻域关于6H次可微且微分E(…)在O 、一,·-。J、一I-·一’U—””””””””—-’”-—”Dg“内连续。 (Q)对任何S>o及 从(S=0 一孔卜 占 口O存在正数牡了(与占相关)使得 tim P[suP n-’K(9)一人(民)}<一k(8)】 1 。、。e。日-川(占) 即4 是弱相容性的,亦即 尸* 4=00 hWOO 定理2.!若上述条件(O)到(Q 成立,且 一o<b<口<o,则9_+ba_<0<0_+arr。的后验概率,即 Q儿一叮厂,;(引入皆…雪XJdo以概率 1趋近于 (2)‘ie‘du,n -+。。 2.将贝叶斯先验选取的一些准则(如共轭分布、不变性原则、最大嫡原则等)看成是一些启发式规则,将贝叶斯启发式方法引入先验分布的选取中,以优化的观点将先验信息与样本信息结合起来,对先验分布的选取进行研究。并引入损失函数和风险函数分别从贝叶斯决策理论与贝叶斯判别分析的角度,对先验分布和合理性与准确性进行评价。推广了现有的MLll方法。证明了在一定的条件下最大后验信念(HPB)先验和二型极大似然(MLll)先验是最合适先验(MSP)。主要有如下一些结论: 定义3*:记ac=stir(XW门,O)m(9)dg,a;=。;DP(尸川卜,a尸”)4o 互互比值 a。a;称为 N(0)对 A(0)后验信念比,。。/。;称为先验信念比。由(3.2), ac h]P(X’”’IN,8)N(8呷 2=——(.3) a ql尸(XU门八,0)八(8)do假设仅有两个可以考虑的先验八汐)和八(的,如果要选用11;(的而不是从(的就必须有闪<;叫,或 一 0 <—— N 定义3.3因子 _DOS柏FlOP 060灯ffiflo *。/al *nDI DPIO厂 *幻l叨 厂口rl口 厂。/兀a,风称为先验贝叶斯因子。先验信念比和后验信念比这两种比率相除,可能会减弱先验信念的影响,突出数据的影响。 定理3.二如果对于各个先验分布的先验信念都相等,则仅得先验似然达到最人值的八用即为最大后验信念先验。 定义 4.lp用称为最合适先验(MSP--Most Suitable Prior MSP卜如果使凡(尸,旬x》达到最小值。 定理4.1如果采用“0.1”损失函数,即 !0,600=A L(A,5k)卜《.二”:’(4,3)

【Abstract】 Bayesian method originated from the well-known Bayes’ theorem and had beendeveloped into a syStematic technology of inference and decision. In 90’s of the 20stcentUry leamable Bayesian netWork was studied and was aPplied tO machine learning.With the boom of date mining, Bayesian netWork has been paid great attention and againin the limeligh< because of the naturai relation betWeen the mathematical statistics anddata mining. Comparing with non-Bnyain methods, it’s prominent featares lay in that itcombines the prior and POsterior information, which avoids the disadVantag ofsubjective bias caused by simply using the prior information only, of blind search causedby the incomplete sample information, of noise affection caused by simply using thesample information only If we choice a suitable prioF, we can conduct the Bayesianleaming effectively, so it fits the problems of data mining and machine leaming thatpossess charaters of probability and statistics, especially when the samples are rare.What is much difficult in Bayesian method is that the detenninations of prior are onlysome guidelines without complete and operationaI theorem, and it’s hard to value thejustice and accuracy of a prior in manyfconditions. In this essay we fOcus on the choosingof suitable prior in Bayesian learningFirst of a1l. we discuss afoundaIional probiem -consistency in 8a)esian learning. Aproperty of consistency is attained, i.e. inference based on consistent leaming is free fromprior. By introducing counterexamples we investigate the existence of inconsistency’which reminds us to carefully choose a priof, proPOse a basic principle in Bayesianleaming - the consist principle. We a1so point out that under certain regular conditionsthe posterior distribution is free from prior and is approkimately normal as the volume ofsamples increases infinitely In this essay we discuss these regular conditions using amethod similar to that of W8lker (l967[2l ]) and Heyde and Johnstone (l979tl4]), but theconditions have been simplified. The followings are some resultS about this aspect:Definition l.l: The pair (8,p.) is said consistent at 00, if for everyneighborhood v of 96, limp’(v) = l almost surely. 1.e. given any n >o and an--.roarbitrary neighborhood V of 90, there exists n 2 n0 suchK(8 E V I X’"’) Z l -- 0 for a1l n Z nolLemma 1.l: Let 90 be an interior POint of e and Kl,K2 be tWO priordensities which are positive and continuous at 00. We assume that the POSterior4(0 I X’) is consistent for i = l,2 thenLT jK, (8 I X(")) -- K,(6 I X’">)k6 = 0Theorem 1.l: if K and Vn are poSterior distributions of 8 and they areconsistent then lfu(m)--+ Jfd(vn) for al1 bounded continuous functions f.nweSuppose that 00is an interior point of e. W6 imPOse the following regulartyconditions throughout:(C l ) The prior density K(0) is continuous and positive at 80.(C2) logf(x I 0) is twice differentiable with respect to 0 in some neighborhoodof 80 and the twice difference E(ap) is continuous in O.b0, ) is continuous inO.(C3) For any 8 > 0 for which N0(8) = {i 8 --00 l< 6} = O there exists apositive number k(5), depend on 5. such thatlimP[ sup n-’{L,,(6) -- Ln(00)} < --k(6)] = lnwt eee--M,(6)this obviously implies that 8. is a weakly consistent, namely thatp lim 6n == 00.n-roTheorem 2.l: suPpose the conditions (C l ) to (C3) hold, then, if -- co < b < a < co,the POsterior probability that 6n + bcr. < 6 < 0n + arr.’ namelyI::Kn (8 l X,..., X.)d0t6nds tO(2.)-1 f.4"duin probability as n - co.Secondly’ we regard some guidelines of choosing prior (such as conjugate prior,Jetws non-infOrmative Prior, maximum entropy prior) as heuriStic aPproaches. Byintroducing the Bayesian hetalstic approaches and combining the prior infOrmation andsample infOrmation, we Stody the choosing of suitable prior in view of oPtimiZaion. wnthe loss function and risk function we aPpraise the justice and accuracy of a prio

  • 【分类号】TP181
  • 【被引频次】13
  • 【下载频次】1707
节点文献中: