节点文献
基于数据分割的受控模型自由变量选择方法
Model-free controlled variable selection via data splitting
【摘要】 从高维数据中识别重要变量并控制错误发现率是一个重要的统计问题.本文在充分降维框架中,通过数据分割的手段针对高维数据提出一种误差可控且模型自由的变量选择方法.该方法首先通过一个响应变换函数将一般化的模型转化为求解最小二乘估计问题.然后构造一系列边际对称的统计量和一个数据驱动的拒绝临界值,来实现错误发现率可控的变量选择.在一些温和的理论条件下,本文证明该方法能精确控制有限样本下的错误发现率,同时也能实现大样本下的错误发现率控制.本文通过数值模拟和高维疾病基因识别的实例分析,验证了该方法相较于其他现有方法能以较高的检验效率且更快更准确地控制错误发现率.
【Abstract】 Addressing the simultaneous identification of contributory variables while controlling the false discovery rate(FDR) in high-dimensional data is a crucial statistical challenge. In this paper, we propose a novel modelfree variable selection procedure in a sufficient dimension reduction framework via a data splitting technique. The variable selection problem is first converted to a least squares procedure with several response transformations. We construct a series of statistics with global symmetry property and leverage the symmetry to derive a data-driven threshold aimed at error rate control. Our approach demonstrates the capability for achieving finite-sample and asymptotic FDR control under mild theoretical conditions. Numerical experiments confirm that our procedure has satisfactory FDR control and higher power compared with existing methods.
【Key words】 data splitting; false discovery rate; model-free; sufficient dimension reduction; symmetry;
- 【文献出处】 中国科学:数学 ,Scientia Sinica(Mathematica) , 编辑部邮箱 ,2025年08期
- 【分类号】O212.1
- 【下载频次】44