节点文献

金融大数据智能决策关键技术研究

Research on Intelligent Decision Making in Financial Big Data

【作者】 周俊;

【导师】 郑小林;

【作者基本信息】 浙江大学 , 电子信息, 2024, 博士

【摘要】 近年来,金融服务领域的智能化需求日益增长,伴随用户增长、风险控制等需求,大数据技术得到了广泛应用。一个包含客户、金融产品与服务及算法平台的金融大数据决策闭环系统渐成体系。客户与金融服务的交互产生了丰富数据,构成了大数据分析的基础;金融机构创新设计产品并应用智能化决策技术以满足客户需求;算法平台则负责数据的收集、处理与分析。当前,金融大数据智能决策领域的研究正受到业界的广泛关注。众多研究致力于将电子商务领域广泛应用的深度学习技术迁移至金融环境。得益于神经网络卓越的表征能力,深度学习算法借助复杂的网络结构和庞大的数据集,极大地推动了金融信息的交互与价值的传递。然而,这一过程中往往忽视了金融领域特有的监管性、透明性、安全性及风险性等关键要素。因此,尽管金融大数据智能决策技术的研究正逐渐展开,该领域仍然面临着众多待解决的挑战。为构建一个稳健的金融智能化决策系统,本文提出的关键技术主要包括:(1)为了缓解该系统中算法黑盒化的问题,提出了可解释的特征归因框架。在金融领域,复杂的算法决策常常缺乏必要的透明度,使得客户对预测逻辑感到困惑。然而,现有研究未能充分考虑生成实例的合理性,这可能导致产生在实际应用场景中根本不可能出现的虚构实例。为提高解释的可信度,本文利用约束扰动和反事实实例技术,并结合一种新的度量方法,准确估算每个特征的重要性。该方法在实际风控场景中表现出明显的优势,尤其在处理不平衡数据集时。为了更好地解释文本和图像数据,本文进一步提出了基于扰动的特征归因框架,利用学到的变分自动编码器将实例映射到潜空间,并应用生成模型进行扰动,以提高效率和解释质量。(2)为了捕捉金融网络中的关系与结构信息,提出了异质图数据的表征学习框架。图关系数据凭借其出色的可用性和丰富的语义表达,在金融领域的应用日益广泛,尤其在促进金融增长和风险控制业务方面发挥着关键作用。然而,传统的处理方法往往依赖于大量的标注数据和专家知识,而对于捕捉数据的高阶连通性和时序性等复杂特征则关注不足。为了高效挖掘这些数据,本文设计了多通道图神经网络和图增强专家网络,以联合建立高阶信息模型和处理多任务推荐,从而提升推荐效果。此外,本文将金融图中的不同角色解耦并构建独立的子图,通过联合学习和时间编码机制,捕捉不同角色之间的交互作用和动态性,改善了金融交易风控的效果。(3)为了在挖掘数据价值的同时保障用户隐私,提出了深度学习隐私保护技术。在处理数据孤岛和高计算复杂性的同时确保隐私安全,限制了神经网络的性能,然而提高模型精度的需求迫在眉睫。现有的隐私保护研究主要包括基于高度安全的密码学技术方法,这些方法虽然提供了强大的安全保障,但计算效率较低;以及分布式机器学习,它通过在参与者之间交换模型更新进行协作训练,提高了效率但也引入了潜在的安全风险。为此,本文结合算法和密码学方法,提出了可扩展且保护隐私的深度神经网络学习框架,并引入安全多方计算技术来保护隐私并保持模型性能。此外,本文还提出了垂直联合图神经网络学习范式,用于隐私保护的节点分类任务。实验证明这些技术在保护隐私的同时,仍能保持良好的神经网络效果。(4)为了解决金融决策中涉及的海量数据与多目标优化问题,提出了大规模决策优化技术。在金融决策场景中,除了追求算法效果,还需考虑成本和风险等因素。现有的求解器受限于单机资源,难以应对实际的中大规模决策问题。为此,本文引入分布式迭代算法并开发具有扩展性和效率的大型分布式框架,以有效解决大规模线性规划问题,优于传统方法,并能扩展到十亿规模的线性规划问题。

【Abstract】 In recent years,the demand for intelligence in the financial services sector has been increasing.With the growth of users and the need for risk control,big data technology has been widely applied.A financial big data decision-making closed-loop system,encompassing customers,financial products and services,and algorithm platforms,has been gradually forming.The interaction between customers and financial services generates a wealth of data,laying the foundation for big data analysis;financial institutions innovate and design products and apply intelligent decision-making technologies to meet customer needs;the algorithm platform is responsible for data collection,processing,and analysis of data.Currently,research in the field of intelligent decision-making in financial big data is receiving widespread attention in the industry.Numerous studies are dedicated to transferring deep learning technologies,widely applied in e-commerce,to the financial environment.Thanks to the superior representational abilities of neural networks,deep learning algorithms have significantly advanced the interaction of financial information and value transfer by leveraging complex network structures and vast datasets.However,this process often overlooks key elements unique to the financial sector,such as regulatory compliance,transparency,security,and risk.Consequently,despite the growing research in financial big data,the field still faces numerous challenges to be addressed.To build a robust intelligent financial decision-making system,the key technologies proposed in this thesis include:(1)To alleviate the black box problem of deep learning algorithms in the system,an interpretable feature attribution framework is proposed.In the financial sector,complex algorithmic decisions often lack the necessary transparency,leaving customers puzzled by the prediction logic.However,existing research has not fully considered the plausibility of generated instances,which may lead to the creation of fictitious instances that are impossible in real-world scenarios.To enhance the credibility of interpretations,this thesis uses constrained perturbation and counterfactual instance techniques,combined with a new measurement method,to accurately estimate the importance of each feature.This method has shown significant advantages in practical risk control scenarios,especially when dealing with unbalanced datasets.To better explain text and image data,this thesis further proposes a perturbation-based feature attribution framework,utilizing learned variational autoencoders to map instances to latent spaces and applying generative models for perturbation to improve efficiency and quality of interpretation.(2)To capture the relationships and structural information within financial networks,representation learning frameworks for heterogeneous graph data are proposed.Graph relational data,with its excellent usability and rich semantic expression,is increasingly applied in the financial sector,playing a key role in promoting financial growth and risk control businesses.However,traditional processing methods often rely on a large amount of labeled data and expert knowledge,while paying insufficient attention to capturing complex features such as the high-order connectivity and temporality of data.To efficiently mine these data,this thesis designs a multi-channel graph neural network and a graph-enhanced expert network to jointly establish a high-order information model and handle multi-task recommendations,thereby improving recommendation effects.Additionally,this thesis decouples different roles in financial graphs and constructs independent subgraphs,capturing the interactions and dynamics between roles through joint learning and timing coding mechanisms,improving the effectiveness of financial transaction risk control.(3)To address the issue of safeguarding user privacy while extracting the value of data,deep learning privacy protection techniques are proposed.Ensuring privacy security while dealing with data isolation and high computational complexity limits the performance of neural networks,yet there is an urgent need to improve model accuracy.Existing research on privacy protection mainly includes methods based on highly secure cryptographic techniques,offering strong security but with lower computational efficiency,and distributed machine learning,which exchanges model updates among participants for collaborative training,improving efficiency but also introducing potential security risks.Therefore,this thesis combines algorithmic and cryptographic methods to propose a scalable and privacy-protecting deep neural network learning framework,incorporating secure multiparty computation technology to protect privacy and maintain model performance.Additionally,a vertical federated graph neural network learning paradigm for privacy-protecting node classification tasks is proposed.Experiments demonstrate that these technologies maintain good neural network effects while protecting privacy.(4)To tackle the challenges associated with the vast amounts of data and multi-objective optimization in financial decision-making,large-scale decision optimization technology is proposed.In financial decision-making scenarios,factors such as cost and risk must be considered in addition to algorithmic effectiveness.Existing solvers,limited by single-machine resources,struggle to cope with practical medium-to-large-scale decision problems.Therefore,this thesis introduces distributed iterative algorithms and develops a scalable and efficient large-scale distributed framework to effectively solve large-scale linear programming problems,outperforming traditional methods and capable of extending to problems of a billion-scale magnitude.

  • 【网络出版投稿人】 浙江大学
  • 【网络出版年期】2025年 04期
  • 【分类号】TP311.13;F832
节点文献中: