节点文献

大模型赋能的可视化与可视分析研究综述

Survey on large model-empowered visualization and visual analytics

  • 推荐 CAJ下载
  • PDF下载
  • 不支持迅雷等下载工具,请取消加速工具后下载。

【作者】 汪云海曹楠陈思明李晨辉曾伟陶钧曾琼王长波张加万

【Author】 Wang Yunhai;Cao Nan;Chen Siming;Li Chenhui;Zeng Wei;Tao Jun;Zeng Qiong;Wang Changbo;Zhang Jiawan;School of Information, Renmin University of China;College of Design and Innovation,Tongji University;School of Data Science, Fudan University;School of Computer Science and Technology, East China Normal University;Thrust of Computational Media and Arts, Hong Kong University of Science and Technology (Guangzhou);School of Computer Science and Engineering, Sun Yat-sen University;School of Computer Science and Technology, Shandong University;School of Data Science and Engineering, East China Normal University;Software College, Tianjin University;

【通讯作者】 张加万;

【机构】 中国人民大学信息学院同济大学设计创意学院复旦大学大数据学院华东师范大学计算机科学与技术学院香港科技大学(广州)计算媒体与艺术学域中山大学数据科学与计算机学院山东大学计算机科学与技术学院华东师范大学数据科学与工程学院天津大学软件学院

【摘要】 数据可视化作为连接人类认知与数据科学研究的重要技术基础,随着大模型的发展迎来范式转型。传统数据可视化主要依赖人工设计的视觉编码规则、图形语法与交互机制,通过显式映射和操作支持数据探索与信息传达。然而,面对日益增长的数据规模、多样化的数据类型以及复杂的分析与决策场景,基于静态图形和参数化交互的传统方法在效率、表达能力和语义支持方面逐渐显现局限。近年来,大规模预训练语言模型、多模态基础模型及数据智能体的兴起,为数据可视化的自动生成、智能分析和交互优化提供了新的技术动力。本文围绕可视化基础理论、可视分析、可视化叙事与可视化评估4个方面,对大模型赋能下的数据可视化研究进展进行系统综述,从基础理论层面看,大模型通过其强大的语义理解与推理能力,推动可视化从低层视觉编码向语义驱动的表达与感知建模演进。在可视分析层面,大模型与数据代理的结合使分析流程从以人为主导的工具操作,转向人—模型—知识协同的混合智能模式。在叙事可视化方面,大模型显著降低了数据叙事的创作门槛,使系统能够自动构建叙事结构、整合图文内容,并根据受众和情境动态调整表达方式。在可视化评估方面,大模型在图形质量评估与设计建议生成方面展现出潜力。本文分析当前面临的关键问题与发展趋势,为大模型时代的数据可视化研究与系统设计提供结构化的理论支撑。

【Abstract】 Data visualization, as a fundamental technology supporting human cognition and data-driven scientific reasoning, is undergoing a profound shift driven by the rapid development of large-scale models and is attracting widespread attention from the academia and industry. For decades, traditional visualization approaches have mainly relied on manually or automatically designed visual encoding rules, graphical grammars, and human-computer interaction techniques. Through explicit visual mappings, graphical composition, and interface operations, these methods enable users to observe patterns, understand structures, and communicate insights from data. However, with the continuous growth of data size, task complexity, and decision-making contexts, statistical visual mappings and parameter-driven interaction are increasingly insufficient to support modern analytical needs. Instead, users now demand more than visual presentation alone; they require systems capable of semantic understanding, task-driven reasoning, and cross-modal information integration. The emergence of large models——particularly large language models, vision-language foundation models, and intelligent agents——has introduced unprecedented technological momentum to the visualization field. These models are not only transforming how visualizations are generated and interacted with but also reshaping the theoretical foundations, analytical workflows, and narrative practices of data visualization. Their strong capabilities in semantic representation, abstraction, and reasoning enable visualization systems to move beyond surface-level depiction toward deeper support for analytical intent and knowledge construction. From a theoretical perspective, visualization research has long been grounded in the Grammar of Graphics and declarative specification languages, which provide highly abstract and formalized descriptions of graphical structures, data mappings, and interaction logic. The introduction of large models does not replace this; instead, it reinforces its importance. Declarative grammars increasingly function as an interpretable “intermediate language” between naturallanguage intent and executable visual representations, enabling large models to translate high-level analytical goals reliably into controlled, consistent visualization specifications. This mediation is critical for ensuring the interpretability, reproducibility, and controllability of automatically generated visualizations. Moreover, the semantic reasoning capabilities of large models allow them to infer the underlying intent and logic behind visual layouts, color encodings, and spatial structures, and pushes visualization research from low-level perceptual encoding toward semantic-driven visual understanding and intent modeling. At the same time, techniques such as differentiable rendering, neural implicit representations, and Gaussian splatting provide new frameworks for high-fidelity scientific rendering, continuous data representation, and optimization in parameterized spaces. These methods enable visualization systems to become differentiable, learnable, and optimizable, and allows tighter integration between visual representation, computational models, and analytical objectives. As a result, visualization is increasingly situated within richer signal and representation spaces, supporting adaptive, expressive, and controllable visual analysis. Beyond theoretical advances, large models are fundamentally changing how users collaborate with visualization systems. Traditionally, effective visualization required users to master visualization grammars, tool-specific operations, and data processing workflows. By contrast, large-model-driven systems allow users to express analytical goals, design preferences, and data-related questions directly through natural language. The system can then interpret user intent, generate appropriate visualizations, restructure data views, or optimize visual encodings accordingly. This fusion of model-based semantic knowledge with formal visualization representations forms the foundation of a new generation of intelligent visualization systems. At the level of visual analytics, large models and agents are driving a transition from human-centered interaction toward a hybrid collaborative paradigm involving humans, agents, and knowledge. Early machine-learning-for-visualization research primarily focused on addressing isolated subtasks, such as view recommendation, feature detection, or layout optimization. By contrast, visualization frameworks based on large language models(LLMs) leverage unified semantic understanding across tasks and modalities to support end-to-end analytical pipelines. These systems can assist with visualization generation, pattern discovery, and reasoning-based explanation, and enable more holistic support for complex analytical processes. As human-machine collaboration evolves from simple question answering toward knowledge generation, large models increasingly function as proactive cognitive collaborators rather than passive assistants. They can produce trend interpretations, support hypothesis testing, and summarize underlying mechanisms while continuously modeling user behavior, analytical state, and task context. This approach enables a shift from reactive assistance to mixed-initiative collaboration, where systems actively guide users through complex, dynamic, and uncertain analytical environments, augmenting human reasoning and sensemaking capabilities. In the domain of narrative visualization, large models mark the beginning of a new era of intelligent content generation. Traditional narrative visualization requires expertise in design, storytelling, writing, and programming, which makes producing high-quality narratives costly and time consuming. Large models significantly lower this barrier by automatically identifying data themes, extracting salient trends, constructing narrative structures, and integrating visual, textual, and multimodal elements into coherent, stylistically consistent narratives. This capability substantially improves efficiency and accessibility and enables end-to-end automation from data to story in applications such as data journalism, educational visualization, science communication, and business reporting. With the maturation of multimodal generation technologies, narrative processes can further incorporate natural-language interaction, dynamically adapt narrative perspectives, and tailor content to audiences with different backgrounds and goals. In visualization evaluation, large models demonstrate promising potential due to their understanding of visual aesthetics, layout quality, perceptual principles, and readability. They can automatically assess visualization quality, detect misleading encodings, propose design improvements, and generate multiple design alternatives for comparison. However, model-based visual judgment is not inherently equivalent to human perception, a situation raising critical research challenges. Ensuring alignment between model assessments and human visual cognition, mitigating hallucinations and biases, and improving transparency, reliability, and interpretability remain central issues in largemodel-driven visualization research. To organize the rapidly evolving landscape of large-model-powered data visualization systematically, this paper presents a comprehensive survey from four perspectives: visualization fundamentals, visual analytics, narrative visualization, and visualization evaluation. This paper reviews representative research advances in each area, analyzes emerging technologies enabled by large models, and discusses key challenges and future research directions. By providing a structured knowledge map and theoretical framework, this survey aims to support future innovation in intelligent visualization systems and contribute to the development of data visualization as a critical bridge between human intelligence and artificial intelligence in the era of large models.

【基金】 国家自然科学基金项目(62132017,62472099,62572191,62572415,62372271)~~
  • 【文献出处】 中国图象图形学报 ,Journal of Image and Graphics , 编辑部邮箱 ,2026年06期
  • 【分类号】TP311.13;TP18
  • 【下载频次】58
节点文献中: 

本文链接的文献网络图示:

本文的引文网络