1.华东师范大学通信与电子工程学院,上海 200241
2.湖南大学机械与运载工程学院,湖南 长沙 410082
3.澳门科技大学创新工程学院,澳门 999078
4.中国科学院自动化研究所多模态人工智能系统全国重点实验室,北京 100190
5.麻省理工学院信息与决策系统实验室,剑桥 02139-4307
[ "李柏(1989- ),男,博士,华东师范大学通信与电子工程学院教授,主要研究方向为自动驾驶轨迹规划、计算最优控制理论、多模态大语言模型与具身智能系统。" ]
[ "郝金第(2002- ),男,湖南大学机械与运载工程学院硕士生,主要研究方向为大语言模型、计算最优控制方法在智慧医疗领域的应用。" ]
[ "孙跃硕(2002- ),男,湖南大学机械与运载工程学院硕士生,主要研究方向为自主无人系统协同运动规划、计算最优控制以及具身智能机器人。" ]
[ "孟雨晴(2001- ),女,华东师范大学通信与电子工程学院硕士生,主要研究方向为多模态大语言模型与具身智能在自动驾驶领域的应用。" ]
[ "黄峻(1998- ),男,澳门科技大学创新工程学院博士生,主要研究方向为平行智能、自动驾驶轨迹预测规划、提示工程以及大语言模型。" ]
[ "田永林(1994- ),男,博士,中国科学院自动化研究所多模态人工智能系统全国重点实验室助理研究员,主要研究方向为平行系统、自动驾驶以及场景工程。" ]
[ "贺正冰(1981- ),男,博士,麻省理工学院信息与决策系统实验室高级研究员,主要研究方向为城市交通、系统工程与人工智能的交叉领域。" ]
收稿:2025-08-02,
修回:2025-09-16,
录用:2025-09-18,
纸质出版:2025-09-15
移动端阅览
李柏,郝金第,孙跃硕等.平行智能范式视角下的视觉-语言-动作模型发展现状与展望[J].智能科学与技术学报,2025,07(03):290-303.
LI Bai,HAO Jindi,SUN Yueshuo,et al.Vision-language-action models under parallel intelligence paradigm: the state of the art and future perspectives[J].Chinese Journal of Intelligent Science and Technology,2025,07(03):290-303.
李柏,郝金第,孙跃硕等.平行智能范式视角下的视觉-语言-动作模型发展现状与展望[J].智能科学与技术学报,2025,07(03):290-303. DOI: 10.11959/j.issn.2096-6652.202536.
LI Bai,HAO Jindi,SUN Yueshuo,et al.Vision-language-action models under parallel intelligence paradigm: the state of the art and future perspectives[J].Chinese Journal of Intelligent Science and Technology,2025,07(03):290-303. DOI: 10.11959/j.issn.2096-6652.202536.
视觉-语言-动作模型是一类面向具身智能的综合性建模方法,它将视觉感知、自然语言理解和动作执行在统一框架下进行表征与学习,旨在实现从环境感知到任务规划再到动作控制的连续闭环。视觉-语言-动作模型的运行逻辑与21世纪初提出的平行智能范式存在呼应。平行智能通过“人工系统、计算实验和平行执行”的三元架构,强调虚拟建模、可复现推演以及虚实交互的闭环机制,这些理念与视觉-语言-动作模型的发展路径在不同阶段形成了对应关系:早期的多模态探索可视为人工系统中的原型实践,随后的大规模模型与跨域训练扩展了计算实验的能力,近年来的分层控制和虚实闭环则体现了平行执行强调的反馈修正与规范指导。在这一框架下,视觉-语言-动作模型呈现出语义与动作深度耦合、虚拟与现实双向循环和可验证性增强等特征,但也存在泛化不足、语义对齐不稳、安全与解释机制薄弱以及部署效率受限等问题。未来研究可聚焦任务语义的契约化表达、长时序规划的可修复性、世界模型的工程化使用、多层次反馈与安全治理,以及跨平台迁移和人机协作等方向。以平行智能为参照重新审视视觉-语言-动作模型,不仅有助于厘清发展脉络,也为其在真实场景中的可信应用提供了方法论支持。
Vision-language-action (VLA) models are a comprehensive modeling approach for embodied intelligence that integrates visual perception
natural language understanding
and action execution within a unified framework
aiming to establish a continuous loop from environmental perception to task planning and action control. Their operational logic corresponds closely to the paradigm of parallel intelligence articulated in the early 21st century. That paradigm comprises artificial systems
computational experiments
and parallel execution
emphasizing virtual modeling
reproducible inference
and closed-loop interaction between the virtual and the real. The initial stage of VLA development
driven by multimodal deep learning
can be regarded as prototypical work within artificial systems
the subsequent stage characterized by large-scale models and cross-domain training expanded the scope of computational experiments
and the more recent focus on hierarchical control and virtual-real closed loops reflects the feedback correction and normative guidance emphasized in parallel execution. VLA models exhibit deep coupling between semantics and action
iterative cycles linking simulation with reality
and steadily improving verifiability. Nonetheless
challenges remain in generalization
semantic alignment
safety and interpretability
and deployment efficiency. Addressing these issues calls for contract-based task semantics
repairable long-horizon hierarchical planning
engineering-oriented use of world models
multi-level feedback and safety governance
and cross-platform transfer with human-machine collaboration. Examining VLA through the lens of parallel intelligence clarifies its developmental logic and provides methodological support for advancing toward trustworthy real-world applications.
SHAO R, LI W, ZHANG L, et al. Large VLM-based vision-language-action models for robotic manipulation: a survey[J]. arXiv preprint, 2025, arXiv: 2508.13073.
SAPKOTA R, CAO Y, ROUMELIOTIS K I, et al. Vision-language-action models: concepts, progress, applications and challenges[J]. arXiv preprint, 2025, arXiv: 2505.04769.
王飞跃. 平行系统方法与复杂系统的管理和控制[J]. 控制与决策, 2004, 19(5): 485-489, 514.
WANG F Y. Parallel system methods for management and control of complex systems[J]. Control and Decision, 2004, 19(5): 485-489,514.
WANG F Y, TANG S M. Artificial societies for integrated and sustainable development of metropolitan systems[J]. IEEE Intelligent Systems, 2004, 19(4): 82-87.
WANG X X, YANG J, LIU Y H, et al. Parallel intelligence in three decades: a historical review and future perspective on ACP and cyber-physical-social systems[J]. Artificial Intelligence Review, 2024, 57(9): 255.
HU X M, LI S, HUANG T Y, et al. How simulation helps autonomous driving: a survey of sim2real, digital twins, and parallel intelligence[J]. IEEE Transactions on Intelligent Vehicles, 2023, 9(1): 593-612.
WANG F Y, TANG S M. A framework for artificial transportation systems: from computer simulations to computational experiments[C]//Proceedings of 2005 IEEE Intelligent Transportation Systems. Piscataway: IEEE Press, 2005: 1130-1134.
吕宜生, 王飞跃, 张宇, 等. 虚实互动的平行城市: 基本框架、方法与应用[J]. 智能科学与技术学报, 2019, 1(3): 311-317.
LYU Y S, WANG F Y, ZHANG Y, et al. Parallel cities: framework, methodology, and application[J]. Chinese Journal of Intelligent Science and Technology, 2019, 1(3): 311-317.
陈龙, 王晓, 杨健健, 等. 平行矿山: 从数字孪生到矿山智能[J]. 自动化学报, 2021, 47(7): 1633-1645.
CHEN L, WANG X, YANG J J, et al. Parallel mining operating systems: from digital twins to mining intelligence[J]. Acta Automatica Sinica, 2021, 47(7): 1633-1645.
周芳, 王瑞. 基于平行系统的网络舆情试验方法[J]. 指挥信息系统与技术, 2013, 4(3): 1-7.
ZHOU F, WANG R. Test method for network opinion based on parallel system[J]. Command Information System and Technology, 2013, 4(3): 1-7.
王惠珍, 张捷, 俞怡, 等. 平行手术室:围术期护理流程与智慧手术平台管理的新模式[J]. 模式识别与人工智能, 2023, 36(10): 867-876.
WANG H Z, ZHANG J, YU Y, et al. Parallel operating rooms: a new model of perioperative nursing process and smart surgical platform management[J]. Pattern Recognition and Artificial Intelligence, 2023, 36(10): 867-876.
李柏, 宋秭函, 李鑫源, 等. 基于代理智能的平行厨师:从AI Agents到智慧数字机器人饮食系统[J]. 模式识别与人工智能, 2025, 38(3): 252-267.
LI B, SONG Z H, LI X Y, et al. Parallel chefs via agentic intelligence: from AI agents to smart digital robotic cuisine systems[J]. Pattern Recognition and Artificial Intelligence, 2025, 38(3): 252-267.
王飞跃. 数字教师与平行教育: 关于ChatGPT之后教学变革的探讨[J]. 智能科学与技术学报, 2023, 5(4): 446-455.
WANG F Y. Digital teachers and parallel education: a paradigm shift in teaching and learning after ChatGPT[J]. Chinese Journal of Intelligent Science and Technology, 2023, 5(4): 446-455.
张腾超, 田永林, 林飞, 等. 平行旅游:基础智能驱动的智慧出游服务[J]. 智能科学与技术学报, 2024, 6(2): 164-178.
ZHANG T C, TIAN Y L, LIN F, et al. Parallel tourism: foundation intelligence driven smart trip services[J]. Chinese Journal of Intelligent Science and Technology, 2024, 6(2): 164-178.
王飞跃. 指控5.0:平行时代的智能指挥与控制体系[J]. 指挥与控制学报, 2015, 1(1): 107-120.
WANG F Y. CC5.0: intelligent command and control systems in the parallel age[J]. Journal of Command and Control, 2015, 1(1): 107-120.
王飞跃, 王雨桐. 数字科学家与平行科学:AI4S和S4AI的本源与目标[J]. 中国科学院院刊, 2024, 39(1): 27-33.
WANG F Y, WANG Y T. Digital scientists and parallel sciences: the origin and goal of AI for science and science for AI[J]. Bulletin of Chinese Academy of Sciences, 2024, 39(1): 27-33.
国务院. 国发〔2025〕11号. 国务院关于深入实施“人工智能+”行动的意见[S]. 2025.
STATE COUNCIL. Guo Fa [2025] No. 11. Opinions of the state council on further implementing the “AI+” initiative[S]. 2025.
王飞跃. 平行控制: 数据驱动的计算控制方法[J]. 自动化学报, 2013, 39(4): 293-302.
WANG F Y. Parallel control: a method for data-driven and computational control[J]. Acta Automatica Sinica, 2013, 39(4): 293-302.
ZITKOVICH B, YU T H, XU S C, et al. Rt-2: vision-language-action models transfer web knowledge to robotic control[C]//Proceedings of the 7th Conference on Robot Learning. Cambridge: PMLR, 2023: 2165-2183.
李浩然, 陈宇辉, 崔文博, 等. 面向具身操作的视觉-语言-动作模型综述[J]. arXiv preprint, 2025, arXiv: 2508.15201.
LI H R, CHEN Y H, CUI W B, et al. Survey of vision-language-action models for embodied manipulation[J]. arXiv preprint, 2025, arXiv: 2508.15201.
KRIZHEVSKY A, SUTSKEVER I, HINTON G E. ImageNet classification with deep convolutional neural networks[J]. Communications of the ACM, 2017, 60(6): 84-90.
DONG S, WANG P, ABBAS K. A survey on deep learning and its applications[J]. Computer Science Review, 2021, 40: 100379.
SILVER D, HUANG A, MADDISON C J, et al. Mastering the game of Go with deep neural networks and tree search[J]. Nature, 2016, 529(7587): 484-489.
TANWANI A K, YAN A, LEE J, et al. Sequential robot imitation learning from observations[J]. The International Journal of Robotics Research, 2021, 40(10/11): 1306-1325.
SHRIDHAR M, MANUELLI L, FOX D. CLIPort: what and where pathways for robotic manipulation[C]//Proceedings of the 5th Conference on Robot Learning. Cambridge: PMLR, 2021: 894-906.
VASWANI A, SHAZEER N, PARMAR N, et al. Attention is all you need[C]//Proceedings of the 31st Conference on Neural Information Processing Systems. Red Hook: Curran Associates Inc, 2017: 5998-6008.
BROHAN A, BROWN N, CARBAJAL J, et al. Rt-1: robotics transformer for real-world control at scale[J]. arXiv preprint, 2022, arXiv: 2212.06817.
KIM M J, PERTSCH K, KARAMCHETI S, et al. OpenVLA: an open-source vision-language-action model[J]. arXiv preprint, 2024, arXiv: 2406.09246.
GE Z, CHEN C, SINHA A, et al. On learning informative trajectory embeddings for imitation, classification and regression[J]. arXiv preprint, 2025, arXiv: 2501.09327.
LIANG J, HUANG W L, XIA F, et al. Code as policies: language model programs for embodied control[J]. arXiv preprint, 2022, arXiv: 2209.07753.
HUANG W L, WANG C, ZHANG R H, et al. VoxPoser: composable 3D value maps for robotic manipulation with language models[J]. arXiv preprint, 2023, arXiv: 2307.05973.
YE S, JANG J, JEON B, et al. Latent action pretraining from videos[J]. arXiv preprint, 2024, arXiv: 2410.11758.
ZAWALSKI M, CHEN W, PERTSCH K, et al. Robotic control via embodied chain-of-thought reasoning[J]. arXiv preprint, 2024, arXiv: 2407.08693.
O'NEILL A, REHMAN A, MADDUKURI A, et al. Open x-embodiment: Robotic learning datasets and rt-x models: Open x-embodiment collaboration[C]//Proceedings of the 2024 IEEE International Conference on Robotics and Automation (ICRA). Piscataway: IEEE Press, 2024: 6892-6903.
KHAZATSKY A, PERTSCH K, NAIR S, et al. DROID: a large-scale in-the-wild robot manipulation dataset[J]. arXiv preprint, 2024, arXiv: 2403.12945.
FANG H S, FANG H J, TANG Z Y, et al. RH20T: a comprehensive robotic dataset for learning diverse skills in one-shot[C]//Proceedings of the 2024 IEEE International Conference on Robotics and Automation (ICRA). Piscataway: IEEE Press, 2024: 653-660.
MEES O, HERMANN L, ROSETE-BEAS E, et al. CALVIN: a benchmark for language-conditioned policy learning for long-horizon robot manipulation tasks[J]. IEEE Robotics and Automation Letters, 2022, 7(3): 7327-7334.
JAMES S, MA Z C, ARROJO D R, et al. RLBench: the robot learning benchmark & learning environment[J]. IEEE Robotics and Automation Letters, 2020, 5(2): 3019-3026.
GUPTA A, KUMAR V, LYNCH C, et al. Relay policy learning: Solving long-horizon tasks via imitation and reinforcement learning[J]. arXiv preprint, 2019, arXiv: 1910.11956.
BLACK K, BROWN N, DRIESS D, et al. π0: a vision-language-action flow model for general robot control[J]. arXiv preprint, 2024, arXiv: 2410.24164.
YANG R H, YU Q X, WU Y C, et al. EgoVLA: learning vision-language-action models from egocentric human videos[J]. arXiv preprint, 2025, arXiv: 2507.12440.
FANG Y, YANG Y, ZHU X H, et al. ReBot: scaling robot learning with real-to-sim-to-real robotic video synthesis[J]. arXiv preprint, 2025, arXiv: 2503.14526.
CHI C, XU Z J, FENG S Y, et al. Diffusion policy: Visuomotor policy learning via action diffusion[J ] . The International Journal of Robotics Research, 2025, 44(10/11): 1684-1704.
MASCARO R, CHLI M. Scene representations for robotic spatial perception[J]. Annual Review of Control, Robotics, and Autonomous Systems, 2025, 8: 351-377
IKEUCHI K, WAKE N, SASABUCHI K, et al. Semantic constraints to represent common sense required in household actions for multimodal learning-from-observation robot[J]. The International Journal of Robotics Research, 2024, 43(2): 134-170.
MIAO R Q, JIA Q X, SUN F C, et al. Semantic representation of robot manipulation with knowledge graph[J]. Entropy, 2023, 25(4): 657.
GARG S, SÜNDERHAUF N, DAYOUB F, et al. Semantics for robotic mapping, perception and interaction: a survey[J]. Foundations and Trends in Robotics, 2020, 8(1-2): 1-224.
ZHAO Z G, CHENG S, DING Y, et al. A survey of optimization-based task and motion planning: From classical to learning approaches[J]. IEEE/ASME Transactions on Mechatronics, 2025, 30(4): 2799-2825.
SZOT A, CLEGG A, UNDERSANDER E, et al. Habitat 2.0: training home assistants to rearrange their habitat[J]. Advances in Neural Information Processing Systems, 2021, 34: 251-266.
HU Y C, GUO Y J, WANG P C, et al. Video prediction policy: a generalist robot policy with predictive visual representations[J]. arXiv preprint, 2024, arXiv: 2412.14803.
VRABIČ R, ŠKULJ G, MALUS A, et al. An architecture for sim-to-real and real-to-sim experimentation in robotic systems[J]. Procedia CIRP, 2021, 104: 336-341.
DEITKE M, VANDERBILT E, HERRASTI A, et al. Procthor: large-scale embodied AI using procedural generation[J]. Advances in Neural Information Processing Systems, 2022, 35: 5982-5994.
AHMED O, TRÄUBLE F, GOYAL A, et al. CausalWorld: a robotic manipulation benchmark for causal structure and transfer learning[J]. arXiv preprint, 2020, arXiv: 2010.04296.
ALT B, PICKLUM M, ARION S, et al. Open, reproducible and trustworthy robot-based experiments with virtual labs and digital-twin-based execution tracing[J]. arXiv preprint, 2025, arXiv: 2508.11406.
LUO J L, XU C, LIU F C, et al. FMB: a functional manipulation benchmark for generalizable robotic learning[J]. International Journal of Robotics Research, 2025, 44(4): 592-606.
LI X L, HSU K, GU J Y, et al. Evaluating real-world robot manipulation policies in simulation[J]. arXiv preprint, 2024, arXiv: 2405.05941.
WILLIBALD C, LEE D. Hierarchical task decomposition for execution monitoring and error recovery: Understanding the rationale behind task demonstrations[J]. arXiv preprint, 2025, arXiv: 2505.04565.
LI Y X, ZHU Y C, WEN J J, et al. WorldEval: world model as real-world robot policies evaluator[J]. arXiv preprint, 2025, arXiv: 2505.19017.
SHI L, XU Y X, WANG S Y, et al. An real-sim-real (RSR) loop framework for generalizable robotic policy transfer with differentiable simulation[J]. arXiv preprint, 2025, arXiv: 2503.10118.
YU Y W, LIU L T. Neural fidelity calibration for informative sim-to-real adaptation[J]. arXiv preprint, 2025, arXiv: 2504.08604.
THODUKA S, HOUBEN S, GALL J, et al. Enhancing video-based robot failure detection using task knowledge[J]. arXiv preprint, 2025, arXiv: 2508.18705.
BAR A, ZHOU G Y, TRAN D, et al. Navigation world models[C]//Proceedings of the 2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). Piscataway: IEEE Press, 2025: 15791-15801.
CUI Z J, BAENA F R Y. Forbidden region dynamic active constraints in robot-assisted minimally invasive surgery[J]. IEEE Robotics and Automation Letters, 2025, 10(3): 2950-2957.
KUBE D, HADWIGER S, MEISEN T. Beyond performance: explaining generalisation failures of robotic foundation models in industrial simulation[J]. Biomimetic Intelligence and Robotics, 2025: doi.org/10.1016/j.birob.2025.100249.
SEO M, PARK H A, YUAN S L, et al. LEGATO: cross-embodiment imitation using a grasping tool[J]. IEEE Robotics and Automation Letters, 2025, 10(3): 2854-2861.
张慧, 梁姝彤, 李明轩, 等. 视觉-语言-动作模型综述: 从前史到前沿[J]. 自动化学报, 2025, 51(9): 1922-1950.
ZHANG H, LIANG S T, LI M X, et al. Vision-language-action models: from the early foundations to the state-of-the-art[J]. Acta Automatica Sinica, 2025, 51(9): 1922-1950.
0
浏览量
422
下载量
0
CSCD
关联资源
相关文章
相关作者
相关机构
京公网安备11010802024621
