Researcher & Founder
Pengxiang Ding 丁鹏翔
I am the founder and CEO of Symbiosis Robotics (共生知行), where we build foundation models for whole-body humanoid intelligence.
My research focuses on vision-language-action models and whole-body control.
I am a Ph.D. student at Zhejiang University in a joint program with Westlake University, advised by Prof. Donglin Wang in the Machine Intelligence Laboratory (MiLAB).
Previously, I received my M.Sc. from the School of Artificial Intelligence at BUPT in 2022, advised by Prof. Jianqin Yin.
News
Scroll for earlier updates- NeurIPS1 paper accepted: Endowing Your Vision-Language-Action Model with a Predictive Mind.
- CoRL1 paper accepted: SAGE-3D.
- ECCV2 papers accepted: Fast-dVLA and Towards More Efficient Decoding for Autoregressive Vision-language-action Models.
- ICML1 paper accepted: Dyn-VPP.
- CVPR2 papers accepted.
- ICRA3 papers accepted.
- ICLR2 papers accepted.
- AAAI4 papers accepted.
- NeurIPS1 paper accepted.
- CoRL1 paper accepted.
- ICCV1 paper accepted.
- ICML3 papers accepted.
Selected publications 21
Google ScholarFirst author or project leader. * Equal contribution † Project leader
SAGE-3D: Semantic-Aware 3D Representations for Generalizable Vision-Language-Action Models
Towards More Efficient Decoding for Autoregressive Vision-language-action Models
HiF-VLA: Hindsight, Insight and Foresight through Motion Representation for Vision-Language-Action Models
CUBic: Coordinated Unified Bimanual Perception and Control Framework
Unified Diffusion VLA: Vision-Language-Action Model via Joint Discrete Denoising Diffusion Process
TrajBooster: Boosting Humanoid Whole-Body Manipulation via Trajectory-Centric Learning
VLA-Adapter: An Effective Paradigm for Tiny-Scale Vision-Language-Action Model
Long-VLA: Unleashing Long-Horizon Capability of Vision Language Action Model for Robot Manipulation
Humanoid-vla: Towards universal humanoid control with visual integration
Accelerating vision-language-action model integrated with action chunking via parallel decoding
GEVRM: Goal-Expressive Video Generation Model For Robust Visual Manipulation
VLAS: Vision-Language-Action Model with Speech Instructions for Customized Robot Manipulation
Other publications 8
2026Fast-dVLA: Accelerating Discrete Diffusion VLA to Real-Time Performance
Dyn-VPP: Video Prediction Policy Optimization for Improved Visual Dynamics
Rethinking the Practicality of Vision-Language-Action Model: A Comprehensive Benchmark and an Improved Baseline
Robust Online Residual Refinement via Koopman-Guided Dynamics Modeling
ReconVLA: Reconstructive Vision-Language-Action Model as Effective Robot Perceiver
Filter, Correlate, Compress: Training-Free Token Reduction for MLLM Acceleration
Rethinking Target Label Conditioning in Adversarial Attacks: A 2D Tensor-Guided Generative Approach
Experience
- –
Xiaomi 小米
Intern
- –
Ant Group · Robbyant
Intern 蚂蚁灵波
- Jan 2025–Jun 2025
DAMO Academy
Research Intern · Machine Intelligence Laboratory
Advisor: Xin Li
- Sep 2022–Mar 2023
RedBook
Research Intern · Intelligent Creation Group
Advisor: Haofan Wang
- Sep 2021–Mar 2022
SenseTime
Research Intern · Smart City Group
Advisor: Dongliang Wang
Academic service
Journal & conference reviewer
ICML, ICLR, NeurIPS, CVPR, ICCV, ACM MM, AAAI, ICRA, IROS, CoRL, TNNLS, TASE, TCSVT
Elsewhere
You can also find me on Redbook .
Talks & discussions 7
-
What Are AI Builders Changing?
Chongli Forum · AI Builders Show · 甲子光年
-
World Models, Robot Intelligence, and Control
WRC 2026 Developer After Party · 北京人形 / Xbotics
-
Advancing VLA Models across Robot Embodiments
EAIRCon 2025 · Shenzhen · 智猩猩 / 智东西
-
VLA Foundations, Reinforcement Learning, and Dexterous Manipulation
Lumina Talk #19 · Westlake MiLAB
-
Full-Stack VLA: From Visuomotor Policies to Robot Foundation Models
3D Vision Workshop · 3D 视觉工坊
-
End-to-End Foundation Models for Quadruped Robots
Shenlan Academy · 深蓝学院
-
Vision-Language-Action Models
Peking University · Hosted by Hao Dong