Selected publications 21
2026First author or project leader. Click a figure to enlarge.
* Equal contribution † Project leader
SAGE-3D: Semantic-Aware 3D Representations for Generalizable Vision-Language-Action Models
Towards More Efficient Decoding for Autoregressive Vision-language-action Models
Consistency distillation and early-exit decoding accelerate autoregressive VLA inference.
HiF-VLA: Hindsight, Insight and Foresight through Motion Representation for Vision-Language-Action Models
Motion-based hindsight and foresight support temporally coherent, long-horizon manipulation.
CUBic: Coordinated Unified Bimanual Perception and Control Framework
A shared perceptual representation connects independent arm observations with coordinated bimanual control.
Unified Diffusion VLA: Vision-Language-Action Model via Joint Discrete Denoising Diffusion Process
Future visual observations and robot actions are refined together through a joint discrete diffusion process.
TrajBooster: Boosting Humanoid Whole-Body Manipulation via Trajectory-Centric Learning
Retargeted end-effector trajectories transfer wheeled-humanoid data to bipedal whole-body manipulation.
VLA-Adapter: An Effective Paradigm for Tiny-Scale Vision-Language-Action Model
Bridge Attention connects a compact vision-language backbone to an action policy.
Earlier selected work13 papers · 2022–2025
Long-VLA: Unleashing Long-Horizon Capability of Vision Language Action Model for Robot Manipulation
CARP: Visuomotor Policy Learning via Coarse-to-Fine Autoregressive Prediction
A coarse-to-fine autoregressive policy progressively refines multi-scale action sequences.
Accelerating vision-language-action model integrated with action chunking via parallel decoding
QUART-Online: Latency-Free Large Multimodal Language Model for Quadruped Robot Learning
Action chunk discretization enables faster multimodal policies for quadruped robots.
VLAS: Vision-Language-Action Model with Speech Instructions for Customized Robot Manipulation
QUAR-VLA: Vision-Language-Action Model for Quadruped Robots
Visual observations and language instructions guide quadruped perception, navigation, and action.