2026
CoRL 2026 Paper Guide | ELAN4D: Learning from Future Robot Motion
Accepted to CoRL 2026 · Vision-Language-Action · 4D Supervision
Paper: ELAN4D: Embodiment-Centric 4D Supervision for Vision-Language-Action Models via Plug-and-Play Adaptation.
Success in a training scene does not guarantee success after changing the background, viewpoint, or object positions. ELAN4D trains an action policy to anticipate the robot’s own future motion, using inexpensive trajectory supervision to improve VLA generalization.
Collaboration highlight: Fan Mo is a coauthor of this work. The author team includes Bowen Yang and Professor Junchi Yan from Shanghai Jiao Tong University. Yan’s faculty appointment is confirmed by SJTU’s School of Artificial Intelligence. Full authorship and paper affiliations are listed below.
ELAN4D in one minute
- Target: Future 3D displacement tracks of robot joints and end-effector keypoints. “4D” means three-dimensional motion over time.
- Construction: Forward kinematics applied to recorded proprioceptive trajectories and a known kinematic chain, without external video trackers for robot tracks.
- Integration: A ControlNet-style action branch, gradient isolation, and zero-initialized residual projection.
- Inference: Remove the track decoder, retain the learned control branch, and preserve the base image–language–proprioception-to-action interface.
The overview below summarizes the trade-off among out-of-distribution robustness, 4D preprocessing cost, and interference with pretrained representations.

Figure 1 | Motivation and overview. Original Figure 1, PDF p. 2. Open full-resolution figure
01 | From current observations to future consequences
VLA policies connect visual observations and language instructions to actions. The paper asks whether reactive action regression sufficiently captures the future spatial changes induced by those actions, especially under unfamiliar views, backgrounds, or layouts.
Predicting whole images can devote capacity to static context and appearance changes, while generating dense scene trajectories may require costly tracking or reconstruction. ELAN4D focuses its auxiliary prediction on the robot embodiment itself.
02 | Building 4D supervision with forward kinematics
For recorded joint state q(t), the known kinematic model gives each keypoint position: pₖ(t) = FKₖ(q(t)). Subtracting the current positions from future positions produces ΔP(t+h) = P(t+h) − P(t). Stacking these displacements gives an H × K × 3 target: horizon, keypoints, and spatial coordinates.
The signal comes from robot state logs and is not directly subject to camera occlusion. The reported setups use 8 keypoints for LIBERO / LIBERO-Plus, 14 for RoboTwin2.0, and 7 for real-world tasks. For an hour of data, the paper compares roughly four GPU-hours of video-based tracking with around one CPU-minute for robot tracks. This comparison concerns supervision preprocessing, not the time required to train the complete model.
03 | A plug-and-play branch for the action expert
ELAN4D is built on π₀ and π₀.₅, which combine a vision-language backbone with a conditional-flow-matching action expert. A residual control branch injects features through a zero-initialized projection, so the initial residual contribution is near zero.
A lightweight decoder uses control features and current robot keypoints to predict future displacements. Training uses L = L_act + λ_track L_track. The action loss supervises the action-generation pathway and control branch. The track loss updates the control branch and track decoder, with stop-gradient preventing its direct propagation into the pretrained VLM and original action branch.

Figure 2 | Action expert, residual control branch, and track decoder. Original Figure 2, PDF p. 4. Open full-resolution figure
What remains at deployment? The track decoder and its dedicated 3D keypoint input are removed, while the residual control branch remains in the action expert. The policy retains the base language, image, and proprioception inputs and action outputs. An unchanged interface does not imply identical inference computation, nor that the backbone is frozen under every training objective.
04 | Simulation: larger gains under distribution shifts
On the nearly saturated original LIBERO benchmark, overall gains are small: π₀ improves from 94.2% to 95.0%, and π₀.₅ from 96.9% to 97.0%. Gains are clearer on the systematically perturbed LIBERO-Plus:
- π₀: 53.6% → 67.6%, a 14.0-percentage-point increase.
- π₀.₅: 73.6% → 78.2%, a 4.6-percentage-point increase.
- Training data: Both LIBERO evaluations use policies trained on original LIBERO demonstrations, without training on LIBERO-Plus augmentations.
The table breaks results down by perturbation and task suite. Higher overall averages do not imply a win on every submetric.

Figure 3 | LIBERO-Plus success rates (%). Original Table 1, PDF p. 7; improvements above are percentage points. Open full-resolution figure
Across eight unseen bimanual settings in RoboTwin2.0, overall success rises from 12% to 15% for π₀ and from 32% to 37% for π₀.₅. Each task is evaluated over 100 trials. Improvements are notable on spatial tasks such as Dump Bin and Lift Pot; these remain benchmark-specific results rather than guarantees for arbitrary robotic tasks.
05 | Real robots: distractors, new positions, and assembly
The real-world suite uses an AgileX Piper arm for fruit placement with unseen distractors, cup stacking at unseen positions, and a two-stage block assembly. Each category uses 50 expert demonstrations and 20 evaluation trials.
- Visual robustness: π₀.₅ improves from 50% to 80%.
- Spatial generalization: 15% to 65%.
- Sequential assembly: 5% to 45%.
The figure reports an overall improvement from approximately 23% to 63%. Within these tested settings, embodiment-centric future-motion supervision helps with distractors and compounding errors; the evidence is limited by the task scope and trial count.

Figure 4 | Real-world tasks and success rates. Original Figure 4, PDF p. 7. Open full-resolution figure
06 | What do the ablations explain?
Adding the control branch without the 4D track loss gives 73.3% on LIBERO-Plus, close to the π₀.₅ baseline’s 73.6%, versus 78.2% for the complete method. This supports a contribution from the supervision rather than simply additional parameters.
Predicting tracks through queries inserted into the VLM gives 66.8%, below the separate control-branch design. CKA analysis also shows more representation drift for the query-token variant, supporting the gradient-isolation design.
Richer whole-scene tracks obtained from simulator ground truth reach 79.3%, 1.1 percentage points above robot-only supervision. Choosing robot keypoints is a trade-off between supervision richness and acquisition cost.

Figure 5 | Supervision placement, prediction targets, and data efficiency. Original Figure 5, PDF p. 8. Open full-resolution figure
07 | Practical takeaways and boundaries
The work turns readily available robot state records into a predictive training signal: demonstrations can supervise both actions and their associated embodiment motion. This offers a direction for improving existing VLA training pipelines.
- Reliable state and kinematics are required: the supervision depends on recorded robot states and a known kinematic chain.
- External dynamics are not directly supervised: sparse robot tracks may be insufficient when external object motion, deformation, or complex contact dominates task success.
- Separate training from inference: future tracks supervise training; inference drops the decoder but retains the control branch.
- Scope remains empirical: the reported experiments do not establish effectiveness for every model, embodiment, or task.
08 | Authors and collaborating institutions
Full author list, in paper order: Zeyuan He, Bowen Yang, Zhirui Fang, Keru Zhou, Lei Jiang, Jingjing Qian, Fan Mo, Junchi Yan, Philip Torr, Xiu Li, Li Jiang, Jialin Yu.
Shanghai Jiao Tong University collaborators: Bowen Yang is one of the equal-contribution first authors. Junchi Yan (严骏驰) is a professor at SJTU’s School of Artificial Intelligence. His official faculty profile lists research in machine learning and interdisciplinary applications, as well as academic service including membership of the ICML Board.
Affiliations below follow the title page of arXiv v1, dated May 28, 2026:
- Torr Vision Group, University of Oxford: Zeyuan He, Philip Torr, Jialin Yu.
- The Chinese University of Hong Kong, Shenzhen: Zeyuan He, Jingjing Qian, Li Jiang.
- Tsinghua University: Zhirui Fang, Keru Zhou, Xiu Li.
- Shanghai Jiao Tong University: Bowen Yang, Junchi Yan.
- University College London: Lei Jiang.
- University of Cambridge: Fan Mo.
The paper marks Zeyuan He, Bowen Yang, Zhirui Fang, and Keru Zhou as equal contributors; Zhirui Fang as project lead; Lei Jiang and Li Jiang as equal supervisors; and Jialin Yu as corresponding author.
Paper and conference information
Status: Accepted to CoRL 2026, as confirmed by the research group. This guide is based on the public arXiv v1; no presentation format is specified.
CoRL, the Conference on Robot Learning, is an annual international conference at the intersection of robotics and machine learning. The CoRL 2026 website lists the main conference for November 9–11, 2026 in Austin, Texas, USA, with workshops on November 12.
arXiv abstract and version history · Read the public PDF · Junchi Yan’s official faculty profile
Figures are taken from the original paper or table, with source numbers and PDF pages identified in the captions. Performance and cost figures should be read in the context of the paper’s reported experimental conditions.