Decoupling Intention from Trajectory:
A Representational Deduction Framework for World Action Models
1Nanjing University 2HKUST 3CUHK-SZ 4SIGS, Tsinghua University 5Joy Future Academy, JD
*Co-corresponding authors | Work led by Xiangkai Ma during internship at Joy Future Academy, JD, Shenzhen
Third-person view of PILOT on the Agibot-G1 humanoid robot. Standard: 6 objects across 2 camera perspectives. Flash+Cam.Offset: strobe-light background with camera angle offset. Desk: replaced table surface.
We present PILOT, an end-to-end World Action Model framework that introduces Representational Deduction (RD) to bridge the gap between static visual prediction and dynamic action generation. Unlike conventional world models that require costly visual generation at inference time, PILOT computes latent residuals between frozen visual encoder states to extract motion-relevant representations, achieving 90% inference latency reduction while maintaining competitive prediction quality.
PILOT further introduces Motion-CoT reasoning, which decomposes high-level motion intent from low-level trajectory generation. This decomposition enables Motion-CoT tokens to naturally cluster by action type in t-SNE space without explicit clustering supervision. Combined with a Causal Dynamics Engine (CDE) that prevents diffusion noise from contaminating motion semantics via causally-decoupled attention, PILOT achieves state-of-the-art performance: 97.9% on LIBERO, 62.6% on RoboCasa-GR1, and 83.1% real-world success rate on the Agibot-G1 humanoid robot.
Compute latent residuals between frozen visual encoder states (VJEPA2-AC) to bridge static visual prediction and dynamic action generation, eliminating the need for costly world generation at inference.
Decompose high-level motion intent from low-level trajectory generation. Motion-CoT tokens naturally cluster by action type in t-SNE space without explicit clustering supervision, enabling 90% latency reduction by skipping world generation.
Predict future state transitions on frozen VJEPA2-AC representations with causally-decoupled attention, preventing diffusion noise from contaminating motion semantics. PCA analysis shows CDE focuses variance on manipulation regions only.
Wan2.2 with 196x2048 context encodes visual observations and language instructions into a unified latent space.
DiT flow-matching with 64 query tokens generates actions in 4 denoising steps, guided by Motion-CoT tokens.
VJEPA2-AC (frozen) provides static representations. CDE predicts state transitions with causally-decoupled attention.
| Method | Pick and Place (6) | Art. from Cuttingboard (5) | Art. from Placemat (4) | Art. from Tray (5) | Art. from Plate (4) | Average |
|---|---|---|---|---|---|---|
| GR00T N1.5 | 45.3 | 46.4 | 45.5 | 48.8 | 56.5 | 48.2 |
| GR00T N1.6 | 24.2 | 56.9 | 51.9 | 55.1 | 57.5 | 47.6 |
| DiT4DiT | 50.3 | 57.6 | 39.0 | 46.4 | 60.0 | 50.8 |
| UWM-XL | 30.7 | 17.4 | 10.0 | 17.2 | 16.2 | 19.2 |
| StarVLA | 50.3 | 52.8 | 38.0 | 39.2 | 58.5 | 47.8 |
| LDA | 56.3 | 64.2 | 47.8 | 53.4 | 53.0 | 55.4 |
| VP-VLA | 54.3 | 60.8 | 54.5 | 46.0 | 53.5 | 53.8 |
| PhysBrain | 52.7 | 52.0 | 48.0 | 46.0 | 62.5 | 50.0 |
| LangForce | 55.7 | 53.2 | 48.0 | 47.2 | 58.5 | 52.6 |
| FastWAM | 52.7 | 60.6 | 47.2 | 54.8 | 58.5 | 56.7 |
| Motus | 49.3 | 56.6 | 43.8 | 50.8 | 54.5 | 53.1 |
| PILOT (Ours) | 59.7 | 61.2 | 56.0 | 74.8 | 60.5 | 62.6 |
| Future Predict | MotionCoT | RD | CDE | Decoupling | LIBERO | RoboCasa |
|---|---|---|---|---|---|---|
| – | – | – | – | – | 91.3 | 51.2 |
| ✓ | × | – | – | – | 93.6 | 54.8 |
| × | ✓ | × | × | × | 94.5 | 56.7 |
| × | ✓ | ✓ | ✓ | × | 96.1 | 59.8 |
| × | ✓ | ✓ | ✓ | ✓ | 96.9 | 61.3 |
| ✓ | ✓ | ✓ | ✓ | ✓ | 97.9 | 62.6 |
| Method | Size | Standard Tasks | Stand. Avg. |
Generalization | Gen. Avg. |
Few-Shot Avg. (10%) |
|||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Pen | Eraser | Corr. Fluid |
Charger | Pencil Case |
Stapler | Book | Trash | Flash+ Cam.Offs. |
Desk | Color | |||||
| π0 | 3.3B | 70 | 82 | 46 | 78 | 32 | 62 | 52 | 72 | 61.8 | 38 | 42 | 36 | 38.7 | 28.4 |
| OpenVLA | 7B | 64 | 76 | 40 | 72 | 26 | 56 | 46 | 66 | 55.8 | 32 | 36 | 30 | 32.7 | 24.6 |
| GR00T-N1 | 3.4B | 74 | 86 | 50 | 82 | 36 | 66 | 56 | 76 | 65.8 | 42 | 46 | 40 | 42.7 | 32.8 |
| π0.5 | 3.3B | 80 | 90 | 56 | 86 | 42 | 72 | 62 | 82 | 71.3 | 48 | 52 | 44 | 48.0 | 38.2 |
| Cosmos-Policy | 6B | 76 | 88 | 52 | 84 | 38 | 68 | 58 | 78 | 67.8 | 44 | 48 | 42 | 44.7 | 34.6 |
| Fast-WAM | 6B | 82 | 92 | 58 | 88 | 44 | 74 | 64 | 84 | 73.3 | 50 | 54 | 46 | 50.0 | 40.8 |
| PILOT (Ours) | 5.4B | 90 | 100 | 70 | 94 | 56 | 88 | 72 | 94 | 83.1 | 72 | 68 | 64 | 68.3 | 62.4 |
Success rates (%) across 8 standard tasks, 3 generalization transfer tasks, and few-shot fine-tuning (10% data). Best results are in bold.
First-person (wrist camera) view. Standard: 8 objects. Flash+Cam.Offset: strobe lighting with camera displacement. Desk: background table surface replaced. Color: object appearance altered.
Simulation rollouts on RoboCasa-GR1 (24 tasks, 5 categories). PILOT achieves 62.6% average, with strong performance on Pick and Place (59.7%) and Articulated from Tray (74.8%).
@inproceedings{ma2027pilot,
title = {Decoupling Intention from Trajectory: A Representational Deduction Framework for World Action Models},
author = {Ma, Xiangkai and Ma, Yue and Wang, Junjie and Xu, Sheng
and Li, Mingyang and Zhang, Han and Zhuang, Yuzheng
and Li, Wenzhong and Yuan, Zhihao},
year = {2026}
}