PILOT

Decoupling Intention from Trajectory:
A Representational Deduction Framework for World Action Models

Xiangkai Ma1Yue Ma2Junjie Wang3Sheng Xu3Mingyang Li4Han Zhang1
Yuzheng Zhuang5Wenzhong Li1*Zhihao Yuan5*

1Nanjing University  2HKUST  3CUHK-SZ  4SIGS, Tsinghua University  5Joy Future Academy, JD

*Co-corresponding authors | Work led by Xiangkai Ma during internship at Joy Future Academy, JD, Shenzhen

PILOT teaser: model comparison, key improvements, and success rates across benchmarks.
PILOT introduces Representational Deduction to bridge visual prediction and action generation, with Motion-CoT reasoning that decomposes high-level motion intent from low-level trajectory generation.

Real-Robot Demonstrations (Third-Person View)

Standard

Pen (Persp. 1)
Pen (Persp. 2)
Eraser (Persp. 1)
Eraser (Persp. 2)
Charger (Persp. 1)
Charger (Persp. 2)
Corr. Fluid (Persp. 1)
Corr. Fluid (Persp. 2)
Book (Persp. 1)
Book (Persp. 2)
Stapler (Persp. 1)
Trash (Persp. 2)

Flash + Cam. Offset

Corr. Fluid (Flash)
Book (Flash)
Stapler (Flash)
Trash (Flash)

Desk Change

Book (Desk)
Charger (Desk)
Eraser (Desk)
Trash (Desk)

Third-person view of PILOT on the Agibot-G1 humanoid robot. Standard: 6 objects across 2 camera perspectives. Flash+Cam.Offset: strobe-light background with camera angle offset. Desk: replaced table surface.

97.9%LIBERO
62.6%RoboCasa-GR1
83.1%Real-World
90%Latency ↓

Overview

We present PILOT, an end-to-end World Action Model framework that introduces Representational Deduction (RD) to bridge the gap between static visual prediction and dynamic action generation. Unlike conventional world models that require costly visual generation at inference time, PILOT computes latent residuals between frozen visual encoder states to extract motion-relevant representations, achieving 90% inference latency reduction while maintaining competitive prediction quality.

PILOT further introduces Motion-CoT reasoning, which decomposes high-level motion intent from low-level trajectory generation. This decomposition enables Motion-CoT tokens to naturally cluster by action type in t-SNE space without explicit clustering supervision. Combined with a Causal Dynamics Engine (CDE) that prevents diffusion noise from contaminating motion semantics via causally-decoupled attention, PILOT achieves state-of-the-art performance: 97.9% on LIBERO, 62.6% on RoboCasa-GR1, and 83.1% real-world success rate on the Agibot-G1 humanoid robot.

Motivation: world model latency vs. PILOT efficiency.
(a) Conventional world models require full visual generation at inference, while PILOT bypasses this via Representational Deduction.
Motivation: Motion-CoT clustering and CDE attention.
(b) Motion-CoT tokens naturally cluster by action type. CDE focuses variance on manipulation regions only.

Method

PILOT framework overview.
PILOT consists of three branches: (a) World-Model (Wan2.2) encodes observations and language instructions; (b) Action-Perceiver uses learnable query tokens with causally-decoupled attention to extract motion semantics; (c) CDE supervises motion semantics by predicting future state transitions on frozen VJEPA2-AC representations.
1

Representational Deduction

Compute latent residuals between frozen visual encoder states (VJEPA2-AC) to bridge static visual prediction and dynamic action generation, eliminating the need for costly world generation at inference.

2

Motion-CoT Reasoning

Decompose high-level motion intent from low-level trajectory generation. Motion-CoT tokens naturally cluster by action type in t-SNE space without explicit clustering supervision, enabling 90% latency reduction by skipping world generation.

3

Causal Dynamics Engine

Predict future state transitions on frozen VJEPA2-AC representations with causally-decoupled attention, preventing diffusion noise from contaminating motion semantics. PCA analysis shows CDE focuses variance on manipulation regions only.

World-Model Branch

Wan2.2 with 196x2048 context encodes visual observations and language instructions into a unified latent space.

Action-Model Branch

DiT flow-matching with 64 query tokens generates actions in 4 denoising steps, guided by Motion-CoT tokens.

RD + CDE Branch

VJEPA2-AC (frozen) provides static representations. CDE predicts state transitions with causally-decoupled attention.

Results in the simulation and the real scenarios

RoboCasa-GR1 Benchmark (24 tasks, 5 categories)

Method Pick and Place (6) Art. from Cuttingboard (5) Art. from Placemat (4) Art. from Tray (5) Art. from Plate (4) Average
GR00T N1.545.346.445.548.856.548.2
GR00T N1.624.256.951.955.157.547.6
DiT4DiT50.357.639.046.460.050.8
UWM-XL30.717.410.017.216.219.2
StarVLA50.352.838.039.258.547.8
LDA56.364.247.853.453.055.4
VP-VLA54.360.854.546.053.553.8
PhysBrain52.752.048.046.062.550.0
LangForce55.753.248.047.258.552.6
FastWAM52.760.647.254.858.556.7
Motus49.356.643.850.854.553.1
PILOT (Ours)59.761.256.074.860.562.6

Cumulative Ablation Study

Future Predict MotionCoT RD CDE Decoupling LIBERO RoboCasa
91.351.2
×93.654.8
××××94.556.7
××96.159.8
×96.961.3
97.962.6

Real-World Evaluation on Agibot-G1

Method Size Standard Tasks Stand.
Avg.
Generalization Gen.
Avg.
Few-Shot
Avg. (10%)
Pen Eraser Corr.
Fluid
Charger Pencil
Case
Stapler Book Trash Flash+
Cam.Offs.
Desk Color
π03.3B708246783262527261.838423638.728.4
OpenVLA7B647640722656466655.832363032.724.6
GR00T-N13.4B748650823666567665.842464042.732.8
π0.53.3B809056864272628271.348524448.038.2
Cosmos-Policy6B768852843868587867.844484244.734.6
Fast-WAM6B829258884474648473.350544650.040.8
PILOT (Ours)5.4B9010070945688729483.172686468.362.4

Success rates (%) across 8 standard tasks, 3 generalization transfer tasks, and few-shot fine-tuning (10% data). Best results are in bold.

Demonstration in the real world

Agibot-G1 humanoid robot and test environment.
The Agibot-G1 humanoid robot and its test environment, featuring 8 daily objects: pen, eraser, correction fluid, charger, pencil case, stapler, book, and trash.
Evaluation tasks and generalization settings.
Evaluation tasks and generalization settings. Standard (83.1%), Flash+Cam.Offset, Desk change, and Color variation (68.3% overall).

First-Person View Videos

Standard

Pen
Eraser
Corr. Fluid
Charger
Pencil Case
Stapler
Book
Trash

Flash + Cam. Offset

Pen (Flash)
Book (Flash)
Stapler (Flash)
Trash (Flash)

Desk Change

Pen (Desk)
Eraser (Desk)
Book (Desk)
Stapler (Desk)

Color Variation

Stapler (Color)
Trash (Color)
Book (Color)
Charger (Color)

First-person (wrist camera) view. Standard: 8 objects. Flash+Cam.Offset: strobe lighting with camera displacement. Desk: background table surface replaced. Color: object appearance altered.

Simulation Results (RoboCasa-GR1)

PnP Close

Bottle to Cabinet
Can to Drawer
Cup to Drawer
Milk to Microwave

Novel From Cuttingboard

To Basket
To Cardboardbox
To Pan
To Pot

Novel From Placemat

To Basket
To Bowl
To Plate
To Tieredshelf

Novel From Tray

To Cardboardbox
To Plate
To Pot
To Tieredbasket

Novel From Plate

To Bowl
To Cardboardbox
To Pan
To Plate

Simulation rollouts on RoboCasa-GR1 (24 tasks, 5 categories). PILOT achieves 62.6% average, with strong performance on Pick and Place (59.7%) and Articulated from Tray (74.8%).

BibTeX

@inproceedings{ma2027pilot,
  title     = {Decoupling Intention from Trajectory: A Representational Deduction Framework for World Action Models},
  author    = {Ma, Xiangkai and Ma, Yue and Wang, Junjie and Xu, Sheng
               and Li, Mingyang and Zhang, Han and Zhuang, Yuzheng
               and Li, Wenzhong and Yuan, Zhihao},
  year      = {2026}
}