Motion Chain of Thought: Disentangling Motion Semantics from Visual Representations for Robotic Manipulation

PILOT learns what should change in the world before refining how the robot should move.

Anonymous submission

Code, model weights, RoboCasa assets and recorded evaluation logs are public. Hugging Face downloads do not require login.

Comparison of conventional visual-language-action and world-action models with PILOT's Motion Chain of Thought.
From visual prediction to motion semantics. Representational Deduction trains learnable Motion-CoT tokens to capture state transitions, leaving the action decoder to refine the robot's trajectory.
58.3%RoboCasa-GR124 tasks, 50 rollouts each
82.5%Real-WorldEight standard tasks

Real-Robot Demonstrations (Third-Person View)

Standard

Pen (Persp. 1)
Pen (Persp. 2)
Eraser (Persp. 1)
Eraser (Persp. 2)
Charger (Persp. 1)
Charger (Persp. 2)
Corr. Fluid (Persp. 1)
Corr. Fluid (Persp. 2)
Book (Persp. 1)
Book (Persp. 2)
Stapler (Persp. 1)
Trash (Persp. 2)

Flash + Cam. Offset

Corr. Fluid (Flash)
Book (Flash)
Stapler (Flash)
Trash (Flash)

Desk Change

Book (Desk)
Charger (Desk)
Eraser (Desk)
Trash (Desk)

Third-person demonstrations of PILOT on Agibot-G1 under standard conditions, strobe lighting with camera offset, and a replaced desk surface. These qualitative videos include multiple tasks; the quantitative generalization results in Experiments are evaluated on Pen only.

Overview

Predicting what a future image looks like is not the same as understanding how an action changes the world. Visual reconstruction can entangle background appearance with motion, forcing the action model to infer high-level state evolution and low-level motor commands from the same representation. PILOT addresses this with Motion Chain of Thought (Motion-CoT): learnable latent tokens that separate motion semantics from fine-grained trajectory generation.

Representational Deduction (RD) supervises these tokens through a Causal Dynamics Engine (CDE), which predicts future states in a frozen VJEPA2-AC representation space. Causally-decoupled attention lets motion semantics guide action decoding without absorbing action diffusion noise. Future observations are supervision targets, not inference inputs; neither future-frame generation nor the CDE is needed to execute the policy.

PILOT reaches 58.3% average success across 24 RoboCasa-GR1 tasks and 82.5% across eight standard real-world tasks. On the Pen task only, it reaches 68.0% mean success across three visual generalization conditions and 78% with only 10% of the Pen demonstrations, compared with its 90% full-data Pen reference.

Method

PILOT framework overview.
The World-Model encodes the current observation and instruction; the Action-Perceiver distills Motion-CoT with learnable queries and uses it to guide flow-matching action decoding. During training, the CDE predicts future VJEPA2-AC representations conditioned on the current representation, motion semantics, and robot state.
1

Learn Motion-Semantic Tokens

A single understanding pass through the Cosmos-Predict2.5 backbone encodes the current observation and instruction. A set of 64 learnable queries distills high-level motion semantics from this context.

2

Supervise State Transitions

Representational Deduction trains the CDE to predict the frozen encoder's future-state representation from the current one, conditioned on Motion-CoT and proprioception. A Smooth L1 objective teaches the queries what changes under an action, rather than how pixels should look.

3

Refine Actions Without Noise Leakage

Causally-decoupled attention allows action tokens to attend to motion semantics, but blocks the reverse flow of action noise into the query tokens. Flow-time conditioning is applied only to action tokens, so the motion context remains stable during trajectory refinement.

World-Model Branch

Cosmos-Predict2.5 provides a 196-token vision-language context. Auxiliary future-frame prediction preserves the pretrained visual and physical prior during training.

Action-Model Branch

A Perceiver-style flow-matching head learns motion semantics and decodes continuous action chunks. At inference, it uses only the current observation, instruction, and robot state.

Training-Only RD Branch

The frozen VJEPA2-AC encoder supplies 256 patch tokens of dimension 1408. The trainable CDE predicts the future representation, and is omitted from the policy's inference path.

Experiments

RoboCasa-GR1 and real-world manipulation on Agibot-G1. Success rates (%) reproduce the manuscript's Tables 3, 4, and 11. Real-world generalization and few-shot evaluation use the Pen task only, separately from the eight-task standard average.

RoboCasa-GR1: 24 Tasks Across Five Categories

MethodPnP Close (6)From Cuttingboard (5)From Placemat (4)From Tray (5)From Plate (4)Overall Average
GR00T N1.545.346.445.548.856.548.2
GR00T N1.624.256.951.955.157.647.6
DiT4DiT50.357.639.046.460.050.8
UWM-XL30.717.410.017.216.319.3
QwenGR00T50.352.838.039.258.547.8
LDA56.364.247.853.453.055.4
VP-VLA54.360.854.546.053.553.8
PhysBrain-4B52.752.048.046.062.552.0
LangForce55.753.248.047.258.552.6
PILOT (Ours)57.062.448.564.857.058.3

Manuscript Table 3. PILOT is evaluated with 50 rollouts per task and reaches 58.3% overall, compared with 55.4% for LDA. Category scores average their constituent tasks; the overall average weights all 24 tasks equally (category weights 6:5:4:5:4). PILOT reaches 57.0% on PnP Close and 64.8% on Tray.

All 24 RoboCasa-GR1 Tasks and Baselines

Manuscript Table 11. PnP = pick and place; Nv = novel; CB = cuttingboard; PM = placemat; TR = tray; PL = plate; CbBox = cardboard box; TBasket = tiered basket; TShelf = tiered shelf; Cab. = cabinet; Micro. = microwave. All reported values and averages are retained. QwenGR00T is the Qwen3-VL-based GR00T-style baseline, not StarVLA-alpha; PhysBrain denotes PhysBrain-4B.

TaskGR00T N1.5GR00T N1.6DiT4DiTUWM-XLQwenGR00TLDAVP-VLAPhysBrain-4BLangForcePILOT (Ours)
PnP Close
PnP Bottle to Cab.5451.54841467654747264
PnP Can to Drawer50137453807172687864
PnP Cup to Drawer388.55212544144424634
PnP Milk to Micro.60145025485274545676
PnP Potato to Micro.3241.53629284134243654
PnP Wine to Cab.3816.54224465748544650
PnP Close (Avg)45.324.250.330.750.356.354.352.755.757.0
PnP Novel From Cuttingboard
NvCB to Basket38585218486566626676
NvCB to CbBox4646.54814406954444042
NvCB to Pan5868.57620687574566886
NvCB to Pot62656225526154584878
NvCB to TBasket2846.55010565156404430
Cuttingboard * (Avg)46.456.957.617.452.864.260.852.053.262.4
PnP Novel From Placemat
NvPM to Basket3058.55016425348425478
NvPM to Bowl6057.55610445574566246
NvPM to Plate56633212485970805244
NvPM to TShelf3628.5182182426142426
Placemat * (Avg)45.551.939.010.038.047.854.548.048.048.5
PnP Novel From Tray
NvTR to CbBox5251.53825386544405076
NvTR to Plate48715618566366665872
NvTR to Pot6064.55425505538526292
NvTR to TBasket52574616365158504454
NvTR to TShelf3231.5382163324222230
Tray * (Avg)48.855.146.417.239.253.446.046.047.264.8
PnP Novel From Plate
NvPL to Bowl5857568605352545442
NvPL to CbBox4443.55810504344504854
NvPL to Pan60516820545556685450
NvPL to Plate6478.75827706162787882
Plate * (Avg)56.557.660.016.358.553.053.562.558.557.0
Overall average48.247.650.819.347.855.453.852.052.658.3
82.5%

Standard Manipulation

Average across eight desktop tasks on Agibot-G1, with 50 rollouts per task.

68.0%

Pen-Only Generalization

Mean across Flash + Camera Offset (72%), Desk (68%), and Color (64%).

78%

Pen-Only Few-Shot

Using 10% of Pen demonstrations, compared with 68% for the strongest baseline, π0.

Agibot-G1: Eight Standard Desk Tasks

MethodPenEraserCorrection
Fluid
ChargerPencil
Case
StaplerBookTrashStandard
Average
π0744642662864527455.8
GR00T-N1.7385230642474808255.5
π0.5886486824070847673.8
PILOT (Ours)909670945688729482.5

Manuscript Table 4, standard evaluation: 50 rollouts per task. The standard average is the arithmetic mean across all eight tasks. PILOT reaches 82.5% overall; its 90% Pen score is the full-data reference for the task-matched generalization and few-shot results below.

Agibot-G1: Pen-Only Generalization and Few-Shot Learning

MethodFlash +
Camera Offset
DeskColorGeneralization
Average
Few-Shot
(10% Pen Data)
π046546254.068
GR00T-N1.718242823.330
π0.562587063.366
PILOT (Ours)72686468.078

Manuscript Table 4, Pen task only. Flash + Camera Offset changes lighting and viewpoint; Desk replaces the surface; Color changes object colors. The generalization average covers these three conditions. Few-shot fine-tuning uses 10% of Pen demonstrations and is a separate setting, not part of that average. PILOT's 78% few-shot success is compared with its 90% full-data Pen success, not the 82.5% average across standard tasks.

Demonstration in the real world

Agibot-G1 humanoid robot and test environment.
The Agibot-G1 dual-arm humanoid and desktop evaluation scenes. The eight tasks involve pen insertion, object grasping and placement, upright book placement, and trash disposal.
Evaluation tasks and generalization settings.
Qualitative book, pen, and trash manipulation sequences under standard, replaced-surface, and flash-plus-camera-offset conditions. These multi-task illustrations are distinct from the Pen-only quantitative generalization evaluation above.

First-Person View Videos

Standard

Pen
Eraser
Corr. Fluid
Charger
Pencil Case
Stapler
Book
Trash

Flash + Cam. Offset

Pen (Flash)
Book (Flash)
Stapler (Flash)
Trash (Flash)

Desk Change

Pen (Desk)
Eraser (Desk)
Book (Desk)
Stapler (Desk)

Color Variation

Stapler (Color)
Trash (Color)
Book (Color)
Charger (Color)

First-person qualitative demonstrations across standard tasks and changed visual conditions. Flash + Camera Offset changes lighting and viewpoint; Desk replaces the table surface; Color changes object appearance. These multi-task videos are illustrative, not the Pen-only quantitative generalization evaluation reported in Table 4.

RoboCasa-GR1 Demonstrations

PnP Close

Bottle to Cabinet
Can to Drawer
Cup to Drawer
Milk to Microwave

Novel From Cuttingboard

To Basket
To Cardboardbox
To Pan
To Pot

Novel From Placemat

To Basket
To Bowl
To Plate
To Tieredshelf

Novel From Tray

To Cardboardbox
To Plate
To Pot
To Tieredbasket

Novel From Plate

To Bowl
To Cardboardbox
To Pan
To Plate

Selected rollouts from the 24-task RoboCasa-GR1 evaluation. PILOT achieves 58.3% reported average success; full category and per-task comparisons are available in the Experiments section.

Code, Weights & Reproducibility

The release selects the original 340000-step RoboCasa-GR1 checkpoint, with its compatible training code, fixed evaluation protocol, customized simulator, and recorded evidence. Other benchmark and real-robot checkpoints are not bundled.

Released checkpoint / repaired-protocol source evaluation 59.75% 717 successes / 1200 episodes · 24 tasks × 50 seeds

This source-run audit is separate from the manuscript's 58.3% result. It is not a new installation benchmark. Success is measured at any executed physical step; successful episodes are not terminated early.

01 / Source

Training & Evaluation Code

The selected model, data interface, policy service, simulator wrapper, CPU tests and reproducibility tools.

Explore GitHub →
02 / Model · Public

Original 340000 Weights

The 23.91 GB checkpoint, SHA-256 verification, metadata and constructor component preparation. The repository ID remains WM4A; the release is PILOT.

Open Hugging Face → Browse checkpoint →
04 / Evaluate

Smoke Test To Full Benchmark

Linux/CUDA/EGL setup, explicit GPU allocation, 24-task commands, scoring semantics, expected outputs and troubleshooting.

Evaluation guide →
05 / Train

Fine-Tune Or Resume

Dataset layout, archived hyperparameters, original-pretrained initialization, weight-only fine-tuning and complete-state resume. Demonstrations are not bundled.

Training guide →

Public access checked September 30, 2026: the Hugging Face repository is public and ungated. Checkpoint and simulator range downloads passed without an account or access token. Browse all model and environment files or follow the installation guide.

Verification scope: a CPU log audit is not action regeneration. Raw per-request streams are not in the log archive. Fresh Linux/CUDA package reproduction, long-run training convergence and multi-rank exact resume remain unverified. Read the validation limits and component-specific licenses.

BibTeX

@unpublished{pilot2026,
  title = {Motion Chain of Thought: Disentangling Motion Semantics from
           Visual Appearance for Robotic Manipulation},
  author = {{Anonymous Authors}},
  year = {2026},
  note = {ICLR 2027 submission, under review}
}