Standard Manipulation
Average across eight desktop tasks on Agibot-G1, with 50 rollouts per task.
PILOT learns what should change in the world before refining how the robot should move.
Code, model weights, RoboCasa assets and recorded evaluation logs are public. Hugging Face downloads do not require login.
Third-person demonstrations of PILOT on Agibot-G1 under standard conditions, strobe lighting with camera offset, and a replaced desk surface. These qualitative videos include multiple tasks; the quantitative generalization results in Experiments are evaluated on Pen only.
Predicting what a future image looks like is not the same as understanding how an action changes the world. Visual reconstruction can entangle background appearance with motion, forcing the action model to infer high-level state evolution and low-level motor commands from the same representation. PILOT addresses this with Motion Chain of Thought (Motion-CoT): learnable latent tokens that separate motion semantics from fine-grained trajectory generation.
Representational Deduction (RD) supervises these tokens through a Causal Dynamics Engine (CDE), which predicts future states in a frozen VJEPA2-AC representation space. Causally-decoupled attention lets motion semantics guide action decoding without absorbing action diffusion noise. Future observations are supervision targets, not inference inputs; neither future-frame generation nor the CDE is needed to execute the policy.
PILOT reaches 58.3% average success across 24 RoboCasa-GR1 tasks and 82.5% across eight standard real-world tasks. On the Pen task only, it reaches 68.0% mean success across three visual generalization conditions and 78% with only 10% of the Pen demonstrations, compared with its 90% full-data Pen reference.
A single understanding pass through the Cosmos-Predict2.5 backbone encodes the current observation and instruction. A set of 64 learnable queries distills high-level motion semantics from this context.
Representational Deduction trains the CDE to predict the frozen encoder's future-state representation from the current one, conditioned on Motion-CoT and proprioception. A Smooth L1 objective teaches the queries what changes under an action, rather than how pixels should look.
Causally-decoupled attention allows action tokens to attend to motion semantics, but blocks the reverse flow of action noise into the query tokens. Flow-time conditioning is applied only to action tokens, so the motion context remains stable during trajectory refinement.
Cosmos-Predict2.5 provides a 196-token vision-language context. Auxiliary future-frame prediction preserves the pretrained visual and physical prior during training.
A Perceiver-style flow-matching head learns motion semantics and decodes continuous action chunks. At inference, it uses only the current observation, instruction, and robot state.
The frozen VJEPA2-AC encoder supplies 256 patch tokens of dimension 1408. The trainable CDE predicts the future representation, and is omitted from the policy's inference path.
RoboCasa-GR1 and real-world manipulation on Agibot-G1. Success rates (%) reproduce the manuscript's Tables 3, 4, and 11. Real-world generalization and few-shot evaluation use the Pen task only, separately from the eight-task standard average.
| Method | PnP Close (6) | From Cuttingboard (5) | From Placemat (4) | From Tray (5) | From Plate (4) | Overall Average |
|---|---|---|---|---|---|---|
| GR00T N1.5 | 45.3 | 46.4 | 45.5 | 48.8 | 56.5 | 48.2 |
| GR00T N1.6 | 24.2 | 56.9 | 51.9 | 55.1 | 57.6 | 47.6 |
| DiT4DiT | 50.3 | 57.6 | 39.0 | 46.4 | 60.0 | 50.8 |
| UWM-XL | 30.7 | 17.4 | 10.0 | 17.2 | 16.3 | 19.3 |
| QwenGR00T | 50.3 | 52.8 | 38.0 | 39.2 | 58.5 | 47.8 |
| LDA | 56.3 | 64.2 | 47.8 | 53.4 | 53.0 | 55.4 |
| VP-VLA | 54.3 | 60.8 | 54.5 | 46.0 | 53.5 | 53.8 |
| PhysBrain-4B | 52.7 | 52.0 | 48.0 | 46.0 | 62.5 | 52.0 |
| LangForce | 55.7 | 53.2 | 48.0 | 47.2 | 58.5 | 52.6 |
| PILOT (Ours) | 57.0 | 62.4 | 48.5 | 64.8 | 57.0 | 58.3 |
Manuscript Table 3. PILOT is evaluated with 50 rollouts per task and reaches 58.3% overall, compared with 55.4% for LDA. Category scores average their constituent tasks; the overall average weights all 24 tasks equally (category weights 6:5:4:5:4). PILOT reaches 57.0% on PnP Close and 64.8% on Tray.
Manuscript Table 11. PnP = pick and place; Nv = novel; CB = cuttingboard; PM = placemat; TR = tray; PL = plate; CbBox = cardboard box; TBasket = tiered basket; TShelf = tiered shelf; Cab. = cabinet; Micro. = microwave. All reported values and averages are retained. QwenGR00T is the Qwen3-VL-based GR00T-style baseline, not StarVLA-alpha; PhysBrain denotes PhysBrain-4B.
| Task | GR00T N1.5 | GR00T N1.6 | DiT4DiT | UWM-XL | QwenGR00T | LDA | VP-VLA | PhysBrain-4B | LangForce | PILOT (Ours) |
|---|---|---|---|---|---|---|---|---|---|---|
| PnP Close | ||||||||||
| PnP Bottle to Cab. | 54 | 51.5 | 48 | 41 | 46 | 76 | 54 | 74 | 72 | 64 |
| PnP Can to Drawer | 50 | 13 | 74 | 53 | 80 | 71 | 72 | 68 | 78 | 64 |
| PnP Cup to Drawer | 38 | 8.5 | 52 | 12 | 54 | 41 | 44 | 42 | 46 | 34 |
| PnP Milk to Micro. | 60 | 14 | 50 | 25 | 48 | 52 | 74 | 54 | 56 | 76 |
| PnP Potato to Micro. | 32 | 41.5 | 36 | 29 | 28 | 41 | 34 | 24 | 36 | 54 |
| PnP Wine to Cab. | 38 | 16.5 | 42 | 24 | 46 | 57 | 48 | 54 | 46 | 50 |
| PnP Close (Avg) | 45.3 | 24.2 | 50.3 | 30.7 | 50.3 | 56.3 | 54.3 | 52.7 | 55.7 | 57.0 |
| PnP Novel From Cuttingboard | ||||||||||
| NvCB to Basket | 38 | 58 | 52 | 18 | 48 | 65 | 66 | 62 | 66 | 76 |
| NvCB to CbBox | 46 | 46.5 | 48 | 14 | 40 | 69 | 54 | 44 | 40 | 42 |
| NvCB to Pan | 58 | 68.5 | 76 | 20 | 68 | 75 | 74 | 56 | 68 | 86 |
| NvCB to Pot | 62 | 65 | 62 | 25 | 52 | 61 | 54 | 58 | 48 | 78 |
| NvCB to TBasket | 28 | 46.5 | 50 | 10 | 56 | 51 | 56 | 40 | 44 | 30 |
| Cuttingboard * (Avg) | 46.4 | 56.9 | 57.6 | 17.4 | 52.8 | 64.2 | 60.8 | 52.0 | 53.2 | 62.4 |
| PnP Novel From Placemat | ||||||||||
| NvPM to Basket | 30 | 58.5 | 50 | 16 | 42 | 53 | 48 | 42 | 54 | 78 |
| NvPM to Bowl | 60 | 57.5 | 56 | 10 | 44 | 55 | 74 | 56 | 62 | 46 |
| NvPM to Plate | 56 | 63 | 32 | 12 | 48 | 59 | 70 | 80 | 52 | 44 |
| NvPM to TShelf | 36 | 28.5 | 18 | 2 | 18 | 24 | 26 | 14 | 24 | 26 |
| Placemat * (Avg) | 45.5 | 51.9 | 39.0 | 10.0 | 38.0 | 47.8 | 54.5 | 48.0 | 48.0 | 48.5 |
| PnP Novel From Tray | ||||||||||
| NvTR to CbBox | 52 | 51.5 | 38 | 25 | 38 | 65 | 44 | 40 | 50 | 76 |
| NvTR to Plate | 48 | 71 | 56 | 18 | 56 | 63 | 66 | 66 | 58 | 72 |
| NvTR to Pot | 60 | 64.5 | 54 | 25 | 50 | 55 | 38 | 52 | 62 | 92 |
| NvTR to TBasket | 52 | 57 | 46 | 16 | 36 | 51 | 58 | 50 | 44 | 54 |
| NvTR to TShelf | 32 | 31.5 | 38 | 2 | 16 | 33 | 24 | 22 | 22 | 30 |
| Tray * (Avg) | 48.8 | 55.1 | 46.4 | 17.2 | 39.2 | 53.4 | 46.0 | 46.0 | 47.2 | 64.8 |
| PnP Novel From Plate | ||||||||||
| NvPL to Bowl | 58 | 57 | 56 | 8 | 60 | 53 | 52 | 54 | 54 | 42 |
| NvPL to CbBox | 44 | 43.5 | 58 | 10 | 50 | 43 | 44 | 50 | 48 | 54 |
| NvPL to Pan | 60 | 51 | 68 | 20 | 54 | 55 | 56 | 68 | 54 | 50 |
| NvPL to Plate | 64 | 78.7 | 58 | 27 | 70 | 61 | 62 | 78 | 78 | 82 |
| Plate * (Avg) | 56.5 | 57.6 | 60.0 | 16.3 | 58.5 | 53.0 | 53.5 | 62.5 | 58.5 | 57.0 |
| Overall average | 48.2 | 47.6 | 50.8 | 19.3 | 47.8 | 55.4 | 53.8 | 52.0 | 52.6 | 58.3 |
Average across eight desktop tasks on Agibot-G1, with 50 rollouts per task.
Mean across Flash + Camera Offset (72%), Desk (68%), and Color (64%).
Using 10% of Pen demonstrations, compared with 68% for the strongest baseline, π0.
| Method | Pen | Eraser | Correction Fluid | Charger | Pencil Case | Stapler | Book | Trash | Standard Average |
|---|---|---|---|---|---|---|---|---|---|
| π0 | 74 | 46 | 42 | 66 | 28 | 64 | 52 | 74 | 55.8 |
| GR00T-N1.7 | 38 | 52 | 30 | 64 | 24 | 74 | 80 | 82 | 55.5 |
| π0.5 | 88 | 64 | 86 | 82 | 40 | 70 | 84 | 76 | 73.8 |
| PILOT (Ours) | 90 | 96 | 70 | 94 | 56 | 88 | 72 | 94 | 82.5 |
Manuscript Table 4, standard evaluation: 50 rollouts per task. The standard average is the arithmetic mean across all eight tasks. PILOT reaches 82.5% overall; its 90% Pen score is the full-data reference for the task-matched generalization and few-shot results below.
| Method | Flash + Camera Offset | Desk | Color | Generalization Average | Few-Shot (10% Pen Data) |
|---|---|---|---|---|---|
| π0 | 46 | 54 | 62 | 54.0 | 68 |
| GR00T-N1.7 | 18 | 24 | 28 | 23.3 | 30 |
| π0.5 | 62 | 58 | 70 | 63.3 | 66 |
| PILOT (Ours) | 72 | 68 | 64 | 68.0 | 78 |
Manuscript Table 4, Pen task only. Flash + Camera Offset changes lighting and viewpoint; Desk replaces the surface; Color changes object colors. The generalization average covers these three conditions. Few-shot fine-tuning uses 10% of Pen demonstrations and is a separate setting, not part of that average. PILOT's 78% few-shot success is compared with its 90% full-data Pen success, not the 82.5% average across standard tasks.
First-person qualitative demonstrations across standard tasks and changed visual conditions. Flash + Camera Offset changes lighting and viewpoint; Desk replaces the table surface; Color changes object appearance. These multi-task videos are illustrative, not the Pen-only quantitative generalization evaluation reported in Table 4.
Selected rollouts from the 24-task RoboCasa-GR1 evaluation. PILOT achieves 58.3% reported average success; full category and per-task comparisons are available in the Experiments section.
The release selects the original 340000-step RoboCasa-GR1 checkpoint, with its compatible training code, fixed evaluation protocol, customized simulator, and recorded evidence. Other benchmark and real-robot checkpoints are not bundled.
This source-run audit is separate from the manuscript's 58.3% result. It is not a new installation benchmark. Success is measured at any executed physical step; successful episodes are not terminated early.
The selected model, data interface, policy service, simulator wrapper, CPU tests and reproducibility tools.
Explore GitHub →The 23.91 GB checkpoint, SHA-256 verification, metadata and constructor component preparation. The repository ID remains WM4A; the release is PILOT.
Open Hugging Face → Browse checkpoint →A 4.53 GB archive with 45,853 simulator and robot resource files, plus separate policy/simulation environment setup.
Installation guide → Simulator archive →Linux/CUDA/EGL setup, explicit GPU allocation, 24-task commands, scoring semantics, expected outputs and troubleshooting.
Evaluation guide →Dataset layout, archived hyperparameters, original-pretrained initialization, weight-only fine-tuning and complete-state resume. Demonstrations are not bundled.
Training guide →24 task logs, 1200 episode diagnostics and 680 historical training metric records. Verify recorded counts without downloading the model.
Evidence and redactions → Download logs (1.93 MB) →Public access checked September 30, 2026: the Hugging Face repository is public and ungated. Checkpoint and simulator range downloads passed without an account or access token. Browse all model and environment files or follow the installation guide.
Verification scope: a CPU log audit is not action regeneration. Raw per-request streams are not in the log archive. Fresh Linux/CUDA package reproduction, long-run training convergence and multi-rank exact resume remain unverified. Read the validation limits and component-specific licenses.
@unpublished{pilot2026,
title = {Motion Chain of Thought: Disentangling Motion Semantics from
Visual Appearance for Robotic Manipulation},
author = {{Anonymous Authors}},
year = {2026},
note = {ICLR 2027 submission, under review}
}