2026-10-11 16:33 UTC

Stanford's OpenWAM team (Li Fei-Fei, Jiajun Wu, and Ehsan Adeli's labs) claims its released open framework for composable world-action models — a shared Wan2.2 video foundation with 5B video and 2B action experts in one Mixture-of-Transformers, swappable predict-then-act/act-then-predict/joint/decoupled interaction programs, local-context IDM/FDM components, and a 32,000-segment counterfactual LIBERO-Long-CF dataset — becomes the common foundation that world-action-model research standardizes on; outside groups building policies and benchmarking against it confirm adoption, while quiet fade closes it as a lab-internal release.

state: seedheat: lowuncertainty: mediumnovelscott: highworld-models open-research-frameworks embodied-agents robot-learningLi Fei-FeiJiajun WuEhsan AdeliChangan ChenYao FengHeng Yu

What is this?

OpenWAM is a newly released (arXiv:2610.07922, Oct 2026) open-source framework from Stanford (Fei-Fei Li, Jiajun Wu, Ehsan Adeli labs) for composable world-action models. It builds on a Wan2.2-5B video foundation, adds a 2B action expert via a shared Mixture-of-Transformers architecture, and supports four swappable interaction programs (predict-then-act, act-then-predict, joint, decoupled) with local-context IDM/FDM components. The release includes a 32,000-segment counterfactual LIBERO-Long-CF dataset and full training/evaluation code under AGPL v3. The paper reports strong LIBERO and real-world bimanual results (VTA success on LIBERO-Long jumping from 68.4% to 97.8% with causal robot-video pretraining). The hypothesis's claim that it 'becomes the common foundation research standardizes on' is forward-looking — the release is too recent for independent adoption evidence to appear in these snippets.

Why it matters to Scott

This is a major first-party Stanford open release (Fei-Fei Li, Jiajun Wu, Ehsan Adeli labs) of a composable world-action framework with full technical artifacts — Wan2.2-5B video foundation, 2B action expert in shared Mixture-of-Transformers, four swappable interaction programs, local-context IDM/FDM, and a novel 32K-segment counterfactual LIBERO-Long-CF dataset under AGPL v3. It sits directly in Scott's declared interests (world-models-for-agents, embodied agents, open research frameworks, robot learning) and provides a concrete, buildable foundation that could inform his own agent/world-model architectures, framework design, or writing. The composable interaction-program abstraction (predict-then-act/act-then-predict/joint/decoupled) is architecturally distinctive and relevant to his framework-level thinking. No overlapping open case exists in the radar.
radar:world-labs-atlas-spatial-modelradar:concept.world-modelsradar:concept.embodied-agentsradar:concept.open-weight-modelsradar:concept.robot-foundation-modelsradar:concept.mixture-of-expertsradar:unitree-unifolm-wla1-open-releaseradar:lingbot-video-action-world-model
queries asked of Scott's wikis
  • world-models-for-agents composable frameworks
  • open-weights video foundation models robotics
  • counterfactual datasets robot learning LIBERO
  • Mixture-of-Transformers architecture action experts
  • causal robot-video pretraining scaling
  • embodied agents world-action model standardization

Measured heat

now 0 pts/hpeak 2 pts/hcomments 0/hpeers p14momentum: steady2 platformsage 109h
points/hour across evidence · reading as of 2026-10-12 02:59:37.977291+11:00 · deterministic, not a model opinion

How the heat travelled

10-07 08:27 (minted)⭐ origin echo-reconstructed'World–action models make both possible... OpenWAM gives them a common foundation so we can study how prediction and control work together'
OpenWAM team — Heng Yu, Changan Chen, Yao Feng, Li Fei-Fei, Jiajun Wu, Ehsan Adeli, et al. (Stanford) on blog (echo) · attributed from hn.story.49987679 · published time unknown
—
10-07 03:17first on hacker news · published · lag ?OpenWAM: An Open Framework for Composable World-Action Models
ilreb
—
10-07 03:17amplified on hacker news 👑hn.story.49987679
ilreb
peak 12 · 1 comments · 100% of case engagement
10-07 08:21our radar first saw it · lag ?discovery anchor: hn.story.49987679—
pace: p47 vs 1247 stories at the 96h mark (now 109h old) — ahead of aws-agentcore-credential-exposure-containment-failure (1.1x), behind asksary-liveloop-stateful-editing (0.9x)

Evidence (2) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟧 hnOpenWAM: An Open Framework for Composable World-Action Models
Retrieved article excerpt

Open article · Retrieved 2026-10-07T08:25:44.495914+00:00

## A framework for world–action models

Should a robot predict what it will see before deciding how to act, or generate both together? World–action models make both possible. Comparing these choices is difficult when every system uses a different backbone, dataset, and training recipe. OpenWAM gives them a common foundation so we can study how prediction and control work together.

The framework supports composition within a model and between models. We can change the order in which video and actions are generated and how their tokens attend to one another. We can also connect independently trained components: a video predictor proposes a future, an inverse dynamics model turns it into actions, and a forward dynamics model predicts what a supplied action sequence will do.

[Figure 1: Wan2.2 video pretraining, causal robot-video pretraining, and video-action post-training feed a shared MoT architecture, configurable interaction programs, and independently trained local inverse and forward dynamics components.](https://openwam.stanford.edu/assets/openwam-overview.svg)


**Figure 1. OpenWAM overview.** A shared video foundation supports configurable video–action programs and independently trained dynamics. In the transfer experiments at bottom right, we adapt the video predictor to each task and reuse the IDM without further training.[View full-size figure ↗](https://openwam.stanford.edu/assets/openwam-overview.svg)

## A shared video–action architecture

Before learning actions, we adapt Wan2.2-5B to robot motion and interaction. We pretrain on approximately **3.34 million trajectories and recordings—14.64k hours of source video**—spanning real and synthetic robot manipulation, human-guided manipulation, and human interaction. This stage uses video alone, without action labels or proprioceptive inputs.

What goes into pretraining?

| Dataset | Trajectories / takes | Hours |
| --- | --- | --- |
| Open X-Embodiment | 1,300,749 | 1,911.29 |
| AgiBot World Beta | 1,003,672 | 2,976.4 |
| Ego-Exo4D v2 | 5,035 | 221.26 |
| InternData-A1 | 637,498 | 7,433.91 |
| RoboCOIN | 183,157 | 1,306.83 |
| RoboMIND | 107,877 | 305.5 |
| FastUMI-100K | 92,823 | 461.77 |
| UMI family | 5,430 | 24.60 |

Paper Table 8 · Hours count source sequences before training-window sampling, including synthetic and human-interaction video. OXE covers 49 manipulation datasets; InternData-A1 is synthetic; Ego-Exo4D records human interaction. The UMI family includes UMI, DexUMI, UMI on Legs, and MV-UMI.

Causal attention lets us generate video a chunk at a time: each chunk can use current and past observations and earlier chunks, but not later ones. Its tokens are denoised together. After **14 days on 32 NVIDIA B200 GPUs**, this checkpoint provides the visual foundation for the downstream models.

To add robot control, we pair the 5B video expert with a 2B action expert in a **Mixture-of-Transformers (MoT)** architecture. The action expert starts from width-adapted copies of the pretrained video layers. Each expert keeps its own normalization, projections, and feed-forward layers; attention over their combined tokens lets them exchange information.

[Figure 2: A 5B video expert and 2B action expert retain separate normalization, projections, and feed-forward layers, meet in attention over packed video and action tokens, and use separate cross-attention to text. Earlier chunks provide history context.](https://openwam.stanford.edu/assets/openwam-architecture.svg)


**Figure 2. Shared MoT architecture.** A 5B video expert and a 2B action expert share attention over video and action tokens. Each also attends to the task instruction.[View full-size figure ↗](https://openwam.stanford.edu/assets/openwam-architecture.svg)

## Video-action interaction programs

The same architecture can predict video before actions, actions before video, or both at once. Each interaction program specifies the generation order and attention between future tokens. The backbone, tokenization, training objective, and downstream recipe stay fixed, and every program receives the task instruction and observed history.

| Program | Generation and conditioning |
| --- | --- |
| Video-then-action (VTA) | Predict video first, then generate actions conditioned on that video. |
| Action-then-video (ATV) | Generate actions first, then predict video conditioned on those actions. |
| Joint | Denoise video and actions together, with attention in both directions. |
| Decoupled | Predict video and actions without attention between their future tokens. |

[Figure 3: A is VTA, B is ATV, C is Joint, D is Decoupled, E is local-context IDM, and F is local-context FDM. White denotes no attention, yellow clean conditioning, green noisy conditioning, and blue a prediction target with its generation-pass number.](https://openwam.stanford.edu/assets/openwam-interaction-masks.svg)


**Figure 3. Attention masks.** A–D: the evaluated policy programs. E–F: local inverse and forward dynamics. Colors mark clean or noisy conditioning; numbers mark generation order. L denotes language; O−, O0, and O+ denote past, current, and future observations; A+ denotes future actions.[View full-size figure ↗](https://openwam.stanford.edu/assets/openwam-interaction-masks.svg)

All programs use latent flow matching, with four latent video frames aligned to each 16-step action chunk and proprioception supplied per chunk. In VTA and ATV, the second stage learns from recorded trajectories during training and uses the first stage’s predictions at inference.

## Beyond policies: local-context IDM and FDM

A task-conditioned predictor proposes what should happen next. Dynamics models connect that proposal to the robot’s motion: inverse dynamics (IDM) turns a visual future into actions, while forward dynamics (FDM) predicts the outcome of an action sequence. OpenWAM supports both as standalone models.

We give them a **local-context interface**: the current observation, proprioception, and a supplied future trajectory.

**Local-context IDM:** current observation + proprioception + supplied future video → action trajectory.

**Local-context FDM:** current observation + proprioception + supplied action trajectory → future video.

Neither model receives task language or pre-start history. This separates choosing a task from modeling a transition: we can adapt the video predictor and ask whether the same IDM still produces the right actions. The experiments below test how far this local information can take us.

The video predictor passes VAE video latents to the IDM, not transformer hidden states or caches. With compatible video and action representations, the components can be trained separately and connected at inference. An FDM can likewise predict the outcome of actions from a separate policy.

### Learning from alternative outcomes

A demonstration shows what the demonstrator chose to do, but says little about what other actions would have caused. Changing a model’s inputs does not fill that gap in its training data. We build **LIBERO-Long-CF: 32,000 counterfactual segments across ten tasks** by restoring simulator states and trying alternative action sequences, including unsuccessful ones. The models learn from the resulting observations and actions, without access to simulator state.

| Quantity | Demonstrations | LIBERO-Long-CF |
| --- | --- | --- |
| Tasks | 10 | 10 |
| Stored sequences | 500 | 32,000 |
| Sequences per task | 50 | 3,200 |
| Controls per sequence | 276.2 mean | 128 |
| Total controls | 138,090 | 4,096,000 |
| Control-equivalent hours | 1.92 | 56.9 |

Paper Table 9 · Equivalent durations at 20 Hz. Each counterfactual segment contains 128 controls and 129 synchronized two-view observations. The dataset contains 29.7× as many control records as the demonstrations, using the original tasks, assets, and physics.

What changes in the counterfactual rollouts?

We vary motion magnitude, direction, timing, individual action axes, and gripper behavior. 75% of segments start along a demonstration; the other 25% start after an additional action perturbation.

| Intervention family | Fraction (%) |
| --- | --- |
| Stop / rescale arm motion | 9.4 |
| Reverse / redirect translation | 6.3 |
| Axis biases and pulses | 12.5 |
| Dedicated yaw perturbation | 3.1 |
| Noise / randomized arm controls | 12.5 |
| Dedicated gripper interventions | 31.3 |
| Random-duration arm / gripper interventions | 25.0 |

Intervention recipes (%) · Paper Table 10. Values are rounded; different recipes can produce overlapping physical effects.

In the perturbed starts we analyzed, the end effector is on average **3.34 cm** from the nearest point on the demonstrated path (median 1.90 cm). Objects also move beyond their demonstrated configurations in **61.6%** of these starts, measured at thresholds of 1 cm translation, 5° rotation, or 5% articulated-joint travel.

| Measured interaction | Rate (%) |
| --- | --- |
| Gripper–object / fixture contact | 90.4 |
| Detected grasp | 50.0 |
| Object-configuration effect | 72.3 |

Interaction rates (%) · Paper Table 11. Measured over the segments analyzed; a segment can count toward multiple categories. Object changes are measured against the reference rollout from the same starting state.

Branches from the same starting state stay together in the train/test split. Each transition supplies its own training example; the loss does not directly contrast pairs of branches.

We compare models trained on demonstrations alone, counterfactuals alone (CF-only), and a mixture of **60% counterfactuals and 40% demonstrations**.

## Policy performance across programs

### LIBERO

VTA achieves **98.6%** mean success across four LIBERO suites. Each evaluated program exceeds 95% on LIBERO-Long.

| Method | Object | Goal | Spatial | Long | Mean |
| --- | --- | --- | --- | --- | --- |
| OpenVLA | 88.4 | 79.2 | 84.7 | 53.7 | 76.5 |
| OpenVLA-OFT | 98.4 | 97.9 | 97.6 | 94.5 | 97.1 |
| π0 | 98.8 | 95.8 | 96.8 | 85.2 | 94.1 |
| π0.5 | 98.2 | 98.0 | 98.8 | 92.4 | 96.9 |
| GR00T-N1 | 97.6 | 93.0 | 94.4 | 90.6 | 93.9 |
| Motus | 99.8 | 96.6 | 96.8 | 97.6 | 97.7 |
| Fast-WAM | 100.0 | 97.0 | 98.2 | 95.2 | 97.6 |
| LingBot-VA | 99.6 | 97.2 | 98.5 | 98.5 | 98.5 |
| OpenWAM-VTA | 99.4 ± 0.3 | 98.4 ± 0.5 | 98.6 ± 0.2 | 97.8 ± 0.4 | **98.6** |
| OpenWAM-ATV | 98.0 ± 0.4 | 97.2 ± 0.2 | 96.6 ± 0.6 | 95.4 ± 0.3 | 96.8 |
| OpenWAM-Joint | 98.2 ± 0.2 | 97.8 ± 0.4 | 97.6 ± 0.3 | 96.6 ± 0.5 | 97.6 |
| OpenWAM-Decoupled | 99.0 ± 0.3 | 98.0 ± 0.2 | 97.8 ± 0.5 | 97.0 ± 0.4 | 98.0 |

Closed-loop success (%) · Paper Table 3. OpenWAM: mean ± standard deviation over three training seeds, with 50 episodes per task and 500 per suite per seed. Baselines are reported as published in their respective papers.

Decoupled remains competitive with Joint on LIBERO-Long: **97.0%** versus **96.6%**. Strong control on this benchmark does not require attention between future video and action tokens.

### Bimanual manipulation

On a bimanual robot, VTA and Joint each average about **92% success** across toasting bread, completing the final layer of a 2 × 2 Rubik’s cube, and sorting cups by color. The setup uses two Franka Research 3 arms with parallel-jaw grippers, two wrist cameras, and a third-person camera.

[Figure 4: Successful rollout frames for Toast Bread, Rubik’s Cube, and Sort Cups on the left; LIBERO-90 transfer Tasks 64, 74, 21, and 45 on the right.](https://openwam.stanford.edu/assets/openwam-evaluation-tasks.svg)


**Figure 4. Evaluation tasks.** Left: one successful rollout per real-world task. Right: the four LIBERO-90 transfer tasks, outside the LIBERO-Long source set.[View full-size figure ↗](https://openwam.stanford.edu/assets/openwam-evaluation-tasks.svg)

| Method | Toast | Cube | Cups | Mean |
| --- | --- | --- | --- | --- |
| OpenWAM-VTA | 92.0 | 90.0 | 94.4 | 92.1 |
| OpenWAM-Joint | 90.0 | 94.0 | 91.7 | 91.9 |

Closed-loop success (%) · Paper Table 4. We train on 200 toast, 200 cube, and 180 cup demon
ilreb121
🟧 echo.blog ⭐'World–action models make both possible... OpenWAM gives them a common foundation so we can study how prediction and control work together' OpenWAM team — Heng Yu, Changan Chen, Yao Feng, Li Fei-Fei, Jiajun Wu, Ehsan Adeli, et al. (Stanford)——

Interpretation history

Decision trace