Controllable World Action Models with
Aligned Intent and Foresight
Overview
World action models jointly predict an agent's future actions (intent) and the video of what it will see (foresight). CtrlWAM trains them so that noised actions and noised video stay physically consistent, and lets a single model predict or follow commands for the ego agent and the agents around it.
World action models (WAMs) jointly predict actions (intent) and visual future (foresight). Standard training adds noise to recorded actions and video simultaneously, but such training paradigms introduce a mismatch: perturbed actions imply counterfactual future visual, while the noised video remains tied to the GT recording. In low-noise regime, the scene geometry and even the dynamic behavior remain clearly visible from the noisy future frames despite the added noise.
We present CtrlWAM, which executes perturbed actions in a simulator and pairs them with their noised visual consequences for joint WAM learning. To accommodate the different denoising requirements of video and actions, we introduce warped video–action noise schedules that aim to keep visual layout responsive as action predictions evolve. We further extend the action interface from ego-only control to a variable number of agent streams, allowing a unified model to represent predicted or commanded futures for multiple agents.
Driving experiments show more accurate action forecasts, closer agreement between generated video and actions, and better following of supplied commands; robotics experiments show stronger motion fidelity and controllability. Matched controls support the benefit of off-path renders for command following and manipulation fidelity. Together, these findings contribute to a more controllable world action model.
Both future streams are denoised jointly; the imagined video should reflect the predicted intent.
Future actions are supplied clean as commands; only video is generated. This is where controllability is measured: does the video follow a rewritten command?
Clean future video, noised actions. A separately trained inverse dynamics model (IDM) also reads trajectories out of generated video for evaluation.
Motivation
Three states, replaying. Perturbing the action changes the implied future — but Gaussian noise on the recorded frames never moves the car.















t₁
A recorded action–video pair specifies how the two modalities agree at σ = 0. It says nothing about how they should deviate together as noise is added.
Methods such as DreamZero corrupt actions and video with independent Gaussian noise. The noised action implies a different ego path and a different visual future — but the corrupted video remains tied to the recording. In the low-noise regime the scene layout and even the dynamics stay clearly visible.
Useful visual structure emerges early in denoising (DIDO: step-1 video features already give 97.7 % vs 98.4 % after four steps on LIBERO). Without alignment, an intermediate action estimate can veer toward an obstacle while the imagined video still depicts a safe pass.
In both driving and manipulation the visual future is recognisable while σ is still high — the explicitly predicted trajectory is still moving. Under a shared schedule the video can therefore commit before the action settles.
Standard training noises video and action with one shared level (red diagonal): both are assumed equally uncertain at every step.
The animation above shows otherwise. The action the partially denoised video implies (blue) converges within the first few steps, while the video itself is still far from clean.
The horizontal gap between the curves is the denoising mismatch: the action token is trained against a video that already fixed the ego path, yet nothing in the pair tells the model the two must agree along the way.
Method
Explore how CtrlWAM constructs physically corresponding inputs, paces video against actions, and represents surrounding agents.
Left: standard noising keeps the recorded future in the video. Right: the simulator executes the noised command, its render is noised, and both branches are still trained toward the recorded pair.
Agent histories store ego-frame positions at t_0 (scaled by 20 m); futures are unicycle controls in each actor's local frame. Actors are selected by visibility in the front camera (≤ 60 m, ≥ 12 px, not occluded), up to K_{\max}=8; a shared agent projection and per-actor spatial rotary positions place them on the latent grid.
Model
A shared transformer processes video, ego-action and optional agent-action tokens, conditioned on observed history. Future tokens are denoised as prediction targets or supplied clean as conditions — the same network serves joint prediction, action-conditioned video and video-conditioned actions.
Tokens flow through the model: noised future tokens from each modality enter the transformer, attend jointly with the clean history condition, and are decoded back into video, ego action and agent trajectories.
Future frames → causal VAE (4× time, 16× space) → latent tokens. Noised in forward dynamics and joint mode, clean for inverse dynamics.
Unicycle controls (acceleration, curvature) or 14-D joint positions; 64 future rows. Noised for prediction, clean when supplied as a command.
0…K agent blocks through a shared agent encoder/decoder; variable length. Backbone: Cosmos 3-Nano (16B), language pathway frozen.
Results · driving
1,156 held-out clips from PhysicalAI-AV. In Joint mode the model predicts video and actions together; in forward dynamics the future actions are supplied as commands. CtrlWAM-S is the ego-only model, CtrlWAM-M the multi-agent model, trained with the same recipe, data and budget.
| Joint mode | Forward dynamics | |||||
|---|---|---|---|---|---|---|
| Model | PSNR (dB) | Self ADE (m) | Rollout ADE (m) | PSNR (dB) | Cmd ADE (m) | Cmd FDE (m) |
| Cosmos 3-Nano (zero-shot) | 19.58 | 6.212 | 8.267 | 20.05 | 1.600 | 4.040 |
| SFT baseline | 20.38 | 0.538 | 2.920 | 21.23 | 1.023 | 2.803 |
| CtrlWAM-S (ego-only) | 20.34 | 0.521 | 2.848 | 21.23 | 1.001 | 2.722 |
| CtrlWAM-M (multi-agent) | 20.42 | 0.513 | 2.804 | 21.28 | 0.981 | 2.671 |
Self ADE compares the IDM read-out of predicted video with the model's own action trajectory (video–action consistency). Rollout ADE compares the generated actions with the recording. Cmd ADE/FDE compare the IDM read-out with the supplied command. Bold = best, underline = second best.
Results · counterfactual ego control
Scene history, text prompt and sampling noise are held fixed; only the ego command changes. Stopping halts forward viewpoint progression, acceleration advances the ego past the parked cars.

Top-down inset: the commanded ego path in the t₀ frame (10 m rings), magenta when rewritten. Generated by the multi-agent CtrlWAM model in forward dynamics; one clip at 10.7 m/s, 5 Hz frames over 6.4 s.
Results · multi-agent control
Scene history, the ego command and sampling noise are held fixed while a selected agent's future trajectory is rewritten. Gemini 3.8 Flash rates agreement between the generated video and the supplied agent command (0–1, reported in %). Only CtrlWAM-M can receive the command; the other models are scored on agreement without conditioning.

Counterfactual control of the lead vehicle. Top: factual agent trajectory. Bottom: the same scene with the agent commanded to stop — the generated video shows it halting in front of the ego. Right: unchanged ego path (grey), factual (green) and rewritten (magenta) agent command. Mean compliance rises from 10.07 % (ego-only) to 18.94 %.
Results · robotic manipulation
RoboTwin 2.0 bimanual manipulation, 1,000 held-out episodes, WorldArena Track-1. Each arm starts from a common 2,000-iteration checkpoint and receives 2,000 more iterations using aligned renders (CtrlWAM), real frames in place of renders (GT-Src) or standard SFT. EWMScore is 100 × the mean of the 15 individual metrics.
| Category | Cosmos 3 zero-shot | Cosmos 3 SFT | GT-Src | CtrlWAM | Oracle |
|---|---|---|---|---|---|
| Visual quality | 0.447 | 0.560 | 0.563 | 0.576 | 0.598 |
| Motion quality | 0.247 | 0.316 | 0.306 | 0.327 | 0.348 |
| Content consistency | 0.413 | 0.650 | 0.620 | 0.665 | 0.662 |
| Physics adherence | 0.258 | 0.456 | 0.480 | 0.540 | 0.782 |
| 3D accuracy | 0.812 | 0.884 | 0.899 | 0.921 | 0.956 |
| Controllability | 0.635 | 0.774 | 0.788 | 0.819 | 0.869 |
| EWMScore | 44.88 | 58.70 | 58.65 | 61.75 | 66.92 |
Bold = best, underline = second best; Oracle scores the recorded videos and is excluded from ranking. Individual metrics do not uniformly favour aligned training (continued SFT has slightly higher image and aesthetic quality); the aggregate gain comes from JEPA similarity, trajectory accuracy, interaction quality and instruction following.
Conclusion
CtrlWAM makes the relationship between an action deviation and the observation it produces part of the denoising supervision. Warped schedules keep visual layout responsive while actions evolve, and multi-agent streams allow surrounding futures to be predicted or supplied as commands.
Limitations. Aligned noising needs a simulator that can execute perturbed actions; the multi-agent extension is shown for driving only; offline metrics do not establish closed-loop safety.
@misc{ctrlwam2026,
title = {CtrlWAM: Controllable World Action Models with Aligned Intent and Foresight},
year = {2026},
note = {Preprint}
}