World Models and World-Action Models

From VLAs to WAMs: what changed by August 2026

Yunsung Lee

WoRV / MaumAI

2026-08-31

Yunsung Lee

  • Team Lead, WoRV at MaumAI: robot foundation models for wheeled platforms
  • Research: world models, robot learning, embodied AI
  • Adjunct professor, DGIST (Physical AI lecture, Fall 2026)
  • Incoming member, TTA PG1004

01

Robot Foundation Models Today

The VLA recipe, and where it strains

A VLA is a vision-language model that outputs robot actions

camera pixels language instruction one network the VLA motor commands human teleop demonstrations imitation
  • Trained by imitation on human teleoperation demonstrations
  • The de facto recipe for robot foundation models since 2023

Three years of the same recipe, scaled

Apr 2023
ALOHA
Jul 2023
RT-2
Jun 2024
OpenVLA
Oct 2024
LAPA
Oct 2024
π0
Feb 2025
Helix
Apr 2025
π0.5
Apr 2026
π0.7

Bigger backbones, more teleop hours, broader embodiments: the axis of progress stayed the same.

Dual-system: a slow VLM thinks, a fast policy acts

Helix: System 1 / System 2Figure AI · Feb 2025 · autonomy claimed · 1x speed
7 to 9 HzSystem 2: VLM reasoning
200 HzSystem 1: visuomotor control

Every VLA stands on human demonstration data

ALOHA teleoperationStanford · Apr 2023 · teleoperated · 1x speed
UMI in-the-wild collectionStanford and Columbia · Feb 2024 · teleoperated · 1x speed

Teleoperation rigs and handheld grippers feed every VLA. The strain: an hour of robot time buys one hour of one embodiment’s data.

2026: VLAs began absorbing world models
π0.7visual subgoals from a lightweight world model at inference
zero-shotexpert-level folding on an unseen UR5e
<200 demosGemini Robotics 2 On-Device adapts to new bi-arm robots
Shirt folding on an unseen UR5ePhysical Intelligence · Apr 2026 · autonomous · 1x speed

02

World Models

Learning the environment itself

A world model learns what happens next

current state an action world model learned dynamics what happens next imagine rollouts, then pick an action before moving
  • Act by imagining futures before moving: the Dreamer lineage, 2019 onward
  • World models complement policies: dynamics, not behavior

Three families of world models

Family Exemplars What it gives you
Latent dynamics Dreamer imagination for planning and RL
Video generation Genie, Cosmos promptable, general visual worlds
Neural simulators DreamDojo, GE-Sim action-conditioned rollouts for robots

Generated worlds became persistent, promptable places
Genie 3real-time interactive generation
minutesof environmental consistency
Genie 3: persistent worldsGoogle DeepMind · Aug 2025 · autonomy not stated · 1x speed

From video models to robot simulators

Cosmos world generationNVIDIA · Jan 2025 · autonomy not stated · 1x speed
  • World models entered robot stacks as data engines and evaluators
  • NVIDIA Cosmos: open world foundation models built for physical AI

Cosmos 3: one omnimodal world model

fully openreasoning, generation, and action in one model
ranked #1among open models in world generation and action policy
Japanprospective members and physical-AI stack builders
  • Adopters: Agile Robots, Doosan Robotics, LG Electronics, Samsung Electronics, Skild AI
Cosmos 3: reasoning to generationNVIDIA · Jun 2026 · autonomy not stated · 1x speed

Announced at NVIDIA GTC Taipei, May 31, 2026; rankings per NVIDIA’s launch announcement.

Everyone is pointing world models at robot training

Who May to Jul 2026 The move
Google DeepMind Project Genie + Street View Street View imagery grounds interactive worlds
World Labs SceniX, real to sim to real scanned scenes into policy-training simulation
GE-Sim 2.0 one world model, three jobs scores rollouts, yields rewards, accelerates them

Interactive world models hit real time
44,000 hof human video pretraining
real timeaction-conditioned rollout
open codegithub.com/NVIDIA/DreamDojo
Real-time interactive world modelNVIDIA · Feb 2026 · autonomy not stated · 1x speed

Two paradigms, one team

DreamZero DreamDojo
Direction joint: video + action forward: action to video
The product a zero-shot policy a neural simulator
Training data 500 h of robot data 44,000 h of human video
Its job acts imagines

Both were released by NVIDIA in February 2026: two halves of one loop.

03

World-Action Models

One model that imagines and acts

A joint-prediction WAM denoises video and action together

noise one joint model DreamZero-style WAM future frames future actions human egocentric video video prediction = world-dynamics supervision
  • Human video supplies world-dynamics supervision
  • Executable actions still need pseudo-labels or robot trajectories
  • DreamZero coined the category: “World Action Model”

DreamZero: the policy is a world model
Elo 1736vs 1609 for π0.5 on RoboArena (Aug 29 snapshot)
zero-shoton unseen tasks and embodiments
500 hof robot data; video-first training
DreamZeroNVIDIA · Feb 2026 · autonomous · 1x speed

Where does action supervision come from?

Latent action pretrainingKAIST and Microsoft · Oct 2024 · autonomous · 1x speed
  • Inverse dynamics models label raw video with pseudo-actions
  • LAPA (2024) learns latent actions from human video, with no robot labels at all

From two stages to one joint model

2024: two stages, trained apart video model separate action decoder two models, two training runs 2026: one joint model one transformer frames + actions, denoised jointly video actions at inference, run the action pathway alone
  • Imagination becomes optional at inference: run the action pathway alone

The architecture consolidated in 2026

Style Idea Exemplar
Joint denoising one flow over video and action DreamZero
Interleaved mixture mixture-of-transformers over frames and actions Dyna-2
Latent world + decoder imagine in latent space, decode actions LAPA lineage

Representation-only Fast-WAM skips video generation at inference and delivers VLA-class latency at about 190 ms.

Dyna-2: 1M hours of human video, zero robot data in pretraining
1,000,000 hof egocentric human video
firsthuman-to-robot scaling law
1.55xthe success of its VLA twin, same data
Dyna-2 hero reelDyna Robotics · Aug 2026 · autonomy claimed · 1x speed

Dyna-2 is deployed, not demoed

Chopping through a blackoutDyna Robotics · Aug 2026 · autonomy claimed · 1x speed
95 vs 35napkins per hour (Dyna-2 vs its VLA twin)
93% vs 75%fold quality
3 daysfrom setup to meeting the production ROI bar

Named customer: Din Tai Fung. The fleet collects about 1 TB per day.

The race is for hours, not parameters

Who Scale What
Dyna Robotics 1,000,000 h egocentric human video, no robot data in pretraining
Generalist AI 500,000+ h manipulation data, about 9,000 end-effectors
Lightwheel EgoSuite 100,000 h, open egocentric video, commercial training allowed
Amazon ABC-130K 3,553 h, open 134,806 bimanual teleop episodes
China 90+ data factories operating or under construction

China: Interact Analysis, Aug 2026. Curse of Precision, Jul 2026: data needs grow super-exponentially near the system limit.

1M

What does a million hours buy?

Six weeks in summer 2026: the in-context learning wave

Four one-shot announcements in six weeks

Date Who Claim
Jul 16 RoboTTT, NVIDIA context length as a new scaling axis; one-shot 6/10 from a human video
Aug 18 Skild S1 one human video runs an unseen 10-minute task; 66% vs 9%
Aug 19 GEN-1.5, Generalist AI physical prompting: one-shot 59%, no gradient updates
Aug 26 Zero-WAM, Robbyant in-context WAM: +29.5 pp over the best baseline

RoboTTT: context became a scaling axis

+62%task completion, 8K vs 1K pretraining context
8K stepsabout 5 minutes at 30 Hz
constantlatency via TTT fast weights

Figure: NVIDIA GEAR, RoboTTT, arXiv:2607.15275.

GEN-1.5: the demo is the program
59%one-shot on unseen tasks, weights frozen
83%few-shot after just 10 gradient steps
emergentno meta-learning, 8+ months of pretraining
Human demo to robot, in contextGeneralist AI · Aug 2026 · autonomy claimed · 1x speed

One human video vs 380 teleop episodes

Prompt: a human makes pour-overSkild AI · Aug 2026 · autonomy not stated · 1x speed
Execution: the robot followsSkild AI · Aug 2026 · autonomous · 4x speed
66% vs 9%video-prompted S1 vs language-prompted VLA
1 demoworth about 380 post-training episodes
frozenno weight updates at deployment

Zero-WAM: the two storylines converge

Prompt: human videoRobbyant Research · Aug 2026 · autonomy not stated · 1x speed
Execution: unseen taskRobbyant Research · Aug 2026 · autonomy claimed · 2x speed
in contexta human video is the task spec
47.0%held-out tasks, +29.5 pp over the best baseline
Aug 26published five days before this talk

04

So What Changed?

VLA, world model, WAM, side by side

Five architectures, one scoreboard

Architecture Scales on Actions come from August 2026 signal
VLA teleop demos imitation head π0.5: RoboArena Elo 1609
VLA + WM demos + video imitation, WM-guided π0.7 visual subgoals
World model video none: it imagines real-time neural simulators
WAM egocentric video joint or video-conditioned action pathway DreamZero: Elo 1736
Hybrid everything both pathways the predicted winner

NVIDIA, Jun 2026: action tuning uses 9 ZFLOPs; full video pretraining uses 50+ ZFLOPs; generative WAM inference takes 3 to 4 times as long as π0.5.

If you remember one slide

VLA World model WAM
Learns behavior dynamics dynamics + action coupling
Scales on teleop hours video video priors + robot trajectories
The 2026 turn absorbing world models converging on robot training one-shot via in-context learning

WoRV at MaumAI: robot foundation models for wheeled platforms
WoRV manipulationMaumAI WoRV · Aug 2026 · autonomy not stated · 1x speed

CostNav: navigation as unit economics

CostNavMaumAI WoRV · May 2026 · autonomy not stated · 1x speed
  • Delivery robots evaluated on cost per delivery, not success rate alone
  • Public research from the WoRV team: arXiv:2511.20216

QR code linking to these slides
alohays.github.io/paper2pr/slides/talks/tta-2026.html

Thank you. Questions welcome: sung@maum.ai

Appendix

Backup detail for Q&A

A1. Dyna-2: the scaling law in detail

cross-embodiment transfer emerges 0 20 40 60 14-task mean normalized score, % 20% 28% 45% 53% 1k h 10k h 100k h 1M h human video hours, log scale, nested subsets
  • Monotonic on-robot scaling across four rungs, 1k to 1M hours
  • Distilled sampler: 10,203 ms to 110 ms, about 90x, one H100

A2. Dyna-2: deployment and infrastructure

65%win rate vs Dyna-1, its VLA twin, same data
87% vs 46%new-site zero-shot success
10% to 50%bottle-cap success after about 10 minutes of data
  • Ingestion 31x faster: 1M hours in under 3 weeks

A3. RoboTTT: the mechanism

538M to 690Maction-head parameters before and after TTT
8K stepscontext at constant per-step latency
79% vs 42%assembly, with vs without TTT
  • One-shot from a human video: 6/10 vs 0/10
  • Bimanual YAM arms; egocentric human video in the mix

A4. GEN-1.5: lineage and evidence

59% ± 10% SDone-shot on unseen tasks
83% ± 9% SDfew-shot after 10 gradient steps
<0.15%weight movement in those ten steps
  • Three releases: GEN-0 270k h, GEN-1 500k+ h, then GEN-1.5
  • Composed prompts interpolate: regrasping and recovery emerge

A5. Skild S1: ground-up in-context learning

Prompt: a human makes a pancakeSkild AI · Aug 2026 · autonomy not stated · 1x speed
Execution under perturbationSkild AI · Aug 2026 · autonomous · 2x speed
1k to 100k hpretraining scale sweep
up to 3xas much degradation: VLA vs S1 at L5

A6. Zero-WAM and the academic lineage

Thread Evidence
HumanGen dataset 74.2K human-robot pairs across 8.6K tasks
RoboTwin 2.0 held-out suite 47.0%, +29.5 pp over LingBot-VA
BPP, Stanford task diversity, not hours, drives ICL ability
MimicDroid, UT Austin the earlier academic thread, ICRA 2026

A7. The data race: sources and fine print

Source The fine print
EgoSuite-Open100K 90k head-view + 10k wrist hours, LeRobot format
ABC-130K 3,553 h, 134,806 episodes, bimanual YAM, fully open
UBTECH 566M CNY of humanoid orders across three cities, 2025
Curse of Precision \log N \propto 1/(P-c): required demos grow super-exponentially as precision P nears limit c

A8. Reading list

Topic Where
DreamZero and DreamDojo 2602.15922, 2602.06949
WAM tutorial, survey, Fast-WAM 2607.00836, 2606.00113, 2603.16666
Cosmos 3; Rise of World-Action Models 2606.02800; NVIDIA blog, Jun 2026
RoboTTT, Zero-WAM, BPP 2607.15275, 2608.26103, 2606.30457