Summary
A concise design-space tutorial that categorizes world models as action-conditioned predictive models and traces the transition to World Action Models (WAMs). The paper compares observation-space vs. state-space world models across visual fidelity, spatial structure, physical interpretability, and control usability, then introduces four representative WAM paradigms: imagine-then-execute, video-feature-conditioned action prediction, joint video-action modeling, and auxiliary video prediction for policy learning.
Key Contributions
- Clean taxonomy of world model design axes (observation-space vs. state-space, predictive vs. generative, low-level vs. high-level) enabling principled comparison across existing methods
- Four WAM paradigm categorization with clear trade-off analysis for each approach relative to policy quality, inference cost, and data efficiency
- Accessible survey of representative recent methods and how they instantiate each paradigm — useful reference for researchers entering the field
Significance
Provides the clearest conceptual roadmap available for understanding the design space from classical world models to the emerging World Action Model paradigm, filling a gap left by longer survey papers in accessibility and clarity.