# The Model That Dreams the World

**Source:** [moe-capital.com — The Model That Dreams the World](https://moe-capital.com/blog-home/the-model-that-dreams-the-world)  
**Authors:** [Henry Yin](https://linkeddata.uriburner.com/describe/?url=https%3A%2F%2Fmoe-capital.com%2Fblog-home%2Fthe-model-that-dreams-the-world%23henryYin) · [Naomi Xia](https://linkeddata.uriburner.com/describe/?url=https%3A%2F%2Fmoe-capital.com%2Fblog-home%2Fthe-model-that-dreams-the-world%23naomiXia)  
**Publisher:** [MoE Capital](https://linkeddata.uriburner.com/describe/?url=https%3A%2F%2Fmoe-capital.com%2Fblog-home%2Fthe-model-that-dreams-the-world%23moeCapital) · May 1, 2026  
**RDF (Turtle):** `world-models-moe-capital-claude-sonnet-1.ttl`  
**RDF (JSON-LD):** `world-models-moe-capital-claude-sonnet-1.jsonld`  
**Base IRI:** `https://moe-capital.com/blog-home/the-model-that-dreams-the-world#`

---

## Overview

[World models](https://linkeddata.uriburner.com/describe/?url=https%3A%2F%2Fmoe-capital.com%2Fblog-home%2Fthe-model-that-dreams-the-world%23worldModel) — learned internal representations that predict future states given actions — are converging from two historically separate threads into a single foundational technology for [Physical AI](https://linkeddata.uriburner.com/describe/?url=https%3A%2F%2Fmoe-capital.com%2Fblog-home%2Fthe-model-that-dreams-the-world%23physicalAI). This deep-dive from [MoE Capital](https://linkeddata.uriburner.com/describe/?url=https%3A%2F%2Fmoe-capital.com%2Fblog-home%2Fthe-model-that-dreams-the-world%23moeCapital) maps the $10B investment wave, [NVIDIA](https://linkeddata.uriburner.com/describe/?url=https%3A%2F%2Fmoe-capital.com%2Fblog-home%2Fthe-model-that-dreams-the-world%23nvidia)'s open-source physical AI stack, and the [JEPA](https://linkeddata.uriburner.com/describe/?url=https%3A%2F%2Fmoe-capital.com%2Fblog-home%2Fthe-model-that-dreams-the-world%23jepa) contrarian bet championed by Turing Award winner [Yann LeCun](https://linkeddata.uriburner.com/describe/?url=https%3A%2F%2Fmoe-capital.com%2Fblog-home%2Fthe-model-that-dreams-the-world%23yannLecun).

---

## Key Statistics

| Metric | Value |
|--------|-------|
| Total capital raised (world models sector) | ~$10B |
| NVIDIA DreamDojo training data | 44,711 hrs egocentric video |
| DreamDojo real-world correlation | r = 0.995 |
| DreamZero VLA generalization improvement | 2× |
| DreamGen: new behaviors from 1 demo | 22 |
| AMI Labs seed round (Yann LeCun) | $1.03B (largest European seed) |
| Wayve Series D | $1.2B |
| Decart (Oasis) raised | $153M |
| DreamerV4 speedup over V3 | 25× |
| Genie 3 resolution / frame rate | 720p · 24 FPS |

---

## The Convergence: Two Threads, One Technology

[World models](https://linkeddata.uriburner.com/describe/?url=https%3A%2F%2Fmoe-capital.com%2Fblog-home%2Fthe-model-that-dreams-the-world%23worldModel) emerged from two independent research traditions that are now merging:

### Thread A — The RL Dream-Machine (1990–2025)

World models as tools for *planning*: an agent learns to simulate its environment internally and trains policies inside that simulation ("dreaming"), dramatically reducing the need for real-world interaction.

| Year | Milestone |
|------|-----------|
| 1990 | Schmidhuber's original world model concept |
| 1991 | Recurrent World Models (Ha & Schmidhuber) |
| 2019 | PlaNet — pure latent imagination planning |
| 2020 | [DreamerV1](https://linkeddata.uriburner.com/describe/?url=https%3A%2F%2Fmoe-capital.com%2Fblog-home%2Fthe-model-that-dreams-the-world%23dreamer) — latent policy training beats model-free RL |
| 2020 | [MuZero](https://linkeddata.uriburner.com/describe/?url=https%3A%2F%2Fmoe-capital.com%2Fblog-home%2Fthe-model-that-dreams-the-world%23muzero) — reward/value world model; mastered Go, chess, Atari |
| 2023 | [DreamerV3](https://linkeddata.uriburner.com/describe/?url=https%3A%2F%2Fmoe-capital.com%2Fblog-home%2Fthe-model-that-dreams-the-world%23dreamer) published in Nature |
| 2025 | [DreamerV4](https://linkeddata.uriburner.com/describe/?url=https%3A%2F%2Fmoe-capital.com%2Fblog-home%2Fthe-model-that-dreams-the-world%23dreamer) — transformer-based, 25× faster |

Key figure: [Danijar Hafner](https://linkeddata.uriburner.com/describe/?url=https%3A%2F%2Fmoe-capital.com%2Fblog-home%2Fthe-model-that-dreams-the-world%23daniarHafner) (creator of the Dreamer series, PlaNet → DreamerV1–V4).

### Thread B — The Video Generation Arc (2016–2025)

[Video world models](https://linkeddata.uriburner.com/describe/?url=https%3A%2F%2Fmoe-capital.com%2Fblog-home%2Fthe-model-that-dreams-the-world%23videoWorldModel) represent state as photorealistic video frames, trained on internet-scale human video. This gives robotics and autonomous systems access to the full richness of the real world.

| Year | Milestone |
|------|-----------|
| 2016 | Generative video as implicit world model |
| 2024 | [Sora](https://linkeddata.uriburner.com/describe/?url=https%3A%2F%2Fmoe-capital.com%2Fblog-home%2Fthe-model-that-dreams-the-world%23sora) (OpenAI, Feb 2024) — bidirectional attention, non-interactive |
| 2024 | [Genie 1](https://linkeddata.uriburner.com/describe/?url=https%3A%2F%2Fmoe-capital.com%2Fblog-home%2Fthe-model-that-dreams-the-world%23genie) (DeepMind) — learns action spaces from unlabeled video |
| 2024 | [Oasis](https://linkeddata.uriburner.com/describe/?url=https%3A%2F%2Fmoe-capital.com%2Fblog-home%2Fthe-model-that-dreams-the-world%23oasis) (Decart) — playable Minecraft-like game at 20 FPS |
| 2025 | [Genie 3](https://linkeddata.uriburner.com/describe/?url=https%3A%2F%2Fmoe-capital.com%2Fblog-home%2Fthe-model-that-dreams-the-world%23genie) — 720p, 24 FPS, fully interactive |
| 2025 | [Cosmos Predict](https://linkeddata.uriburner.com/describe/?url=https%3A%2F%2Fmoe-capital.com%2Fblog-home%2Fthe-model-that-dreams-the-world%23cosmos) (NVIDIA) — 14B params, 200M video clips, Apache 2.0 |
| 2026 | [Pi-0.7](https://linkeddata.uriburner.com/describe/?url=https%3A%2F%2Fmoe-capital.com%2Fblog-home%2Fthe-model-that-dreams-the-world%23pi07) (Physical Intelligence) — VLA with world model for subgoal planning |
| 2026 | OpenAI shuts down Sora app (March), pivots to robotics simulation |

---

## Five Properties of a World Model

[Xun Huang](https://linkeddata.uriburner.com/describe/?url=https%3A%2F%2Fmoe-capital.com%2Fblog-home%2Fthe-model-that-dreams-the-world%23xunHuang) (co-developer of AR-DiT and Self Forcing) proposed five defining properties, distinguishing *hard requirements* (binary) from *soft* ones (spectrum):

| Property | Type | Description |
|----------|------|-------------|
| **Causal** | Hard | Must model cause-and-effect, not just correlations |
| **Interactive** | Hard | Must respond to agent actions in real time |
| **Persistent** | Soft | State must persist coherently across timesteps |
| **Real-time** | Soft | Must generate fast enough to be useful for control |
| **Physically accurate** | Soft | Must respect physical laws to varying degrees |

The key insight: Sora fails on *interactive* (bidirectional attention, not causal/real-time). MuZero fails on *physically accurate* (reward-only, never generating pixels). True world models for physical AI must satisfy at minimum the two hard constraints.

---

## Use Cases & Deployment Maturity

| Domain | Representative System | World Model Role | Maturity |
|--------|----------------------|-----------------|----------|
| Robotics policy evaluation | [DreamDojo](https://linkeddata.uriburner.com/describe/?url=https%3A%2F%2Fmoe-capital.com%2Fblog-home%2Fthe-model-that-dreams-the-world%23dreamDojo) (NVIDIA) | Evaluates policies before real deployment; r=0.995 with real outcomes | Production-ready |
| Robotics generalization | [DreamZero](https://linkeddata.uriburner.com/describe/?url=https%3A%2F%2Fmoe-capital.com%2Fblog-home%2Fthe-model-that-dreams-the-world%23dreamZero) (NVIDIA) | Joint video + action prediction; 2× VLA generalization | Early production |
| Synthetic robot data | [DreamGen](https://linkeddata.uriburner.com/describe/?url=https%3A%2F%2Fmoe-capital.com%2Fblog-home%2Fthe-model-that-dreams-the-world%23dreamGen) (NVIDIA) | Generates 22 new behaviors from a single teleoperation demo | Research/production |
| Subgoal planning | [Pi-0.7](https://linkeddata.uriburner.com/describe/?url=https%3A%2F%2Fmoe-capital.com%2Fblog-home%2Fthe-model-that-dreams-the-world%23pi07) (Physical Intelligence) | World model for compositional VLA planning | Early production |
| Autonomous vehicles | GAIA ([Wayve](https://linkeddata.uriburner.com/describe/?url=https%3A%2F%2Fmoe-capital.com%2Fblog-home%2Fthe-model-that-dreams-the-world%23wayve)) | Driving simulation and scenario generation | Production |
| RL training | [Dreamer V4](https://linkeddata.uriburner.com/describe/?url=https%3A%2F%2Fmoe-capital.com%2Fblog-home%2Fthe-model-that-dreams-the-world%23dreamer) (Hafner) | Latent-space policy training; 25× faster than V3 | Research |
| Interactive games | [Oasis](https://linkeddata.uriburner.com/describe/?url=https%3A%2F%2Fmoe-capital.com%2Fblog-home%2Fthe-model-that-dreams-the-world%23oasis) (Decart) | Playable Minecraft-like game on video world model at 20 FPS | Demo/product |

---

## The $10B Capital Landscape

| Company | Focus | Capital Raised |
|---------|-------|---------------|
| [AMI Labs](https://linkeddata.uriburner.com/describe/?url=https%3A%2F%2Fmoe-capital.com%2Fblog-home%2Fthe-model-that-dreams-the-world%23amiLabs) (Yann LeCun) | JEPA-based world models | $1.03B seed |
| [Wayve](https://linkeddata.uriburner.com/describe/?url=https%3A%2F%2Fmoe-capital.com%2Fblog-home%2Fthe-model-that-dreams-the-world%23wayve) | AV world model (GAIA) | $1.2B Series D |
| [Physical Intelligence](https://linkeddata.uriburner.com/describe/?url=https%3A%2F%2Fmoe-capital.com%2Fblog-home%2Fthe-model-that-dreams-the-world%23physicalIntelligence) | Pi-0 VLA series | Undisclosed |
| [Decart](https://linkeddata.uriburner.com/describe/?url=https%3A%2F%2Fmoe-capital.com%2Fblog-home%2Fthe-model-that-dreams-the-world%23decart) | Video world models (Oasis) | $153M |
| [NVIDIA](https://linkeddata.uriburner.com/describe/?url=https%3A%2F%2Fmoe-capital.com%2Fblog-home%2Fthe-model-that-dreams-the-world%23nvidia) | Full physical AI platform | Public / strategic |
| [Google DeepMind](https://linkeddata.uriburner.com/describe/?url=https%3A%2F%2Fmoe-capital.com%2Fblog-home%2Fthe-model-that-dreams-the-world%23googleDeepMind) | Genie series + MuZero | Public / strategic |
| [OpenAI](https://linkeddata.uriburner.com/describe/?url=https%3A%2F%2Fmoe-capital.com%2Fblog-home%2Fthe-model-that-dreams-the-world%23openAI) | Robotics simulation | Public / strategic |

The overall sector has attracted approximately **$10B** in investment as world models transition from research curiosities to production infrastructure.

---

## NVIDIA's Open Physical AI Stack

[NVIDIA](https://linkeddata.uriburner.com/describe/?url=https%3A%2F%2Fmoe-capital.com%2Fblog-home%2Fthe-model-that-dreams-the-world%23nvidia)'s strategy mirrors their CUDA play: own the platform, open-source the stack, and let the ecosystem drive adoption. All four layers are Apache 2.0:

```
┌─────────────────────────────────────────────┐
│  DreamDojo — Policy Evaluator               │
│  44,711 hrs egocentric video · r=0.995      │
├─────────────────────────────────────────────┤
│  DreamZero — World Action Model             │
│  Joint video + motor action · 2× VLA gen.  │
├─────────────────────────────────────────────┤
│  DreamGen — Synthetic Data Generator        │
│  22 new behaviors from 1 teleoperation demo │
├─────────────────────────────────────────────┤
│  Cosmos Predict — Video Foundation Model    │
│  14B params · 200M video clips · Apache 2.0 │
└─────────────────────────────────────────────┘
```

[Jim Fan](https://linkeddata.uriburner.com/describe/?url=https%3A%2F%2Fmoe-capital.com%2Fblog-home%2Fthe-model-that-dreams-the-world%23jimFan) (Head of NVIDIA GEAR Lab) called this "The Great Parallel" — robotics replicating the LLM playbook at scale.

The **EgoScale law** underpins the stack: NVIDIA found an R²=0.9983 scaling relationship between hours of human egocentric video used in training and real-world robot performance. This gives world models the same data-flywheel advantage LLMs have with internet text.

---

## The JEPA Contrarian Bet

[JEPA](https://linkeddata.uriburner.com/describe/?url=https%3A%2F%2Fmoe-capital.com%2Fblog-home%2Fthe-model-that-dreams-the-world%23jepa) (Joint Embedding Predictive Architecture), championed by [Yann LeCun](https://linkeddata.uriburner.com/describe/?url=https%3A%2F%2Fmoe-capital.com%2Fblog-home%2Fthe-model-that-dreams-the-world%23yannLecun) at his new $1.03B venture [AMI Labs](https://linkeddata.uriburner.com/describe/?url=https%3A%2F%2Fmoe-capital.com%2Fblog-home%2Fthe-model-that-dreams-the-world%23amiLabs), bets against the [video world model](https://linkeddata.uriburner.com/describe/?url=https%3A%2F%2Fmoe-capital.com%2Fblog-home%2Fthe-model-that-dreams-the-world%23videoWorldModel) consensus.

**The core argument:**
- Video world models waste compute generating every pixel — most of which is irrelevant for decision-making
- JEPA predicts in *abstract representation space* only, never generating pixels
- This is more efficient and more robust — closer to how biological intelligence works
- LeCun's thesis: autoregressive video prediction will plateau; abstract prediction will scale further

**The counterargument:**
- Internet-scale video gives video world models an enormous pre-training advantage
- Physical realism matters for robot sim-to-real transfer — pixels carry physical information
- Genie 3, DreamDojo, and DreamZero suggest the video paradigm is already at production quality

The JEPA vs. video world model debate is the central technical schism of the 2026 world model moment.

---

## The Great Parallel

[Jim Fan](https://linkeddata.uriburner.com/describe/?url=https%3A%2F%2Fmoe-capital.com%2Fblog-home%2Fthe-model-that-dreams-the-world%23jimFan) of NVIDIA GEAR Lab coined "[The Great Parallel](https://linkeddata.uriburner.com/describe/?url=https%3A%2F%2Fmoe-capital.com%2Fblog-home%2Fthe-model-that-dreams-the-world%23greatParallel)" — the observation that robotics is replicating the exact LLM playbook:

```
LLMs:     Text pretraining → Instruction fine-tuning → RLHF
Robotics: Video pretraining → Action fine-tuning    → RL
```

Just as LLMs benefited from the internet's text corpus, robot foundation models benefit from internet video of humans performing tasks. The world model is the bridge between passive video observation and active physical control.

---

## FAQ

**[Q1. What is a world model?](https://linkeddata.uriburner.com/describe/?url=https%3A%2F%2Fmoe-capital.com%2Fblog-home%2Fthe-model-that-dreams-the-world%23faq1)**  
A [world model](https://linkeddata.uriburner.com/describe/?url=https%3A%2F%2Fmoe-capital.com%2Fblog-home%2Fthe-model-that-dreams-the-world%23worldModel) is a learned internal representation of an environment that predicts future states given actions. Rather than acting blindly, an agent with a world model can *simulate* possible futures before committing to an action — enabling planning without real-world trial and error.

**[Q2. What's the difference between RL world models and video world models?](https://linkeddata.uriburner.com/describe/?url=https%3A%2F%2Fmoe-capital.com%2Fblog-home%2Fthe-model-that-dreams-the-world%23faq2)**  
RL world models (Thread A: [Dreamer](https://linkeddata.uriburner.com/describe/?url=https%3A%2F%2Fmoe-capital.com%2Fblog-home%2Fthe-model-that-dreams-the-world%23dreamer), [MuZero](https://linkeddata.uriburner.com/describe/?url=https%3A%2F%2Fmoe-capital.com%2Fblog-home%2Fthe-model-that-dreams-the-world%23muzero)) operate in compact latent spaces optimized for reward prediction — they never generate pixels. [Video world models](https://linkeddata.uriburner.com/describe/?url=https%3A%2F%2Fmoe-capital.com%2Fblog-home%2Fthe-model-that-dreams-the-world%23videoWorldModel) (Thread B: [Genie](https://linkeddata.uriburner.com/describe/?url=https%3A%2F%2Fmoe-capital.com%2Fblog-home%2Fthe-model-that-dreams-the-world%23genie), [Cosmos](https://linkeddata.uriburner.com/describe/?url=https%3A%2F%2Fmoe-capital.com%2Fblog-home%2Fthe-model-that-dreams-the-world%23cosmos)) represent state as photorealistic video frames, enabling training on internet-scale human video.

**[Q3. What is JEPA and why is Yann LeCun betting on it?](https://linkeddata.uriburner.com/describe/?url=https%3A%2F%2Fmoe-capital.com%2Fblog-home%2Fthe-model-that-dreams-the-world%23faq3)**  
[JEPA](https://linkeddata.uriburner.com/describe/?url=https%3A%2F%2Fmoe-capital.com%2Fblog-home%2Fthe-model-that-dreams-the-world%23jepa) (Joint Embedding Predictive Architecture) predicts in abstract representation space rather than pixel space. [Yann LeCun](https://linkeddata.uriburner.com/describe/?url=https%3A%2F%2Fmoe-capital.com%2Fblog-home%2Fthe-model-that-dreams-the-world%23yannLecun) argues this is more compute-efficient and better aligned with biological cognition. His $1.03B [AMI Labs](https://linkeddata.uriburner.com/describe/?url=https%3A%2F%2Fmoe-capital.com%2Fblog-home%2Fthe-model-that-dreams-the-world%23amiLabs) is the primary bet on this architecture.

**[Q4. How does NVIDIA's physical AI stack work?](https://linkeddata.uriburner.com/describe/?url=https%3A%2F%2Fmoe-capital.com%2Fblog-home%2Fthe-model-that-dreams-the-world%23faq4)**  
[NVIDIA](https://linkeddata.uriburner.com/describe/?url=https%3A%2F%2Fmoe-capital.com%2Fblog-home%2Fthe-model-that-dreams-the-world%23nvidia) open-sourced (Apache 2.0) a four-layer stack: [Cosmos Predict](https://linkeddata.uriburner.com/describe/?url=https%3A%2F%2Fmoe-capital.com%2Fblog-home%2Fthe-model-that-dreams-the-world%23cosmos) (video foundation model) → [DreamGen](https://linkeddata.uriburner.com/describe/?url=https%3A%2F%2Fmoe-capital.com%2Fblog-home%2Fthe-model-that-dreams-the-world%23dreamGen) (synthetic data) → [DreamZero](https://linkeddata.uriburner.com/describe/?url=https%3A%2F%2Fmoe-capital.com%2Fblog-home%2Fthe-model-that-dreams-the-world%23dreamZero) (world action model) → [DreamDojo](https://linkeddata.uriburner.com/describe/?url=https%3A%2F%2Fmoe-capital.com%2Fblog-home%2Fthe-model-that-dreams-the-world%23dreamDojo) (policy evaluator). Each layer composes with the next to form an end-to-end physical AI pipeline.

**[Q5. What are Xun Huang's five properties of a world model?](https://linkeddata.uriburner.com/describe/?url=https%3A%2F%2Fmoe-capital.com%2Fblog-home%2Fthe-model-that-dreams-the-world%23faq5)**  
[Xun Huang](https://linkeddata.uriburner.com/describe/?url=https%3A%2F%2Fmoe-capital.com%2Fblog-home%2Fthe-model-that-dreams-the-world%23xunHuang) defined: (1) **Causal** (hard) — models cause and effect; (2) **Interactive** (hard) — responds to agent actions; (3) **Persistent** (soft) — maintains coherent state over time; (4) **Real-time** (soft) — generates fast enough for control; (5) **Physically accurate** (soft) — respects physical laws. The two hard constraints rule out both Sora (non-interactive) and MuZero (non-pixel-generating).

**[Q6. What is "The Great Parallel"?](https://linkeddata.uriburner.com/describe/?url=https%3A%2F%2Fmoe-capital.com%2Fblog-home%2Fthe-model-that-dreams-the-world%23faq6)**  
"[The Great Parallel](https://linkeddata.uriburner.com/describe/?url=https%3A%2F%2Fmoe-capital.com%2Fblog-home%2Fthe-model-that-dreams-the-world%23greatParallel)" is [Jim Fan](https://linkeddata.uriburner.com/describe/?url=https%3A%2F%2Fmoe-capital.com%2Fblog-home%2Fthe-model-that-dreams-the-world%23jimFan)'s analogy: robotics is replicating the LLM playbook — video pretraining → action fine-tuning → RL — just as LLMs went from text pretraining → instruction fine-tuning → RLHF. The world model is the equivalent of the transformer — the core generative architecture everything else depends on.

**[Q7. Why did OpenAI shut down Sora in March 2026?](https://linkeddata.uriburner.com/describe/?url=https%3A%2F%2Fmoe-capital.com%2Fblog-home%2Fthe-model-that-dreams-the-world%23faq7)**  
[Sora](https://linkeddata.uriburner.com/describe/?url=https%3A%2F%2Fmoe-capital.com%2Fblog-home%2Fthe-model-that-dreams-the-world%23sora) used bidirectional attention — it could generate beautiful video but was fundamentally non-interactive, violating Xun Huang's *interactive* hard constraint. [OpenAI](https://linkeddata.uriburner.com/describe/?url=https%3A%2F%2Fmoe-capital.com%2Fblog-home%2Fthe-model-that-dreams-the-world%23openAI) pivoted the underlying technology toward robotics world simulation, where causal, interactive generation is essential.

---

## Glossary

| Term | Definition |
|------|-----------|
| [World Model](https://linkeddata.uriburner.com/describe/?url=https%3A%2F%2Fmoe-capital.com%2Fblog-home%2Fthe-model-that-dreams-the-world%23worldModel) | A learned internal model of an environment that predicts future states given actions, enabling planning by imagining futures. |
| [Video World Model](https://linkeddata.uriburner.com/describe/?url=https%3A%2F%2Fmoe-capital.com%2Fblog-home%2Fthe-model-that-dreams-the-world%23videoWorldModel) | A world model representing state as video frames — enabling photorealistic simulation from internet-scale human video. |
| [JEPA](https://linkeddata.uriburner.com/describe/?url=https%3A%2F%2Fmoe-capital.com%2Fblog-home%2Fthe-model-that-dreams-the-world%23jepa) | Joint Embedding Predictive Architecture — predicts in abstract representation space, never generating pixels. LeCun's contrarian bet. |
| [VLA](https://linkeddata.uriburner.com/describe/?url=https%3A%2F%2Fmoe-capital.com%2Fblog-home%2Fthe-model-that-dreams-the-world%23vla) | Vision-Language-Action Model — maps visual observations and language instructions to motor actions. Dominant production robotics approach. |
| [Physical AI](https://linkeddata.uriburner.com/describe/?url=https%3A%2F%2Fmoe-capital.com%2Fblog-home%2Fthe-model-that-dreams-the-world%23physicalAI) | AI that understands and acts in the physical world — encompassing world models, VLAs, and robot foundation models. |
| [The Great Parallel](https://linkeddata.uriburner.com/describe/?url=https%3A%2F%2Fmoe-capital.com%2Fblog-home%2Fthe-model-that-dreams-the-world%23greatParallel) | Jim Fan's analogy: robotics replicates the LLM playbook — video pretraining → action fine-tuning → RL. |
| EgoScale | NVIDIA's empirical scaling law: R²=0.9983 between hours of human egocentric video training data and real-world robot performance. |
| Latent Space | A compressed learned representation of environment state used by RL-based world models (Dreamer, JEPA) for efficient simulation. |

---

## Entity Index

### Persons

| Entity | Role | Resolver |
|--------|------|---------|
| [Henry Yin](https://linkeddata.uriburner.com/describe/?url=https%3A%2F%2Fmoe-capital.com%2Fblog-home%2Fthe-model-that-dreams-the-world%23henryYin) | Analyst / Author, MoE Capital | [🔗 Explore](https://linkeddata.uriburner.com/describe/?url=https%3A%2F%2Fmoe-capital.com%2Fblog-home%2Fthe-model-that-dreams-the-world%23henryYin) |
| [Naomi Xia](https://linkeddata.uriburner.com/describe/?url=https%3A%2F%2Fmoe-capital.com%2Fblog-home%2Fthe-model-that-dreams-the-world%23naomiXia) | Analyst / Author, MoE Capital | [🔗 Explore](https://linkeddata.uriburner.com/describe/?url=https%3A%2F%2Fmoe-capital.com%2Fblog-home%2Fthe-model-that-dreams-the-world%23naomiXia) |
| [Yann LeCun](https://linkeddata.uriburner.com/describe/?url=https%3A%2F%2Fmoe-capital.com%2Fblog-home%2Fthe-model-that-dreams-the-world%23yannLecun) | Turing Award winner; founder AMI Labs; JEPA champion | [🔗 Explore](https://linkeddata.uriburner.com/describe/?url=https%3A%2F%2Fmoe-capital.com%2Fblog-home%2Fthe-model-that-dreams-the-world%23yannLecun) |
| [Danijar Hafner](https://linkeddata.uriburner.com/describe/?url=https%3A%2F%2Fmoe-capital.com%2Fblog-home%2Fthe-model-that-dreams-the-world%23daniarHafner) | Creator of Dreamer series (PlaNet → DreamerV1–V4) | [🔗 Explore](https://linkeddata.uriburner.com/describe/?url=https%3A%2F%2Fmoe-capital.com%2Fblog-home%2Fthe-model-that-dreams-the-world%23daniarHafner) |
| [Xun Huang](https://linkeddata.uriburner.com/describe/?url=https%3A%2F%2Fmoe-capital.com%2Fblog-home%2Fthe-model-that-dreams-the-world%23xunHuang) | AR-DiT / Self Forcing; Five Properties framework | [🔗 Explore](https://linkeddata.uriburner.com/describe/?url=https%3A%2F%2Fmoe-capital.com%2Fblog-home%2Fthe-model-that-dreams-the-world%23xunHuang) |
| [Jim Fan](https://linkeddata.uriburner.com/describe/?url=https%3A%2F%2Fmoe-capital.com%2Fblog-home%2Fthe-model-that-dreams-the-world%23jimFan) | Head of NVIDIA GEAR Lab; coined "The Great Parallel" | [🔗 Explore](https://linkeddata.uriburner.com/describe/?url=https%3A%2F%2Fmoe-capital.com%2Fblog-home%2Fthe-model-that-dreams-the-world%23jimFan) |

### Organizations

| Entity | Description | Resolver |
|--------|-------------|---------|
| [MoE Capital](https://linkeddata.uriburner.com/describe/?url=https%3A%2F%2Fmoe-capital.com%2Fblog-home%2Fthe-model-that-dreams-the-world%23moeCapital) | Early-stage VC — agentic infrastructure & AI for Science | [🔗 Explore](https://linkeddata.uriburner.com/describe/?url=https%3A%2F%2Fmoe-capital.com%2Fblog-home%2Fthe-model-that-dreams-the-world%23moeCapital) |
| [NVIDIA](https://linkeddata.uriburner.com/describe/?url=https%3A%2F%2Fmoe-capital.com%2Fblog-home%2Fthe-model-that-dreams-the-world%23nvidia) | Dominant physical AI platform; open-sourced full stack | [🔗 Explore](https://linkeddata.uriburner.com/describe/?url=https%3A%2F%2Fmoe-capital.com%2Fblog-home%2Fthe-model-that-dreams-the-world%23nvidia) |
| [Google DeepMind](https://linkeddata.uriburner.com/describe/?url=https%3A%2F%2Fmoe-capital.com%2Fblog-home%2Fthe-model-that-dreams-the-world%23googleDeepMind) | Genie series + MuZero | [🔗 Explore](https://linkeddata.uriburner.com/describe/?url=https%3A%2F%2Fmoe-capital.com%2Fblog-home%2Fthe-model-that-dreams-the-world%23googleDeepMind) |
| [OpenAI](https://linkeddata.uriburner.com/describe/?url=https%3A%2F%2Fmoe-capital.com%2Fblog-home%2Fthe-model-that-dreams-the-world%23openAI) | Sora (Feb 2024); shut Sora app Mar 2026, pivots to robotics | [🔗 Explore](https://linkeddata.uriburner.com/describe/?url=https%3A%2F%2Fmoe-capital.com%2Fblog-home%2Fthe-model-that-dreams-the-world%23openAI) |
| [Physical Intelligence](https://linkeddata.uriburner.com/describe/?url=https%3A%2F%2Fmoe-capital.com%2Fblog-home%2Fthe-model-that-dreams-the-world%23physicalIntelligence) | Pi-0 VLA series | [🔗 Explore](https://linkeddata.uriburner.com/describe/?url=https%3A%2F%2Fmoe-capital.com%2Fblog-home%2Fthe-model-that-dreams-the-world%23physicalIntelligence) |
| [AMI Labs](https://linkeddata.uriburner.com/describe/?url=https%3A%2F%2Fmoe-capital.com%2Fblog-home%2Fthe-model-that-dreams-the-world%23amiLabs) | LeCun's JEPA lab; $1.03B seed (largest European seed) | [🔗 Explore](https://linkeddata.uriburner.com/describe/?url=https%3A%2F%2Fmoe-capital.com%2Fblog-home%2Fthe-model-that-dreams-the-world%23amiLabs) |
| [Decart](https://linkeddata.uriburner.com/describe/?url=https%3A%2F%2Fmoe-capital.com%2Fblog-home%2Fthe-model-that-dreams-the-world%23decart) | Oasis interactive video world model; $153M | [🔗 Explore](https://linkeddata.uriburner.com/describe/?url=https%3A%2F%2Fmoe-capital.com%2Fblog-home%2Fthe-model-that-dreams-the-world%23decart) |
| [Wayve](https://linkeddata.uriburner.com/describe/?url=https%3A%2F%2Fmoe-capital.com%2Fblog-home%2Fthe-model-that-dreams-the-world%23wayve) | AV — GAIA world model for driving simulation; $1.2B Series D | [🔗 Explore](https://linkeddata.uriburner.com/describe/?url=https%3A%2F%2Fmoe-capital.com%2Fblog-home%2Fthe-model-that-dreams-the-world%23wayve) |

### AI Systems

| Entity | Creator | Description | Resolver |
|--------|---------|-------------|---------|
| [Dreamer V1–V4](https://linkeddata.uriburner.com/describe/?url=https%3A%2F%2Fmoe-capital.com%2Fblog-home%2Fthe-model-that-dreams-the-world%23dreamer) | Hafner | RL latent-space world model series; V3 in Nature 2025; V4 25× faster | [🔗](https://linkeddata.uriburner.com/describe/?url=https%3A%2F%2Fmoe-capital.com%2Fblog-home%2Fthe-model-that-dreams-the-world%23dreamer) |
| [Genie 1–3](https://linkeddata.uriburner.com/describe/?url=https%3A%2F%2Fmoe-capital.com%2Fblog-home%2Fthe-model-that-dreams-the-world%23genie) | DeepMind | Interactive video world model; learns actions from unlabeled video | [🔗](https://linkeddata.uriburner.com/describe/?url=https%3A%2F%2Fmoe-capital.com%2Fblog-home%2Fthe-model-that-dreams-the-world%23genie) |
| [DreamDojo](https://linkeddata.uriburner.com/describe/?url=https%3A%2F%2Fmoe-capital.com%2Fblog-home%2Fthe-model-that-dreams-the-world%23dreamDojo) | NVIDIA | Policy evaluator; 44,711 hrs video; r=0.995 | [🔗](https://linkeddata.uriburner.com/describe/?url=https%3A%2F%2Fmoe-capital.com%2Fblog-home%2Fthe-model-that-dreams-the-world%23dreamDojo) |
| [DreamZero](https://linkeddata.uriburner.com/describe/?url=https%3A%2F%2Fmoe-capital.com%2Fblog-home%2Fthe-model-that-dreams-the-world%23dreamZero) | NVIDIA | World action model; joint video + motor prediction; 2× VLA gen. | [🔗](https://linkeddata.uriburner.com/describe/?url=https%3A%2F%2Fmoe-capital.com%2Fblog-home%2Fthe-model-that-dreams-the-world%23dreamZero) |
| [DreamGen](https://linkeddata.uriburner.com/describe/?url=https%3A%2F%2Fmoe-capital.com%2Fblog-home%2Fthe-model-that-dreams-the-world%23dreamGen) | NVIDIA | Synthetic robot data; 22 behaviors from 1 teleoperation demo | [🔗](https://linkeddata.uriburner.com/describe/?url=https%3A%2F%2Fmoe-capital.com%2Fblog-home%2Fthe-model-that-dreams-the-world%23dreamGen) |
| [Cosmos Predict](https://linkeddata.uriburner.com/describe/?url=https%3A%2F%2Fmoe-capital.com%2Fblog-home%2Fthe-model-that-dreams-the-world%23cosmos) | NVIDIA | Video foundation model; 14B params; 200M video clips; Apache 2.0 | [🔗](https://linkeddata.uriburner.com/describe/?url=https%3A%2F%2Fmoe-capital.com%2Fblog-home%2Fthe-model-that-dreams-the-world%23cosmos) |
| [Sora](https://linkeddata.uriburner.com/describe/?url=https%3A%2F%2Fmoe-capital.com%2Fblog-home%2Fthe-model-that-dreams-the-world%23sora) | OpenAI | Video model Feb 2024; bidirectional (non-interactive); shut Mar 2026 | [🔗](https://linkeddata.uriburner.com/describe/?url=https%3A%2F%2Fmoe-capital.com%2Fblog-home%2Fthe-model-that-dreams-the-world%23sora) |
| [Pi-0.7](https://linkeddata.uriburner.com/describe/?url=https%3A%2F%2Fmoe-capital.com%2Fblog-home%2Fthe-model-that-dreams-the-world%23pi07) | Physical Intelligence | VLA with world model for subgoal planning (Apr 2026) | [🔗](https://linkeddata.uriburner.com/describe/?url=https%3A%2F%2Fmoe-capital.com%2Fblog-home%2Fthe-model-that-dreams-the-world%23pi07) |
| [Oasis](https://linkeddata.uriburner.com/describe/?url=https%3A%2F%2Fmoe-capital.com%2Fblog-home%2Fthe-model-that-dreams-the-world%23oasis) | Decart | Playable Minecraft-like game on video world model at 20 FPS | [🔗](https://linkeddata.uriburner.com/describe/?url=https%3A%2F%2Fmoe-capital.com%2Fblog-home%2Fthe-model-that-dreams-the-world%23oasis) |
| [MuZero](https://linkeddata.uriburner.com/describe/?url=https%3A%2F%2Fmoe-capital.com%2Fblog-home%2Fthe-model-that-dreams-the-world%23muzero) | DeepMind | Reward/value world model (2020); mastered Go, chess, Atari | [🔗](https://linkeddata.uriburner.com/describe/?url=https%3A%2F%2Fmoe-capital.com%2Fblog-home%2Fthe-model-that-dreams-the-world%23muzero) |

---

## Provenance

| Role | Tool / System | Link |
|------|--------------|------|
| Knowledge Extraction | KG Generator Skill | [🔗 Visit](https://github.com/OpenLinkSoftware/ai-agent-skills/tree/main/kg-generator) |
| HTML Infographic | RDF Infographic Skill v1.1 | [🔗 Visit](https://github.com/OpenLinkSoftware/ai-agent-skills/tree/main/rdf-infographic-skill) |
| LLM Reasoning | Claude Sonnet (claude-sonnet-4-6) | [🔗 Visit](https://www.anthropic.com/claude) |
| Execution Environment | Cowork Desktop Environment | [🔗 Visit](https://claude.ai/download) |
| Linked Data Resolver | URIBurner | [🔗 Visit](https://linkeddata.uriburner.com/) |
| RDF Quad Store | OpenLink Virtuoso Server | [🔗 Visit](https://virtuoso.openlinksw.com/) |

---

*Generated from source article: [The Model That Dreams the World](https://moe-capital.com/blog-home/the-model-that-dreams-the-world) · MoE Capital · May 1, 2026*
