Model Releases
Learning causality from internet videos in latent space first, and then using RL to teach the foundation model how to act. This approach is …
Learning causality from internet videos in latent space first, and then using RL to teach the foundation model how to act. This approach is 30× cheaper than Gemini 3.1 Flash on pretraining and achieve
Learning causality from internet videos in latent space first, and then using RL to teach the foundation model how to act. This approach is 30× cheaper than Gemini 3.1 Flash on pretraining and achieves a better result. JEPA is all you need! However, actions are still learned during post-training. Figuring out how to automatically learn actions without action labels is still very important! We’re introducing imagination models: a new foundation model architecture that unlocks learning from internet-scale video. Our first imagination model, Photon-1, learned to use a computer by watching 18 years of screen recording video without action labels.
Related
- Gemini Omni Flash can swap objects, change environments, and make targeted edits to existing videos using simple text prompts. To try this w…
- FOUND-IT: Foundation-model-first Task-driven 3D Scene Graphs with Granularity on Demand
- CLASP: Language-Driven Robot Skill Selection and Composition using Task-Parameterized Learning
- Attacking the Trusted Imagination: Oracle-Level Integrity Attacks on Imagine-then-Act World Models
Source: Yann LeCun (X) | 2026-07-24