A product discussed on Latent Space.

Inside xAI: Building Grok Imagine in 3 Months, Videogen vs World Models, and Video Agents— Ethan He
Jun 1, 2026 · 1:44:43
Ethan He, former xAI and NVIDIA Cosmos researcher, explains how xAI built its first image and video models (Grok Imagine 0.9) from zero to one in three months, attributing rapid iteration to small teams with minimal meetings and strong infra that enabled fixing tiny data and training bugs for biggest quality gains. He argues that most improvements in video generation now come from language models and agents rather than diffusion technology, predicting that by end of 2025 video agents will produce production-grade content for ads. He defines world models as real-time, interactive, long-horizon videos, and details challenges like temporal compression, context management, and the high cost of storing and moving video data (e.g., tens of petabytes for a billion videos). Ethan also shares why he left xAI to focus on language model research, believing the next frontier is models that manage their own context length, similar to solutions already being explored in video generation.

SAM 3: The Eyes for AI — Nikhila & Pengchuan (Meta Superintelligence), ft. Joseph Nelson (Roboflow)
Dec 18, 2025 · 1:15:04
Meta's SAM 3, introduced by Nikhila Ravi and Pengchuan Zhang alongside Roboflow CEO Joseph Nelson, unifies interactive segmentation, open-vocabulary detection, and video tracking into a single model that runs in 30ms on images and scales to real-time video on multi-GPU setups. The model uses concept prompts like "yellow school bus" to detect and segment every instance, separating recognition from localization via a presence token. Its data engine automated exhaustive annotation from two minutes per image down to 25 seconds using AI verifiers fine-tuned on Llama, while the new SACO benchmark contains over 200,000 unique concepts versus previous 1.2k. For video, decoupling the detector and tracker preserves object identity, and SAM 3 agents pair with multimodal LLMs like Gemini to handle complex visual reasoning. The real-world impact includes 106 million smart polygons created on Roboflow, saving an estimated 130+ years of labeling time across fields from cancer research to underwater trash cleanup.
Powered by PodHood