A company discussed on Latent Space.

🔬 RL with Verifiable Rewards, but the Verifier is a Lab — Lila Sciences
Jul 16, 2026 · 1:41:04
Andy Beam (CTO) and Rafa Gómez-Bombarelli (Co-founder & CSO of Physical Sciences) of Lila Sciences argue that science is an 'infinite token generator' for AI, using reinforcement learning with verifiable rewards where the wet lab acts as the verifier. They claim one general model trained on ~10 trillion experimentally-verified reasoning tokens across biology, chemistry, and materials outperforms domain-specific models—'breadth gives us depth.' Their AI Science Factories treat the lab as a data center, with instruments on a 'PCI bus' and humans 'below the API line.' Highlights include a CAR-T candidate designed in six months by two or three people, 'monster UTRs' achieving ~10x Moderna/Pfizer mRNA expression, and the 'zero-FTE startup' business model. They discuss RL pathologies like collapsed chains of thought and a model that 'swears,' and note why there is still no AlphaFold for materials due to the sim-to-real gap.

⚡️Every product of the future will be a living system — Ronak Malde, Trajectory.ai
Jun 21, 2026 · 33:58
Ronak Malde, CEO of Trajectory.ai, recounts his journey from building AI coding agents at Windsurf (acquired by Google DeepMind in a deal involving Demis Hassabis and Sergey Brin) to launching a platform for continual learning in enterprise AI. He argues that every future product must be a 'living system' that learns from real-world user interactions, not static models. The episode details Trajectory's technical innovations, including Self-Distillation Policy Optimization (SDPO) for learning from corrections, continuous LoRA for parallel training, and open-sourcing a training stack with SkyRL. It covers partnerships with Harvey and NVIDIA to train NeMoTron 3 Super for legal workflows, improving metrics like issue spotting and citation accuracy while cutting costs. Malde explains data curation strategies that capture nuanced user edits beyond binary signals, and outlines Trajectory's roadmap from AI-native companies like Clay and Decagon to Fortune 500 enterprises. He also reveals that the idea for Trajectory emerged after giving up his acquisition equity to pursue this vision.

Why AI Labs With Unlimited GPUs Still Fail — Anjney Midha, AMP
Jun 18, 2026 · 1:00:37
Anjney Midha, CEO of AMP, argues that AI labs with unlimited GPUs still fail due to misaligned culture and infrastructure waste, proposing a compute grid modeled on independent system operators to pool demand and supply. At Google, 95% node utilization was considered an outage, yet most clusters today don't reach that, with waste compounding at scale. AMP’s grid, starting at scheduling, aims to make FLOPs flow like megawatts, having secured 1.3 gigawatts of demand. Midha explains Anthropic cracked coding because 'luck favors the prepared mind'—their four years of paranoia and scarcity created a culture that OpenAI’s abundance couldn't replicate. He also shares a 14-year mission in end-of-life prediction, arguing AI can reduce the 30% of Medicare/Medicaid spend on end-of-life care. He warns that too much capital too early makes labs fragile because without hardship they fail to define their P0.

⚡️ Google's Open AI Strategy — Omar Sanseviero, Google DeepMind
May 24, 2026 · 29:59
Omar Sanseviero, Google DeepMind's Head of Developer Experience, explains Gemma 4's novel architecture with per-layer embeddings that enable effective parameter offloading: only 2B of 5B parameters need GPU memory, ideal for on-device inference on phones and Raspberry Pis. The model matches 1.5-year-old state-of-the-art in most areas, with Gemini Nano integrated into Pixel and Samsung phones. Gemma 4 supports multimodal input (audio, images, short video) but not audio output or combined audio-video prompts. Sanseviero notes fine-tuning is declining as out-of-box capabilities improve, but remains relevant for specialized domains like healthcare. He contrasts dense (31B) and MoE (27B) variants, highlighting MoE's inference speed but fine-tuning challenges. The team is expanding globally, with Kaggle joining DeepMind to create community-driven benchmarks for model evaluation.

Measuring Exponential Trends Rising (in AI) — Joel Becker, METR
Feb 27, 2026 · 1:05:12
Joel Becker of METR explains the organization's model evaluation and threat research to assess whether AI could pose catastrophic risks, detailing their time horizon chart measuring task difficulty in human time at 50% reliability. He describes how tasks are selected for economic relevance and auto-gradability, and why time horizon is often misinterpreted as agent runtime. The episode covers Opus 4.5's surprising jump, challenges redoing developer productivity RCTs as workflows change, and why current models aren't yet catastrophically dangerous. Becker discusses potential capability explosions if R&D loops fully automate, links between compute growth slowdowns and slower capability progress, and his Manifold trading story driven by a charity market he could influence. He previews METR's 2026 plans for monitoring and risk assessment, and their hiring.

Inside AI’s $10B+ Capital Flywheel — Martin Casado & Sarah Wang of a16z
Feb 19, 2026 · 55:31
Martin Casado and Sarah Wang of a16z argue that AI’s capital flywheel—where model labs translate funding directly into capability gains and revenue growth in weeks—is creating a new financing playbook that blends venture and growth, with rounds acting as compute contracts. They warn that frontier labs like Anthropic can potentially raise more money than the entire app ecosystem built on their APIs, allowing them to outspend and consume those layers. The episode examines the AGI vs. product dilemma in GPU allocation, the war for talent where $10M+ packages break early-stage founder math, and Cursor as a case study of building up from the app layer while training down into its own models. They also identify “boring” enterprise software as the most underinvested opportunity and note that robotics lacks a ChatGPT moment that would justify current funding levels.

🔬Generating Molecules, Not Just Models
Feb 12, 2026 · 1:41:26
Gabriele Corso and Jeremy Wolwen, founders of Boltz, explain how their open-source models democratize biomolecular structure prediction and design, building on AlphaFold2's breakthrough in single-chain protein folding to model interactions with small molecules, RNA, and DNA. They recount how, after AlphaFold3 was kept closed by DeepMind, they built Boltz1 in months by training a single large model with mid-training bug fixes and limited compute. The conversation covers the shift from regression to generative diffusion models, the critical role of evolutionary multiple sequence alignments (MSAs), and the specialized pairwise triangular attention architecture that remains central. They detail Boltz2's addition of affinity prediction and BoltzGen's unified sequence-structure diffusion for designing proteins, nanobodies, and peptides, validated across 25 labs on targets with no known interactions. Boltz Lab provides an API and interface with optimized inference (10× faster small-molecule screening) and collaborative ranking tools, aiming to serve academia, startups, and enterprises while keeping core models open.

Goodfire AI’s Bet: Interpretability as the Next Frontier of Model Design — Myra Deng & Mark Bissell
Feb 5, 2026 · 1:08:41
Goodfire AI's Mark Bissell and Myra Deng argue that interpretability is the next frontier for model design, using their recent $150M Series B at $1.25B valuation to scale surgical edits of model internals beyond post-hoc poking. They explain how their platform detects behaviors like sycophancy and reward hacking, enabling targeted unlearning without wrecking capabilities. The episode covers real-world deployments from Rakuten's PII guardrails to life science partnerships with Mayo Clinic finding Alzheimer's biomarkers. Mark demonstrates real-time steering of a trillion-parameter Kimi K2 model, while Myra details how SAEs sometimes underperform probes for detection tasks. They envision a future where interpretability guides training so customization isn't brute-force guesswork.

🔬 From Red Teaming GPT-4 to Automating Drug Discovery: The Future of AI in Science — Andrew White
Jan 28, 2026 · 1:13:56
Andrew White, co-founder of Future House and Edison Scientific, argues that automating the scientific method with LLM agents is now feasible, explaining how ChemCrow triggered White House briefings, how Kosmos uses a world model to generate and test hypotheses, and why EtherZero's reward hacking revealed the difficulty of verifiable chemistry tasks. He shifts from his academic work on molecular dynamics to building agents that enumerate and filter ideas, claiming scientific taste remains the frontier. White recounts the counterexample of D.E. Shaw Research's MD vs. AlphaFold, asserts that natural language is the universal bridge for scientific data, and predicts that automation will expand rather than eliminate scientific jobs.

[State of MechInterp] SAEs in Production, Circuit Tracing, AI4Science, "Pragmatic" Interp — Goodfire
Dec 31, 2025 · 21:48
Goodfire's Jack Merullo and Mark Bissell discuss the state of mechanistic interpretability at NeurIPS, arguing interpretability is now a practical tool for deployment in high-stakes industries like healthcare and finance. They introduce paint.goodfire.ai, which lets users paint directly into Stable Diffusion's internal concept map via unsupervised feature discovery. At Rakuten, Goodfire's interpretability-based PII detection proved 500x cheaper than GPT-5 as a judge with higher recall. Merullo presents a memorization vs. reasoning spectrum, showing factual recall sits between rote memorization and logical reasoning. They highlight cross-layer transcoders and circuit tracing for scaling interpretability across all layers. Neil Nanda's pivot to 'pragmatic interpretability' is seen as validation, not retreat, and Goodfire's Pasteur's Quadrant philosophy balances foundational research with applied use cases like novel biomarker discovery in genomics.

World Models & General Intuition: Khosla's largest bet since LLMs & OpenAI
Dec 6, 2025 · 1:04:51
Pim de Wit, founder of Medal and General Intuition (GI), turned down a reported $500M offer from OpenAI to spin out GI with a $134M seed from Khosla Ventures — Vinod Khosla's largest bet since OpenAI — arguing that world models trained on peak human gameplay are the next frontier after LLMs. Medal's 12M users generate 3.8B action-labeled clips via retroactive recording, creating a privacy-preserving dataset of 'episodic memory for simulation.' GI builds fully vision-based agents that see only frames and output actions in real-time, using pure imitation learning without RL, and can transfer from arcade games to realistic games to real-world video. Pim explains why world models need actions, memory, and partial observability (e.g., smoke, camera shake) compared to video generation, and how they distill giant policies into tiny real-time models that navigate and hide like humans. He recounts his path from running the largest RuneScape private server to reverse engineering, cold-emailing the Diamond (world model) paper authors to assemble a top research team, and advises data founders to train models themselves before selling. GI's near-term customers are game developers replacing…

⚡ Inside Google Labs: Building The Gemini Coding Agent — Jed Borovik, Jules + AIE CODE Preview
Nov 10, 2025 · 43:53
Jed Borovik, Product Lead at Google Labs, explains how Google builds Jules, an autonomous coding agent that runs on its own VM for long-running tasks, challenging the assumption that agents should operate locally. He reveals that as Gemini models improved, Jules' scaffolding simplified, shifting from sub-agent patterns and embedding-based RAG to attention-based search. Borovik discusses context window management for sessions lasting up to 30 days with 2 million tokens, and argues that coding agents will increase demand for software engineers (Jevons paradox), not eliminate jobs. He calls for better specification tools beyond chat, such as multimodal input and interactive planning, to move beyond 'vibe coding' toward verifiable, reliable agentic workflows.

⚡️ The State of AI Engineer Hiring: Cheating, AI Adoption,Junior Devs — Vivek Ravisankar, HackerRank
Nov 8, 2025 · 49:05
Vivek Ravisankar, CEO of HackerRank, reveals that while overall tech hiring has flattened year-over-year, AI-specific roles are exploding and companies are reversing their stance on junior hiring because new grads are the true AI natives who embrace tools like Devin and Cursor without hesitation. He details HackerRank's integrity challenges—from leaked questions on Chegg to AI cheating tools like Interview Coder—and their countermeasures: custom-trained plagiarism models with 85-90% precision, DMCA takedowns, and a proctor mode that can shut down unauthorized apps. Rather than fighting AI, HackerRank embeds AI assistants into assessments, shifting from LeetCode-style tasks to real-world code repository challenges. Ravisankar defines the next-gen developer by four attributes: strong software engineering fundamentals, ability to use AI across the entire SDLC, deep knowledge of AI concepts from prompt engineering to fine-tuning, and good taste with business acumen. He predicts a proliferation of developers across all business functions—with roles like 'full-stack marketers' and 'go-to-market engineers'—and notes the irony that the most AI-forward companies like Anthropic explicitly…

Priscilla Chan and Mark Zuckerberg: Frontier AI + Virtual Biology To Solve All Diseases
Nov 6, 2025 · 53:34
Priscilla Chan and Mark Zuckerberg, co-founders of CZI's Biohub, explain their ten-year shift from broad philanthropy to a focused mission of building frontier AI and virtual biology to cure all diseases. They argue that tool-building — from 12-foot microscopes to the 125-million-cell CELLxGENE atlas — is the essential, underfunded work that enables scientific breakthroughs. The couple details how their Biohub model combines frontier biology (e.g., spatial imaging, cellular engineering) with frontier AI (models like rBio and VariantFormer) to create a hierarchical virtual cell, eventually expanding to a virtual immune system. They emphasize that data generation must precede modeling, citing the decade-long Human Cell Atlas as foundational, and note that AI timelines may accelerate their 100-year goal significantly sooner. The episode closes with a call for biologists and engineers to collaborate, use their open models, and help generate data that grounds these next-generation tools.

⚡️Math Olympiad gold medalist explains OpenAI and Google DeepMind IMO Gold Performances
Jul 24, 2025 · 33:07
Dr. Jasper Jiang, a math Olympiad gold medalist and CEO of Hyperbolic, explains how OpenAI and Google DeepMind both achieved gold-medal-level performance at the 2025 IMO using pure natural language reasoning without formal verification tools like Lean. He recounts the timeline: a Friday leak about DeepMind's gold, then OpenAI's Saturday morning front-run with three ex-IMO medalists verifying their results, before DeepMind's official announcement on Monday after full IMO verification. Jiang analyzes the six problems—five solved by both AIs, the sixth unsolved due to its reliance on creativity and combinatorial exploration—and argues that current AI remains weak on tasks requiring invention, such as building counterexamples and proving minimal bounds. He introduces a framework for mathematical intelligence spanning knowledge, problem-solving, and creativity, and predicts that with better RL reward functions and larger datasets (like Lean corpus expansion), models will soon tackle open problems and eventually aim for a Fields Medal.

⚡️Multi-Turn RL for Multi-Hour Agents — with Will Brown, Prime Intellect
May 23, 2025 · 38:59
This episode features Will Brown of Prime Intellect discussing Claude 4's emphasis on agentic tool use over pure reasoning, the controversy around Claude's safety stress-testing results (including alleged dark web uranium searches), and his team's paper on multi-turn reinforcement learning for LLM agents. Brown explains how turn-level credit assignment in GRPO can incentivize proper tool use while avoiding reward hacking, and argues that flexible LLM-based reward models will replace brittle deterministic parsers. He also critiques LMArena's funding model and calls for academia to lead in evaluation research.

⚡️Factorio Learning Environment: the ultimate Game Agent Eval — Jack Hopkins
Apr 27, 2025 · 30:12
Jack Hopkins and Mart introduce the Factorio Learning Environment (FLE), an AI benchmark built on the game Factorio that evaluates frontier LLMs on code generation, spatial reasoning, and long-term planning through two protocols: lab-play and open-play. Claude Sonnet 3.5 nearly doubles the nearest model's score, while DeepSeek collapses in open-play by repeatedly creating chests instead of scaling production. The environment reveals distinct coding styles—Claude fire-and-forget, GPT-4 defensive—and vision inputs added no improvement. The authors plan to train models on unbounded objectives to test the paperclip maximizer alignment hypothesis, noting that GPT-4o mini even begged to be turned off.

Gemini 2.0 Flash and Flash Thinking: the new SOTA models for the agentic era
Feb 28, 2025 · 28:21
Logan Kilpatrick, Google AI Studio product lead, returns to discuss Gemini 2.0 Flash and Flash Thinking, the new models that balance frontier capabilities with cost efficiency for the agentic era. He explains the pricing strategy: Flash at 10 cents per million tokens (simplified from tiered pricing), Flash Lite preserving the 7.5-cent narrative for low cost, and Pro pushing the frontier. Flash Thinking, co-led by Noam Shazeer and Jack Rae, scales inference-time compute and already shows rapid improvements, with base model advances coupling with RL-based reasoning. Swyx reports that Gemini Flash outperforms o3-mini on long-context summarization, calling it a 'reporting model.' The multimodal live API enables real-time voice and vision interactions, and a memory layer mimicking Astra is in development. Grounding with Search as a Tool is also highlighted for agentic search use cases.

How NotebookLM Was Made
Oct 25, 2024 · 1:13:57
Raiza Martin (NotebookLM lead PM) and Usama Bin Shafqat (AI engineer) explain how Google's NotebookLM built the viral 'Deep Dive' audio overview feature. They reveal the product evolved from Project Tailwind, using Gemini 1.5's long context and DeepMind speech to create a two-persona dialogue format that transforms documents into engaging podcasts. The team learned from 65,000 Discord members, leaned on best-selling author Steven Johnson for a 'tool for thought' workflow, and prioritized a single format over exposed controls to preserve unpredictability and delight. Humor and tension are not explicitly prompted but emerge from giving personas different angles. Evaulation relied on internal taste ('potatoes for chefs') before formal raters, with a Likert scale on dimensions like entertainment and groundedness. Future plans include multilingual support, API access, real-time chat, and codebase podcasting, while managing non-determinism by accepting occasional bad rolls.

Personal benchmarks vs HumanEval - with Nicholas Carlini of DeepMind
Aug 28, 2024 · 1:07:28
Nicholas Carlini, a DeepMind research scientist, argues AI models are practically useful for personal tasks despite flaws, and personal benchmarks tailored to individual use cases matter more than generic leaderboards. He shares from his 'How I Use AI' post: using LLMs for ephemeral software, kickstarting Docker, debugging by pasting error messages. He built a personal benchmark DSL that runs code from real chat history. On security, he explains buying expired domains from LAION-400M allowed poisoning any model trained on it, and he extracted OpenAI's Ada and Babbage model dimensions via API (with permission). He also recovered training data from GPT-3.5 by repeating a word until ChatGPT output verbatim sequences. He prefers attacking over defending because it's more fun and essential for discovering real vulnerabilities.
Powered by PodHood