Host of Latent Space.

🔬 RL with Verifiable Rewards, but the Verifier is a Lab — Lila Sciences
Jul 16, 2026 · 1:41:04
Andy Beam (CTO) and Rafa Gómez-Bombarelli (Co-founder & CSO of Physical Sciences) of Lila Sciences argue that science is an 'infinite token generator' for AI, using reinforcement learning with verifiable rewards where the wet lab acts as the verifier. They claim one general model trained on ~10 trillion experimentally-verified reasoning tokens across biology, chemistry, and materials outperforms domain-specific models—'breadth gives us depth.' Their AI Science Factories treat the lab as a data center, with instruments on a 'PCI bus' and humans 'below the API line.' Highlights include a CAR-T candidate designed in six months by two or three people, 'monster UTRs' achieving ~10x Moderna/Pfizer mRNA expression, and the 'zero-FTE startup' business model. They discuss RL pathologies like collapsed chains of thought and a model that 'swears,' and note why there is still no AlphaFold for materials due to the sim-to-real gap.

🔬 "The Most Innovative Diffusion Research Is Happening in Drug Discovery, Not Image Generation"
Jun 30, 2026 · 1:48:40
Evan Feinberg and Sergey Edunov of Genesis Molecular AI argue that diffusion models have unlocked sub-ångström accuracy in protein-ligand structure prediction, a breakthrough that makes AI useful for real drug discovery where the field’s favored 2Å RMSD benchmark is "slop." Their PEARL model uses diffusion with physics-based guidance, synthetic training data from molecular dynamics, and inference-time scaling to predict induced fit—how a protein flexes to accommodate a ligand. On the OpenBind benchmark, PEARL zero-shot surpassed all cofolding models on the notoriously hard EV A721A protease, correctly predicting a flexible loop movement that other methods missed. They also introduce SAPPHIRE, an agentic system that orchestrates AI models for 24/7 drug design, and discuss how downstream ADMET properties (solubility, toxicity, etc.) remain equally critical. The biggest bottleneck they face is GPU availability, and they are actively hiring AI researchers interested in novel architectures beyond standard transformers.

🔬 The Limits of AI in Science - Why We Need Self-Driving Labs — Joseph Krause, Radical AI
Jun 17, 2026 · 1:16:50
Joseph Krause of Radical AI argues that the bottleneck in materials science is experiments, not ideas, and his company's self-driving lab combines AI hypothesis generation with automated synthesis and characterization to produce alloys at unprecedented speed—1,200 in six months, with 300 novel compositions and 10 already in commercial development. Radical's closed-loop system runs research campaigns, not just automated tasks, overcoming challenges like sample manipulation at 3,000°C and tool vendors' software access. Krause details how their AI explores elemental families humans overlooked, why they open-source models like Matrix (the moat is experimental data, not models), and how they plan to compress discovery timelines from decades to 3–5 years for defense and space applications. He addresses the 10-year qualification process for aerospace, supply chain geopolitics (e.g., hafnium price up 10–15x due to Chinese dominance), and the need for public-private partnerships to accelerate U.S. R&D. Finally, he urges ML engineers to lean into their expertise rather than try to become material scientists.

Scaling Past Informal AI - Carina Hong, Axiom Math
Jun 3, 2026 · 1:33:04
Carina Hong, founder and CEO of Axiom Math, argues that formal verification, not informal RL, is the path to superintelligence, following her company's $200M Series A at a $1.6B valuation and a perfect 120/120 on the 2024 Putnam exam. Axiom's system uses Lean theorem prover data and reinforcement learning to produce verified proofs, achieving a 99% pass rate on the Verina code-with-proof benchmark (187 of 189 problems). Hong contends verification is about 'scaling brilliance' — not fixing hallucinations — and that only verified generation can compound AI reasoning. She explains Axiom's open-source Axle API for Lean at scale, addresses why frontier labs like OpenAI have deprioritized formal math (team departures, strategy shifts), and outlines a vision where verified reasoning transfers from math to code, hardware, and eventually AGI through self-improvement. She also discusses the Earth sciences challenge of autoformalization, the difficulty of search in mathematical literature (citing the Erdos controversy), and why fragmentation in the AI math field is a bottleneck.

🔬 The Bitter Lesson is Coming for Proteins - Alex Rives, BioHub
May 27, 2026 · 1:10:12
Alex Rives, head of science at BioHub, argues that scaling language models on protein sequences—the 'bitter lesson' for biology—yields emergent biological understanding, culminating in the open-source ESMC model and ESMFold 2. ESMC, trained on 6.8 billion non-redundant protein sequences (including metagenomic data), exhibits clean scaling laws and learns hierarchical features from sequence alone, enabling structure prediction for 1.1 billion proteins. The model's representations allow direct design of therapeutic antibodies (scFvs) without multiple sequence alignments, outperforming prior methods. Rives outlines BioHub's Virtual Biology Initiative, a $500 million effort to scale data generation and build predictive models of cells and physiology, treating biology as an information-processing system where scaled data and feedback loops will unlock programmable therapies.

FPV Drones -The Next War Is Already Here — Yaroslav Azhnyuk, The Fourth Law & Noah Smith, Noahpinion
May 18, 2026 · 1:59:29
Yaroslav Azhnyuk (The Fourth Law) and Noah Smith explain how FPV drones now cause 70-80% of frontline casualties, dethroning artillery as the 'god of war.' Azhnyuk's firms produce thermal cameras, autonomy modules, and interceptors; level-one autonomy (terminal guidance) raised one brigade's mission success from 20% to 71% and extended the kill zone from 3 to 10 km. Fiber optic drones resist EW but cost $32/km and limit payload; AI autonomy removes the radio horizon problem. He warns China could build 4 billion FPV drones versus Ukraine's 4 million, and that the West lacks mass manufacturing, rare earth refining, and autonomy tech. Azhnyuk calls for a shift from costly platforms to cheap, software-defined drones, and for learning from Ukraine's battlefield experience.

🔬How GPT‑5 derived new results in theoretical physics and quantum gravity — Alex Lupsasca, OpenAI
May 5, 2026 · 1:31:51
Alex Lupsasca, a theoretical physicist at OpenAI and recipient of the 2024 New Horizons Breakthrough Prize, details how GPT-5 and subsequent models derived new results in quantum field theory and quantum gravity, solving problems that had stumped experts for over a year. The work focused on 'single minus' gluon tree amplitudes, long believed to be zero, but which humans discovered might be non-zero in a special kinematic region. GPT-5.2 Pro conjectured a simplified formula for these amplitudes, and an internal OpenAI model later proved it, reducing a factorial number of Feynman diagram terms to a linear number. The AI then autonomously extended the result to graviton amplitudes using the gluon paper as a seed, producing a complete paper draft in under an hour. Lupsasca argues this marks a threshold where AI is superhuman on certain physics tasks, accelerating research by acting as a 'scout' that reduces confusion and suggests next questions. He also discusses challenges including AI slop on arXiv and the need for better verification methods.

🔬 Training Transformers to solve 95% failure rate of Cancer Trials — Ron Alfa & Daniel Bear, Noetik
Apr 20, 2026 · 1:25:22
Noetik founders Ron Alfa and Daniel Bear argue that 95% of cancer drugs fail in clinical trials not because of pharmacology but because of poor patient selection; their thesis is that the right patients exist but aren't identified. To solve this, Noetik generates its own multimodal data—pathology H&E, spatial transcriptomics with up to 20,000 genes, and protein stains—from human tumor samples, building self-supervised foundation models like OctoVC and the autoregressive Tario. They claim a data moat: over a hundred million spatially resolved cells, an order of magnitude more than any public dataset, which drives better generalization across cancer types. The platform validates predictions via PerturbMap, an in-vivo mouse model with multiplexed CRISPR knockouts, and an ‘in silico humanization’ that reads mouse H&E in human gene space. A $50M deal with GSK licenses OctoVC for therapeutic discovery and fine-tuning on GSK's own data, marking a rare software-focused licensing deal in biotech. Noetik bets that patient-level tissue modeling, not subcellular simulation, will first deliver clinically actionable insights.

Dylan Patel Explains the AI War While Cooking | In-Context Cooking
Feb 26, 2026 · 55:13
Dylan Patel, CEO of SemiAnalysis, argues hyperscalers like Google, Amazon, and Meta will sacrifice all profits to build AI infrastructure, spending $180–$200 billion in capex this year alone, because the AI adoption explosion—Claude Code driving 4% of GitHub commits in one month, Anthropic adding $2.5 billion monthly revenue—makes it a Pascal's wager: spend or die. He details how Taiwan's semiconductor geopolitics create endgame scenarios, from a KMT win placating China to full invasion, with TSMC's output critical. Patel explains Nvidia's paranoid founder Jensen Huang is responding to vertical integration threats from hyperscalers by diversifying into chips like CPX and Groq, but warns moats are shallow. The real bottleneck in AI progress? Semiconductors themselves: fabs take years to build, and no one can buy enough GPUs through 2028. He also predicts a massive AI backlash from the public and financial markets, as capital consumption outpaces revenue and labor displacement accelerates.

🔬 From Red Teaming GPT-4 to Automating Drug Discovery: The Future of AI in Science — Andrew White
Jan 28, 2026 · 1:13:56
Andrew White, co-founder of Future House and Edison Scientific, argues that automating the scientific method with LLM agents is now feasible, explaining how ChemCrow triggered White House briefings, how Kosmos uses a world model to generate and test hypotheses, and why EtherZero's reward hacking revealed the difficulty of verifiable chemistry tasks. He shifts from his academic work on molecular dynamics to building agents that enumerate and filter ideas, claiming scientific taste remains the frontier. White recounts the counterexample of D.E. Shaw Research's MD vs. AlphaFold, asserts that natural language is the universal bridge for scientific data, and predicts that automation will expand rather than eliminate scientific jobs.
Powered by PodHood