Episodes from Latent Space about Biology.

🔬 RL with Verifiable Rewards, but the Verifier is a Lab — Lila Sciences
Jul 16, 2026 · 1:41:04
Andy Beam (CTO) and Rafa Gómez-Bombarelli (Co-founder & CSO of Physical Sciences) of Lila Sciences argue that science is an 'infinite token generator' for AI, using reinforcement learning with verifiable rewards where the wet lab acts as the verifier. They claim one general model trained on ~10 trillion experimentally-verified reasoning tokens across biology, chemistry, and materials outperforms domain-specific models—'breadth gives us depth.' Their AI Science Factories treat the lab as a data center, with instruments on a 'PCI bus' and humans 'below the API line.' Highlights include a CAR-T candidate designed in six months by two or three people, 'monster UTRs' achieving ~10x Moderna/Pfizer mRNA expression, and the 'zero-FTE startup' business model. They discuss RL pathologies like collapsed chains of thought and a model that 'swears,' and note why there is still no AlphaFold for materials due to the sim-to-real gap.

🔬 "The Most Innovative Diffusion Research Is Happening in Drug Discovery, Not Image Generation"
Jun 30, 2026 · 1:48:40
Evan Feinberg and Sergey Edunov of Genesis Molecular AI argue that diffusion models have unlocked sub-ångström accuracy in protein-ligand structure prediction, a breakthrough that makes AI useful for real drug discovery where the field’s favored 2Å RMSD benchmark is "slop." Their PEARL model uses diffusion with physics-based guidance, synthetic training data from molecular dynamics, and inference-time scaling to predict induced fit—how a protein flexes to accommodate a ligand. On the OpenBind benchmark, PEARL zero-shot surpassed all cofolding models on the notoriously hard EV A721A protease, correctly predicting a flexible loop movement that other methods missed. They also introduce SAPPHIRE, an agentic system that orchestrates AI models for 24/7 drug design, and discuss how downstream ADMET properties (solubility, toxicity, etc.) remain equally critical. The biggest bottleneck they face is GPU availability, and they are actively hiring AI researchers interested in novel architectures beyond standard transformers.

🔬 The Bitter Lesson is Coming for Proteins - Alex Rives, BioHub
May 27, 2026 · 1:10:12
Alex Rives, head of science at BioHub, argues that scaling language models on protein sequences—the 'bitter lesson' for biology—yields emergent biological understanding, culminating in the open-source ESMC model and ESMFold 2. ESMC, trained on 6.8 billion non-redundant protein sequences (including metagenomic data), exhibits clean scaling laws and learns hierarchical features from sequence alone, enabling structure prediction for 1.1 billion proteins. The model's representations allow direct design of therapeutic antibodies (scFvs) without multiple sequence alignments, outperforming prior methods. Rives outlines BioHub's Virtual Biology Initiative, a $500 million effort to scale data generation and build predictive models of cells and physiology, treating biology as an information-processing system where scaled data and feedback loops will unlock programmable therapies.

🔬 Training Transformers to solve 95% failure rate of Cancer Trials — Ron Alfa & Daniel Bear, Noetik
Apr 20, 2026 · 1:25:22
Noetik founders Ron Alfa and Daniel Bear argue that 95% of cancer drugs fail in clinical trials not because of pharmacology but because of poor patient selection; their thesis is that the right patients exist but aren't identified. To solve this, Noetik generates its own multimodal data—pathology H&E, spatial transcriptomics with up to 20,000 genes, and protein stains—from human tumor samples, building self-supervised foundation models like OctoVC and the autoregressive Tario. They claim a data moat: over a hundred million spatially resolved cells, an order of magnitude more than any public dataset, which drives better generalization across cancer types. The platform validates predictions via PerturbMap, an in-vivo mouse model with multiplexed CRISPR knockouts, and an ‘in silico humanization’ that reads mouse H&E in human gene space. A $50M deal with GSK licenses OctoVC for therapeutic discovery and fine-tuning on GSK's own data, marking a rare software-focused licensing deal in biotech. Noetik bets that patient-level tissue modeling, not subcellular simulation, will first deliver clinically actionable insights.

🔬Generating Molecules, Not Just Models
Feb 12, 2026 · 1:41:26
Gabriele Corso and Jeremy Wolwen, founders of Boltz, explain how their open-source models democratize biomolecular structure prediction and design, building on AlphaFold2's breakthrough in single-chain protein folding to model interactions with small molecules, RNA, and DNA. They recount how, after AlphaFold3 was kept closed by DeepMind, they built Boltz1 in months by training a single large model with mid-training bug fixes and limited compute. The conversation covers the shift from regression to generative diffusion models, the critical role of evolutionary multiple sequence alignments (MSAs), and the specialized pairwise triangular attention architecture that remains central. They detail Boltz2's addition of affinity prediction and BoltzGen's unified sequence-structure diffusion for designing proteins, nanobodies, and peptides, validated across 25 labs on targets with no known interactions. Boltz Lab provides an API and interface with optimized inference (10× faster small-molecule screening) and collaborative ranking tools, aiming to serve academia, startups, and enterprises while keeping core models open.

🔬 From Red Teaming GPT-4 to Automating Drug Discovery: The Future of AI in Science — Andrew White
Jan 28, 2026 · 1:13:56
Andrew White, co-founder of Future House and Edison Scientific, argues that automating the scientific method with LLM agents is now feasible, explaining how ChemCrow triggered White House briefings, how Kosmos uses a world model to generate and test hypotheses, and why EtherZero's reward hacking revealed the difficulty of verifiable chemistry tasks. He shifts from his academic work on molecular dynamics to building agents that enumerate and filter ideas, claiming scientific taste remains the frontier. White recounts the counterexample of D.E. Shaw Research's MD vs. AlphaFold, asserts that natural language is the universal bridge for scientific data, and predicts that automation will expand rather than eliminate scientific jobs.

Priscilla Chan and Mark Zuckerberg: Frontier AI + Virtual Biology To Solve All Diseases
Nov 6, 2025 · 53:34
Priscilla Chan and Mark Zuckerberg, co-founders of CZI's Biohub, explain their ten-year shift from broad philanthropy to a focused mission of building frontier AI and virtual biology to cure all diseases. They argue that tool-building — from 12-foot microscopes to the 125-million-cell CELLxGENE atlas — is the essential, underfunded work that enables scientific breakthroughs. The couple details how their Biohub model combines frontier biology (e.g., spatial imaging, cellular engineering) with frontier AI (models like rBio and VariantFormer) to create a hierarchical virtual cell, eventually expanding to a virtual immune system. They emphasize that data generation must precede modeling, citing the decade-long Human Cell Atlas as foundational, and note that AI timelines may accelerate their 100-year goal significantly sooner. The episode closes with a call for biologists and engineers to collaborate, use their open models, and help generate data that grounds these next-generation tools.
Powered by PodHood