Andy Beam (CTO) and Rafa Gómez-Bombarelli (Co-founder & CSO of Physical Sciences) of Lila Sciences argue that science is an 'infinite token generator' for AI, using reinforcement learning with verifiable rewards where the wet lab acts as the verifier. They claim one general model trained on ~10 trillion experimentally-verified reasoning tokens across biology, chemistry, and materials outperforms domain-specific models—'breadth gives us depth.' Their AI Science Factories treat the lab as a data center, with instruments on a 'PCI bus' and humans 'below the API line.' Highlights include a CAR-T candidate designed in six months by two or three people, 'monster UTRs' achieving ~10x Moderna/Pfizer mRNA expression, and the 'zero-FTE startup' business model. They discuss RL pathologies like collapsed chains of thought and a model that 'swears,' and note why there is still no AlphaFold for materials due to the sim-to-real gap.
Andy Beam (CTO) and Rafa Gómez-Bombarelli (Co-founder & CSO of Physical Sciences) of Lila Sciences join us to talk about building scientific superintelligence. Andy makes the case that the internet is a spent resource ("we have but one internet. It's the fossil fuel. We fracked"), and that the next internet-scale dataset comes from running the scientific method as reinforcement learning, with the wet lab as verifier. Science becomes an "infinite token generator" — the lab isn't the product, the model is. The counterintuitive result: one general model trained on ~10 trillion experimentally-verified reasoning tokens across biology, chemistry, and materials beats the domain-specific ones — "breadth gives us depth." Great blog post from Escalante Bio referenced in the episode: "Your Experiment has a Runtime" (https://blog.escalante.bio/your-experiment-has-a-runtime/) Highlights: - The lab as data center: instruments on "a PCI bus," humans "below the API line" - A CAR-T candidate designed in six months by two or three people - "Monster UTRs" hitting ~10x Moderna/Pfizer mRNA expression - The "zero-FTE startup" business model - "You can't have scientific superintelligence if you're just a good test taker" - Rafa's "bittersweet lesson": "only the things that you can scale matter" - Why there's still no AlphaFold for materials - RL pathologies: collapsed chains of thought, a model that "swears" - A vision-language model driving a Windows 95 instrument - "The world's largest collection of voided warranties in biology" Links: Andy Beam: https://www.linkedin.com/in/andrew-beam-01a6aa295/ Andy Beam (Lila): https://www.lila.ai/team/andrew-beam Rafa Gómez-Bombarelli: https://www.linkedin.com/in/rgbombarelli/ Rafa Gómez-Bombarelli (Lila): https://www.lila.ai/team/rafael-gomez-bombarelli Lila Sciences: https://www.lila.ai/ Lila Sciences (LinkedIn): https://www.linkedin.com/company/lila-sciences Chapters: 0:00 "We have but one internet" 0:46 Intro & guest backgrounds 5:36 The thesis: the bitter lesson & the infinite token generator 10:01 Inside the AI Science Factory: the "PCI bus" & the API line 14:34 Safety, security & scientific rigor 24:39 RL, reward hacking & chain-of-thought pathologies 28:16 Why Lila isn't a biotech: the model is the product 32:36 10 trillion tokens & why the general model wins 35:25 Not just TechBio: materials, quantum dots & MOFs 41:42 Scaling & the "bittersweet lesson" of materials 44:12 The in-vivo CAR-T proof point 49:13 The "zero-FTE startup" model 52:56 Clinical translation & loading the die 59:40 Ken Stanley & open-endedness 1:01:07 Lab video walkthrough & the lab as a data center 1:07:07 Orchestration, scaling & faster assays 1:14:54 Instrument onboarding & the 10T-token dataset 1:24:22 Lila & the Flagship ecosystem 1:31:33 What's harder: materials or biology? 1:35:53 Bottlenecks, MFU & closing thoughts
