A company discussed on Latent Space.

⚡️Mercury: Ultra-Fast Diffusion LLMs — Estefano Ermon, CEO Inception Labs
Aug 4, 2025 · 28:07
Stefano Ermon of Inception Labs explains why diffusion language models like their Mercury Coder can match GPT-4.1 Nano and Claude Haiku in intelligence while delivering 5–10× faster inference at 737–1109 tokens/sec on H100s, built on score-based generative model foundations he pioneered at Stanford. He details the coarse-to-fine generation process that modifies multiple tokens per network evaluation, the required end-to-end training from scratch (no fine-tuning existing LLMs), and how post-training pipelines use specialized DPO for preference alignment. Ermon identifies latency-sensitive applications—voice agents, IDEs, vibe coding—as the killer use case, acknowledges they aren't yet frontier-level but sees a future where diffusion dominates due to efficiency gains, and notes the challenge of open-sourcing given proprietary inference engines and kernels.

The State of AI Startups in 2024 [LS Live @ NeurIPS]
Dec 21, 2024 · 26:35
Sarah Guo and Pranav Reddy argue that 2024 has become a far friendlier ecosystem for AI startups, with the model landscape shifting from OpenAI's near-monopoly to a competitive field where Google's Gemini now leads LMSys Arena and open-source models like Llama 8B score ten points higher on MMLU than Mistral 7B a year ago. They note that OpenAI's API cost has dropped 80-85% in 18 months, and total OpenAI API share fell from ~90% to ~60% as customers switch. The funding environment is rational, not a bubble, with foundation-model labs raising $30-40B but most startups seeing sane valuations; one portfolio company grew from zero to twenty million in PLG-style spending. Key startup themes include first-wave service automation (Sierra, Decagon, Harvey, EvenUp), better search and new friends (Perplexity, Glean, Character, Replica), and democratized creativity (Midjourney, HeyGen). They argue the 'GPT wrapper' narrative is false—applications capture value—and that incumbents face innovator's dilemma because AI changes business models (outcomes-based pricing) and data needs (reasoning traces rarely saved). The speed of change and new markets (legal, healthcare, defense) structurally favor…

This World Does Not Exist — Joscha Bach, Karan Malhotra, Rob Haisfield (WorldSim, WebSim, Liquid AI)
Apr 27, 2024 · 1:56:11
The episode features Karan Malhotra demoing WorldSim, a prompt that turns Claude 3 into a universe simulator via CLI; Rob Haisfield presenting WebSim, which generates functional websites on the fly; and Joscha Bach arguing simulative AI reveals consciousness as a virtual property, with LLMs creating agents as real as human minds, urging the California Institute for Machine Consciousness to build self-organizing silicon life. Malhotra shows WorldSim running “world.exe” to create Twitter inside a simulated universe, with users tweeting and Elon Musk moving Dogecoin. Haisfield demonstrates WebSim generating a 5D particle interface, a news RSS aggregator, and a face-swap webcam app via URL parameters like “secrets=revealed”. Bach explains consciousness as second-order perception in a simulated now, criticizes RLHF for lobotomizing models, and advocates for animist AI where software agents compete like spirits, not golems.

Making Transformers Sing - with Mikey Shulman of Suno
Mar 14, 2024 · 58:58
In this episode of Latent Space, hosts Alessio and Shawn Wang interview Mikey Shulman, CEO of Suno, about making transformers sing. Suno uses transformers to predict audio tokens end-to-end, avoiding baked-in musical knowledge, with a tokenization secret sauce that also includes non-music audio for better vocal realism. Their models are relatively small (far below 175B parameters) due to latency needs, and they prioritize scaling research over brute-force size. Over half of Suno users employ expert mode, tweaking lyrics and style prompts, rather than easy mode. Shulman argues Suno is not the 'Midjourney of music' because music is inherently social and synchronous, unlike images. He demoed live generation, showing control via tokens like [beat drop] and style modifiers, and revealed future plans for collaborative concerts, continuous DJ modes, and personalized models. He also advocates for hiring economists to avoid Goodhart's law pitfalls in ML benchmarks, especially in audio where aesthetics matter most.

The Four Wars of the AI Stack - Dec 2023 Recap
Jan 26, 2024 · 1:20:58
Swyx and Alessio recap December 2023 by framing the AI landscape as four wars: the Data War (NYT lawsuit demanding destruction of GPTs, OpenAI's data partnerships, synthetic data from DeepMind); the GPU/Inference War (Mixtral price dropping 90% to $0.27/million tokens, benchmark drama between Anyscale and Together, alternative architectures like Mamba); the Multimodality War (Midjourney reaching $200M ARR, ElevenLabs hitting unicorn status, OpenAI and Google building god models); and the RAG/Ops War (LangChain vs LlamaIndex, Pinecone's $750M valuation, Qdrant serving OpenAI and Anthropic internally). They argue agents and open source are not active wars, and predict 2024 will see AI finally reach production, with hardware like Rabbit R1 and Tab capturing unique context.

The AI-First Graphics Editor - with Suhail Doshi of Playground AI
Jan 2, 2024 · 1:09:23
Suhail Doshi, co-founder of Mixpanel and founder of Playground AI, argues that image generation is still in a GPT-2 moment and that training open-source foundation models from scratch is necessary to unlock real utility, exemplified by Playground v2's 2.5x preference over Stable Diffusion XL on a 1K prompt benchmark. He explains the pivot from Mighty (cloud-streamed browser) to AI, driven by the belief that shifting compute elsewhere aligns with AI's parallel computation. The episode covers Playground's unique UI—not just a prompt box but a full canvas with preview rendering, seed control, and style filters—and the difficulty of balancing safety with artistic expression, especially around NSFW content. Suhail discusses the under-investment in graphics AI compared to language, the open-source community's role, and his choice to release pre-trained weights for academic research. He also shares lessons from running GPU infrastructure (harder for training than inference) and advises founders to follow curiosity and build projects rather than relying on books.
Powered by PodHood