A company discussed on Latent Space.

Artificial Analysis: The Independent LLM Analysis House — with George Cameron and Micah Hill-Smith
Jan 9, 2026 · 1:18:15
Artificial Analysis founders George Cameron and Micah Hill-Smith explain how their independent benchmarking platform became the gold standard by running their own evals with a mystery shopper policy to prevent labs from manipulating results. They launched in January 2024 after building it as a side project in Sydney, going viral after Swyx's retweet. The Intelligence Index V3 synthesizes 10 datasets with 95% confidence intervals, while the Omniscience Index measures hallucination rates from -100 to +100 (Claude models lead). Their GDP Val AA benchmark tests 44 white-collar tasks, and they open-sourced their agentic harness Stirrup. They also introduced an Openness Index scoring models out of 18 points. The episode covers how they make money through enterprise benchmarking subscriptions and custom work, and why the cost of GPT-4-level intelligence has dropped over 100× while total inference spend rises due to reasoning and agentic workflows.

[State of Evals] LMArena's $1.7B Vision — Anastasios Angelopoulos, LMArena
Dec 31, 2025 · 24:02
Anastasios Angelopoulos, founder of LMArena (now Arena), discusses the platform's $100M raise at a $1.7B valuation, its spin-out from Berkeley incubation by a16z's Anjney Midha, and its mission to be the industry's north star for real-world AI evaluation. He defends against the 'leaderboard illusion' paper, citing factual errors and reaffirming that models cannot pay to be on or off the public leaderboard. Arena funds inference costs for millions of monthly users, with 25% of its 5M+ users in software, and is expanding into occupational verticals (medicine, legal, creative) and multimodal video arenas. Key challenges include consumer retention, which improved with sign-in and persistent history, and moving off Gradio to React for better development. Angelopoulos calls for top talent in ML, product, and go-to-market to join the high-performance team.

The Great Evals Debate — Ankur Goyal & Malte Ubl
Dec 7, 2025 · 34:33
Ankur Goyal (Braintrust) and Malte Ubl (Vercel) debate whether offline evals are essential infrastructure or premature optimization for AI coding agents, arguing that the best teams deliberately invest in multiple feedback loops—offline evals, A/B tests, and vibe checks—to build effective AI products. They explain that modern evals are not about manufacturing golden datasets but pulling real user failures from production logs into eval suites to iterate faster. The conversation highlights how evals provide a 'first derivative' that enables aggressive shipping without regression fears, akin to unit tests in traditional software. Coding evals are uniquely verifiable (e.g., 'does it compile?') yet underutilized; Vercel uses them in RL pipelines to fine-tune models that fix trivial errors 100x faster than agentic loops. They also discuss how product managers encode domain expertise through rubrics and LLM-as-judge scoring, and why proprietary evals are competitive moats while public benchmarks serve marketing. The debate concludes that RL environments are a promising frontier for computer-use agents but require specialized expertise to avoid reward hacking.

⚡️Raising $1.1b to build the fastest LLM Chips on Earth — Andrew Feldman, Cerebras
Oct 1, 2025 · 29:14
Andrew Feldman, CEO of Cerebras, joins Latent Space to discuss their $1.1B fundraise at an $8.1B valuation and their wafer-scale chip that delivers 20x faster inference than NVIDIA's B200 GPUs. He explains how their architecture uses SRAM instead of HBM, providing 2,625x more memory bandwidth by eliminating the narrow straw between compute and memory. Feldman details the decision to accelerate sparse linear algebra rather than specialized convolutions, enabling support for transformers and diffusion models unseen during design. He discusses the explosive growth in AI inference demand, the shift from closed-source to fast open-source models, and the importance of speed—citing Paul Graham's observation that ChatGPT's slowness drives users away. The conversation covers enterprise trends in the 10-30B parameter space, the complexity of building data centers that pull gigawatts of power, and the often-overlooked routing and caching systems that make AI work seamlessly.

⚡️Mercury: Ultra-Fast Diffusion LLMs — Estefano Ermon, CEO Inception Labs
Aug 4, 2025 · 28:07
Stefano Ermon of Inception Labs explains why diffusion language models like their Mercury Coder can match GPT-4.1 Nano and Claude Haiku in intelligence while delivering 5–10× faster inference at 737–1109 tokens/sec on H100s, built on score-based generative model foundations he pioneered at Stanford. He details the coarse-to-fine generation process that modifies multiple tokens per network evaluation, the required end-to-end training from scratch (no fine-tuning existing LLMs), and how post-training pipelines use specialized DPO for preference alignment. Ermon identifies latency-sensitive applications—voice agents, IDEs, vibe coding—as the killer use case, acknowledges they aren't yet frontier-level but sees a future where diffusion dominates due to efficiency gains, and notes the challenge of open-sourcing given proprietary inference engines and kernels.
Powered by PodHood