# ⚡️Math Olympiad gold medalist explains OpenAI and Google DeepMind IMO Gold Performances

Latent Space · 2025-07-24

<https://addtry.com/8892bad7-5f95-4ee4-9a27-61fe106641cd>

Dr. Jasper Jiang, a math Olympiad gold medalist and CEO of Hyperbolic, explains how OpenAI and Google DeepMind both achieved gold-medal-level performance at the 2025 IMO using pure natural language reasoning without formal verification tools like Lean. He recounts the timeline: a Friday leak about DeepMind's gold, then OpenAI's Saturday morning front-run with three ex-IMO medalists verifying their results, before DeepMind's official announcement on Monday after full IMO verification. Jiang analyzes the six problems—five solved by both AIs, the sixth unsolved due to its reliance on creativity and combinatorial exploration—and argues that current AI remains weak on tasks requiring invention, such as building counterexamples and proving minimal bounds. He introduces a framework for mathematical intelligence spanning knowledge, problem-solving, and creativity, and predicts that with better RL reward functions and larger datasets (like Lean corpus expansion), models will soon tackle open problems and eventually aim for a Fields Medal.

## Questions this episode answers

### How did OpenAI and Google DeepMind's verification of their IMO gold medal performances differ?

According to Jasper Jiang, OpenAI did not involve the IMO officially; they used the competition problems and had three former IMO medalists review the model's solutions. In contrast, Google DeepMind waited for official IMO verification, which required extra time. This led to OpenAI releasing results first, while DeepMind's were later confirmed as fully verified by the IMO.

[3:20](https://addtry.com/8892bad7-5f95-4ee4-9a27-61fe106641cd?t=200000)

### Why couldn't the AI models solve the last combinatorics problem in the IMO?

Jasper Jiang, a former IMO gold medalist, explains that the final combinatorics problem requires creativity—generating examples and proving they represent minimal or maximal bounds—which AI currently lacks. While the first five problems can be decomposed into logical steps, the last demands experimental exploration and creative insight, making it unsolved by any AI.

[14:01](https://addtry.com/8892bad7-5f95-4ee4-9a27-61fe106641cd?t=841000)

### What new math benchmark is Jasper Jiang developing, and what skills does it aim to measure?

Jasper Jiang is developing a holistic math benchmark with PhDs and AI researchers, breaking down mathematical skills into three categories: mathematical knowledge and understanding (definitions, theorems, calculations); problem-solving and communication (strategies, logic, presentation); and learning meta-skills and creativity (abstract thinking, transfer learning, intuition, and inventing new methods). He believes current benchmarks like IMO test only narrow aspects of math.

[16:09](https://addtry.com/8892bad7-5f95-4ee4-9a27-61fe106641cd?t=969000)

## Key moments

- **[0:00] Intro**
- **[1:18] IMO Gold**
  - [1:49] Jasper Zhang reveals DeepMind's IMO gold on X; OpenAI announces its own unverified gold result hours later, stealing the spotlight.
- **[4:16] Natural Language**
  - [4:16] IMO 2025: AI models reason in natural language without formal verification, a leap from 2024's Lean-dependent silver medal.
- **[7:34] Taste & Skills**
  - [9:21] "It kind of surprised me that AI can come up with this idea" — Jasper Zhang on AI reducing IMO problem 1 to any N to 3.
  - [10:14] Q: Does IMO require research-level creativity? A: IMO problems are brain teasers that test problem-solving, not conjectures or advanced math, says Jasper Zhang.
  - [11:37] It's harder to make China's IMO national team than to win a gold medal at IMO, says Jasper Zhang.
- **[12:26] Problem Analysis**
  - [13:18] Combinatorics problems like the final IMO problem require exploring examples and creative leaps, an area where current AI struggles.
- **[15:09] New Benchmark**
  - [15:10] Q: How can AI solve creative math? A: RL with auto-evaluation and mathematician-designed reward functions will be key, says Jasper Zhang.
  - [16:15] Jasper Zhang is developing a new math benchmark that rates AI on knowledge, problem-solving, and creativity, going beyond competition-style problems.
  - [19:23] For logical reasoning, using Lean as a verifier for RL training will quickly produce powerful math reasoning models.
- **[19:53] Creativity & RL**
  - [20:39] Progress in creative math AI requires more mathematicians collaborating with AI researchers to design reward functions for creativity.
- **[23:24] Lean vs Language**
  - [23:24] Q: Should AI math rely on formal proofs like Lean? A: Natural language is easier for data, Lean offers trustworthiness; both approaches will be needed, says Jasper Zhang.
  - [26:03] Mathematician Kevin Buzzard is formalizing the proof of Fermat's Last Theorem in Lean, aiming to complete it in two to three years.
- **[27:08] Breakthroughs**
  - [27:30] OpenAI and DeepMind's IMO success likely came from scaling parallel thinking and multi-step RL, not a single algorithmic breakthrough.
  - [29:41] DeepMind submitted a second Gemini model without in-context learning that also achieved IMO gold, confirming genuine reasoning.
- **[31:20] Future Milestones**
  - [31:34] Jasper Zhang predicts Math AGI will be achieved when an AI solves a long-standing open problem and wins a Fields Medal.
- **[32:36] Outro**

## Speakers

- **Alessio** (host)
- **Swyx** (host)
- **Jasper Jiang** (guest)

## Topics

Reasoning, Benchmarks

## Mentioned

DeepMind (company), Hyperbolic (company), OpenAI (company), FrontierMath (product), Gemini (product), Lean (product)

## Transcript

### Intro

**Alessio** [0:03]
Hey, everyone. Welcome back to another Latent Space Lightning Pod. This is Alessio, partner and CTO at Decibel, and I'm joined by Swyx, founder of smol.ai.

**Swyx** [0:12]
Hi. Hi. Hi. We are very happy to have-- finally have Dr. Jasper Jiang of Hyperbolic, uh, on the pod. Welcome.

**Jasper Jiang** [0:19]
Thanks for inviting me. Great to be here.

**Swyx** [0:22]
Yeah. Um, the story of Hyperbolic, uh, you know, I feel like we, we will do a, a full podcast on Hyperbolic at, at, uh, at another time, but maybe like a quick two seconds on, uh, what's Hyperbolic, and then we can go into the IMO stuff.

**Jasper Jiang** [0:35]
Yeah, sounds good. Yeah. So I'm, I'm the CEO and co-founder of Hyperbolic, and, um, I faced like a lot of pain when I trying to ac-get access to GPUs, uh, during my PhD time and also my serial time.

So I decided I want to make sure that any AI team who want to build or use AI in the future, they won't go-- they won't get constrained by compute. And so, uh, we decided to build Hyperbolic, which is the on-demand AI cloud made for developers, so everyone can get fast, affordable access to compute and inference.

**Swyx** [1:06]
Yeah. Um, by the time that we, uh, release this episode, uh, we should be publishing, uh, Jasper's talk at AIE as well, so you can go over to his talk and, and see more about, uh, the, the story there.

Uh, but we're actually coming here today to talk about something slightly different, uh, something that you also have, uh, experience in, which is Math Olympiads. Um, and, uh, maybe, yeah, yeah, uh, you know, the, the, the sort of headline, you know, not to bury the lead is that, you know, um, both DeepMind and OpenAI have claimed to have, um, gold le-- gold medal level performance at the IMO.

### IMO Gold

**Swyx** [1:39]
Uh, and then you started being like super involved in like sort of breaking down, uh, what the story is. Maybe can you sort of recap what happened, and then we can go from there?

**Jasper Jiang** [1:49]
Sounds good. Yeah. So, uh, on Friday-- So I, I, I, like, participate in different Math Olympiads in the past, won a few gold medals, and, uh, I have a lot of friends who are actively involved in IMO. And so I was just chatting with my friends on Friday, um, and around like three p.m., uh, one friend told me that, "Wow, uh, China team won the IMO again, uh, uh, and, and also, uh, DeepMind also won the gold medal."

And, uh, so I just started posting on X. I mean, I try to let everyone knows the news. So uh, I posted on X, uh, some people get excited, but then never confirmed by DeepMind. And so it's a little bit, a little bit weird because I thought like DeepMind was, was the-- would be the first one to kind of announce that news.

Uh, and then, uh, at, at one a.m. Saturday, OpenAI starting, um, starting posting-- A lot of OpenAI people starting posting about like them winning the IMO, uh, IMO gold medal and kind of like stole the, the spotlight. Um, and so I initially thought Google was just slow due to their marketing approval, like the process.

So I posted another one to just like share my thoughts, uh, and then I got pinged from Google, like someone, someone from Google just told me that it's actually like IMO and Google themselves need extra time to verify.

Um, and so like it kind of raised a question to me, like why? Why DeepMind need to verify, but then, uh, OpenAI didn't, didn't need to do that, they can just release? And then it turns out like OpenAI, uh, actually didn't involve officially with IMO.

They just like used the, the problems, but and then, um, just like used their model to test the, the results and asked like three previous, uh, IMO medalists to review them. Um, I think I, I read, I read the pr-proof, the, uh, it's still correct, uh, but it's just like less official and that's why, uh, people kind of like kind of-- OpenAI re-re-receive a few backlash, uh, over the weekend and on Monday DeepMind, uh, officially confirmed, uh, they have the gold medal, uh, and, uh, also fully verified by IMO, so.

**Swyx** [4:16]
Yeah. And, uh, I mean, so congrats. I, I, I think, you know, like the relationship between OpenAI and IMO and all that, I, I... honestly, it's, uh, that's between them. Uh, the-- what matters is that they have a model, both, both labs independently achieved a model that reasons in natural language, uh, without internet, without tools, uh, which I think, you know, this time last year it was-- we didn't think it was possible.

### Natural Language

**Jasper Jiang** [4:40]
Yeah. So, uh, I have a feeling that this year, um, there, there at least like an AI model can win, um, can win the IMO gold medal because, uh, last year, uh, by using like formal verification, formal language, uh, Google DeepMind can-- uh, already won a silver medal.

And, um, so maybe I should share some background about IMO. So usually, uh, IMO is basically the, uh, the top-tier competition, uh, for high school students, right? Uh, it doesn't require you to really understand like the advanced math.

It's just basically it's problems about, uh, simple knowledge like combinatorics, uh, number theory, geometry, and algebra. Uh, but then it kind of require you to have very strong problem-solving skills. You need to know how to do a contradiction, you need to know how to create like counterexamples, you need to do like observation, et cetera.

And so, uh, people treat this as like the, uh, one of the, the bar for AI to hit, uh, before they become the real mathematicians in the future. Um, so last year Google already won the silver medal and, uh, seem-seems to me that they already figured out a way to, to, to make it work.

Uh, it just like need a little more time, a little more resources, a little more data, a little more compute. And so, uh, naturally, I, I would say, uh, Google will win gold medal. What surprised me is, is this time they don't use formal language But instead they just use LLM.

And so last year when they tried to do, uh, the IMO, they need like a, uh, like a human to kind of translate the natural language problems to Lean, and then they use Lean to kind of prove the, the, the problems, uh, to prove the, uh, solve the problems.

And, and this year, uh, no longer need, need that. And the same for OpenAI too.

**Swyx** [6:39]
Yeah. I, I, I think that this, uh, uh, dropping the, the need for Lean and, you know, other, other sort of specialized systems like, um, you know, I, I was, um, covering this last year, uh, with the, uh, the, the AlphaProof and AlphaGeometry team.

Um, and you know, it, it seemed like a very complicated system. Some of these solutions took 60 hours, which is way longer than the 4.5 hours that, uh, IMO, um, uh, limits, uh, have for humans. Um-

**Jasper Jiang** [7:06]
Um.

**Swyx** [7:06]
And so yeah, these, these, uh, some, some systems took up to three days. And, uh, you know, I think that's, that's a huge step. You know, it's, it's, it's, uh... I, I don't even think it's like, um, incremental, like, "Oh, scale this system more."

They, they just have to throw away the whole thing and do something else. Um, and then I think also, like I think people can also just, um, look up the problems. Um, you know, I think y-yeah, what's, what's cool about you is that you can actually analyze and, and study the problems.

### Taste & Skills

**Jasper Jiang** [7:34]
Problems.

**Swyx** [7:34]
I think that there's some different ratings or different pass rates of different problems. I think probably either problem two or problem three was like the, quote-unquote, "easiest." And, uh, you know, it's like, it's like just a few paragraphs.

It's not, not, not a, not a lot, but like it, it involves a lot of creative thinking, right? Or what, what are they testing basically?

**Jasper Jiang** [7:54]
So, um, I would say, uh, different problems testing different skills. For example, like the, the, uh, if you look at the problem one, we can start from-

**Swyx** [8:05]
Yeah. Yeah

**Jasper Jiang** [8:05]
... problem one is like a combinatorial geometry. So it, it kind of, uh, involves like integer... You need to kind of, uh, it, it kind of asks you a question about the all the possibilities for different cases, like for different N, uh, for the any, any positive integer that are greater than or greater or equal to three.

Uh, and so it require you to find a key lemma, which is like, uh, for any N greater than three, like let's say N, N equal to five, you can reduce the problem to N equal to four, and then reduce the problem to N equal to three.

**Swyx** [8:48]
Mm-hmm. Mm-hmm.

**Jasper Jiang** [8:48]
And so this is, uh... There are different ways to do this like induction. Um, but, uh, like Google has one method, and then, uh, OpenAI, uh, AI also have another pro-- uh, method. Uh, but both kind of works.

And basically you just reduced any N to three, and then you just do case-by-case analysis for N equal to three. And so I'll say, um, this is probably not that challenging if you are very good, uh, like Math Olympiad pr- uh, participants.

But, uh, it kind of surprised me that AI can, can kind of come up with this, uh, like ideas too.

**Swyx** [9:30]
Yeah. It's-

**Jasper Jiang** [9:30]
Yeah

**Swyx** [9:31]
... it's crazy.

**Jasper Jiang** [9:31]
It's crazy.

**Swyx** [9:31]
Um-

**Alessio** [9:32]
And, uh-

**Swyx** [9:32]
Yeah, go ahead.

**Alessio** [9:33]
How do you think about rating, like the taste? I actually used to do, uh, Math Olympics too, but I got third. I never made it to IMO-

**Swyx** [9:41]
What?

**Alessio** [9:41]
... for Italy. Um, uh, and, and so it's hard for me today to judge. But i-is it-- How important is taste in your mind where, hey, you can actually take then some of these results and maybe build on top of it, versus sometimes it's kind of competition for competition sake?

You know what I mean? It's kind of like with code, it's like, "Hey, look, you just want it to compile." Like is taste created a lot in the competition, and is it useful, or do you think we're better off having models that can just brute force some of these things and we can focus on more interesting stuff?

**Jasper Jiang** [10:14]
Yeah. So, uh, so I think, um, what we bas-- I, I basically think like... I basically look at my, my journey from, from like high school, like high school to PhD, right? And I, I realized that IMO kind of like it's a good exam to test certain aspect of math skills, which is like problem-solving or you need to have like really, uh, complete chain of thoughts, uh, like logic, et cetera.

Uh, but it doesn't test like your creativity in terms of like research. It doesn't test your, uh, your, your capabilities of doing conjectures. Uh, it doesn't test your, uh, capabilities of learning like advanced knowledge or like applying advanced knowledge.

And so I think, uh, we probably should just treat this as like one step stone for AI to become the real mathematicians in the future. Um, and I would say most of the problems here are more like brain teaser style.

So like you can solve them, uh, but it doesn't mean that you are a good mathematician. Um, and, uh, and, and yeah, I, I mean, I mean, I can, I can feel the AGI vibe from the solutions, but I think we're still, we're still not there.

**Swyx** [11:33]
Uh, yeah. I mean, uh, you, you did the CMO, right?

**Jasper Jiang** [11:37]
Yeah.

**Swyx** [11:37]
Is that, uh-

**Jasper Jiang** [11:38]
That's it

**Swyx** [11:38]
... your, your, your background? Um, how does it compare, by the way, CMO versus IMO?

**Jasper Jiang** [11:42]
Uh, so, uh, I mean, CMO is Chinese Math Olympiad. And so, uh, in terms of difficulty, uh, CMO is, uh, kind of similar to IMO. Um, I would say it's super competitive in China. So there is a saying, like if you, if you, if you get...

It's harder to get selected for the national-- to become the national team instead of- ... uh, than like getting, getting the gold medal at IMO. And so, so it's similar. Um, and I, I mean, like there, there are many like, uh- It's, it's all about probability.

Like, there are so many people can compete at CMO, and so, like, a lot of students are pretty good, but just, like, six people, um, were selected.

### Problem Analysis

**Swyx** [12:26]
Yeah. Yeah. Um, I mean, that, that's, uh, that's also very common in terms of, uh, any eli-elite institution. It's harder to get in than it is to, uh, come out of it. Um, okay. Uh, any other, like, any other overviews of the IMO problems?

Uh, I, I saw some-- I saw a minor piece of discussion around how the, the five problems that both AI solved were of the same kind or, or the sixth problem was, like, a clear discontinuity. Um-

**Jasper Jiang** [12:54]
Hmm.

**Swyx** [12:54]
Which I think is normal, right? Like, IMO always has a sixth problem that is, like, challenging. Um, this is part of, like, the Eliezer Yudkowsky bet. Um, but, like, I just wanted to, you know, you, you, you, you prepped-- you, you were kind enough to prep some show notes which, which we're gonna publish.

Um, but you, you actually had some analysis of, like, the, the, the problems and maybe you wanna sort of talk about, uh, what you thought about the problems.

**Jasper Jiang** [13:18]
Yeah. So I'll say, um, so in-- for IMO, there are, like, four, four categories, right? Algebra, combinatorics, geometry, number theory. Um, usually combinatorics, uh, is kind of more-- the most challenging and creative, uh, especially for AI. So, like, at one time, I, I was at, like, this conference, which is a Frontiermath Symposium.

So basically, they invited, like, uh, thirty to forty mathematicians to kind of assess the capabilities of, uh, of OpenAI and many other advanced AIs. Uh, we realized that if, if it's like, uh, textbook stuff or like, uh, step-by-step, uh, problem, then, then it's-- uh, AI can solve it.

But however, if it require, like, creativity, uh, especially like in combinatorics, you kind of need to create, uh, some example and then try to prove them, uh, that it's the minimal bound or like the upper bound. Uh, it's the lower bound or the upper bound, then AI kind of sucks at it.

And so I think, uh, this is why no AI can solve the last problem. Uh, but it also is super hard for human as well. Uh- And, and like, if you look at the first five problems, it-- you can basically, when you look at a problem, you already know how to decompose the, uh, how to decompose a problem into several steps, and then y-you know that if you can solve each step, then you can solve the whole problem.

But then for the last problem, you kind of need to do a lot of, uh, experiments or, like, you need to look at the sm-smaller examples and try to create some example-- try to create some solutions for that small examples and try to figure out if you can prove it, uh, that is the, it's the minimal bound, minimal number.

Yeah.

### New Benchmark

**Swyx** [15:09]
How-

**Jasper Jiang** [15:10]
Yeah.

**Swyx** [15:10]
How do you think, um, we can get the models to do that? Do you have any intuition on how to get there or AGI? Uh, do, do, do you have an idea of, like, how can we get the models to solve those problems?

Is it just a matter of better RL? Is it something fundamental about what data they need to get? Um, yeah.

**Jasper Jiang** [15:29]
Yeah. So are you talking about the combinatorics problem or like the real research problems?

**Swyx** [15:35]
No, like these problems that you're saying, uh, the format... Uh, uh, you're basically saying, look, if the type of problem is this type, the models are not good. Do you know why that is or how, uh, we're gonna get them to solve them?

**Jasper Jiang** [15:49]
Yeah. So, uh, this kind, kind of come to, uh, one thing that I'm just, like, doing as a hobby project, which is, uh, like, uh, we, uh, we think we need a better benchmark for math, right? Like, uh, most of the benchmark that people, uh, are using right now is like AMM, IMO, USAMO, all these, like, competition math problems.

But, like, they're-- they just test like a certain aspect of other math skills that a mathematician have. And so what we're doing here is, like, we really want to think from first principles, like, what makes a human a mathematician, and then we break down all the skills that are required for a human to become mathematicians, and then we kind of create problems or, like, benchmark based on those categories.

And so we kind of, uh, right now it's like still, uh, in progress, but we're working with, like, several PhDs, professors, and, like, uh, AI researchers in different AI labs together to kind of, uh, to, to produce like this more holistic, uh, benchmark.

Uh, but basically, we kind of organize the skills into three category. So one is mathematical knowledge and understanding. And so you need to know all the math knowledge, uh, like definition, theorems, lemmas, et cetera, and then you also need to understand theories, right?

Like how to apply them, why, why they're true, what are the history of those theory, and you also need to understand how to do calculations. You know, how to do like analytical skills, et cetera. And then the second is, like, problem-solving and, uh, communication.

So you need to kind of have a holistic problem-solving framework. Basically, you want to understand the problems, you need to ex-- you need, need to know how to explore examples, and you need to know how to apply key strategies like contradiction, decomposition, and analogy, et cetera.

And the same, uh, and then also you need to have logical thinking, and you need to be very good at writing and presentation. Then I would say for IMO, uh, problem-solving and communication probably is the, the most skills that you require.

Uh, and then lastly is, like, learning meta skills and, uh, creativity. So you need to know how to learn new knowledge. You need to know how to, uh, you need to how, how to think abstractly and then transfer tools across domains.

For example, like you want to use probabilistic method in combinatorics, then you kind of know how to do this transfer learning stuff. And, uh, and in mathematical modeling, all right, you need to how to Translate real world scenarios into formal mathematical models.

Uh, and then generalization, which is you look at a few examples and you bas-- you immediately come up with an idea. Okay, I can generate this to a more powerful theorem and see if I can prove that. Uh, and, and one, one, one very important piece of a mathematician is like creativity.

So you need to be very, very creative. Like when you look at, look at, uh, one object or one example, then you basically say, "Okay, now I think, uh, I can invent some new methods, and you can invent some new theory."

Uh, and lastly is like intuition. So you look at the, for example, the geometry. You look at the geometry problems, and you kind of already know, okay, these, these two line are parallel, or these two line are perpendicular.

Or, uh, even just look at algebra problems and, and you know, like, okay, this, this group has this structure. And so those are like the main pieces, uh, I think will define a mathematician. Uh, and then in terms of how we are-- we can get there, um, I think it really depends on different skills.

So, for example, uh, for the logical thinking and reasoning, I think we can... Uh, a lot of people are trying to use Lean, uh, because like RL works super well for verified ML. And so if we can just use Lean to kind of become the verifier, then, uh, it's, it's easy.

So I think, uh, we probably will see very powerful reasoning models, uh, in, in ter-- in math, uh, very soon. Uh, but then, and then for creativity, this is a little bit hard because-

**Swyx** [19:53]
Mm. Mm.

### Creativity & RL

**Jasper Jiang** [19:54]
Uh, there is no measurement to-

**Swyx** [19:57]
Increase temperature.

Increase temperature.

**Jasper Jiang** [19:59]
Yeah. This decrease temp... No, we want, we want it to be more, uh, creative.

**Swyx** [20:04]
Increase, increase. Increase.

**Jasper Jiang** [20:05]
Increase. Uh, but then you also need to, uh, have the, the taste, like, right? You need to know, like, which direction is right. Like, this is a common, uh, problem in RL. And so I would say, uh, you need a new way, like you need to define a new reward function to kind of capture how good this model is, uh, in terms of creativity and, uh, and just like you need times-- you need like specific domain knowledge to kind of create this reward.

Uh, and I, I, I think I started seeing more and more mathematicians working together with AI scientists, researchers, and I think this is a good combo because they're, they're... If you are just like AI researchers, you actually don't know what makes a mathematician good, then there's no way that you can come up with a reward function well.

Uh, but then if you're just a mathematician, you maybe know the, the concept, but you don't know how to apply that to, to AI research. And so, um, so basically I think we'll see-- we will start seeing more and more mathematicians working together with AI scientists.

And, uh, my, my guess, I mean, different people have different opinions, and then my guess is I believe in RL. So, uh, if for each category, we can figure out the-- a way to auto evaluate the results, then after that it will just be compute and data.

**Swyx** [21:33]
Yeah. Uh, I just wanna, uh, offer a few links for listeners and viewers. Uh, this is the symposium that you went to. It was two months ago, and these are the people involved. Uh, and the, this is obviously leaning on the FrontierMath work that, uh, was done by Epoch and OpenAI.

Um, and this is some of the, some of the example about the, the kinds of problems that they're working on, which is like not even English to me. Um, uh, but like, uh, yeah, I mean, uh, it's... I think it's very nice how you broke down the different levels of reasoning.

I feel like this is very cool in a sense of, um, th-yes, this is particularly applied to math, but it does generalize. It does generalize to like, this is how we... These are, these are like the higher level skills of intelligence other than predicting the next token, um, that, that, you know, that, that people want.

And, and, uh, s- so you, you have a framework of like, you know, knowledge and then, uh, going all the way down to communication, which is actually really cool to me. Um, I'm gonna think about that a lot more.

**Jasper Jiang** [22:33]
Thanks. Yeah. Uh, stay tuned. We probably will publish the first draft, first version in, in a few weeks.

**Swyx** [22:39]
This is different work from, from the FrontierMath people?

**Jasper Jiang** [22:42]
Yeah.

**Swyx** [22:43]
Okay. Okay. Interesting.

**Jasper Jiang** [22:44]
It's, uh... So FrontierMath is still problem focused, but then I, I think we need, uh, more skills, uh, breakdown. Um-

**Swyx** [22:51]
Yeah.

**Jasper Jiang** [22:52]
Yeah.

**Swyx** [22:52]
And is this like-- Is, is this related to Hyperbolic at all, or is it like a personal project?

**Jasper Jiang** [22:56]
So, uh, I mean, Hyperbolic, the goal of Hyperbolic- ... is that we want to democratize AI innovation and try to accelerate AI evolution. And so, um, I als- I would position this as a hobby project, but-

**Swyx** [23:09]
Okay. Okay

**Jasper Jiang** [23:10]
... uh, we, we will probably give people some credits if they actually can solve- ... some-- We, we'll, we'll give some researchers some GPU credits if they can actually solve, uh, some problems in the benchmark.

**Swyx** [23:23]
Yeah.

**Jasper Jiang** [23:24]
Yeah.

### Lean vs Language

**Swyx** [23:24]
Uh, and then I think the, the last follow-up I had was about a comment that you made, uh, maybe like three comments into your, your comments, uh, which is, uh, about, about the usage of Lean and v- uh, and like sort of verifiability.

Um, I, I saw some mathematicians actually say... Like, so, okay, on the one hand, if you're very bitter lesson pilled, and if you're very pro scaling AI and all that, it's a win to reason purely in natural language, purely within your weights.

No-

**Jasper Jiang** [23:53]
No.

**Swyx** [23:53]
No tools, no, no nothing. Um, you know, the OpenAI output was very raw. It could have been fine-tuned for human reliability, but like it was still, it was still fine. Like, it still got the right answer. GDM, uh, was a bit more polished.

Um, but ultimately, it, um, the, the, the answers were hard to verify. You had to have... You know, you had to hire like three former IMO gold medalists to go and read it, and then say like, "Yeah, this looks right," instead of like, "No, they didn't get lucky."

Or, or like, or like, "They just got lucky." Right. And I think like the-- I saw some arguments from mathematicians that are saying like, "Actually using Lean is good because then we can trust that that is, uh, good reasoning."

You know, do you have any perspectives on that stuff?

**Jasper Jiang** [24:34]
Yeah, I think, um, the different approach have different challenges. Like, uh, natural language is easier to gather data, but it's harder to verify. Um, but it, it's more human readable, right? Uh, for Lean, um, there, there are not enough data for Lean, right?

There are like probably one million tokens, uh, in, uh... about Lean right now. And, and I know a lot of like professors and, uh, PhD, uh, researchers are trying to build the biggest, uh, Lean dataset in the future.

Uh, we-- if we reach, can reach like one trillion token, I think then you can start doing some cool AI training on top of that. Um, and but, uh, but I think that's like a, a more... Personally, I feel like it's probably a more clear approach.

I was super surprised when I saw OpenAI and Google DeepMind didn't, didn't even use Lean and can solve that. And as, as many might know, like, uh, AI neural networks is basically a black box, and it's like a experimental science.

So I, I can't make any prediction whether we will start s- we'll first see like a mathematician AI built on top of Lean or built on top of natural language. But, um, but generally, I feel like both approaches might-- should work, and I, I think, uh, mathematicians, uh, will like both of them.

Um, and besides that, uh, uh, anecdote, I know, uh, some professors at, uh... in, in London are trying to using, um, using Lean to write the proof for Fermat's last theorem. And so that's, that's pretty, uh, impressive if you can, you can do it.

Yeah.

**Swyx** [26:21]
Do you think, do you think it's possible? Uh, I, I mean, I, I knew that, uh, Fermat's last theorems has been proven, but-

**Jasper Jiang** [26:28]
Yes.

**Swyx** [26:29]
Uh, also, I, I don't know if Lean is, you know, complete enough to, to get to that level.

**Jasper Jiang** [26:36]
So I would say, uh, it's, uh, I think definitely possible. Um-

**Swyx** [26:40]
Yeah.

**Jasper Jiang** [26:40]
Like he, he... So the professor is Kevin Buzzard, and he got like a grant, and now he's just like focused on writing the proof for Lean. Like he's-

**Swyx** [26:50]
Wow.

**Jasper Jiang** [26:51]
He's hoping to finish that in like two or three years. Uh, and, uh, and then basically like if Lean is not complete, then he can just build more-

**Swyx** [26:59]
Yeah. Yeah.

**Jasper Jiang** [27:00]
... more libraries on top of that. So, so everything is doable. It's just time. And I think he got time, and he got, uh, resources, so I'm very excited to see that happen.

### Breakthroughs

**Swyx** [27:08]
Okay, cool. Well, looking forward to the, the skills breakdown. Um, we'll, you know, share it w-whenever it comes out. Um, any other, uh, thoughts on, on just this IMO chapter of, of, uh, I, I guess, the twenty twenty-five edition of the IMO?

Um, you know, any speculations on like what breakthroughs either lab has done?

**Jasper Jiang** [27:30]
Can you, can you say that again? What do you mean by breakthrough?

**Swyx** [27:32]
Any, any... Do you want to speculate-- Uh, is there, is there any credible speculation on what has been the breakthrough that-

**Jasper Jiang** [27:38]
That-

**Swyx** [27:39]
... um, that either OpenAI or GDM have done in order to achieve this? Like, um, you know, they, they... all they say is like, "Oh, we scaled test time scaling more." But like I don't think that's enough.

**Jasper Jiang** [27:50]
You don't think that's enough? You don't think like, uh... I, I... To be honest, I feel like a lot of, uh, AI, uh, like, a, a lot of AI models built in industry is just like a bigger scale of, uh, of like, uh, res- in, uh, like research results.

They, um, they... probably they will use like existing research, but it's just more, more like larger scale, more systematic, um, and more data. Um, I mean, I know, um, I know, uh, uh, when I read like the Google DeepMind report, they said they can, they can now do like parallel thinking.

Basically, the, the model can, um, can like parallelly try different new ideas. Uh, they try to like leverage more multi-step reasoning. Uh, uh, sorry. They, they try to like figure out new reinforcement learning techniques that can leverage multi-step reasoning.

I think that's probably like a, a very, uh, important, important step. Uh, I know, uh, AL is pretty bad at multi-step. Uh, but now seems like they, they figured out a way to kind of, uh, train, uh, using AL for multi-step reasoning, and I think that's like, uh, pretty impressive in terms of the, the solution.

Uh, and, uh... But then I, I would say mo- none of this will probably worth like the best paper in, at NeurIPS, but it's more about like you just be the best in terms of each technique and just get more data, get more compute.

**Swyx** [29:32]
Yeah, a lot of, a lot of, uh, people are focusing on this sentence where they say, "We provided Gemini with access to a curated corpus of solutions to math problems." And they were like, "Oh, okay, like you, you trained on tests."

Um.

**Jasper Jiang** [29:42]
It's not.

**Swyx** [29:43]
Uh, but no. But, uh, I think a smaller covered story that people may have missed is that there was a second Gemini model that was also submitted to IMO, uh, with no, no pre, uh, in-context learning, and they also did...

got the same result. So, uh, yeah, that's, that's a, that's a big win.

**Jasper Jiang** [29:58]
Wow. Um.

**Swyx** [29:59]
Yeah, I think it's reasonably, uh, known that Deep Think and o3 pro, um, and, uh, Grok 4 Heavy are all like parallel thinking, right? Like they, they all just run like ten, uh, different var-variations of the, the, the same model, and then they have a, uh, a judge model choose the best result.

Um, I... no, yeah, I'm not sure. Like I, I, I think the, the existence proof that there's been something qualitatively different is the, is the wall clock time going from sixty hours to, to, you know, under four point five.

**Jasper Jiang** [30:29]
Four point five.

**Swyx** [30:29]
That's, that's an order of magnitude change, and it is not that-- typically, that is not achieved by scaling up, um, existing ideas. You have to just do new ideas. Because it's, it's a, it's a efficiency, uh, issue.

**Jasper Jiang** [30:42]
Can we... Like, I mean, GPU getting better, now you get B200- ... it's faster. Uh, like even-

**Swyx** [30:50]
Is it one order of magnitude better? I don't think so. It's like two times better.

**Jasper Jiang** [30:52]
Uh, I think 2X, and then inference, inference speed also improve a lot, right? Like, uh, if you look at if you want to run like Llama 3 and, and like now compared to now it's like s- 3X, 2X better.

So, so kind of-

**Swyx** [31:09]
Yeah.

**Jasper Jiang** [31:10]
I-

**Swyx** [31:10]
Yeah.

**Jasper Jiang** [31:10]
Yeah.

**Swyx** [31:11]
Maybe. Maybe.

**Jasper Jiang** [31:12]
Maybe.

**Swyx** [31:13]
Maybe.

**Jasper Jiang** [31:13]
But I definitely agree. It's, uh, it's also like, um, the, the model is more powerful, and you can solve the problems faster too.

### Future Milestones

**Swyx** [31:20]
Uh, cool. Any, uh... What, what's next? Uh, what, what other, like IMO, uh, what other math like, uh... You know, you have like Frontier Math, you have, uh, your benchmark that you're putting out. Uh, any other like milestones that you're looking for in, in, in this part of the world?

**Jasper Jiang** [31:34]
Yes. Uh, so I would say, uh, a lot of benchmark is just like, it's just like exams, right? Like you take... The purpose of becoming a mathematician is not to take exams. It can give you like, uh... It's, it's the same for any, any, any job too.

It kind of give you a, like a benchmark or like a, uh, like a metrics to see how good these model are. But you really need to put them into the battlefield, right? And, uh, the goal of a mathematician is you want to push forward the frontier of math.

And so I would say the, the next step is trying to solve some open problems that's like open for 30 years, 50 years. And then, uh, I think the ultimate goal is to solve like a very challenging problem, like create a new theory and then win a Fields Medal.

That's how we , how I define a Math AGI in the, in the future. Yeah.

**Swyx** [32:30]
Fields Medal.

**Jasper Jiang** [32:31]
Fields Medal.

**Swyx** [32:31]
Okay, let's go.

**Jasper Jiang** [32:32]
Yeah. That's my dream. That's my dream.

**Swyx** [32:36]
Uh, cool. Um, Alessio, any, uh, parting thoughts, questions?

### Outro

**Alessio** [32:39]
No, that was it. Maybe I should go and try and back and do the problems this afternoon and see if I get-

**Jasper Jiang** [32:45]
Let's, uh, let's solve some, uh, IMO problems together.

**Alessio** [32:48]
Yeah. Awesome, man. Thanks for joining us.

**Jasper Jiang** [32:52]
Yeah. Thanks for inviting me.

---

This library is powered by PodHood (https://podhood.com), the podcast website platform.
