LALatent SpaceFeb 26, 2025· 17:43

S1: the $6 DeepSeek R1 Competitor (ft. Entropix)

Tim Kellogg joins hosts Alessio and Swyx to break down his viral blog post on S1, showing how Stanford researchers cloned DeepSeek R1's reasoning for $6 by inserting a 'think token' and fine-tuning on just 1,000 examples. Kellogg argues the real innovation isn't the cost but the technique's simplicity—making advanced inference-time scaling accessible. He explains Entropix, a dynamic sampler that uses model entropy and varentropy to adjust generation on the fly, and notes its creators are starting a company. The conversation contrasts supervised fine-tuning's targeted results with RL's broader reasoning growth, using an analogy: 'RL causes a lot of growth, then SFT trims it.' Kellogg predicts that embedding introspection signals into RL training could solve agent 'doom loops' and unlock higher agency—a key challenge behind products like OpenAI's Deep Research.

Transcript

Intro0:00

Alessio0:03

Hey, everyone. Welcome back to another Lightning Pod. This is Alessio, partner and CTO at Decibel, and I'm joined by Swyx, founder of Osmo AI.

Swyx0:12

Hello. Uh, DeepSeek is all the rage. It was the biggest news story in the world. Uh, DeepSeek was the number one downloaded app in the US App Store in January, which is crazy. Uh, and we wind up-- We're looking for a guest, and Tim has been blogging about R1, S1, um, and, and all the other stuff.

So welcome, Tim.

Tim Kellogg0:33

Thanks.

Swyx0:34

Uh, Tim, what do you do?

Tim Kellogg0:36

Uh, I am a software engineer. Uh, I, I've been the CEO of, uh, Dentropy, which is, you know, two-person startup. We're winding that down and kind of transitioning into a different, different role.

Swyx0:48

More consulting and, and, uh, and all that. Um, so what got you interested in... You know, you blogged about, uh, like an R1 explainer, uh, you know, pretty early on and, you know. How, how have you been sort of exploring DeepSeek?

Tim Kellogg1:02

Uh, yeah. So I, I think I've just been, uh, kind of listening quietly on, on Twitter, and I hear, I hear little rumbles of different papers. And so every now and then, something strikes me as like, "Oh, this is really important," but people aren't, you know, talking loudly enough about it.

Um, so the R1 thing, um, that one came the same way. I just heard something. People were talking about it. Didn't seem like enough people were talking about it, so I wrote this up, wrote it really plainly, and, uh, I think it resonated with a lot of people.

And the same thing with S1.

Swyx1:34

Yeah.

Tim Kellogg1:34

People were, like, basically vague posting about it on, on Twitter. And I was like, "No, no, this is a pretty big deal, I think."

Swyx1:43

Yeah. It, it seemed like a... You know, so there's a meta question about... Look, I'm looking at your recent posts, right? You didn't really wor- You know, weren't really like, you know... Y-y-you, you explain, you explain stuff, and you sort of learn in, in public.

Um, but like, you know, h- what is your blogging strategy? Like, shou-should I do this? Like, we, we've debated having a DeepSeek blog. Alessio actually even wrote one. But I was like, "I don't think I'm, we're saying anything new here."

I don't know, like...

Tim Kellogg2:08

Um, so I think-- I don't know if I have, like, a overarching strategy aside from the ones that do well are the ones that, um, I know people want ahead of time.

Swyx2:19

Yeah.

Tim Kellogg2:19

Um, so some of the ones like the AI engineering primer, um, that one ca- I was, I was mentoring somebody, and he was asking a ton of questions. I'm like, "Man, I should just write this down." Um, so that one, that one did okay.

Um, R1, uh, I was posting on it about, about it on, on Blue Sky, and then it was just doing really well. I was like, "I'll just, I'll just consolidate this into a blog."

Swyx2:44

Yeah.

Tim Kellogg2:45

Uh, and then S1 the same way. I think, I think if I test out the topic, um, see if it's people are actually interested in it and then, uh, consolidate pretty much.

S12:45

Swyx2:55

Okay, cool. Um, should we double-click more on, on R1, or j- or do we, should we go into S1? I'll, I'll give a fork in the road here for, for the two of you.

Tim Kellogg3:03

Uh, S, S1. Let's go to S1.

Swyx3:04

Let's go to S1. All right. Um, so S1 came out more recently, and, um, it, we cloned DeepSeek for $6 is, is, uh, what the, the headline is.

Tim Kellogg3:13

Yeah. You need something to grab people's attention, you know? Well, it is important. It is important. The, the cost reduction is important. I just don't-- I'm not sure it's really, like, the, the major thing here, but it is, it is definitely something that is eye-catching, for sure.

Swyx3:30

Um, yeah. So, so, you know, I've, I've, like, vaguely, you know, I've read through the paper. I think the, the main claim, uh... There, there's sort of two claims, which I think, I think, uh, you, you, you explained, you explained a little bit.

Uh, main claim is that, um, DeepSeek had some inference time scaling, but they didn't, they didn't exactly replicate the test time scaling chart that, um, that o1 had. And, uh, also that you could basically replicate it with by inserting a weight token, uh, or a think token.

Um-

Tim Kellogg4:01

Mm-hmm

Swyx4:01

... and, and that was a big part of it. The other part was that it was very sort of-- Uh, there was a little bit of fine-tuning with, like, a thousand examples, basically. That's why it's $6. Like, otherwise you don't really need $6.

Um, yeah. That, that's my TLDR. I don't know if you would disagree or modify that in any way.

Tim Kellogg4:17

Yeah. No, absolutely. Yeah. I, I was, I was just thinking about the R1 paper, or actually the S1-- uh, the o1 paper. So many letters and numbers.

Um, uh, yeah, like they had that, that, uh, that graph, the graph, uh, on a log scale, the, uh, accuracy versus the, the compute that they put into it. And I'm thinking about it and like, you could just put like, you know, har- like harder questions are gonna deserve more, more thought naturally.

But that's kind of a wonky way to graph things out, 'cause if you, if you-- it's all, if that's all you did, then what you-- the graph would be saying is that the harder questions you get, the more accurate they are.

And that's, that seems a little weird. So they're, they're clearly doing something else, right? That's where, where-- When I read S1, I was like, "Oh, that's how they're doing it," right?

Swyx5:08

Yeah. Yeah. Um, uh, I'm, I'm also kind of curious about the distillation that bas- So basically, a lot of people are saying that, uh, DeepSeek, um, you know, distilled or st- or stole even, if you wanna use very politically charged words, uh, from OpenAI.

Um, it seems like you can distill from, um, you know, like just the outputs, which is interesting, because you don't need like the process rewards. Um, but the... I mean, here they had, they also had like really interesting, um, I guess the-- some kind of fil-filtration process, right?

Or it's like, yeah, data frugality, you said.

Tim Kellogg5:43

Yeah.

Swyx5:43

Yeah.

Tim Kellogg5:43

Yeah. I mean, it's kinda crazy that if, if you, if you know what you're looking for, um, ahead of time, you only need 1,000 examples to get like that, that kind of boost. Uh, which is enough, like one person could do that if they're, if they're, you know, focused.

And I, I was thinking about this, um- Like, that's, that's one of the advantages that OpenAI, OpenAI has, 'cause the three criteria they're looking for, so data quality, which is just basic, you know, the, the text is complete.

Um, they're also looking for diversity, um, like different topics, and then they're also looking for, like, different how, how hard a question is. Um, when-- so when they're sitting there listening to your ChatGPT conversations, you're automatically giving them diversity because you're interested in lots of different things, and you're also...

They can just look at how long, um, the, the LLM takes to respond to figure out the, the difficulty. You're, you're giving them basically O four , O five, O six just by, just by using their, their web app, which is kind of crazy.

Swyx6:49

Yeah. Like, uh, I mean, I'm also trying to wonder, like, what is it-- what, what goes from here, right? Like, is, um, do we... Like, is, is the, is the, is the goal to just copy this for every open model in the world?

Um, how, you know, how, how should, how should people think about, like, the implications of S1?

Tim Kellogg7:10

Um, I mean, I think copying is interesting 'cause then you, you get, um, smaller and smaller models, but it's not very interesting. It's just, it's more of a, a curiosity. I think one reason why I think this, this particular blog did well is because it, it, it seemed easy all of a sudden.

Like, all this AI stuff seems really difficult most of the time. Then it's like, "Oh, you just put like a weight token in. Got it. Okay." Um, I think that's, uh, really just it s-suddenly it seemed like, "Oh, this is accessible."

Um, where, where this goes, I think, I think this is where, like, the-- I mentioned Entropix in there. Um, Entropix is basically automating this. Um, so what they did in the S1 paper is they just looked at the length of the output so far, and if like, if they wanted to generate more, like to replicate the graph, they would just stick another weight token and go a little longer and keep on going, stick another, stick it in again.

Entropix7:43

Tim Kellogg8:09

Um, Entropix uses the internal signals from the model. Um, so, uh, there's been a few papers saying that the, the entropy of the, the model will, uh, indicate the, the model's confidence. So if it's, if it's feeling unsure about the answer, you'll, you should see it in the entropy and also the varentropy as well.

Um, the varentropy is just like the slope of the entropy in a sense. Uh, you, you can just think about it as like on, on the blog, uh, or sorry, on the, on the GitHub repository, they, they, they have some kind of flowery language to describe entropy versus varentropy.

Um, it's like, uh, entropy is like the, the clarity of the, um, like if you're looking at the horizon, how clearly can you-- how far can you see? And then varentropy is like the patchiness, like if there's clouds.

But really you can think of varentropy as kind of like the how, how, how much, how much the entropy is changing. So if it's changing a lot, that's potentially that there's, there's multiple valid paths to take that are, that seem clear.

Um, I guess one good example is like if you have, uh, let's see, um, James as the first token, the next token could be like, um, who, who are some, uh, people named James? Like, uh-

Swyx9:23

Bond. Uh-

Tim Kellogg9:24

James Bond. There's one, right? And then, uh, some other person named James, right?

Swyx9:28

Hammer.

Tim Kellogg9:29

Those two seem like very clear, so the varentropy is gonna be high, but the, um, the entropy is gonna feel, feel pretty low in general.

Swyx9:35

Yeah.

Tim Kellogg9:35

So there's two... Uh, so like the Entropix are trying to, um, uh, listen to the internal entropy of the model to, uh, figure out which path to take. And so, um-

Swyx9:45

Is it just from the, the vector of logits? Like, that's all you need? Um, that's it, right? Like-

Tim Kellogg9:52

So that's the core of, that's the core of what Entropix does. They added in also, um, the, the attention as well. But the, the, the core, the core concept is just the logits. That's it. Yeah.

Swyx10:04

That's pretty simple. Uh, why-- what happened to Entropix? Why did it stop? It looks like it's dead.

Tim Kellogg10:09

I'm not part of it , so I can't speak.

Swyx10:10

Yeah.

Tim Kellogg10:11

But what they've been saying online is basically the, um-

Swyx10:15

Frog died

Tim Kellogg10:15

... the, the scope grew a lot. Um, which, I mean, it happens to a lot of projects. Like, it's you get this idea that gets bigger and bigger and bigger, and they just announced like today, I think, that they're starting a new company.

Swyx10:26

Yeah.

Tim Kellogg10:27

Maybe yesterday.

Swyx10:27

Yeah.

Tim Kellogg10:28

So they're starting a new company. So it's, it's, it's grown. It's, it's a big count, uh, idea now.

Swyx10:33

Yeah, but I mean, they could have come out with like E1 and then

you know, at least, at least made somewhat of a splash. This, this like random academic, uh, they kind of took like the b- the most basic implementation and, you know, made much more of a splash. Um, interesting. Uh, so, so like you, you like-- we think that Entropix is still like the right...

Like using varentropy is actually the right way to do it. This is just the dumbest possible way of doing it, right?

Tim Kellogg11:00

So I think-

Swyx11:01

It works pretty well

Tim Kellogg11:02

... I think they're going, I think they're going further. Like, so Entropix is just a sampler. It's a dynamic sampler, so it's extra cool. But-

Swyx11:07

Yeah

Tim Kellogg11:07

... um, if you, if you think about-

Swyx11:09

Sorry, what's dynamic about it?

Tim Kellogg11:11

What was that?

Swyx11:12

What's dynamic about it?

Tim Kellogg11:14

Um, so it's listening to-- So, um, it, it instead of... Okay, so like a, a normal sampler would be, uh-

Swyx11:21

Top K.

Tim Kellogg11:22

Like looking at the temperature. You set, set one temperature for the entire generation-

Swyx11:26

Yeah

Tim Kellogg11:26

... um, and then taking just the probability of-

Swyx11:28

I see.

Tim Kellogg11:28

Right. Versus dy-- uh, uh, Entropix is changing that threshold, changing the top K, changing the top P-

Swyx11:35

Yeah

Tim Kellogg11:35

... um, throughout the generation. Um, and also potentially resampling entirely or forking, doing two different, um-

Swyx11:42

Yeah. Okay

Tim Kellogg11:42

... branches.

Swyx11:42

Yeah. Like fundamentally not too complicated. It's just part of the sort of, you know, inference process.

Tim Kellogg11:46

Yeah. So it's just, it's just doing it dynamically though.

Swyx11:48

Okay.

Tim Kellogg11:48

Um, but if you think about it, all this signal they're going off of is coming from within the model, and if you can do that at runtime, you should be able to train the model to do it on its own.

So I think the direction they're going right now is to like em-embed this into like a RL, uh, reward function and get the model to Essentially start introspecting on itself, being able to tell itself, "Oh, I'm unsure. I should probably, you know, clarify this," right?

Which is-

Swyx12:17

Yeah

Tim Kellogg12:17

... it's interesting. It'd be, be curious to see what kind of, uh, behavior emerges when we start training on, uh, introspection, introspection-

Swyx12:26

And on the-

Tim Kellogg12:27

Queries

Swyx12:27

... on the RL point, uh, one of the conclusions on, on S1 is, like, that supervised fine-tuning also gets, uh, good results. Um, but then now you mentioned that maybe RL is still, like, a, a good way to go.

Do you have any ideas on, like, is the S1 path of taking RL models and bringing them into SFT just a way to get smaller models to perform better and then have the advantages of, like, small inference, but then on kinda like the cutting-edge side you should still do RL, or maybe SFT is also still a good path for cutting edge but just we need a different approach for data generation?

RL vs SFT12:40

Tim Kellogg13:04

Uh, I don't know. I think, I think you need both. I mean, R1, R1 did both. R- there s- there seemed to be something kind of magical about, um, RL, um, that, uh... I mean, I- to be, to be totally frank, a lot of, a lot of these machine learning concepts, um, I, I, I just watch my children growing up and I'm like, "Oh, I kinda see the same thing."

And so you, you kinda see, like, just by just giving them small feedback how, how quickly they learn. Like, I think the same sort of thing is gonna happen with AI as well, and, and reinforcement learning is pretty good at that.

The supervised fine-tuning, um, it obviously works, right? But you have to- it's a much more hands-on approach if you wanna, uh... Like, that one, like they're, they're targeting specific, you know... Like th- it was just math. That was it, right?

Versus when, um, R1 did the R10, um, it developed not just math, but an, an, an, an... like, the whole reasoning ability and be able to, uh, think as long as it needed to. It developed that entire ability, right?

So RL seems to have, like, this... It's almost like it...

Okay, I guess if you, if you, if you view it as, like, a plant, I feel like RL causes a lot of growth, then SFT, you trim it, you trim it down to, like, the shape you want it to be so it's more useful.

Um, I, I, I think, I think that might be the right way to think about it, but I'm not sure. I'm not a researcher, so what can I say?

Doom Loops14:27

Swyx14:33

Yeah. Like, uh, what do you... Like, uh, I, I see you're also in the game of, uh, making predictions. Um, what else are you looking for this year in terms of progress? Like, there- so we've already had a lot and we're, you know, a month and a half into 2025.

Tim Kellogg14:48

I don't, I don't even know what to say.

Swyx14:50

Yeah. What are, like, your, your hot topics, your pet topics, like, stuff you're, you're, you're like, "Okay," like, "this is the, you know, this is something I'm really paying attention to"?

Tim Kellogg14:57

So, um, honestly, the, the thing on top of my mind is really just, uh, the, um, like, the doom, doom loops. I don't know if you've heard that term, but, like, this, like, the idea of, like, you just-

Swyx15:10

Oh

Tim Kellogg15:10

... that these a- the agents kinda get stuck in a loop.

Swyx15:13

Yep.

Tim Kellogg15:14

Um-

Swyx15:14

That's what, that's what I was hoping that VarEntropy, you know, gets you out of

Tim Kellogg15:18

I kinda think it will. I don't know. We'll see. Um-

Swyx15:21

Why?

Tim Kellogg15:21

... I, I don't think, I don't think the runtime, um, Entropix was really the right answer, but maybe if you do it during training time, it might evolve that-

Swyx15:30

Yeah

Tim Kellogg15:30

... capability.

Swyx15:31

Okay.

Tim Kellogg15:32

Um, that, that seems like it unlocks a whole lot because then, then you get a higher level of just agency in general in these models, but you also get that, the ability to introspect and say like, "Oh, actually I am not confident what I'm currently doing, therefore I should do something that creates my- increases my confidence.

I should go use a tool to get information or, you know, whatever." Like, that's, that's where I think, I think that's where a lot of, uh, capabilities start, start turning on.

Swyx16:02

Yeah.

Tim Kellogg16:02

Um, it feels like a linchpin that we haven't really talked about a whole lot.

Swyx16:08

Yeah. Interesting. Um, I, I, I think that's partially what Deep Research had to solve, uh, when OpenAI released it. Uh, obviously Gemini also had, uh, Deep Research. Um, it seemed like... Yeah, it's a, it's a key part, quite like Operator as well.

You know, all the, all the agents that OpenAI is releasing, they're gonna have to resolve doom loops, and they probably have already worked on it, like, uh, maybe not that difficult. Um, uh, but there's no paper that we can point to that says, "This is how people do it."

Like, it's just-

Tim Kellogg16:41

Mm-hmm

Swyx16:41

... guessing and theory.

Tim Kellogg16:46

Yeah.

Outro16:47

Swyx16:47

Yeah. Cool. Um, yeah, that's all the questions I had. Any other, uh, pet topics or calls to action? Any, any, uh, anything else that you're, you want, uh, people to check out?

Tim Kellogg16:56

Uh, no, I don't, I don't think I have anything right now.

Swyx17:00

All right. Well, thanks for jumping on, uh, for a quick pod. Uh, we wanted to talk about S1 and, you know, your blog post is a top Hacker News, so thank you for, uh, sharing your knowledge and doing the write-ups, honestly.

You, you do a better- You, you make me inspired to write more because, you know- ... like, people, people should. Like, even though, even though, like, you know, this, this, this can be out there, like, um, I think, I think it should be explained well.

I think you explained it very well.

Tim Kellogg17:24

Awesome. Thank you.

Swyx17:26

Yeah. Cool. Thank you, Tim.