# Answer.ai & AI Magic with Jeremy Howard

Latent Space · 2024-08-17

<https://addtry.com/0beadac0-82e1-4e61-85da-89bb84e09edd>

Jeremy Howard of Answer.AI argues that continuous pre-training should be treated as a continuum, not separate phases, and demonstrates how FSDP+QLoRA enables training a 70B model on just two NVIDIA 4090s. He reveals Answer’s non-hierarchical, manager-free R&D lab model that recruited unusual talent like Benjamin Warner and Ben Claviez, who independently launched BERT 24 to revive encoder-only architectures. Howard introduces FastHTML, a pure-Python web framework built on HTMX and Starlette, for creating modern SPAs without JavaScript. He previews “AI Magic,” a dialogue engineering system that moves beyond teletype-style chat interfaces and code editors, aiming to make AI-assisted development more interactive. The episode also covers Answer’s Public Benefit Corporation structure designed to resist hostile takeovers, and critiques decoder-only hype while advocating for encoder-decoder and state-space models.

## Questions this episode answers

### How is it possible to fine-tune a 70 billion parameter language model on consumer GPUs like the NVIDIA RTX 4090?

Jeremy Howard and his team at Answer.AI developed a system combining FSDP (Fully Sharded Data Parallel) with QLoRA (quantized low-rank adaptation) that enables fine-tuning a 70B model on two NVIDIA 4090 GPUs. This required extensive engineering to overcome undocumented library issues and ensure accuracy. They later upgraded to DoRA for even better performance.

[38:58](https://addtry.com/0beadac0-82e1-4e61-85da-89bb84e09edd?t=2338000)

### How does Answer.AI's corporate structure as a Public Benefit Corporation help ensure its long-term mission alignment?

Jeremy Howard explains that Answer.AI is a Delaware C Corp registered as a Public Benefit Corporation (PBC). This legal structure allows the company to prioritize its stated public benefit—maximizing societal good through AI—over short-term profits. As a PBC, Answer.AI can reject buyout offers that conflict with its mission, preventing it from being forced to pivot into harmful uses.

[20:20](https://addtry.com/0beadac0-82e1-4e61-85da-89bb84e09edd?t=1220000)

## Key moments

- **[0:00] Intro**
- **[1:07] Continuous Pre-Training**
  - [1:14] Continued pre-training is replacing distinct fine-tuning steps, says Jeremy Howard.
- **[4:48] Optimizer Schedules**
  - [4:48] Jeremy Howard prefers scheduled optimizers over schedule-free, valuing more control over hyperparameters.
- **[6:08] AI Governance**
  - [6:58] Jeremy Howard predicted OpenAI's governance collapse two days before Sam Altman was fired.
  - [8:38] Jeremy Howard: 'Companies are sociopathic, by design ... the alignment problem as it relates to companies has not been solved.'
  - [12:32] Jeremy Howard explains how a Public Benefit Corporation legally resists acquisition offers that violate its mission.
- **[13:32] Answer.ai Structure**
  - [14:41] Jeremy Howard argues that top AI researchers often have unconventional backgrounds, not just elite institutions.
  - [15:34] Jeremy Howard says 80% of the time, people who produce really high-quality work have unusual backgrounds.
  - [17:24] At Answer AI, nearly every team member admits to imposter syndrome, says Jeremy Howard.
  - [19:28] Answer AI operates with zero managers and no hierarchy, says Jeremy Howard.
  - [20:35] Ben Claviez initiated BERT 24, a collaboration to build a new BERT, without direction from Jeremy Howard.
  - [22:32] Kerem updated Answer AI's FSDP QLoRA to use DoRA, which outperforms fine-tuning, says Jeremy Howard.
- **[27:04] Recruiting Talent**
- **[32:34] BERT Revival**
  - [34:00] Jeremy Howard explains why encoder-decoder models are superior for tasks with a fixed input context.
  - [36:22] Jeremy Howard: 'Everybody should look at the thing that has shown signs of being useful in the past, but nobody really followed up with properly.'
- **[37:10] QLoRA Innovations**
  - [38:16] Jeremy Howard describes creating FSDP QLoRA as extremely unpleasant, fiddly work requiring tenacious janitorial engineering.
  - [41:41] Jeremy Howard warns that validating open-source performance improvements requires rigorous benchmarks, as many claims don't hold up.
- **[43:42] Inference Optimization**
  - [45:36] Jeremy Howard predicts the future is quantized models with adapters, eliminating merged model downloads.
- **[47:48] FastHTML**
  - [50:34] Jeremy Howard announces FastHTML, a Python framework to build full web apps in a single file.
  - [57:07] Jeremy Howard reveals AI Magic, a new dialogue engineering approach beyond prompt engineering.
  - [57:56] Jeremy Howard compares ChatGPT's UX to a 1970s teletype, calling it suboptimal.
- **[1:01:16] Dialogue Engineering**
  - [1:02:43] Jeremy Howard plans to launch AI Magic via a course called 'How to Solve It With Code'.
- **[1:04:11] AI Wishlist**
  - [1:05:14] Jeremy Howard predicts BERT 24 will spark a renewed interest in encoder-only architectures.
  - [1:06:04] Jeremy Howard calls for KV cache saving as a new class of technique, noting Gemini's upcoming support.
  - [1:08:35] Jeremy Howard sees diffusion models fitting in generative pipelines as a pre-token generation planning step.

## Speakers

- **Alessio** (host)
- **Swyx** (host)
- **Jeremy Howard** (guest)

## Topics

Language Models, Hardware, State Space Models

## Mentioned

Answer.ai (company), Google DeepMind (company), Hugging Face (company), Meta (company), OpenAI (company), fast.ai (company), AI Magic (product), BERT24 (product), Bits and Bytes (product), ChatGPT (product), Claude (product), Claudette (product), FSDP QLoRA (product), FastHTML (product), Flash Attention (product), Gemini (product), HQQ (product), Llama (product), PEFT (product), vLLM (product)

## Transcript

### Intro

**Alessio** [0:01]
Hey everyone. Welcome to the Latent Space Podcast. This is Alessio, partner and CTO in residence at Decibel Partners, and I'm joined by my co-host Swyx, founder of Smol.ai

**Swyx** [0:10]
And today we're back, uh, with Jeremy Howard. I think your third appearance on Latent Space. Welcome.

**Jeremy Howard** [0:16]
Wait, third? Second.

**Swyx** [0:18]
Well-

**Jeremy Howard** [0:18]
Second?

**Swyx** [0:18]
... I grabbed you in Europe, so and, and we had a-

**Jeremy Howard** [0:20]
I see.

**Swyx** [0:21]
... very fun-

**Jeremy Howard** [0:22]
Okay, fair enough

**Swyx** [0:23]
... uh, standing outside street episode of just, yeah.

**Jeremy Howard** [0:24]
I never heard that, by the way. You gotta send me a link. I gotta hear what it sounded like.

**Swyx** [0:27]
Yeah, yeah.

**Alessio** [0:28]
Right.

**Swyx** [0:28]
It's one of the, the Europe ones.

**Alessio** [0:29]
I, I, I think the two episodes are six hours, so there, there's plenty to, to listen. We'll make sure to send it over.

**Swyx** [0:35]
Yeah, we're trying this thing where, uh, the major ML conferences, we, you know, do a little audio tour of, uh, the conference and, uh, give people a sense of what it's like. Um, but the last time you were on, you declared the end of fine-tuning.

Uh, I hope that-- I, I know that I, you know, I, I, I sort of editorialized the title a little bit, and I know you were slightly uncomfortable-

**Jeremy Howard** [0:55]
Yeah, yeah

**Swyx** [0:55]
... with it, but you sort-- you just own it anyway. Uh, I think you're very good at the hot takes. Um- ... and we were just discussing in our pre-show that, um, things have-- it's really happening, that, uh, the continued pre-training is, is really happening.

### Continuous Pre-Training

**Jeremy Howard** [1:08]
Yeah, absolutely. Um, I think people are starting to understand that

treating the three ULMFiT steps of like pre-training, you know, and then the kind of like what people would now call instruction tuning, and then I don't know if we've got a general term for this DPO, RLHF-y step, you know, but, you know, the task training.

They're not actually as separate as we originally suggested they were in our paper. And when you treat it more as a continuum, and that you make sure that you have, you know, more of kind of the original dataset incorporated into the later stages, and that, you know, we've also seen with like Llama 3, this idea that those later stages can be done for a lot longer.

These are all of the things I was kind of trying to describe there. It wasn't like, yeah, it wasn't the end of pre-training-- Sorry, it wasn't the end of fine-tuning, but more that we should treat it as a continuum, and we should, we should have much higher expectations of how much you can do with a already trained model.

You can really add a lot of behavior to it. You can change its behavior. You can, you know, you can do a lot. So a lot of our research has been around trying to figure out how to modify the model by a larger amount, rather than starting from random weights, 'cause I, I get very offended at the idea of starting from random weights.

**Swyx** [2:36]
Yeah. I saw that, um, in iClear in Vienna, there was a, there was a outstanding paper about starting transformers from data-driven priors. I don't know if you've, uh, saw that one. It, they called it sort of never train from scratch, and, uh, I think it was kind of rebelling against like the, the sort of random initialization, uh, of it.

Um-

**Jeremy Howard** [2:54]
Yeah, I've-- You know, that's been our kind of continuous message since we started fast.ai is, if you're training for random weights, you better have a really good reason, you know, 'cause it seems so unlikely to me that nobody has ever trained on data that has any similarity whatsoever to the general class of data you're working with, and that's the only situation in which I think starting from random weights makes sense.

**Swyx** [3:20]
Yeah. The other trends, uh, since our last pod that I would point people to is, um, I'm seeing a rise in multi-phase pre-training. Um, so Snowflake released a, a large model called, uh, Snowflake Arctic, where they detailed, uh, three phases of training, where they, they had like a different mixture of like-- There was like 75% web in the first m- uh, first instance, and then they reduced the percentage of the web text by 10% each time and increased the amount of code, um, in each phase.

Um-

**Jeremy Howard** [3:51]
Yeah

**Swyx** [3:51]
... and I feel like multi-phase is being called out in papers more. I feel like it's always been a thing. Like changing data mix is not something new, but calling it a distinct phase is new, and, uh, I wonder if there's something-

**Jeremy Howard** [4:05]
Yeah

**Swyx** [4:05]
... that you're seeing on, on your end.

**Jeremy Howard** [4:06]
Well, so they're getting there, right? So the point at which they're doing proper continued pre-training is the point at which that becomes a continuum rather than a phase. So the only difference with what I was describing last time is to say like, oh, there should-- you know, there's a, a function or whatever which is happening every batch.

And it doesn't like-- It's not a huge difference, but it's like way back, you know, I always used to get offended when people had learning rates that like jumped. And so one of the things I started doing early on in fast.ai was to say to people like, "No, you should actually have-- Your learning rate schedule should be a function, not a list of numbers."

So now I'm trying to give the same idea about, um, training mix.

### Optimizer Schedules

**Swyx** [4:48]
There's been pretty public work from Meta on schedule-free optimizers. I don't know if you've been following Aaron Defazio-

**Jeremy Howard** [4:54]
Mm-hmm

**Swyx** [4:54]
... and what he's doing. Uh-

**Jeremy Howard** [4:56]
Sure

**Swyx** [4:56]
... it's j-just because you mentioned learning rate schedules. Uh, you know, what if you didn't have a schedule?

**Jeremy Howard** [5:01]
I mean, I don't, I don't care very much, honestly. Like I don't think that schedule-free optimizer's that exciting. It's fine. Um, yeah, we've had non-scheduled optimizers for ages, like, um, Les Wright, who's now at Meta, who was part of the fast.ai community there, created something called the Ranger Optimizer.

Um, you know, the, um-- I actually like having more hyperparameters. You know, as soon as you say schedule-free, then like, well, now I don't get to choose, and there isn't really a mathematically correct way of-- Like I actually try to schedule more parameters rather than less.

So like I like scheduling my epsilon in my, in my Adam, for example. I schedule all the things. Um, so, um, but then the other thing we always did with the fast.ai library was make it so you don't have to, you don't have to set any schedules.

So fast.ai always supported like not-- You didn't even have to pass a learning rate. Like it would always just try to have good defaults and do the right thing. Um, but to me, I like to have more parameters I can play with if I want to, but that you don't have to.

### AI Governance

**Alessio** [6:09]
And then the more less technical side, I guess, of your, um, issue, I guess, uh, with the, with the market was some of the large research labs taking all this innovation kind of behind closed doors and whether or not that's good, which it isn't.

Um, and how we could maybe make it more available to people. And then after a month, a month after we released the, the episode, there was the whole Sam Altman drama and like all the OpenAI governance issues. Um, and maybe people started to think more, "Okay, what happens if some of these kind of labs, um, you know, start to break from within, so to speak?"

And the, uh, the alignment of the humans is probably gonna evolve before the alignment of the models. Um, so I'm curious, like if you have any new thoughts, and maybe we can also tie in some of the, the way that we've been building Answer as like a public benefit corp and, um, some of those aspects.

**Jeremy Howard** [6:58]
Sure. So yeah, I mean, it was kind of uncomfortable because two days before Altman got fired, I did a small public video interview in which I said I'm, uh, quite sure that OpenAI's covenant, current governance structure can't continue, um, and that it was definitely gonna fall apart.

And it fell apart two days later, and a bunch of people were like, "What did you know, Jeremy?" It's like-

**Alessio** [7:27]
What did Jeremy see?

**Jeremy Howard** [7:28]
I didn't see anything. It's just obviously true. Um, and so yeah. So my friend Eric Ries and I spoke a lot before that about, you know, Eric's... I think probably most people would agree the top expert in the world on kind of startup and AI governance.

And,

you know, we could both clearly see that this didn't make sense to have like a so-called non-profit, where then there are people working at a commercial company that's owned by or controlled nominally by the non-profit, where the people in the company are being given the equivalent of stock options.

Like everybody there was working there with expecting to make money largely from their equity. So the idea that then a board could exercise control by saying like, "Oh, we're worried about safety issues, and so we're gonna do something that decreases the profit of the company," when every stakeholder in the company, their remuneration pretty much is tied to their profit, it obviously couldn't work.

So I mean, that was a huge oversight there by someone, and I guess it's like... I guess part of the problem is that the kind of people who work at non-profits, you know, and in this case the board, you know, who are kind of academics and, you know, people who...

the kind of true believers, I think it's hard for them to realize that 99.999% of the world is driven very heavily by money, especially huge amounts of money. So, so yeah, Eric and I had been talking for a long time before that about like, well, h- what could be done differently.

Because also companies are sociopathic, like by design. And so the alignment problem as it relates to companies has not been solved. Like companies become huge. They devour their founders, they devour their communities, and they do things where even the CEOs, you know, often of big companies tell me like, "I, I wish our company didn't do that thing.

Uh, but, you know, I know that if I didn't do it, then I would just get fired and the board would put in somebody else." And the board knows if they don't do it, then their shareholders can sue them 'cause they're not maximizing profitability or whatever.

So, um, what Eric's spent a lot of time doing is trying to think about like how do we make companies less sociopathic. You know, how do... or more, you know, maybe a better way to think of it is like how do we make it so that the, you know, the founders of companies can ensure that their companies continue to actually do the things they want them to do.

Um, so,

you know, when we started a company, w- you know, like well, A, we very explicitly decided we're gonna start a company, not a academic lab, not a non-profit, you know. We created a Delaware C corp, you know, the most company kind of company.

Um, but

when we did so, we told everybody, you know, including our first investors, which was you, Alessio.

**Alessio** [10:57]
They sound great.

**Jeremy Howard** [11:00]
We are gonna run this company on the basis of maximizing long-term value, you know. Um, uh, so, you know, uh, in, and in fact, so when we did our, our second round, which was an angel round, we had everybody invest through a long-term SPV which we set up, where everybody had to agree to vote in line with long-term value principles.

Um, so like it's not just... it's n- it's never enough just to say to people like, "Okay, we're trying to create long-term value here for society as well as for ourselves," and everybody's like, "Oh, yeah, yeah. I totally agree with that."

But when it comes to like, "Okay, well, here's a specific decision we have to make which will not maximize short-term value," people suddenly change their mind. So, you know, it has to be written into the legal documents of everybody so that it, it...

there's no question that that's the way the company has to be managed. So then you mentioned the PBC aspect, Public Benefit Corporation, which I never quite understood, uh, previously, and turns out it's incredibly simple. Like it took- You know, like one paragraph added to our corporate documents to become a PBC.

It was cheap, it was easy, but it's got this huge benefit, which is if you're not a public benefit corporation, then somebody can come along and offer to buy you

w- with a stated description of, like, turning your company into the thing you most hate, right? And if they offer you more than the market value of your company and you don't accept it, then you are not necessarily meeting the, kind of your fiduciary responsibilities.

**Alessio** [12:48]
Mm-hmm.

**Jeremy Howard** [12:49]
So the way, like, Eric always described it to me, you know, is like, uh, if Philip Morris came along and said, "You've got great technology for marketing cigarettes to children, so we're gonna pivot your company to do that entirely, and we're gonna pay you 50% more than the market value," you're gonna have to say yes.

If you have a PBC, then you are more than welcome to say no if that offer is not in line with your stated public benefit. So our stated public benefit is to maximize this, this, you know, the benefit to society through using AI.

So given that more children smoking doesn't do that, then we can say like, "No, we're not selling to you."

### Answer.ai Structure

**Alessio** [13:33]
Yep. And I, I was looking back at, um, some of our emails. Um, you, you sent me an email on November 13th, uh, about talking, and then on four- on the 14th, I sent you an email, uh, working together to free AI was the, the subject line.

Um, and then that was kind of the, the start of the, the seed round. And then two days later, Sam Altman got fired. So, uh, th- this was, like, not even... The, you know, you were having these thoughts even before we had, like, a public example of, like, why some of the current structures didn't work.

So, um, yeah, you were very ahead of, uh, ahead of the curve, so to speak.

**Jeremy Howard** [14:07]
Mm.

**Alessio** [14:07]
I would love just to, you know... People, people can read your awesome introduction blog on Answer and the idea of having a R&D lab versus, um, R lab-

**Jeremy Howard** [14:17]
Mm

**Alessio** [14:17]
... and then a D lab-

**Jeremy Howard** [14:18]
Mm

**Alessio** [14:18]
... uh, somewhere else. Uh, I think to me the most interesting thing has been hiring-

**Jeremy Howard** [14:22]
Mm

**Alessio** [14:23]
... and some of the awesome people that you've been bringing on that maybe don't fit the central casting of Silicon Valley, so to speak.

**Jeremy Howard** [14:29]
Mm.

**Alessio** [14:29]
Like, sometimes I call it, like, playing baseball cards, you know, and people are like, "Oh, what teams was this person on? Where did they work?"

**Jeremy Howard** [14:35]
I know.

**Alessio** [14:35]
Versus focusing on ability. So I would love to, for you to give a shout-out to, to some of the awesome folks-

**Jeremy Howard** [14:40]
Yeah

**Alessio** [14:40]
... that you have on the team.

**Jeremy Howard** [14:41]
So, you know, there's, like, a graphic going around describing, like, the people at xAI, you know, the Elon Musk thing, and, like, they are all connected to, like, you know, multiple of Stanford, Meta, DeepMind, OpenAI, Berkeley, Oxford. It's just...

Look, these are all great institutions, and they have good people, and I'm definitely not at all against that. But damn, there's so many other people. And one of the things I found really interesting is, um,

kind of anytime I... almost anytime I see something which I think, like, this is really high-quality work, and it's, like, something I don't think would have been built if that person hadn't built the thing right now, I nearly always reach out to them and ask to chat.

And I tend to dig in to find out, like, "Okay, you know, why did you do that thing? Everybody else has done this other thing. Your thing's much better, but it's not what other people are working on." And, like, 80% of the time, I find out the person has a really unusual background.

So, like, often they'll have, like, either they, like, came from poverty and, like, didn't get an opportunity to go to a good school, or they, like, you know, had dyslexia and, you know, got kicked out of school in year 11.

Or, you know, or they had a health issue that meant they couldn't go to university, or something happened in their past and they ended up out of the mainstream, and then they kind of succeeded anyway. And those are the people that throughout my career I've tended to kind of accidentally hire more of.

But I... It's not exactly accidentally. It's like when I see somebody who's done... Two people who have done extremely well. One of them did extremely well in exactly the normal way, from the background to k- entirely pointing in that direction, and they achieved all the hurdles to get there.

And like, okay, that's quite impressive, you know. But another person who did just as well, despite lots of constraints and doing things in really unusual ways and came up with different approaches, like, that's normally the person I'm likely to find useful to work with.

'Cause they're often, like, risk-takers, they're often creative, they're often extremely tenacious, um, they're often very open-minded. So that's the kind of folks we, you know... I tend to find myself hiring. And I think, like... So now at Answer AI, um,

it's a group of people that are strong enough that nearly every one of them has independently come to me in the past few weeks and said... and told me that they have imposter syndrome and they're not convinced that they're good enough to be here.

You know? And I kind of heard it at the point where I was like, "Okay, I don't think it's possible that all of you are so far behind your peers that you shouldn't get to be here." But I think part of the problem is, like, as an R&D lab, the great developers look at the great researchers and they're like, "Wow, these big-brained crazy research people with all their math and shit, they're too cool for me.

Oh my God." And then the researchers look at the developers and they're like, "Oh, they're killing it, making all this stuff with all these people using it and talking on Twitter about how great it is." And I think they're both a bit intimidated by each other, you know?

And so I have to kind of remind them, like, "Okay-" There are lots of things in this world where you suck compared to lots of other people in this company, but also vice versa, you know, for all things.

And the reason you came here is because you wanted to learn about those other things from those other people and have an opportunity to, like, bring them all together into a single unit. Um, so, you know, it's not reasonable to expect you're gonna be better at everything than everybody else.

Even although, like, I guess the other part of it is for nearly all of the people in the company, to be honest, they have nearly always been better than everybody else at nearly everything they're doing, nearly everywhere they've been.

So it's kinda weird to be in this situation now where it's like, "Gee, I can clearly see that I suck at this thing that I'm meant to be able to do compared to these other people," or I'm like the worst in the company at this thing for some things.

So I think that's a healthy place to be, you know, uh, as long as you keep reminding each other about that's actually why we're here. Um, and it's been really nice to s- like, it's all a bit of an experiment.

Like, um, we don't have any managers. Uh, we don't have any hierarchy from that point of view. So for example, I'm not a manager, which means I, I don't get to tell people what to do or how to do it or when to do it.

Um, and it's been a... Yeah, it's been a, been an experiment to see how that would work out, and it's been great. Like, um, so for instance, um, Ben Claviez, who you might have come across, he's the author of Raggtui, he's the author of ReRankers, super strong information retrieval guy.

And a few weeks ago, he was-- he... You know, this additional channel appeared on Discord, on our, on our private Discord called BERT 24. Like, these people started appearing, as in our collab sections. We have a collab section for, like, collaborating with outsiders.

And these people started appearing, they're all these names that I recognize, like BERT 24, and they're all talking about, like, the next generation of BERT, and I start following along, and it's like, okay, Ben decided that, I think quite rightly, we need a new BERT.

Um, 'cause everybody... Like, so many people are still using BERT, and it's still the best at so many things.

**Guest** [20:40]
Mm-hmm.

**Jeremy Howard** [20:40]
But it actually doesn't take advantage of lots of best practices. And so he just went out and found basically everybody who's created better BERTs in the last four or five years, brought them all together. Suddenly, there's this huge collaboration going on.

So yeah, I didn't tell him to do that. He didn't ask my permission to do that. Um, and then, like, Benjamin Warner dived in, and he's like, "Oh, I created a

whole transformers from scratch implementation designed to be maximally hackable." Um, he originally did it largely as a teaching exercise to show other people, but he was like, "I could, you know, use that to create a really hackable BERT implementation."

Um, in fact, he didn't say that. He said, "I just did do that," you know. Uh, "And I created a repo," and then everybody's, like, starts using it. They're like, "Oh my God, this is amazing. I can now implement all these other BERT things," you know.

Um, and it's not just Answer.AI guys there. You know, there's lots of folks, you know, who have, like, contributed new dataset mixes and blah, blah, blah. So, I mean, I can help in the same way that other people can help.

So, like, then Ben Claviez reached out to me at one point and said, like, "Okay, can you help me... Like, what have you learnt over time about how to manage, you know, intimidatingly capable and large groups of people who you're nominally meant to be leading?"

Um, and so I, you know, I, like, I try to help, but I don't direct. Um, another great example was, uh, Kerem, um, uh, who, um, after our FSDP QLoRA work,

decided quite correctly that it didn't really make sense to use LoRA in today's world. You wanna use the normalized version, which is called DoRA. And like two or three weeks after we did FSDP QLoRA, he just popped up and said, "Okay, I've just converted the whole thing to DoRA, and I've also created these vLLM extensions, and I've got all these benchmarks, and, you know, now I've got, um, training of quantized models with adapters that are as fast as LoRA and as...

actually better than, weirdly, fine-tuning." It's just like, "Okay. That's great," you know. Um, and yeah, so the things we've done to try to help make these things happen as well, it's like we have... So we don't have any required meetings, you know, but we do have a meeting for each pair of major time zones that everybody's invited to.

And, you know, people see their colleagues doing stuff that looks really cool and say like, "Oh, how can I help?" You know, or, "How can I learn?" Or whatever. So another example is, uh, Austin, who, you know, amazing background.

He ran AI at Fidelity, he ran AI at Pfizer, he ran browsing and retrieval for, for Google's DeepMind stuff, um, created Gemma.cpp, and he's been working on a new, uh, system to make it easier to do WebGPU programming.

'Cause again, he quite correctly identified, like, you know, this is a way that not everybody has to use CUDA, not everybody has to use NVIDIA. You can do stuff on your own computer, optionally through the browser. We need to make this easier to do.

And so I- Yeah. So I said to him, like, "Okay, I, I wanna learn about that." Not an area that I have much expertise in, so yeah, he's gonna-

**Alessio** [24:36]
Yeah

**Jeremy Howard** [24:36]
... show me what he's working on and teach me a bit about it, and hopefully I can help contribute. I think one of the key things that's happened in all of these is everybody understands the, um, what Eric Gilliam, who wrote the second blog post in our series, the, the R&D historian, describes as everybody has total flexibility to do what they want, but we all understand, like, kind of roughly why we're here.

You know, we all have the same i- you know, we agree with the premises around, like, you know, everything's too expensive, everything's too complicated. You know, people are building too many vanity foundation models rather than taking better advantage of fine-tuning.

Like, there's this kind of general, like, sense of, like, we're all on the same wavelength about,

you know, all the ways in which r- current research is fucked up and-

**Alessio** [25:34]
Mm-hmm

**Jeremy Howard** [25:34]
... you know, all the ways in which, you know, we kind of try... You know, are worried about centralization, and we, you know, um, we all care a lot about not just research for the point of citations, but research that actually wouldn't have happened otherwise and actually is gonna relate to real world outcomes.

And so, yeah, with this kind of like shared vision, people understand, like, you know, so when they-- Then I say like, "Oh, well, you know, tell me, Ben, about BERT24. What's that about?" And he's like, you know, like, "Oh, well, you know, you can see from an accessibility point of view, or you can see from a kind of a actual practical impact point of view, there's far too much focus on, um, decoder-only models and, you know, like BERT's used in all of these different places and industry.

And so I can see, like in terms of our basic principles, what we're trying to achieve, this seems like something important."

**Alessio** [26:28]
Mm-hmm.

**Jeremy Howard** [26:28]
And so I think that's like a really helpful that we have that kind of shared perspective-

**Alessio** [26:35]
Mm-hmm

**Jeremy Howard** [26:35]
... you know.

**Alessio** [26:36]
Yeah. And before we maybe talk about some of the specific research, when you're like reaching out to people, interviewing them, what are some, some of the traits... Like, h- how do these things come out, you know, usually? Is it working on side projects that you, you know, you're already familiar with?

Is there anything like in the interview process that like helps you screen for people that are more, uh, less pragmatic and more research-driven versus some of these folks that are like, are just gonna do it, you know, they're not waiting for like the-

**Jeremy Howard** [27:03]
So-

**Alessio** [27:03]
... perfect process?

### Recruiting Talent

**Jeremy Howard** [27:05]
Everybody who comes through the recruiting is interviewed by everybody in the company. Um,

you know, our, our, our goal is 12 people, so it's not an unreasonable amount. And like, the way I... So the other thing to say is everybody so far who's come into the recruiting pipeline, everybody bar one, has been hired.

So which is to say our original curation has been good. Um, and that's actually pretty easy 'cause nearly everybody who's coming through the recruiting pipeline are people I, I know pretty well. So, you know, Jono Whittaker and I, you know, he worked on the stable diffusion course we did.

Um, he's outrageously creative and talented, and he's just super like enthusiastic tinkerer, um, just likes making things. And, um, you know, Benjamin was one of the strongest parts of the Fast.ai community, which is now the alumni is like hundreds of thousands of people.

Um, and you know, again, like they're not people who a normal interview process would pick up, right? So Benjamin doesn't have any qualifications in math or computer science. Um, you know, Jono was living in Zimbabwe. He was not...

You know, he was working on like helping some African startups, you know, but not FAANG kind of credentials. Um, but yeah, I mean, when you actually see people doing real work and they stand out above, you know, the, the...

We've got lots of Stanford graduates and OpenAI people and whatever in our alumni community as well. You know, when you stand out above all of, above all of those people anyway, obviously you've got something going for you. Um, you know, Austin, uh, him and I worked together on the, um, masks study we did in the Proceeding of the National Academy of Science.

Uh, so you know, we had worked together, and again, that was a group of like basically the 18 or 19 top experts in the world on public health and epidemiology and, um, uh, research design and so forth, and Austin was, you know, one of the strongest people in that collaboration.

So yeah, you know, like I've been lucky enough to have had opportunities to work with some people who are great and, you know, I'm a very open-minded person, so I kind of am always happy to try working with pretty much anybody, and some people stand out.

You know, there have been some exceptions, people I haven't previously known, like Ben Claviez, actually, I didn't know before. But, you know, w- w- with him, like, j- you just read his code and I'm like, "Oh, that's really well-written code."

Like I... And like it's not written exactly the same way as everybody else's code, and it's not written to do exactly the same thing as everybody else's code. So yeah. And then when I chatted to him, it's just like, I don't know, it felt like we'd known each other for years.

Like we just were on the same wavelength and... But I could pretty much tell that was gonna happen just by reading his code. I think you express a lot- In the code you choose to write and how you choose to write it, I guess.

Um, you know, or another example, uh, is a guy named Vic who was previously the CEO of Dataquest. Um, and like in that case, like he's, you know, he's created a really successful startup. He's like-- He won the, the, the first basically Kaggle NLP competition, which was automatic essay grading.

Um, he's got the current state-of-the-art OCR system, Syria. Um, again, he's just a guy who obviously just builds stuff. You know, he doesn't ask for permission. He doesn't need any like external resources. Um, actually, Kerem's another great example of this.

I mean, I already knew Kerem very well because he was my best ever master's student. But it wasn't my-- it wasn't a surprise to me then when he then went off to create the world state-of-the-art language model in Turkish on his own in his spare time with no budget, you know, from scratch.

This is not fine-tuning or whatever. He like went back to Common Crawl and did everything. So yeah, it's kind of-- I don't know what I'd describe that process as, but it's-

**Swyx** [31:58]
Assemble the inventors

**Jeremy Howard** [31:58]
... not at all based on credentials.

**Swyx** [32:01]
Assemble based on talent, yeah. Um, we wanted to, uh, dive in a little bit more on, you know, turn- turning from the people side of things into the technical bets that you're making. Uh, just a little bit more on BERT.

Uh, I, I was actually-- We just did an interview with Yi Tay from Reca. Uh, I don't know if you're familiar with, uh-

**Jeremy Howard** [32:19]
Yep. Excellent

**Swyx** [32:19]
... his work, but also another encoder-decoder bet. And, um, one of his arguments was actually people kind of over-index on the decoder-only GPT-3 type, uh, paradigm.

**Jeremy Howard** [32:30]
Yeah. Definitely.

**Swyx** [32:31]
I, I wonder if you have, if you have thoughts there that it's maybe non-consensus as well.

### BERT Revival

**Jeremy Howard** [32:35]
Yeah. No, absolutely. So I think it's a great example. So one of the people we're collaborating with a little bit with BERT 24 is, um, Colin Raffel, who is the guy behind-

**Swyx** [32:43]
Oh, GPT-5. Yeah.

**Jeremy Howard** [32:44]
Yeah, most of that stuff. Um, you know, between that and UL2, there's a lot of really interesting work. And so one of the things I've been encouraging the BERT group to do, and Colin has as well, is to consider using a T5 pre-trained, uh, encoder backbone as a thing you fine-tune, which I think would be really cool.

Um, but although he was saying, you know, Colin was also saying actually just use encoder-decoder as your BERT you know, why don't you like use that as a baseline, which I also think is a good idea. Yeah. Look-

**Swyx** [33:25]
But like, you know, what technical arguments are, you know, are people under, under weighting?

**Jeremy Howard** [33:30]
I think Colin would be able to describe this much better than I can, but I'll, I'll, I'll give my slightly non-expert attempt. Look, I mean, think about like diffusion models, right? Like in stable diffusion. Like we use things like U-Net.

We, you know, you, you, you have this kind of downward path, and then in the upward path you have the cross connections, which you-- it's, it's not a tension, but it's like a similar idea, right? You're, you're, you're, you're, you're inputting the original encoding path into your decoding path.

It's, it's critical to make it work, right? 'Cause otherwise, in the decoding part, the model has to like do so much kind of from scratch, right? So like if you're doing translation, like that's a classic kind of encoder-decoder example.

If it's decoder only, you never get the opportunity to find the right, you know, feature, uh, engineering, the right feature encoding for the original sentence. Um, and it kind of means then on every token that you generate, you have to recreate the whole, the whole thing, you know.

So if you have an encoder, it's basically saying like, okay, this is your opportunity model to create a really useful feature representation for your, for your input information. Um,

so I think there's really strong arguments for encoder-decoder models anywhere that there is this kind of like context or source thing, you know. Um, and then why encoder only? Well, because like so much of the time, what we actually care about is like, you know, a classification.

You know, it's like an output. It's like we're not generating an arbitrary length sequence of tokens. So anytime you're not generating an arbitrary length sequence of tokens, uh, decoder models don't seem to make much sense to me. Now, the interesting thing is you see on like Kaggle competitions that decoder models still are at least competitive with things like DeBERTa V3.

Um, but they have to be way bigger to be competitive with things like DeBERTa V3. Um, and the only reason they are competitive is because people have put a lot more time and money and effort into training the decoder only ones.

You know, there, there isn't a recent DeBERTa. There isn't a recent BERT. So yeah, it's a whole part of the world that people have slept on a little bit, and this is just what happens. This is how trends happen rather than like-- To me, everybody should be like, "Oh, let's look at the thing that has shown signs of being useful in the past, but nobody really followed up with properly."

That's, that's the more interesting path, you know. But people tend to be like, "Oh, I, I need to get citations. So what's everybody else doing? Can I make it point one percent better, you know, or point one percent faster?"

That's what everybody tends to do. Yeah. So I think it's like Yi Tay's work commercially now is interesting because here's like a whole- Here's a whole model that's been trained in a different way, so there's probably a whole lot of tasks it's probably better at than, um, you know, GPT and Gemini and Claude.

Um, so that should be a good commercial opportunity for them if they can figure out what those tasks are.

**Swyx** [37:00]
Well, if rumors are to be believed, uh, and he didn't comment on this, but, you know, Snowflake may, uh, may figure out the commercialization for them, so we'll see.

**Jeremy Howard** [37:09]
Good day.

### QLoRA Innovations

**Alessio** [37:11]
Let's talk about FSDP, QLoRA, QDoRA, and all of that awesome stuff. One, one of the things we talked about last time, some of these models are meant to run on systems that nobody can really own, no single person.

Um, and then you were like, "Well, what if you could fine-tune a 70B model on like a 4090?" And I was like, "No, that sounds great, Jeremy," but like can, can we actually do it? Um, and then obviously, you all figured it out.

**Jeremy Howard** [37:38]
Mm-hmm.

**Alessio** [37:38]
Um, can you maybe tell some of the war stories behind that, like, uh, the, the idea behind FSDP, which is kind of taking, uh, you know, sharded, um, data parallel, uh, computation, then QLoRA, which is do not touch all the weights, just go, uh, uh, at the quant- quantize some of the model, and then ju- within the quantized model, only do certain layers instead of doing everything.

**Jeremy Howard** [38:00]
Well, do the adapters, yeah.

**Alessio** [38:02]
Yeah, yeah, to the, to, to the adapters.

**Jeremy Howard** [38:04]
Mm-hmm.

**Alessio** [38:05]
Um, yeah, I, I will leave the floor to you. I think before you published it, nobody thought this was like a short-term thing that we're just gonna have, and now it's like, oh, obviously you can do it, but it's not that easy.

**Jeremy Howard** [38:16]
Yeah. I mean, to be honest, it was extremely unpleasant work to do. Um, it's like not at all enjoyable. It's, um... So I, I kind of did version 0.1 of it myself before we had launched the company. Um, or at least the kind of like the, the, the pieces, which is I just...

They're just, they're all pieces that are difficult to work with, right? So for, for the quantization, you know, I chatted to Tim Dettmers quite a bit and, you know, he very much encouraged me by saying like, "Yeah, it's possible."

He actually thought it'd be easy. It probably would be easy for him, but I'm not Tim Dettmers. You know, so st- so he wrote Bits and Bytes, which is his quantization library and, um, you know, he wrote that for a paper.

Um, he didn't write that to be production like code. It's now like everybody's using it.

**Alessio** [39:07]
He wrote it in one night apparently. Yeah.

**Jeremy Howard** [39:09]
Yeah. So like it's not

particularly well structured. There's lots of code paths that never get used. There's lots of s- you know, multiple versions of the same thing. You have to try to figure it out. So trying to get my head around that was hard and, you know, because it, like, the interesting bits are all written in CUDA, it's hard to like just step through it and see what's happening.

Um, and then, you know, FSDP is this very complicated library in PyTorch, which not particularly well documented, so the only really, really way to understand it properly is, again, just read the code and step through the code. And then, um, like Bits and Bytes doesn't really work in practice unless it's used with PEFT, the Hugging Face library, and PEFT doesn't really work in practice unless you use it with other things.

And there's a lot of coupling in the Hugging Face ecosystem where like none of it works separately. You have to use it all together, which I don't love. Um, so yeah, trying to just get a minimal example that I can play with was really hard, and so I ended up having to rewrite a lot of it myself, um, to kind of create this like minimal script.

One thing that helped a lot was, um, Meta had this Llama recipes repo that came out just a little bit before I started working on that, and like they had a kind of role model example of like, here's how to train FSDP, LoRA.

Didn't work with QLoRA on Llama. Actually, a lot of that had been put together, like a lot of the stuff I discovered, the interesting stuff had been put together by Les Wright, who's, uh... He was actually the, the guy in the Fast.ai community I mentioned who created the Ranger Optimizer, so he's doing a lot of great stuff at Meta now.

Um,

so yeah, I kind of... That helped get some minimum stuff going, and then it was great once Benjamin and Jono joined full time, and so we basically hacked at that together, and then Kerem joined like a month later or something.

Um, but gee, it was just a lot of like fiddly detailed engineering on like barely documented bits of obscure internals. Uh, so my focus was to see if it kind of could work, and I kind of got a bit of a proof of concept working, and then the rest of the guys actually did all the work to make it work properly.

And you know, every time we thought we had something, we, you know, we needed to have good benchmarks, right? So we'd like, we'd, we'd... It's very easy to convince yourself you've done the work when you haven't, you know.

So then we'd actually try lots of things and be like, "Oh, in these like really important cases, the, the memory use is higher," you know, or it's actually slower. And we'd go in and we'd just find like all these things that were nothing to do with our library-

**Alessio** [42:05]
Mm-hmm

**Jeremy Howard** [42:06]
... that just didn't work properly, and nobody had noticed they hadn't worked properly because nobody had really benchmarked it properly. So we ended up, you know, trying to fix a whole lot of different things. And even as we did so, new regressions were appearing in like transformers and stuff that Benjamin then had to go away and figure out like, "Oh, how come flash attention doesn't work in this version of transformers anymore with this set of models?"

And we'd go, "It turns out they accidentally changed this thing so it doesn't work." You know, there's just, there's not a lot of, uh, really good

performance type evals going on in the open source ecosystem, so there's an extraordinary amount of like things where people say like, "Oh, we built this thing and it has this result," and when you actually check it, it, it doesn't.

So yeah, there's a shitload of war stories from From getting that thing to work, and it did require a particularly like tenacious group of people and a group of people who don't mind doing a whole lot of kind of like really janitorial work, to be honest, um, to get the details right, to check them.

**Alessio** [43:10]
Yeah. Yeah, we had, uh, TreeDAO on the podcast, and, uh, we talked about how a lot of it is like systems work-

**Jeremy Howard** [43:17]
Mm-hmm

**Alessio** [43:17]
... to make some of these things work. It's not just like beautiful, pure math that you do on a blackboard. It's like how do you get into the-

**Jeremy Howard** [43:23]
Absolutely

**Alessio** [43:23]
... the nitty-gritty of it. Um-

**Jeremy Howard** [43:25]
I mean, Flash Attention's a great example of that. Like it's-- it basically is just like, "Oh, let's just take the attention and just do the tiled version of it." Which sounds simple enough, you know, but then implementing that is challenging at lots of levels.

**Alessio** [43:41]
Yeah. Uh, what about, uh, inference? You know, obviously you've done all this amazing work on fine-tuning. Um, do you have any research you've been doing on the inference side, how to make local inference really fast on these models too?

### Inference Optimization

**Jeremy Howard** [43:53]
We're doing quite a bit on that at the moment. We haven't released too much there yet. Um, but, uh, one of the things I've been trying to do is also just to help other people. Um, and one of the nice things that's happened is that, um, a couple of folks at, at Meta, including Mark Seraphim, have done a nice job of creating this CUDA mode community of people working on like CUDA kernels or learning about that, and I tried to help get that going well as well and did some lessons to help people get into it.

Um,

so there's a lot going on in both inference and fine-tuning performance, and a lot of it's actually happening kind of related to that. Also, the PyTorch team have created this, uh, Torch AO project on quantization. Um, and so there's a, yeah, big overlap now between kind of the fastai and Answer.AI and CUDA mode communities of people, um,

working on stuff for both inference and fine-tuning. But, um, we're getting close now. You know, our goal is that nobody should be merging models. Nobody should be downloading merged models. Everybody should be using basically quantized plus adapters for almost everything, and just downloading the adapters, um, and that should be much faster.

So that's kind of the place we're trying to get to. It's difficult, you know, because like Karen's been doing a lot of work with, with vLLM, for example. The-- these, these inference engines are pretty complex bits of code.

They have a whole lot of custom kernel stuff going on as well as do the quantization libraries. So we've been working on... We're also quite a bit of collaborating with the folks who do HQQ, HQQ, which is a really great quantization library and works super well.

Um, so yeah, there's a lot of other people outside Answer.AI that we're working with a lot who are, who are really helping on, on all this performance optimization stuff, open source.

**Swyx** [46:03]
Just to follow up on, on merging models, uh, I picked up there that you said nobody should be merging models. That's interesting because, uh, you know, obviously a lot of people are experimenting with this and finding interesting results.

I would say in defense of merging models, you can do it without data. That, that's probably the only thing that's going for it.

**Jeremy Howard** [46:26]
Um, to explain, it's not that you shouldn't merge models, it should... You shouldn't be distributing a merged model. You should distribute a merged adapter, um, 99% of the time. And actually often, one of the best things happening in the model merging world is actually that of- often merging adapters works better.

Uh, the point is, Shawn, that, that once you've got your new model, if you distribute it as an adapter that sits on top of a quantized model that somebody's already downloaded, then it's a much smaller download for them, and also the inference should be much faster because you're not having to transfer FP16 weights from FP-- from HBM memory at all, or, or ever load them off disk.

Um, you know, all the main weights are quantized, and the only floating point weights are in the adapters, so that should make both inference and fine-tuning faster.

**Swyx** [47:25]
Got it. Got it. Okay, perfect. Um, we're moving on a little bit to the rest of the fast universe. Um, I had... I would have thought that, uh, you know, once you started Answer.AI that the sort of fast universe would be kind of on hold.

Uh, and then today you just dropped FastLight, and it looks like, uh, you know, there's, there's more activity going on in sort of fast land.

**Jeremy Howard** [47:47]
Yeah. So fast land and An- Answer land are not really distinct things. Answer land is kind of like the fast land grown up and funded. Um, they both have the same mission, which is to maximize the societal benefit of AI broadly.

### FastHTML

**Jeremy Howard** [48:07]
We want to create thousands of

commercially successful products at Answer.AI, uh, and we want to do that with like 12 people. So that's means we need a pretty efficient stack, you know. Like quite a few orders of magnitude more efficient, not just for creation, but for deployment and maintenance than anything that currently exists.

Um,

people often forget about the D part of our R&D firm. So we've got to be extremely good at, you know, creating, deploying, and maintaining applications, not just models. Much to my, you know, horror, the story around creating web applications is Much worse now than it was 10 or 15 years ago in terms of like if I say to a data scientist, "Here's how to create and deploy a web application,"

you know, either you have to learn JavaScript or TypeScript and about all the complex like libraries like React and stuff, and all the complex like details around security and web protocol stuff around how you then talk to a backend, and then all the details about creating the backend.

You know, if that's your job, you know, and you're, you know, you have specialists who work in just one of those areas, it is possible to... for that to all work. But

compared to like, oh, write a PHP script and put it in the home directory that you get when you sign up to this shell provider, which is what it was like in the '90s. You know, here are the 25 lines of code, um, and you're done, and now you can pass that URL around to all your friends.

You know? Or put this, you know, .po file inside the cgi bin directory that you got when you signed up to this web host. Um, so

yeah, the thing I've been mainly working on the last two weeks is fixing all that, and I, I think I fixed it. Um, so I've created this thing called Fast-

**Swyx** [50:24]
Are we, are we gonna announce it? Yeah. Sorry.

**Jeremy Howard** [50:26]
Uh, I don't know if this is an announcement, but I can... I tell you guys, so yeah. There's this thing called FastHTML, um, which

basically lets you create a complete web application in a single Python file. Um, unlike excellent projects like Streamlit and Gradio, you're not working on top of a highly abstracted thing that's got nothing to do with web foundations. You're working with web foundations directly, but you're able to do it by u- using pure Python.

There's no template, there's no Jinja, there's no separate like CSS and JavaScript files. Um, it looks and behaves like a modern SPA web application. Um,

and you can create components for like Daisy UI or Bootstrap or Shoelace or whatever fancy JavaScript and/or CSS, Tailwind, et cetera, library you like. Um, but you can write it all in Python. You can pip install somebody else's set of components and use them entirely from Python.

You can develop and prototype it all in a Jupyter Notebook if you want to. It all displays correctly, um, so you can like interactively do that. And then you mentioned FastLite, so, um, specifically now if you're using SQLite in particular, it's like ridiculously easy to have that persistence, you know, uh, and you can basically all of your handlers will be passed database-ready objects automatically, um, that you can just call .delete, .update, .insert on.

Um, yeah, you get session, you get security, you get all that. So it's, it's, it's again, like with most of everything I do, it's very little code. It's mainly tying together really cool stuff that other people have written, so.

Um, um, you don't have to use it, but a lot of the best stuff comes from its incorporation of, uh, HTMX, um, which to me is basically the thing that changes your browser to make it work the way it always should've.

So it's a... it just does four small things, but those four small things are the things that are basically un- unnecessary constraints that HTML should never have had, so it removes the constraints. Um, uh, it sits on top of Starlette, which is a very nice, you know, kind of lower level platform for building these kind of web applications.

The, um, the actual interface matches as closely as possible to FastAPI, which is a really nice system for creating the kind of classic Jav- JavaScript type applications. And, uh, Sebastian, who wrote FastAPI, has been kind enough to help me think through some of these design decisions and so forth.

Um, I mean, everybody involved has been super helpful. Actually, I chatted to Carson, who created HTMX, you know, also about it. Chatted to some of the folks involved in Django. Um, like ev- everybody in the community I've spoken to definitely realizes there's a big gap to be filled around like

highly scalable web foundation-based, you know, pure Python framework, um, with a minimum of fuss. So yeah, getting a lot of support and trying to make sure that Fast- FastHTML works well for people.

**Swyx** [54:20]
Uh, yeah. I, I would say when I heard about this, I, I, I just... I just... I texted Alexio, "I think this is gonna be pretty huge." Uh, you know, like, um, people consider Streamlit and Gradio to be the state of the art, but I think there's so much to improve and, uh, you know, having a...

having sort of what you say, what you call web found- web foundations or web fundamentals at the core of it, I think it would be really helpful. Um-

**Jeremy Howard** [54:40]
I mean, it's based on 25 years of thinking and work for me. So like- ... FastMail was built on a system much like this one, um, but that was of Hell. And so I spent, you know, 10 years working on that.

We had millions of people using that every day, really pushing it hard, and I really always enjoyed working in that. So, you know, and obviously lots of other people have done like great stuff, and particularly HTMX, you know.

So I've been thinking about like, yeah, how do I pull together the best of the fra- web framework I created for FastMail with HTMX? There's also things like PicoCSS, which, um, is the CSS system which by default FastHTML comes with.

Although, as I say, you can pip install anything you want to, but it makes it like super easy to... You know, so we try to make it so that just out of the box, you don't have any choices to make, you know, if you don't want to.

**Swyx** [55:40]
Yeah.

**Jeremy Howard** [55:40]
You can make choices, but if for most people, you just, you know, it's like the PHP in your home directory thing. You just start typing, and just by default you'll get something which looks and feels, you know, pretty okay.

And if you wanna then write a version of Gradio or Streamlit on top of that, you totally can. And then the nice thing is, if you then write it in kind of the Gradio equivalent, which will be, you know, I imagine we'll create some kind of pip installable thing for that.

Once you've outgrown or if you outgrow that, it's not like, okay, throw that all away and start again in this like whole separate language, but it's like this kind of smooth, gentle path that you can take step by step, 'cause it's all just standard web foundations all the way, you know.

**Swyx** [56:32]
Yeah. Got it. Um, well, so, uh, you know, just, just to wrap up the sort of open source, um, work that you're doing. Um, you know, you're, you're, you're aiming to create thousands of projects with a, with a very, very small team.

Um, and I haven't heard you mention once AI agents or AI developer tooling or AI-

**Jeremy Howard** [56:51]
Sorry

**Swyx** [56:51]
... code maintenance. Um, you know, please, I, I know you're very productive, but, you know, what is the role of AI in your own work?

**Jeremy Howard** [57:00]
So I'm making something.

**Swyx** [57:02]
Ooh.

**Jeremy Howard** [57:02]
I'm not sure how much I wanna say just yet. Uh, okay.

**Swyx** [57:06]
Give us a nibble.

**Jeremy Howard** [57:07]
All right. I'll give you the key thing. So I've created a new, uh, approach. It's not called prompt engineering, it's called dialogue engineering. Um, and I'm creating a system for doing dialogue engineering. Um, um, it's currently called AI Magic.

Um, I'm doing most of my work in this system, and it's making me much more productive than I was before I used it. So I always just build stuff for myself and hope that it'll be useful for somebody else.

Um,

think about ChatGPT with code in- code interpreter, right? Um,

the, the basic UX is the same as a 1970s teletype, right? So if you wrote APL on a teletype in the 1970s, you typed onto a thing, your words appeared at the bottom of a sheet of paper, and you'd like hit Enter and it would scroll up, and then the answer from APL would, would be printed out and it would scroll up, and then you would type the next thing and like...

Uh, which is also the way, for example, um, a shell works, like Bash or ZSH, whatever. Um, it's, it's not terrible, you know. Like, we all get a lot done in these like very, very basic teletype style REPL environments.

Um, but I've never felt like it's optimal, you know. And to me, um, you know, so... And, and everybody else has just copied ChatGPT. So it's also the way Bard and Gemini work. It's also the way the Claude web app works.

And then you add code interpreter, and the most you can do is to like plead with ChatGPT to write the kind of code I want.

**Swyx** [59:05]
Uh-

**Jeremy Howard** [59:05]
It's pretty good for very, very, very beginner users who like can't code at all. Like by default now the code's even hidden away, so you never even have to see it ever happened. But for somebody who's like wanting to learn to code or who already knows a bit of code or whatever, it's, it seems really not ideal.

So okay, that's one end of the spectrum. The other end of the spectrum, which is where Shawn's work comes in, is, um, oh, you wanna do more than ChatGPT? No worries. Here is Visual Studio Code. I run it.

There's an empty screen with a flashing cursor. Okay, start coding. You know? And it's like, okay, you can use systems like Shawn's or like Cursor or whatever to be like, okay, uh, Apple-K in Cursor to like, uh, create a form that blah, blah, blah.

But it's... In the end, it's like a convenience over the top of this incredibly complicated system that full-time sophisticated software engineers have designed over the past few decades in a totally different environment as a way to build software, you know?

And so we're trying to like shoehorn in AI into that, and it's, it's not easy to do, and I think there are like much better ways of thinking about the craft of software development in a language model world to be much more interactive, you know.

So the thing that I'm building is, is neither of those things. It's something between the two, and it's built around this idea of crafting a dialogue, you know, where the outcome of the dialogue is, you know, uh, the, the artifacts that you want, whether it be a piece of analysis or whether it be a Python library or whether it be a technical blog post or whatever.

So as part of building that, I've created something called Claudette, which is a library for Cl- for Claude. I've created something called Cosette, which is a library for OpenAI. Um, they're libraries which are designed to make those APIs much more usable, much easier to use, much more concise.

### Dialogue Engineering

**Jeremy Howard** [1:01:26]
Um, and then I've written AI Magic on top of those. Um, and that's been an interesting exercise because, uh, I did Claudette first, and rather than trying to, like... I was looking at what Simon Willison did with his fantastic LLM library, and his library is designed around, like, let's make something that supports all the LLM inference engines and commercial providers.

And I thought, "Okay, what if I did something different, which is like make something that's as Claude-friendly as possible and forget everything else?" So that's what Claudette was. So for example, one of the really nice things in Claude is prefill.

Uh, so by telling the assistant that this is what your response started with, there's a lot of powerful things you can take advantage of. Um, so yeah, I created Claudette to be as Claude-friendly as possible. And then after I did that, um, and then with Claude, with...

particularly with GPT-4o coming out, I kind of thought, "Okay, now let's create something that's as OpenAI-friendly as possible." And then I tried to look to see, well, where are the similarities and where are the differences, and how can I make, make them compatible in places where it makes sense for them to be compatible without losing out on the things that make each one special for what they are.

Um, so yeah, those are some of the things I've been working on in that space, and I'm thinking we might launch AI Magic via a, a course called How to Solve It With Code. Uh, the name is based on the classic Polya book, if you know How to Solve It, which is, um, you know, one of the classic math books of all time.

Um,

where we're basically gonna try to show people how to solve challenging problems that they didn't think they could solve without doing a full computer science course, by taking advantage of a bit of AI and a bit of, like,

practical skills. Um, and it's particularly for this, like, whole generation of people who are learning to code with and because of ChatGPT. Like, I know a lot, know a lot of people who didn't really know how to code, but they've created things because they use ChatGPT, but they don't really know how to maintain them or fix them or add things to them that ChatGPT can't do, because they don't really know how to code.

So this course will be designed to show you how you can like, you know, either become a developer who can, like, supercharge their capabilities by using language models, or become a language model-first developer who can supercharge their capabilities by understanding a bit about process and fundamentals.

Um, so yeah.

### AI Wishlist

**Alessio** [1:04:13]
Nice. That's a, that's a great spoiler, you know. I guess the fourth time you're gonna be on Learning Space, we're gonna talk about AI Magic. Uh, Jeremy, before, before we wrap, um, th- this was a, just a great run through everything.

Um, what are the things that when you next come on the podcast in 9, 12 months, we're gonna be like, "Man, Jeremy was like really ahead of it." Like, is there anything that you see in the space that maybe people are not talking enough?

Um, you know, what's the next company that's gonna fall, like a drama internally? Anything, anything like-

**Jeremy Howard** [1:04:43]
I think, you know, hopefully we'll be talking a lot about FastHTML and hopefully the international community that at that point has come up around that, and also about AI Magic and about dialogue engineering. Um, hopefully dialogue engineering catches on, because I think it's the right way to think about a lot of this stuff.

What else? Um, just trying to think about more on the research side. Yeah, I th- you know, I mean, we've talked about a lot of it, like I think encoder-decoder architectures, encoder ar- only architectures. Hopefully, we'll be talking about like the whole re-interest in BERT that BERT 24 stimulated.

Um...

**Swyx** [1:05:15]
There's a state-space model that came out today that, uh, might be interesting for just general discussion. Um, one thing that stood out to me with Cartesia's blog post-

**Jeremy Howard** [1:05:25]
Mm

**Swyx** [1:05:25]
... was that they were talking about real-time ingestion of billions and trillions of tokens, um, and keeping that context obviously in the, in the state space that they have. Uh-

**Jeremy Howard** [1:05:35]
Yeah

**Swyx** [1:05:35]
... I'm wondering what your thoughts are because you, you've been entirely-

**Jeremy Howard** [1:05:38]
Yeah

**Swyx** [1:05:38]
... transformers, uh, the whole time.

**Jeremy Howard** [1:05:39]
Yeah. No, so obviously my background is RNNs and LSTMs. Um-

**Swyx** [1:05:44]
Of course.

**Jeremy Howard** [1:05:46]
... and I'm still a believer in the idea that state is something you can update, you know. So obviously, um, Sep Pokhrel came up... came out with, uh, xLSTM recently. Um,

oh my God, okay, another whole thing we haven't talked about, which is somewhat related. Uh, I've gone... I've been going crazy for like a long time about, like, why can I not pay anybody to save my KV cache f- you know, for like, I just ingested The Great Gatsby or the documentation for Starlette or whatever.

You know, I'm sending it as my prompt context. Why are you redoing it every time, you know? So Gemini is about to finally come out with, um, KV caching, and this is something that Austin actually in Gemma.cpp had had on his roadmap for years.

Um, well, not years, months, long time, um, is q- is, is that. The idea that the KV cache is like a thing that, like, it's, it's a third thing, right? So there's RAG, you know, there's in-context learning, you know, and prompt engineering, and there's, uh, KV cache creation.

Um, I think it creates like a whole new class almost of applications or of, of, of, of, of techniques where- You know, for me, for example, I very often, like, I very often work with really new libraries, or I've created my own library that I'm now writing with rather than on.

So I want all the docs to my new library to be there all the time. So yeah, I want to upload them once and then all of-- have a whole discussion about building this application using FastHTML. Well, nobody's got FastHTML in their, in their, uh, language model yet.

I don't want to send all the FastHTML docs across every time. So one of the things I'm looking at doing in AI Magic actually is taking advantage of some of these ideas so that you can, um,

have the documentation of the libraries you're working on be kind of always available. So there'll be ways to, uh, you know, something over the next 12 months people will be spending time thinking about is how to like, where to use RAG, where to use fine-tuning, where to use KV cache storage, you know, and, and how to use state.

Um, because in state models and xLSTM, again, state is something you, you update. So how do we combine the best of all of these worlds?

**Alessio** [1:08:35]
And Jeremy, j-- I know before you talked about how some of the autoregressive models are not maybe a great fit for agents. Any other thoughts on like JEPA, Diffusion for Text, any interesting thing that you've seen pop up?

**Jeremy Howard** [1:08:46]
In the same way that like we probably ought to have state that you can update, i.e. xLSTM and state models, in the same way, um, a lot of things probably should have an encoder. JEPA and Diffusion both seem like the right conceptual mapping for a lot of things we probably want to do.

So

the idea of like

there, there should be a, a, a piece of the generative pipeline, which is like thinking about the answer and coming up with a sketch of to what the answer looks like before you start outputting tokens. That's where it kind of feels like Diffusion ought to fit.

You know, and Diffusion is, because it's not autoregressive-

**Alessio** [1:09:37]
Mm-hmm

**Jeremy Howard** [1:09:37]
... it's like, let's try to like gradually de-blur the picture of how to solve this. So this is also where dialogue engineering fits in, by the way. So with dialogue engineering, one of the reasons it's working so well for me is I use it to kind of like craft the thought process before I generate the code, you know.

**Alessio** [1:10:02]
Mm.

**Jeremy Howard** [1:10:03]
Um, so yeah, there's a lot of different pieces here and,

uh, I don't know how they'll all kind of exactly fit together. I don't know if JEPA is going to actually end up working in the text world. I don't know if Diffusion will end up working in the text world.

But they seem to be like trying to solve a class of problem which is currently unsolved.

**Alessio** [1:10:22]
Awesome, Jeremy. This was great, um, as usual. Um, thanks again for coming back on the pod.

**Jeremy Howard** [1:10:28]
Thank you, Alessio. Thank you, Shawn.

**Alessio** [1:10:28]
Um, and thank you all for listening.

**Swyx** [1:10:31]
Yeah, that was fantastic.

---

This library is powered by PodHood (https://podhood.com), the podcast website platform.
