# The End of Finetuning — with Jeremy Howard of Fast.ai

Latent Space · 2023-10-20

<https://addtry.com/d5d5c49e-3f22-41df-b931-56721bc214a3>

Jeremy Howard of Fast.ai argues that fine-tuning language models is essentially continued pre-training, not a separate process, and that the common practice of fine-tuning on a single task causes catastrophic forgetting. He recounts the discovery of single-shot memorization in LLMs, where models memorize entire datasets after one epoch, a phenomenon many practitioners ignore. Howard criticizes the current focus on zero-shot and few-shot learning, advocating for transfer learning and small models. He shares his journey from philosophy to founding Fast.ai, the creation of ULMFit (which inspired GPT), and his ongoing work on making AI accessible. He also discusses his involvement with Modular's Mojo language and the importance of democratizing AI technology.

## Questions this episode answers

### What is Jeremy Howard's view on fine-tuning language models and why does he think it should be replaced with continued pre-training?

Jeremy Howard argues that the common three-step fine-tuning process (pretrain, domain fine-tune, task fine-tune) is flawed and leads to catastrophic forgetting. He now believes that 'there's no such thing as fine-tuning; there's only continued pre-training.' The right approach is to include all desired data types from the start and gradually curate them, never discarding any data entirely. This prevents models like Code Llama from forgetting general knowledge while specializing.

[0:44](https://addtry.com/d5d5c49e-3f22-41df-b931-56721bc214a3?t=44000)

### What is the single-example memorization phenomenon that Jeremy Howard observed when fine-tuning language models?

When fine-tuning a language model, Jeremy Howard and Jono Whitaker noticed a sharp drop in validation loss at the end of each epoch, even after seeing the training data only once. They hypothesized and confirmed through experiments that the model was memorizing the dataset after a single pass, causing a loss drop of about .3. This challenges assumed training dynamics and raises questions about the proper way to fine-tune.

[0:37](https://addtry.com/d5d5c49e-3f22-41df-b931-56721bc214a3?t=37000)

### How did Jeremy Howard's ULMFit model influence the development of GPT and modern NLP?

Jeremy Howard created ULMFit, a three-step process for NLP: pre-train a language model on a large corpus, fine-tune on a domain corpus, then fine-tune on a task. At the time, experts told him it wouldn't work, but it achieved state-of-the-art results and directly inspired Alec Radford to develop GPT. ULMFit demonstrated that transfer learning could revolutionize NLP, much like ImageNet did for vision.

[0:06](https://addtry.com/d5d5c49e-3f22-41df-b931-56721bc214a3?t=6000)

## Key moments

- **[0:00] Intro**
  - [1:12] Jeremy Howard barely attended university lectures while working 80-100 hour weeks at McKinsey & Company.
- **[4:05] ULMFit**
  - [4:05] fast.ai started in 2016 to make deep learning accessible, which was widely considered a stupid idea.
  - [5:28] If deep learning is only accessible to a handful of computer science PhDs, it's both a waste and dangerous.
  - [6:15] Q: How small was the ULMFit model? A: It had only 24 to 33 million parameters, trained on Wikipedia.
  - [9:50] Jeremy Howard's 30-year-old idea from the Chinese room experiment led him to believe a language model could appear intelligent.
  - [14:04] ULMFit's three-step process — pre-train, fine-tune on domain, fine-tune on task — is still used by ChatGPT today.
  - [14:24] Within hours of trying ULMFit on IMDb, Howard achieved a new state-of-the-art result.
- **[18:12] Validation Loss**
  - [19:20] After GPT's zero-shot success, Howard toured to promote fine-tuning but found no interest, feeling the field went backwards for years.
- **[22:36] Motivation**
  - [25:03] "I was determined to not waste more years of my life doing things which I could not be reasonably sure would have a lot of value."
  - [26:11] Howard and Rachel Thomas started fast.ai to prevent AI from being controlled by a small group of elites, aiming to make it accessible.
- **[29:10] Library**
- **[33:37] DAWNBench**
  - [34:25] Using progressive resizing and other techniques, fast.ai's team of students and Rachel Thomas won DAWNBench with only 10 days of work.
  - [35:35] "All of my research is always how do we do more with less? Less data, less compute, less complexity, less education."
  - [35:44] Q: What did Howard's recent paper on single-example learning find? A: LLMs can memorize a dataset after a single epoch, causing sudden loss drops.
- **[38:20] Memorization**
  - [38:42] Other researchers dismissed the validation loss clunks as normal fine-tuning behavior, but Howard insisted on finding the cause.
  - [40:37] Experiments by Howard and Whitaker supported the hypothesis that LLMs can memorize a large dataset after seeing it just once.
- **[43:15] Pre-training**
  - [43:15] Q: Does fine-tuning cause catastrophic forgetting? A: Yes, as seen in Code Llama, which lost general knowledge after fine-tuning on mostly code.
  - [44:37] Howard now believes that fine-tuning as a separate phase is wrong; all training should be seen as continued pre-training with a diverse data mix.
  - [46:36] Howard predicts that the ULMFit three-step fine-tuning approach will be replaced by a gradual continued pre-training strategy that never discards data types.
  - [47:07] Howard argues that academic incentives reward incremental improvements over real innovation, as seen in GANs vs. better alternatives.
- **[48:38] Incentives & Communities**
- **[54:51] Mojo**
  - [55:24] Howard approached Chris Lattner at the first TensorFlow Dev Summit and spent 30 minutes explaining all the ways AI development was broken.
  - [59:20] Howard and Lattner built a fastai-like library for Swift for TensorFlow, but Google eventually discontinued the project.
  - [1:00:30] "You've gotta be your own boss, man, 'cause only you've got the ideas."
  - [1:00:57] A few years later, Lattner called Howard to say he had started a company to build the language they had discussed, which became Mojo.
- **[1:02:47] Small Models**
  - [1:03:10] Q: What AI areas is Howard most excited about? A: Fine-tuning/transfer learning is still hugely underappreciated, especially compared to RAG.
  - [1:05:15] Q: Can you fine-tune new knowledge into a language model? A: Of course, because fine-tuning is just continued pre-training, which is how knowledge got in originally.
  - [1:06:28] Howard highlights BTLM 3B as a small model that achieves 7B-level quality, showing that efficient small models are possible.
  - [1:09:27] Q: What should people interested in deep learning focus on in 2024? A: Coding is now more accessible, and fine-tuning small models for specific tasks is a huge opportunity.
  - [1:11:50] Despite rapid progress, Howard says no one really knows how to train, fine-tune, or prompt LLMs optimally; there's massive technical debt.
  - [1:12:55] A user created a 6,000-line Python prompting strategy that made GPT-4 play chess at an Elo of 3400, near the best engines.
- **[1:17:08] Lightning Round**
  - [1:19:09] Q: What is the most interesting unsolved question in AI? A: How do language models learn? Understanding training dynamics and loss surfaces.
  - [1:21:43] "Humanity on net is a marvelous species... we should do everything we can to enable more of them to contribute."

## Speakers

- **Alessio** (host)
- **Swyx** (host)
- **Jeremy Howard** (guest)

## Topics

Language Models

## Mentioned

Modular (company), fast.ai (company), BTLM 3B (product), Code Llama (product), Fast Progress (product), GPT (product), Hugging Face Trainer (product), JAX (product), Kaggle (product), Llama (product), Mojo (product), Phi-1 (product), Phi-1.5 (product), Swift for TensorFlow (product), TensorFlow (product), ULMFiT (product)

## Transcript

### Intro

**Alessio** [0:11]
Hey, everyone. Welcome to the Latent Space Podcast. This is Alessio, partner and CTO in residence at Decibel Partners, and I'm joined by my co-host Swyx, founder of Small AI.

**Swyx** [0:21]
Hey, and today we have in the remote studio Jeremy Howard from all the way from Australia. Good morning.

**Jeremy Howard** [0:28]
The remote studio, also known as my house. Good morning. Uh, nice to see you both.

**Swyx** [0:33]
Nice to see you too. Um, a- I, I'm actually very used to seeing you, uh, in your mask as, as a message to people, but today we're mostly audio. Uh, but, uh, thank you for doing the, the very important public service of, uh, COVID awareness.

Um-

**Jeremy Howard** [0:47]
I won't say it was a pleasure. It was, uh, all very annoying and frustrating and tedious, but somebody had to do it, so-

**Swyx** [0:52]
Somebody had to do it

**Jeremy Howard** [0:53]
... such is life

**Swyx** [0:53]
... especially somebody with your profile, I think, um, uh, really drives home the message. Um, so I tend-- So we tend to really, uh, we, we tend to introduce people for them and then, uh, ask people to fill in the blanks on the personal side.

You, uh, something I did not know about you was that you graduated with a BA in philosophy, uh, from the University of Melbourne. Um-

**Jeremy Howard** [1:12]
Mm-hmm

**Swyx** [1:13]
... I, I assumed you had a PhD.

**Jeremy Howard** [1:15]
No, I mean, I barely got through my BA because I was working 80 to 100-hour weeks at McKinsey & Company from 19 years old onwards. So I-

**Swyx** [1:29]
Yeah

**Jeremy Howard** [1:29]
... I actually didn't attend any lectures in second and third-year university.

**Swyx** [1:36]
Well, you d- I guess you didn't need it or you're, you're very sort of self-driven and self-motivated.

**Jeremy Howard** [1:40]
I just, um, I took two weeks off, yeah, before each exam period, um, uh, when I was working at McKinsey, and then... I mean, I can't believe I got away with this in hindsight. I would go to all my professors and say, "Oh, I was meant to be in your class this semester, and I didn't quite turn up.

Were there any assignments I was meant to have done?" Whatever. And- ... I can't believe all of them let me basically have... They basically always would say like, "Okay. Well, if you can have this written by tomorrow, I'll accept it."

So yeah, stressful way to get through university, but...

**Swyx** [2:13]
Well, yeah, it, it shows that, I guess, you, you, uh, min-maxed, uh, the opportunities. That definitely was a precursor-

**Jeremy Howard** [2:19]
I should mention, though, I mean, funnily, like, in as much as I... You know, in, in philosophy, the things I found interesting and focused on in the little bit of time I did spend on it was, um, was ethics and cognitive science.

**Swyx** [2:31]
Mm.

**Jeremy Howard** [2:31]
And it's kind of really amazing that come, it's now come back around and those are actually genuinely useful things to know about, which I never thought would happen.

**Swyx** [2:39]
A lot of, uh, yeah, a lot of relevant conversations there. Um, so, uh, you were a consultant for a while, and then in the magical month of June 1989, you, you founded both Optimal Decisions and FastMail, uh, which I also-

**Jeremy Howard** [2:51]
Yes

**Swyx** [2:52]
... briefly used, so thank you for that.

**Jeremy Howard** [2:53]
Oh, good for you. Yeah, 'cause I had read the statistics, which is at, like, 90% or something of small businesses fail, so I thought if I start two businesses-

**Swyx** [3:02]
Do two

**Jeremy Howard** [3:02]
... I have a higher chance. In hindsight, I was thinking of it as some kind of stochastic thing I didn't have control over, but it's a bit odd. But anyway.

**Swyx** [3:10]
Um, and then you were president and chief scientist at Kaggle, which, uh-

**Jeremy Howard** [3:14]
Uh-huh

**Swyx** [3:14]
... obviously is, is, uh, the, uh, mission, um, uh, sort of composition platform, uh, of, of, uh, machine learning. Um, and then Enlitic, uh, where, uh, you were working on using deep learning to improve medical diagnostics and clinical decisions.

**Jeremy Howard** [3:28]
Yeah, that was actually the first company to use deep learning in medicine, so kind of founded the field.

**Swyx** [3:33]
A- and, and even now that's still, like, a pretty early, uh, phase. And, uh, I, I, I actually heard you on your new podcast with Tanishq-

**Jeremy Howard** [3:40]
Mm. Mm-hmm

**Swyx** [3:41]
... where you went-

**Jeremy Howard** [3:41]
Yeah

**Swyx** [3:41]
... uh, very, very deep into, uh, the stuff, the kind of work that he's doing. Uh, such a young prodigy at, at his age.

**Jeremy Howard** [3:47]
Oh, maybe he's too old to be called a prodigy now. A ex-prodigy.

**Swyx** [3:51]
No, I think he still counts. Um, and anyway, just to, just to round out the bio, you have a lot more other, uh, cr- credentials obviously, but, um, uh, most recently you started fast.ai, which is, which is still, I guess, your, your primary identity, uh, with Rachel Thomas.

### ULMFit

**Swyx** [4:05]
Um, so welcome-

**Jeremy Howard** [4:06]
Yeah, that's my wife

**Swyx** [4:06]
... thanks.

**Jeremy Howard** [4:07]
Thank you.

**Swyx** [4:07]
Yeah. Doing a lot of, uh, public service there, uh, with, like, getting people involved in, in AI and, and, I can't imagine a better way to describe it than fast. fast.ai is you teach people-

**Jeremy Howard** [4:18]
Yeah

**Swyx** [4:18]
... from nothing to stable diffusion in, you know, seven weeks or something, and, and that's amazing.

**Jeremy Howard** [4:22]
Yeah, yeah. I mean, it's funny. You know, when we started that... What was that, like, 2016 or something? The idea that deep learning was something that you could make more accessible was generally considered stupid. Um, but everybody knew that deep learning was a thing that you got a math or a computer science PhD.

You know, there was one of five labs that could give you the appropriate skills, and then you would join. Yeah. Basically from one of those labs you might be able to write some papers. So yeah, the idea that normal people could use that technology to do good work was considered kind of ridiculous when we started it.

And we weren't sure if it was possible either, but we kind of felt like we had to give it a go- ... 'cause the alternative was we were pretty sure that deep learning was on its way to becoming, you know, the most or one of the most technol- you know, important technologies in human history, and if the only people that could use it were a handful of computer science PhDs, that seemed like, A, a big waste, and B, kind of dangerous.

**Swyx** [5:28]
Yep. And, you know, well, I, I just wanted to know one thing on your bio that at Kaggle you were also the top-ranked participant in both 2010 and 2011. So sometimes you see a lot of founders running companies that are not really in touch with the problem, but- ...

you were clearly building something that you knew a lot about, which is awesome. Um, and even, yeah, talking about deep learning, you created, published a paper on ULMFit, which was kind of the predecessor to multitask learning and a lot of the groundwork that then went into transformers.

**Alessio** [6:00]
I read back on the paper and you trained this model AWD-LSTM, which I-I mean, I did the math and it was like 24 to 33 million parameters depending on what training dataset you use. Today, that's kind of like not even small.

It's like super small. Um, what were some of the kind of like contrarian takes that you had at the time and maybe set the stage a little bit for the rest of the audience on what was kind of like the state of the art-

**Jeremy Howard** [6:29]
Sure

**Alessio** [6:29]
... so to speak at the time and what people-

**Jeremy Howard** [6:31]
The whole thing-

**Alessio** [6:31]
... were working towards.

**Jeremy Howard** [6:32]
Yeah, the whole thing was a contrarian take, you know. So, okay, so we started fastai, um, my wife and I, and we th-- yeah, so we're trying to think, okay, how do we make it more accessible? So, so 20-- well, when we started thinking about it, it was probably 2015, and then 2016 we started doing something about it.

Why is it inaccessible? Okay, well, A, no one knows how to do it, um, other than a few number of people. And then when we ask those few number of people, "Well, how do you actually get good results?"

They would say like: "Oh, it's like, you know, a box of tricks that aren't published, so you have to join one of the, you know, labs and learn the tricks." So bunch of unpublished tricks. Um, not much software around, but, you know, thankfully there was Theano, um, and, you know, wrappers and particularly Lasagne, the wrapper.

Um, but yeah, not, not much software around, not much in the way of datasets. You know, uh, very hard to get started in terms of the compute, like how do you get that set up? So, you know, everything was kind of inaccessible.

And, you know, as we started looking into it, we had a key insight, which was like, you know what? Um, most of the compute and data for image recognition, for example, we don't need to do it. You know, there's this thing which nobody knows about, nobody talks about called transfer learning, where you take somebody else's model where they already figured out like how to detect edges and gradients and corners and text and whatever else, and then you can fine-tune it to do the thing you wanna do.

And we thought that's the key. That's the key to, to becoming more accessible in terms of compute and data requirements. So when we started fastai, we focused from day one on transfer learning. Lesson one, in fact, was transfer learning, literally lesson one.

Uh, it was something not normally even mentioned in... I mean, there wasn't much in the way of courses. Um, you know, the qu- basically the-- really the courses out there were PhD programs that had happened to have recorded their lessons.

They would rarely mention it at all. We wanted to show how to do four things that seemed really useful. You know, work with vision, work with tables of data, work with kind of recommendation systems and collaborative filtering, and work with text.

'Cause we felt like those four kind of modalities covered a lot of the stuff that, you know, are useful in real life. And no one was doing anything much useful with text. Everybody was talking about Word2Vec, you know, like king plus queen minus woman and blah, blah, blah.

And it's like cool experiments, but nobody's doing anything like useful with it. NLP was all like lemmatization and stop words and topic models-

**Alessio** [9:27]
Mm-hmm

**Jeremy Howard** [9:28]
... and bigrams and SVMs, and it was really academic and not practical. But yeah, and to be honest, I'd been thinking about this crazy idea for nearly 30 years since I had done cognitive science at university, where we talked a lot about, um, the, the Sel's Chinese room experiment.

This idea of like, what if there was somebody that could kind of like knew all of the symbolic manipulations required to answer questions in Chinese, but they didn't speak Chinese. They were kind of inside a room with no other way to talk to the outside world other than taking in slips of paper with Chinese written on them, and then they do all their rules, and then they s- pass back a piece of paper with Chinese back.

And this room with a person in is actually fantastically good at answering any question you give them written in Chinese. You know, do they, do they understand Chinese? Um, and, uh, is this, you know, something that's intelligently working with Chinese?

Ever since that time, I'd say the, the, the most thought-- to me, the most thoughtful and compelling philosophical response is yes. Um, you know, uh, intuitively it feels like no, because b- but that's just because we can't imagine such a large kind of system.

Um, but, you know, if it looks like a duck and acts like a duck, it's, it's a duck, you know, or to all intents and purposes. And so I always kind of thought, you know... So this is basically a, a s- a, a kind of an as- analysis of the limits of, of text.

And I kind of felt like, yeah, if something could ingest enough text and could use the patterns it saw to then generate text in response to text, it could appear to be intelligent, you know. Um, and whether that means it is intelligent or not is a different discussion and not one I find very interesting.

Yeah. And then when I came across neural nets when I was about 20, you know, and I learned about the universal approximation theorem and stuff, and I started thinking like: Oh, I wonder if like a neural net could ever get big enough and take in enough data to be a, a Chinese room experiment.

You know, with that background and this kind of like interest in transfer learning, you know, I'd been thinking about this thing for kind of 30 years, and I thought like: Oh, I, I wonder if we're there yet, you know.

'Cause we have a lot of text. Like, I can literally download Wikipedia, which is a lot of text. And I thought, you know, how would something learn to Kind of answer questions or, you know, respond text. And I thought, "Well, what if we used a language model?"

So language models are already a thing. You know, they were not a popular or well-known thing, but they were a thing. But language models exist at this idea that you could train a model to fill in the gaps, or actually in those days it wasn't fill in the gaps, it was finish a string.

And in fact, uh, Andrej Karpathy did his fantastic, uh, RNN demonstration from this at a similar time, where he showed like you can have it ingest Shakespeare and it will generate something that looks a bit like Shakespeare.

**Alessio** [12:35]
Mm-hmm.

**Jeremy Howard** [12:37]
I thought, "Okay, so if I do this at a much bigger scale using, using all of Wikipedia, what would it need to be able to do to finish a sentence in Wikipedia, uh, effectively, to do it quite accurately quite often?"

I thought, "Geez, it, it would actually have to know a lot about the world. You know, it'd have to know that there is a world and that there are objects and that objects relate to each other through time and cause each other to react in ways and that causes precede effects and that, you know, and there are animals and there are people and that people can be in certain positions during certain time frames."

And then you could... You know, all that together you can then finish a sentence like, "This was signed into law in 2016 by US President X," and it would fill in the middle.

**Alessio** [13:23]
Mm-hmm.

**Jeremy Howard** [13:24]
You know? So that's why I tried to create a, what in those days was considered a big language model trained on the entirety on Wikipedia, which is-- That was b- you know, a bit unheard of. And my interest was not in,

you know, just having a language model. My interest was in like what latent capabilities would such a, a system have that would allow it to finish those kind of sentences. Because I was pretty sure based on our work with transfer learning and vision that I could then suck out those latent capabilities by transfer learning, you know, by fine-tuning it on a task dataset or whatever.

So we generated this three-step system. So step one was train a language model on a big corpus. Step two was fine-tune a language model on a more curated corpus. And step three was further fine-tune that model on a task.

Um, and of course that's what-

**Alessio** [14:20]
Got it

**Jeremy Howard** [14:20]
... everybody still does today, right? That's what ChatGPT is. And so-

**Alessio** [14:24]
Mm-hmm

**Jeremy Howard** [14:25]
... the first time I tried it, within hours I had a new state-of-the-art academic result on IMDb. And I was like, "Holy shit, it does work." And so you asked to what degree was this kind of like pushing against the, you know, established wisdom.

You know, every way. Like, the reason it took me so long to try it was 'cause I asked all my friends in NLP if this could work, and everybody said, "No, it definitely won't work." It wasn't like a maybe.

Everybody was like, "It definitely won't work. NLP is much more complicated than vision. Language is a much more vastly complicated domain. You know, and you've got problems like the grounding problem. We know from like philosophy and theory of mind that it's actually impossible for it to work.

So, um, yeah, so don't waste your time."

**Alessio** [15:11]
Jeremy-

**Jeremy Howard** [15:11]
Yeah

**Alessio** [15:11]
... had people not tried because it was like too complicated to actually get the data and like set up the training, or like were people just lazy and kind of like, "Hey, this is just not gonna work"?

**Jeremy Howard** [15:21]
No, everyone was lazy. So like, so the person I thought at that time who... You know, there were two people I thought at that time actually who were the strongest at language models were, um, Stephen Merity and Alec Radford.

And, um, at the time I didn't know Alec, but, um, I... after we had both-- after I'd released ULMFit and he had released GPT, um, I organized a chat for both of us with Cade Metz of The New York Times, and Cade Metz answered...

So and, and Alec answered this question for Cade, and Cade was like, "So how did, you know, GPT come about?" And he said, "Well, I was pretty sure that pre-training on a general large corpus wouldn't work, so I hadn't tried it.

And then I read ULMFit, and turns out it did work. And so I did it, you know, bigger and it worked even better." And similar with, with Stephen. You know, I asked Stephen Merity like, "Why don't we just find-- you know, take your AWDS2LM and like train it on all of Wikipedia and fine-tune it?"

And he was kinda like, "Eh. I don't think that's gonna really fly." Like two years before, I d- did a very popular talk at KDD, the conference, where everybody in NLP was in the audience. I recognized half the faces, you know.

And I told them all this, "I'm sure transfer learning is the key. I'm sure ImageNet, you know, is gonna be an NLP thing as well." And you know, everybody was interested, and, um, people asked me questions afterwards. But n- not just...

Yeah, nobody followed up because everybody knew that it didn't work. I mean-

**Alessio** [17:02]
Yeah

**Jeremy Howard** [17:02]
... even like... So we were scooped-

**Alessio** [17:05]
Mm-hmm

**Jeremy Howard** [17:05]
... a little bit by, um, Dai and Le, Quoc Le at, at Google. They had, they had, uh, already... I didn't even realize this, which is a bit embarrassing. They had already done a large language model and fine-tuned it.

But, um, again, they didn't create a general purpose large language model on a general purpose corpus. They only ever tested a domain-specific corpus. And I haven't spoken to Quoc actually about that, but I assume that the reason was the same.

It probably just didn't occur to them-

**Alessio** [17:36]
Mm-hmm

**Jeremy Howard** [17:36]
... that the, the general approach could work.

**Alessio** [17:39]
Mm-hmm.

**Jeremy Howard** [17:39]
So maybe it was that kind of 30 years of mulling over the, this old Chinese room experiment that had convinced me that it probably would work. I don't know.

**Alessio** [17:48]
Yeah. Interesting. I just, uh, dug up a, Alec, um, announcement tweet from 2018. He said, "Inspired by CoVE, ELMo, and ULMFit, we showed a single transformer language model can be fine-tuned to a wide variety." Um, it's interesting because, you know, today people think of AI as the leader, kind of like- Kind of like the research lab pushing forward the field.

What was that at the time? You know, like, kinda like going back-

### Validation Loss

**Jeremy Howard** [18:12]
Yeah

**Alessio** [18:12]
... five years. People think-

**Jeremy Howard** [18:14]
Yeah

**Alessio** [18:14]
... OpenAI as an overnight success, but obviously it took a while.

**Jeremy Howard** [18:17]
Yeah, yeah. No, uh, I mean, absolutely. And I'll say, like, you know, it's interesting that it mentioned ELMo because in some ways that was kind of diametrically opposed to, to ULMFit. You know, there was these kind of like...

So there was a lot of, um, there was a lot of activity at the same time as ULMFit's release. So there was, um... So before it, as Brian McCann, I think, at Salesforce had come out with this neat model that did a kind of multitask learning, but again, they didn't create a general fine-tune language-

**Alessio** [18:50]
Mm-hmm

**Jeremy Howard** [18:50]
... model first. There was ELMo, um, which I think was a lit- you know, actually quite a few months after the first ULMFit example, I think. Um, but yeah, there was a bit of this stuff going on, and the problem was everybody was doing...

And particularly after GPT came out, then everybody wanted to focus on zero-shot and few-shot learning. You know, everybody hated fine-tuning. Everybody hated transfer learning. And like I literally did tours trying to get people to start doing transfer learning.

And peop- you know, nobody was interested, particularly after GPT showed such good results with, with few, zero-shot and few-shot learning. And so I actually feel like we kinda went backwards for years. And, and to be honest, I mean, I'm a bit sad about this now, but I, I kind of got so disappointed and dissuaded by, like...

It felt like these bigger lab, much bigger labs, you know, like fastai had only ever been just me and Rachel, were getting all of this attention for an approach I thought was the wrong way to do it. You know, I was convinced was the wrong way to do it.

And so, yeah, for years, people were really focused on getting better at zero-shot and few-shot. And it wasn't until, you know, this key idea of like, whoa, let's take the ULMFit approach, but for step two, rather than fine-tuning on a kind of a domain corpus, let's fine-tune on an instruction corpus.

And then in step three, rather than fine-tuning on a reasonably specific task classification, let's fine-tune on a, on a RLHF task classification. And so that was really, that was really key, you know. So I was kind of like out of the NLP field for a few years there because, yeah, it just felt like, I don't know, pushing uphill against this-

**Alessio** [20:39]
Right

**Jeremy Howard** [20:39]
... vast tide, uh, which I was convinced was not the right direction, but who's gonna listen to me, you know? 'Cause I, as you said, I don't have a PhD, not at a university, or at least I wasn't then.

I don't have a big set of computers to fine-tune huge transformer models.

**Alessio** [20:56]
Yeah.

**Jeremy Howard** [20:56]
So yeah, it was definitely difficult. It's always been hard. You know, it's always been hard. Like, I've always been somebody who does not wanna build stuff on lots of big computers because most people don't have lots of big computers.

**Alessio** [21:11]
Mm-hmm.

**Jeremy Howard** [21:11]
And I hate creating stuff that most people can't use, you know. And also stuff that's created on lots of big computers has always been, like, much more media friendly. So like, it might seem like a recent thing, but actually throughout my 30 years in data science, the attention's always been on, you know, the big iron results.

So when I first started, everybody was talking about data warehouses, and it was all about Teradata, and it'd be like, oh, this big bank has this huge room full of computers, and they have like terabytes of data available, you know, at the press of a button.

And yeah, that's always what, what people wanna talk about, what people wanna write about. And then, of course, students coming out of their PhDs and stuff, that's where they wanna go work 'cause that's where they read about. And to me, it's a huge distraction, you know, because, like I say, most people don't have unlimited compute.

**Alessio** [22:10]
Mm-hmm.

**Jeremy Howard** [22:10]
And I wanna help most people-

**Alessio** [22:12]
Mm-hmm

**Jeremy Howard** [22:12]
... not the small subset of the most well-off people.

**Alessio** [22:17]
Yeah. That's awesome. And it- it's great to hear, you know, you do such a great job educating, that a lot of times you're not telling your own story, you know. So I love, I love this conversation. Um, and the other thing before we jum- jump into fastai, actually, you know, a lot of people that I know, they run across a new architecture and what not.

### Motivation

**Alessio** [22:36]
They're like, "I gotta start a company and raise a bunch of money and do all of this stuff." Instead, you were like, "I want everybody to have access to this." The, the... Why was that the case for you?

Was it because you already had like a successful, you know, venture in like FastMail, and you were more interested in that? What was the, the reasoning?

**Jeremy Howard** [22:53]
That's a really good question. So I guess the answer's yes. It is. That's the reason why. Um, so when I was a teenager, I thought it would be really cool to like have my own company. You know, I didn't know the word startup.

I didn't know the word entrepreneur. I didn't know the word VC. And I didn't really know what any of those things were really until after we started Kaggle, to be honest. Even though the way it started to what we'd now call startups, I just thought they were just small businesses.

You know, they were just companies. Um, so yeah. So those two companies were FastMail and Optimal Decisions. FastMail was the first kind of, uh, synchronized email provider for non-businesses. So something you can get your same email at home, on your laptop, at work, on your phone, whatever.

And then, uh, Optimal Decisions, um, uh, invented a new approach to insurance pricing, something called, uh, profit-optimized insurance pricing. So I sold both of those companies, um, you know, after 10 years. And at that point, I had achieved the thing that as a teenager I had wanted to do.

You know. It took a lot longer than it should have 'cause I spent way longer in management consulting than I should have 'cause I got caught up in that stupid rat race. But, you know, eventually I got there.

And I remember my mom saying to me, "Oh, you must be so proud." You know, 'cause she remembered my, my dream. She's like, "You've done it." And I kind of reflected and I was like, "I'm not. I'm not proud at all."

You know? Like, people quite liked FastMail. You know, it's quite nice to have synchronized email. It probably would've happened anyway. Um, yeah, I'm certainly not proud that I've helped some insurance companies suck more money out of their customers.

Yeah, no, I'm not proud. You know? It's, it's actually... I haven't really helped the world very much. Um, for, you know, maybe in the insurance case I've made it a little bit worse. I don't know. So yeah, I was determined to not waste more years of my life doing things, working hard to do things which I could not be reasonably sure would have a lot of value.

So, uh, you know, I took some time off. I wasn't sure if I'd ever work again, actually. I didn't particularly want to, 'cause it felt like, yeah, it felt like such a disappointment. Um, and but you know, and I didn't need to.

I had enough money. Like, I wasn't super rich, but I had enough money. I didn't need to work. And I certainly recognized that amongst the other people I knew who had enough money that they didn't need to work, they all worked ridiculously hard, you know, and constantly put themselves in extremely stressful situations.

And I thought, "I don't wanna be one of those idiots who's tied to, you know, buying a bigger plane than the next guy or whatever." You know, Kaggle came along, and I mainly kind of did that just 'cause it was fun and interesting and got to hang out with interesting people.

But, you know, with, with fastai in particular, you know, Rachel and I had a very explicit, you know, long series of conversations over a long period of time about like, well, how can we be the most helpful to society as a whole, and particularly to those people who maybe need more help, you know?

And so we definitely saw the world going in a potentially pretty dystopian direction if the world's most powerful technology was c- controlled by a s- small group of elites. Um, so we thought, yeah, we should focus on trying to help that not happen.

Um, you know, sadly, it looks like it still is likely to happen, but, I mean, I feel like we've, we've helped make it a little bit less likely, so we've done our bit.

**Swyx** [26:40]
You've shown that it's possible, and I think, uh, and, and I think y- uh, your constant advocacy, your courses, um, your research that you publish, you know... You know, just the other day you published, uh, a, a finding on, on, um, you know, learning that, that I think is, is still something that people are still talking about, uh, quite a lot.

I think that that is, um, the, the origin story of a lot of people who are gonna be, you know, little Jeremy Howards furthering your mission with, uh... You know, you don't have to do everything y- by yourself is what I'm saying.

**Jeremy Howard** [27:10]
No, definitely. Definitely. You know, that was a, that was a big takeaway from, like, Enlytic, was at Enlytic it definitely felt like we had to do everything ourselves, and I kind of... I wanted to solve medicine. I said, "Yeah, okay, solving medicine's actually quite difficult, and, uh, I can't do it on my own, and there's a lot of other things I'd like to solve, and I can't do those either."

So that was, that was definitely the other piece, was like, yeah, you know, can we create an, an army of, of, of passionate domain experts who can change their little part of the world? And that's definitely happened. Like, I find nowadays at least half the time, probably quite a bit more, that I get in contact with somebody who's done really interesting work in some domain.

Most of the time I'd say they say, "Yeah, I got my start with fastai." Um, so it's definitely... I can, I can see that. And I, I also know from talking to folks at places like Amazon and Adobe and stuff, which, you know, there's lots of alumni there, and they say, "Oh my God, I got here, and, like, half of the people are fastai alumni."

Um, so-

**Swyx** [28:13]
It's fantastic

**Jeremy Howard** [28:14]
... uh... Yeah, actually Andrej Karpathy grabbed me when I saw him at Europe a few years ago, and he's like, "I have to tell you, thanks to the fastai courses, when people come to Tesla and they need to know more about deep learning, we always send them to your course."

And the OpenAI Scholars program was doing the same thing. So it's kind of like, yeah, it's had a, a surprising impact. You know, uh, that's, that's just one of like three things we do is the course, you know?

**Swyx** [28:41]
Yes.

**Jeremy Howard** [28:41]
And it's, it's, it's only ever been at most two people, either me and Rachel or me and Sylvere. Nowadays it's just me. Uh, so yeah, I think it shows you don't necessarily need a huge amount of money and a huge team of people to, to make an impact.

**Swyx** [28:58]
Yep. Uh, so just to in- reintroduce fastai for, uh, people who may not have, uh, dived into it much, um, there is the, the courses that you do. There is the library that is, uh-

**Jeremy Howard** [29:10]
Yep

**Swyx** [29:10]
... that is, um, very, uh, well loved. And, uh, I kind of think it, think of it as a, a nicer layer on top of PyTorch that-

### Library

**Jeremy Howard** [29:17]
Mm-hmm

**Swyx** [29:17]
... people should start with by default, um, and-

**Jeremy Howard** [29:19]
Yeah

**Swyx** [29:19]
... use it as the basis for a lot of your courses. Um-

**Jeremy Howard** [29:22]
Yep

**Swyx** [29:22]
... and then you have, you have, like, uh, NBDev, which I, I don't know. Is, is that the third one?

**Jeremy Howard** [29:27]
Oh, so the three areas were, uh, research, software, and, and courses.

**Swyx** [29:33]
Oh, sorry. I was going by-

**Jeremy Howard** [29:34]
So then in s-

**Swyx** [29:34]
In terms of software

**Jeremy Howard** [29:35]
... yeah, so then in software-

**Swyx** [29:36]
Yeah

**Jeremy Howard** [29:36]
... you know, um, fastai is the main thing, but NBDev is not far behind. But then there's also things like FastCore, GHAPI, I mean, dozens of open source projects that I've created, and some of them have, um, been pretty popular, and some of them are still a little bit hidden, actually.

I should... Some of them I should try to do a better job of telling people about.

**Swyx** [30:01]
What are you, what are you thinking about? Yeah, what, what's on the horizon? Just like when-

**Jeremy Howard** [30:04]
Oh, I don't know. Just like little things. Like, for example, for working with EC2 and AWS, I created a fast EC2 library, which I think is, like, way more convenient and nice to use than anything else out there.

**Swyx** [30:14]
Yeah. Okay.

**Jeremy Howard** [30:15]
And it's literally got a whole auto-complete, dynamic auto-complete that works both on the command line and in notebooks. It'll, like, auto-fit your instance names and- And everything like that. You know, just little things like that. I, I try to make like...

When I work with some domain, I try to make it like, I want to make it as enjoyable as possible for me to do that. So I always try to kind of like, like with GHAPI, for example, um, I think that GitHub API is incredibly powerful, but I didn't find it good to work with 'cause I didn't particularly like the libraries that were out there.

So like GHAPI, like FastEC2, it like auto-completes both at the command line or in a notebook or whatever, like literally the entire GitHub API. Um, the entire thing is like, I think it's like less than 100K of code because it, um, actually is, as far as I know, the only one that, that grabs it directly from the official OpenAPI spec that GitHub produces.

**Swyx** [31:14]
Interesting.

**Jeremy Howard** [31:14]
And like if you're in GitHub and you just type @api, um, you know, auto-complete API, uh, method and hit enter, it prints out the docs or the six brief docs and then gives you a link to the actual documentation page.

Um, you know, GitHub actions I can write now in Python, which is just so much easier than writing them in TypeScript and stuff. So, you know, just little things like that.

**Swyx** [31:41]
I think that's a approach that more, uh, wish more developers, uh, took to publish some of their work along the way. Um, is... You described the, the third arm of fastai as research. It's not something I d- I see often.

Obviously, obviously you do, you do do some research and, um, um, how do you, how do you run your research? Uh, what, what are your research interests?

**Jeremy Howard** [32:00]
Yeah. So research is what I spend the vast majority of my time on. Um, and the artifacts that come out of that are largely software and courses, you know. So to me, the main artifact shouldn't be papers, 'cause papers are things read by a small exclusive group of people.

You know, to me, the main artifacts should be like something teaching you people, "Here's how to use this insight and here's software you can use that builds it in." Um, so I think I've only ever done three first-person papers in my life, you know, and they were...

And none of those are ones I wanted to do, you know. They were all ones that... Like, so one was ULMFit, where Sebastian Ruder reached out to me after seeing the course and said like, "You have to publish this as a paper."

You know? And, uh, and he said, "I'll write it." I was like, "Oh," you know. And he said, "I want to write it, 'cause if I do, I can put it on my PhD, and that would be great."

And it's like, "Okay, well, I want to help you with your PhD, and I... That sounds great." So like, you know, one was the masks paper, which just had to exist and nobody else was writing it. And then the third was, um, the fastai library paper, which again, um, um, somebody reached out and said, "Please, please write this.

We will waive the fee for the journal and everything and-"

**Swyx** [33:20]
Excellent

**Jeremy Howard** [33:21]
... help you get it through publishing and stuff." So yeah, so I, I, I don't... Other than that, I've never written a first author paper. So the research is like... Well, so for example, you know, DAWNBench was a competition which Stanford ran a few years ago.

### DAWNBench

**Jeremy Howard** [33:37]
Um, it was kind of the first big competition of like who can train neural nets the fastest rather than the most accurate. And, um, and, uh, specifically it was who can train ImageNet the fastest. And again, this was like one of these things where it was created by necessity.

So Google had just released their TPUs, and so I heard from my friends at T- at Google that they'd put together this big team to, to smash DAWNBench so that they could prove to people that they had to use Google Cloud and use their TPUs and show how good their TPUs were.

And we kind of thought, "Oh shit, this would be a disaster if they do that, because then everybody's gonna be like, 'Oh, deep learning's not accessible.' You know, to actually be good at it, you have to be Google and you have to use special silicon and so."

So, you know, we, we only found out about this 10 days before the competition finished. Um, but, you know, we basically got together an emergency bunch of our students, and Rachel and I, and sat for the next 10 days and just tried to crunch through and, um, try to use all of our best ideas, um, that had come from our research.

And so particularly progressive resizing, which is basically train mainly on small things, um, train on non-square things, um, you know, stuff like that. And so yeah, we ended up winning, um, thank God. And so, you know, we turned it around from being like, uh, like, "Oh shit, you know, this is gonna show that you have to be Google and have TPUs," to being like, "Oh my God, even the little guy can, can do deep learning."

Um, so that's an example of the kind of like research artifacts we do. And, uh, yeah, so all of my research is always how do we do more with less? You know, so how do we get better results with less data, with less com- less compute, with less complexity, with less education, you know, stuff like that.

So ULMFit's obviously a good example of-

**Swyx** [35:37]
Yes

**Jeremy Howard** [35:37]
... of that.

**Swyx** [35:39]
Uh, and most recently you published, uh, Can LLMs Learn From a Single Example? Um-

**Jeremy Howard** [35:44]
Hmm

**Swyx** [35:44]
... w- maybe could you tell the story a little bit behind that? And maybe that goes a little bit too far into the learning on very low resource- The literature.

**Jeremy Howard** [35:53]
Yeah. Yeah. So, um, me and my friend Jono Whitaker, um, basically had been, um, playing around with this fun Kaggle competition, which is actually still running as we speak, which is, um, can you create a model which can answer multiple choice questions about anything that's in Wikipedia?

And the thing that makes it interesting is that your model has to run on Kaggle within nine hours, and Kaggle's very, very limited. So you've only got 14 gig RAM, only two CPUs, and a small, very old GPU.

Um, so this is cool, you know, if you could do well at this, and this is a good example of like, oh, you can do more with less. Um, so yeah, John O and I were playing around with fine-tuning, of course, transfer learning, um, uh, pre-trained language models, and we saw this like...

So we always, you know, plot our losses as we go. So here's another thing we created, or actually, um, Sylvain Gugger, when he worked with us, created called Fast Progress, which is kind of like TQDM, but we think a lot better.

So we look at our fast progress curves, and they kind of go down, down, down, down, down, down, down a little bit, little bit, little bit, and then suddenly go clunk, and they drop, and then down, down, down, down, down a little bit, and then suddenly clunk, they drop.

We're like, "What the hell? These clunks are occurring at the end of each epoch." So normally in deep learning, this would be... This... You know, I've seen this before, and it's always been a bug. It's always turned out that like, oh, we accidentally forgot to turn on eval mode during the validation set, so it was actually learning then.

Or, oh, acci- acci- we accidentally were calculating moving average statistics throughout the epoch. So, you know, so it's reset the moving average or whatever. And so we were using Hugging Face trainer. So, you know, I, I did not give my friends at Hugging Face the benefit of the doubt.

I thought, "Oh, they've fucked up Hugging Face trainer." You know, idiots. Well, you all use the fastai trainer instead, so we switched over to Learner. We still saw the clunks, and you know, that's... Yeah. It, it shouldn't really happen because semantically speaking, an epoch isn't like...

It's not a thing. You know, like nothing happens. Well, nothing's meant to happen when you go from ending one epoch to starting the next one.

**Swyx** [38:19]
Mm.

**Jeremy Howard** [38:20]
Um, so there shouldn't be a clunk. You know, so I kind of asked around on the open source Discords, and I was like, "What's going on here?" And everybody was just like, "Oh, that's just what, that's just what these training curves look like.

### Memorization

**Jeremy Howard** [38:33]
Ours all look like that. Don't worry about it." And I was like, "Oh, are you all using Trainer?" "Yes." "Oh, well, must... There must be some bug with Trainer." And I was like, "Well, we, we also saw it in Learner."

And somebody else was like, "No, we've got our own trainer. We get it as well." They're just like, "Don't worry about it. It's just something we see. It's just normal." And I can't do that. I can't just be like, "Here's something that's like in the previous 30 years of neural networks, nobody ever saw it, and now suddenly we see it, so don't worry about it."

Like, I just, I have to know why. You know?

**Swyx** [39:03]
Can, can I clarify, this is all-

**Jeremy Howard** [39:04]
Yeah

**Swyx** [39:04]
... was everyone that you're talking to, were they all seeing it for the same dataset or in different datasets?

**Jeremy Howard** [39:09]
Data... Different datasets, different trainers. They're just like, "No, this is just, this is just what it looks like when you fine-tune language models. Don't worry about it." You know. As I say, I've been-

**Swyx** [39:18]
You've never seen this before. Yeah.

**Jeremy Howard** [39:19]
I hadn't seen it before, but I'd been kind of like... As I say, I, you know, I kept working on them for a couple of years after ULMFit, and then I kind of moved onto other things, partly out of frustration.

So I hadn't been fine-tuning, you know, um... I mean, Llama's only been out for a few months, right? But I, I, I, I, I, I wasn't one of those people who jumped straight into it, you know. So I was relatively new to the kind of Llama fine-tuning world, where else these guys had been, you know, doing it since day one.

Um, which is only a few months ago, but it's still quite a bit of time. So, so yeah, they're just like, "No, this is all what we see. Don't worry about it." So yeah, I, I've got a very kind of like, I don't know how, I've just got this brain where I have to know why things are.

And so I kind of, I ask people like, "Well, why, why do you think it's happening?" And they'd be like, "Oh, pretty obviously, 'cause it's like memorized the dataset." It's just like, that can't be right. It's only seen it once.

Like, look at this, the, the loss has dropped by .3. .3. Which is like, basically it, it knows the answer. Um, they're like, "No, no, it's just... It, it is. It's, it's memorized the dataset." So yeah, so look, John O and I did not discover this, and John O and I did not come up with a hypothesis.

You know, I guess we were just the ones, I guess, who had been around for long enough to recognize that like this, this isn't how it's meant to work. And so we, we, you know, and so we went back and we're like, "Okay, let's just run some experiments," you know, 'cause nobody since we've actually published anything about this.

Um, well, not quite true. Some people have published things, but nobody ever actually stepped back and said like, "What the hell?" You know. "How can this be possible? Is it possible? Is it what's happening?" And so yeah, we created a bunch of experiments where we basically predicted ahead of time.

It's like, okay, if this hypothesis is correct, that it's memorizing the training set, then we ought to see blah under conditions blah, but not under these conditions. And so we ran a bunch of experiments, and all of them supported the hypothesis that it was memorizing the dataset in a single...

seeing it once. Um, and it's a pretty big dataset, you know. Um, which in hindsight, it's not totally surprising because the theory, remember, of the ULMFit theory was like, well, it's kind of creating all these latent capabilities to make it easier for it to predict the next token.

So if it's got all this kind of latent capability, it ought to also be really good at compressing new tokens because it can immediately recognize it as like, oh, that's just a version of this. So-

**Swyx** [41:55]
Oh

**Jeremy Howard** [41:56]
... it's, it's not so crazy, you know. But it is, um... It requires us to rethink everything because like... And nobody knows. Like, okay, so how do we fine-tune these things? Because like does it even matter? Like maybe it's fine.

Like maybe it's fine that it's memorized the dataset after one go, and you do a second go, and okay, the validation loss is terrible because it's now really overconfident. That's fine. Don't, you know, don't... I keep telling people, "Don't track validation loss, track validation accuracy," um, 'cause at least that, that will still be useful.

Um, there's another thing that's got lost since ULMFit. Nobody tracks accuracy of language models anymore. Um, but you know, it'll still, it'll still keep learning, and it does. It does keep improving. But- Is it worse? You know, like, is it like now that it's kind of memorized it, it's probably getting a less strong signal, you know?

Um, I don't know. So I still don't know how to fine-tune language models properly, and I haven't found anybody who feels like they do. Like, nobody really knows whether this memorization thing is... It's probably a feature in some ways.

There's probably some things that you can do usefully with it. It's probably-

**Swyx** [43:10]
It, it, um-

**Jeremy Howard** [43:11]
But, um, yeah, I have a feeling it's messing up training dynamics as well.

**Swyx** [43:15]
And does it come at the cost of catastrophic forgetting as well, right? Like, which is the other side of the coin.

### Pre-training

**Jeremy Howard** [43:20]
Um, it does, um, to some extent. Like we know it does. Like look at Code Llama, for example. So Code Llama was a... I think it was like a 500 billion token fine-tuning of Llama 2 using code, and also pros about code, uh, that Meta did.

And, um, honestly, they kind of blew it because, uh, Code Llama is good at coding but is bad at everything else, you know.

**Swyx** [43:45]
Yeah. Okay.

**Jeremy Howard** [43:45]
And it used to be good. Yeah, I was pretty sure it was... Like, every-- before they released it, me and lots of people in the open source Discords were all like, "Oh my God, you know, we know this is coming.

Jan Lehnardt's saying it's coming. I sh- I hope they kept at least like 50% non-code data, 'cause otherwise it's gonna forget everything else." And they didn't. Only like .3 of their, .3% of their epochs were non-code data. So it did, it forgot everything else.

So now it's good at code and it's bad at everything else.

**Swyx** [44:12]
Ugh.

**Jeremy Howard** [44:13]
So we definitely-

**Swyx** [44:14]
It's fixable

**Jeremy Howard** [44:14]
... have catastrophic forgetting. It's fixable, just somebody has to do-

**Swyx** [44:18]
Yeah

**Jeremy Howard** [44:18]
... you know, somebody, somebody has to spend their time training a model on a, a good mix of data. Like, so, okay, so here's the thing. Um, even though I originally created the three-step approach that everybody now does, my view is it's actually wrong and we shouldn't use it.

Um,

um, and that's because people are using it in a way different to why I created it. You know, I created it thinking that the task-specific models would be more specific. You know, it's like, oh, this is like a, a sentiment classifier was an example of a task, you know.

Uh, uh, but the tasks now are like a, um, you know, RLHF, which is basically like answer questions that make people feel happy about your answer. So that's a much more general task, and it's a really cool approach.

And so we see, for example, RLHF also breaks models like, you know, like GPT-4, RLHFed. We, we know from kind of the, the work that Microsoft did, you know, the, the pre-- the, the earlier less aligned version was better.

Um, and these are all kind of examples of catastrophic forgetting. And so to me, the right way to do this is to fine sha- fine-tune language models, is to actually throw away the idea of fine-tuning. There's no such thing.

There's only continued pre-training. Uh, and pre-training is something where from the very start, you in- try to include all the kinds of data that you care about, all the kinds of problems that you care about: instructions, exercises, code, general purpose document completion, whatever.

And then as you train, you gradually curate that. You know, you gradually make that higher and higher quality and more and more specific to the kinds of tasks you want it to do. Um, but you never throw away any data.

You always keep all of the data types there in reasonably high quantities. Um, you know, maybe the quality filter, you stop training on low-quality data, 'cause that's probably fine to forget how to write badly maybe. Um, so yeah, that's now my view, is I think ULMFit is the wrong approach.

**Swyx** [46:36]
Mm-hmm.

**Jeremy Howard** [46:36]
Um, and that's why we're seeing a lot of these, uh, you know, so-called alignment tasks and this view of like, oh, a model can't both code and do other things. And, you know, I, I think it's actually 'cause people are training them wrong.

**Swyx** [46:49]
Hmm. Well, I think you have a clear anti-laziness approach. I think other people are not as, um, good-hearted, you know. They're like, "Hey, they told me this thing works, and if I release a model this way, people will appreciate it, I'll get promoted, and I'll kind of make, make more money."

Uh-

**Jeremy Howard** [47:07]
Oh, absolutely

**Swyx** [47:08]
... which is how humans are.

**Jeremy Howard** [47:08]
Yeah.

**Swyx** [47:09]
Yeah.

**Jeremy Howard** [47:09]
And it's not just money. It's like this is how citations work most-

**Swyx** [47:12]
Mm.

**Jeremy Howard** [47:12]
... most badly, you know. So if you wanna get cited, you need to write a paper that people in your field recognize as an advancement-

**Swyx** [47:19]
Mm-hmm

**Jeremy Howard** [47:20]
... on things that we know are good. And so we've seen this happen again and again. So like I say, like zero-shot and few-shot learning, everybody was writing about that. Or, you know, with, um, image generation, every- everybody just was writing about GANs, you know.

And I was trying to say like, "No, GANs are not the right approach," you know, and I showed, again, through research that we demonstrated in our videos that you can do better with than GANs much faster and with much less data.

And nobody cared because, again, like if you wanna get published, uh, you write a GAN paper that slightly improves this part of GANs and this tiny field, you'll, you'll get published, you know. So it's, yeah, it's not set up for real innovation.

Um, it's, you know, it's-- again, it's really helpful for, for me. You know, I, I have my own research lab with nobody telling me what to do, and I don't even publish, so it doesn't matter if I get citations.

So I just write what I think actually matters. Um, I wish there was... And, you know, and actually places like OpenAI, you know, the researchers there can do that as well. It's a shame, you know. I wish there was more academic open venues in which people can focus on, like, genuine innovation.

### Incentives & Communities

**Swyx** [48:39]
Uh, Twitter, which is- ... uh, unironically has, has become a, a little bit of that forum. Uh, I, I wanted to follow up on one thing that you mentioned, uh, which is that you checked around the open source Discords.

Uh, I don't know if it's, uh, too, uh, I don't know if it's, uh, kosher to ask, like what Discords are lively, uh, or useful right now. Um, I think that Something I, I definitely felt like I missed out on was the early days of Eleuther AI, which, where-

**Jeremy Howard** [49:05]
Yeah, yeah

**Swyx** [49:06]
... which is a fair hard bit. And-

**Jeremy Howard** [49:08]
Yeah

**Swyx** [49:08]
... uh, you know, like, what is the new Eleuther? And you would-- you actually shouted out the Alignment Lab AI Discord-

**Jeremy Howard** [49:13]
Mm-hmm

**Swyx** [49:13]
... in your blog post, and that was the first time I even knew. Like, I saw them on Twitter, and never knew they had a Discord, never knew that there was actually substantive discussions going on in there, and that you were an active member of it.

**Jeremy Howard** [49:23]
Okay, yeah. And then even then, if you do know about that, and you go there, it'll look like it's totally dead. And that's because unfortunately, nearly all the Discords, nearly all of the conversation happens in private channels, you know?

**Swyx** [49:33]
Ah, yeah.

**Jeremy Howard** [49:34]
Um, and that's I guess-

**Swyx** [49:36]
So how does, how does, how does someone get into that world? 'Cause it's obviously very, very, um, in- instructive, right?

**Jeremy Howard** [49:43]
Well, you could just come to the first AI Discord, which I'll be honest with you, it's less bustling than some of the others, but it's, it's not terrible. And so, like, at least... Yeah, to be fair, one of our most bustling channels is private.

I guess... So I'm just thinking why, like-

**Swyx** [50:02]
That's just the nature of quality discussion, right? It, it-

**Jeremy Howard** [50:04]
Yeah

**Swyx** [50:05]
... you want to have it in a closed space

**Jeremy Howard** [50:05]
... I guess when I think about it, like, I didn't have any private discussions on our Discord for years. Um, but there was a lot of people who came in with like, "Oh, I just had this amazing idea for AGI.

How- if you just thought about, like, if you imagine the AI as a brain, then we..." You know, that's just... I don't even wanna talk about it. You know, I don't wanna like... But you don't wanna be dismissive, whatever, and it's like, "Oh, well, that's an interesting comment, but maybe you should, like, try training some models first to see if that aligns with your intuition."

Like, "Oh, but how could I possibly learn?" It's like, "Well, we have a course. Just actually spend time learning." Like, uh, yeah. Anyway, and then it's like, okay, I know the people who always have good answers there, and so I created a private channel and put them all in it, and I got to admit, I- that's where I post more often 'cause there's much less, you know, flight of fancy views about how we could solve AGI, blah, blah, blah.

So there is a bit of that, but having said that, like, I think the bar is pretty low. Like, if you join a Discord, and you can hit the, like, participants or community or whatever button, you can see who's in it, and then you'll see at the top who, who the admins or moderators or people in the dev role are.

And, uh, just DM one of them and say like, "Oh, I... Here's my GitHub," or, "Here's some blog posts I wrote," you know. "I'm interested in talking about this," you know. "Can I join the private channels?" And-

**Swyx** [51:33]
Right

**Jeremy Howard** [51:33]
... I've never heard of anybody saying no. I will say, you know, uh, Eleuther's all pretty open, so you can do the Eleuther Discord still. You know, one problem with the Eleuther Discord is it's been going on for so long that it's like, it's very inside baseball.

**Swyx** [51:50]
Yeah. It's hard to join as a newcomer.

**Jeremy Howard** [51:51]
It's quite hard to get started.

**Swyx** [51:53]
Yeah.

**Jeremy Howard** [51:53]
Um, Carpa AI looks-

**Swyx** [51:58]
Oh, okay

**Jeremy Howard** [51:58]
... I think it's all open.

**Swyx** [52:00]
This just left, uh, Stability.

**Jeremy Howard** [52:02]
That's more accessible.

**Swyx** [52:04]
Yeah.

**Jeremy Howard** [52:04]
Um, uh, there's also, um, just recently, uh, Nous Research that does, like, the Hermes models and dataset-

**Swyx** [52:13]
Yes

**Jeremy Howard** [52:13]
... just, just opened. They've got some private channels, but it's pretty open, I think. Uh, you mentioned Alignment Lab. That one, it's all the interesting stuff is on private channels, so just ask. Um, if, if you know me, ask me, 'cause I've got admin on that one.

There's also, yeah, uh, OS Skunk Works. OS Skunk Works AI is a good Discord, which I think it's open. So they're... Yeah, they're all pretty good.

**Swyx** [52:41]
I, I don't want you to leak any, any, uh, you know, uh, Discords that don't want any publicity, but-

**Jeremy Howard** [52:46]
No, I mean, we all like-

**Swyx** [52:46]
... this is all helpful

**Jeremy Howard** [52:47]
... we all want people. Like, we all want people. We just, we just want people-

**Swyx** [52:51]
Quality discussion

**Jeremy Howard** [52:52]
... who, like, wanna build stuff.

**Swyx** [52:53]
Serious. Exactly.

**Jeremy Howard** [52:54]
You know?

**Swyx** [52:54]
Yeah.

**Jeremy Howard** [52:54]
Um, rather than people who... And, like, it's fine to not know anything as well, but if you don't know anything, but you wanna tell everybody else what to do and how to do it, that's annoying. If you don't know anything and wanna be told, like, "Here's a really small kind of task that as somebody who doesn't know anything, it's gonna take you a really long time to do, but it would still be helpful," then, and then you go and do it, that would be great.

The truth is, yeah, like, I don't know, maybe 5% of people who come in with great enthusiasm and saying that they wanna learn and they'll do anything, and then somebody says like, "Okay, here's some work you can do," almost nobody does that work.

So if you're somebody who actually does the work and follows up, you will massively stand out. That's an extreme rarity, and everybody will then want to help you do more work. So yeah. So just, um, yeah, just do work and people will want to support you.

**Swyx** [53:48]
Yeah. Our, our Discord used to be referral only for a long time.

**Jeremy Howard** [53:51]
Oh.

**Swyx** [53:51]
We didn't have a, a public invite, and then-

**Jeremy Howard** [53:54]
Oh

**Swyx** [53:54]
... we, we opened it and then kind of like channel gating. Um, yeah, a lot of people just wanna do... I remember it used to be like, you know, a forum moderator. It's like people just wanna do, like, drive-by posting, you know?

And like, they don't wanna help the community. They just wanna-

**Jeremy Howard** [54:07]
Yeah

**Swyx** [54:07]
... get their question answered.

**Jeremy Howard** [54:08]
I mean, the funny thing is our, um, our forum community does not have any of that garbage. You know? There's something specific about the low latency thing where people, like, expect an instant answer and... Yeah, where somehow in a, in a forum thread where they know it's, like, there forever, people are a bit more thoughtful.

But then the forums

are less active than they used to be because Discord has got more popular, you know? So it's, it's all a bit of a compromise. You know, running a healthy community is, yeah, it's, it's always a bit of a challenge.

**Swyx** [54:47]
All right, we got so many more things we wanna dive in, but I don't wanna keep you here four hours.

**Alessio** [54:51]
Uh, this is not the, the Lex Fridman podcast, we always like to say. Uh, one topic I would love to maybe chat a bit about is Mojo Modular. You know, Chris Lattner nominated-

### Mojo

**Jeremy Howard** [55:01]
Oh, yeah

**Alessio** [55:01]
... you on the podcast, so, uh, we wanna spend a little time there. You recently did a hacker's guide to language models, and you ran through everything from quantized model to, like, smaller models, larger models, and all of that.

Um-

**Jeremy Howard** [55:14]
Yeah, that's right

**Alessio** [55:14]
... but obviously modular has taken its own approach. Uh-

**Jeremy Howard** [55:18]
Mm-hmm

**Alessio** [55:18]
... yeah, what got you excited? I know you and Chris have been talking about this for, like, years and-

**Jeremy Howard** [55:22]
Oh, yeah

**Alessio** [55:22]
... a lot of the ideas you had, so.

**Jeremy Howard** [55:24]
Yeah, yeah, yeah, yeah. No, absolutely. So I met Chris, I think it was at the first TensorFlow Dev Summit, and I don't think he had even, like... I'm not sure if he'd even officially started his employment with Google at that point, so I, I don't know.

You know, certainly nothing had been mentioned. So I, you know, I admired him from afar with LLVM and Swift and whatever, and so I saw him walk into the courtyard at, at Google, just like, "Oh, shit, man. That's Chris Lattner.

I wonder if he would lower his standards enough to talk to me." Well, it's worth a try. So I got up my courage because, like, he... nobody was talking to him. He looked a bit lost and I, I wandered over and it's like, "Oh, you're Chris Lattner, right?"

It's like, "What are you doing here?" And he was like, "Yeah, yeah, I am." I was like, "Oh, and Jeremy Howard." He's like, "Oh, yeah, do you do some of this AI stuff?" And I was like, "Yeah, yeah, I, I like this AI stuff.

Uh, are you doing AI stuff?" He's like, "Well, I'm thinking about starting to do some AI stuff, yeah. I think it's gonna be cool." And I said, "Oh." So, like, I spent the next half hour just basically brain dumping all the ways in which AI was stupid to him, and he listened patiently, and I thought he probably wouldn't even remember or care or whatever.

But, um, yeah, then I kind of, like, I guess I re-caught up with him a few months later, and he's like, "I've been thinking about everything you said in that conversation," and he, like, narrated back his response to every part of it, the projects he was planning to do, and it's just like, "Oh, this dude follows up.

Holy shit." Uh, and I was like, "Wow, okay." And he was like, "Yeah, so we're gonna create this new thing called Swift for TensorFlow, and it's gonna be like... it's gonna be a compiler with auto differentiation built in, and blah, blah, blah."

And I say, "Uh, wait, why would that help? You know, why would you..." And he was like, "Okay, with a compiler during the forward pass, you don't have to worry about saving context, you know, 'cause it'll ought to be optimized in the backward..."

But I was like, "Oh my God," 'cause I didn't really know much about compilers. Just, just that, you know, I spent... enough to kind of, like, understand the ideas, but it hadn't occurred to me that a compiler basically solves a lot of the problems we have as end users.

I was like, "Wow, that's amazing. Okay. You do know, right, that nobody's gonna use this unless it's, like, usable?" And it's like, "Yeah, I know, right? So I was thinking you should create, like, a fastai for this." So, "Okay.

I... But I don't even know Swift." And he was like, "Well, um, why don't you start learning it, and if you have any questions, ask me." It's just like, holy shit. Like, not only has Chris Lattner lowered his standards enough to talk to me- ...

but he's offering me personal tutoring-

**Alessio** [58:00]
Personal guidance

**Jeremy Howard** [58:01]
... in the programming language that he made. So I was just like, "I'm not gonna let him down." So I spent, like, the next two months, like, just nerding out on Swift, and it was just before Christmas that I kind of, like, started writing down what I'd learned.

And so I wrote a couple of blog posts on, like, "Okay, this is, like, my attempt to do numeric programming in Swift, and these are all the challenges I had, and these are some of the issues I had with, like, making things properly performant, and here are some libraries I wrote."

And I sent it to Chris, and I was like, "I hope he's not too disappointed with me," you know, 'cause that would be the worst. I was like, I, you know... And I was also like... I was like, "I hope he doesn't dislike the fact that I've...

you know, didn't love everything."

And, uh, yeah, he was like, "Oh, thanks for sending me that. Let's get on a call and talk about it." And we spoke, and he was like, "This is amazing. I can't believe that you made this. This is exactly what Swift needs."

And he was like... And so, like, somebody set up, like, a new Swift, uh, I can't remember what they call them, the equivalent of a PIP, you know, kind of RFC thing of like, "Oh, you know, let's look at how we can implement Jeremy's ideas in the language."

And it's just like, oh wow. And, um, so yeah, you know, um,

so, you know, and then we ended up, like, literally teaching some lessons together about Swift for TensorFlow, and we built a, a fastai, uh, kind of equivalent, um, with him and his team. It was so much fun. And then in the end, you know, Google didn't follow through, which is fair enough.

Like, asking everybody to, you know, le- to learn a new programming language is gonna be tough. But, like, it was very obvious, very, very obvious at that time that TensorFlow 2 is gonna be a failure, you know. And so this felt like, okay, I...

You know. Well, you know, what are you gonna do? Like, um, you can't focus on TensorFlow 2 'cause it's not gonna... Like, it's not working. It's never gonna work, you know. Nobody at Google's using it, um, internally. Um, so, you know, in the end, Chris left.

You know, um, Swift for TensorFlow got archived. There was no backup plan, so it kind of felt like Google was kind of screwed, you know. And, uh, Chris went and did something else. But we kept talking, and I was like, "Look, Chris, you know, you've gotta be your own boss, man, 'cause, like, you know, you've got the ideas.

You know, like, only you've got the ideas, you know. And if, if your ideas are implemented, we'd all be so much better off," 'cause, like, Python's the best of a whole bunch of shit, you know. Like, I would...

It's amazing, but it's awful, you know, compared to what it could be. And, um, anyway, so eventually a few years later, he called me up and he was like, "Jeremy, I've taken your advice." I've started a company. So that's like, oh my God.

So we're gonna create a new language, we're gonna create a new infrastructure. It's gonna build... It's gonna have all the stuff we've talked about. And it's like, oh, wow. So that's, that's, uh, that's what Modula is. And so Mojo is like, uh, you know, building on all the stuff that Chris has figured out over...

I mean, really from when he did his PhD thesis, which developed LLVM onwards, you know, and Swift and MLIR, you know, the TensorFlow runtime engine, which is very good. You know, that was something that he, he built and has lasted.

Um, so yeah, I'm, I'm, I'm, I'm pumped about that. I mean, it's very speculative. Creating a whole new language is tough. I mean, Chris has done it before, and, um, and he's created a whole C++ compiler, amongst other things.

Um, looking pretty hopeful. I mean, and I hope it works because, uh, you know, I mean-

**Alessio** [1:01:54]
You told him to quit his job, so...

**Jeremy Howard** [1:01:58]
But I mean, in the meantime, I will say, you know, Google now does have a backup plan. You know, they have JAX, which was never a strategy. It was just a bunch of people who also recognized TensorFlow 2 as shit, and they just decided to build something else.

And for years, my friends in that team were like, "Don't tell anybody about us 'cause we don't, you know, we don't want to be anything but a research project." So now these poor guys, suddenly they're the great white hope for Google's future.

Um, and so JAX is, you know, also not terrible, um, but it's still, still written in Python.

**Alessio** [1:02:30]
Mm-hmm.

**Jeremy Howard** [1:02:30]
Like it would be cool if we had all the benefits of JAX, but in a language that was designed for those kind of purposes. Um, so you know, fingers crossed that, uh, yeah, that, that Mojo turns out great.

**Alessio** [1:02:47]
Yeah. Any other thoughts on when, where people should be spending their time? So that's more the kind of language framework level. Then you have the, you know, GGML, some of these other like quantization-focused, uh, kind of model level things.

### Small Models

**Jeremy Howard** [1:03:02]
Yeah.

**Alessio** [1:03:02]
Then you got the hardware people. It's like a whole other bucket. Um, yeah, what are some of the exciting stuff that you're excited about?

**Jeremy Howard** [1:03:10]
Well, you won't be surprised to hear me say this, but I think fine-tuning transfer learning is, um, still a hugely underappreciated area. So today's zero-shot, few-shot learning equivalent is retrieval augmented generation, you know? Rag. Uh, which is like, just like few-shot learning is a thing.

Like it's a real thing. It's a useful thing. It's not a thing anybody would want to ignore. Why are people not spending at least as much effort on fine-tuning, you know? 'Cause, you know, RAG is like such a inefficient hack really, isn't it?

It's like, you know, segment up my data in some somewhat arbitrary way, embed it, ask questions about that, you know, hope that my embedding quest- you know, model embeds questions in the same embedding space as the paragraphs, which obviously is not going to...

If your question is like, if I've got a whole bunch of archive papers embeddings and I asked like, "What are all the ways in which we can make inference more efficient?" Like the only paragraphs it'll find is like if there's a review paper that says, "Here's a list of ways to make, you know, um-

**Alessio** [1:04:26]
Right

**Jeremy Howard** [1:04:26]
... inference more efficient."

**Alessio** [1:04:27]
Doesn't have any of the specifics. Yeah.

**Jeremy Howard** [1:04:29]
No, it's not gonna be like, oh, here's one way, here's one way, here's a different way in different papers, you know? Um, yeah, if you fine-tune a model then, then all of that information is getting directly incorporated into the, into the weights of your model in a much more efficient and nuanced way.

Um, and then you can use RAG on top of that. So I think that that's one area that's definitely like underappreciated. Um, and, and also the kind of like the, the confluence of like, okay, how do you combine RAG and fine-tuning, for example.

**Alessio** [1:05:01]
Something that I think a lot of people are uncertain about, and I, I don't expect you to know either, is that whether or not you can fine-tune new information in, um, and, and, and I think that that is the focus of, uh-

**Jeremy Howard** [1:05:15]
Of course you can

**Alessio** [1:05:15]
... some of your open questions and research.

**Jeremy Howard** [1:05:17]
But of course you can, right? Like-

**Alessio** [1:05:18]
Because it's additional pre-training

**Jeremy Howard** [1:05:19]
... like obviously you can. Because there's no such thing as fine... There's no such thing as fine-tuning. There's only continued-

**Alessio** [1:05:24]
AI

**Jeremy Howard** [1:05:24]
... pre-training. Continued... So fine-tuning is pre-training, like they're literally the same thing.

**Alessio** [1:05:30]
Yeah.

**Jeremy Howard** [1:05:30]
Um, so the knowledge got in there in the first place through pre-training, so how could like continuing to pre-train not put more knowledge in? Like it's, it's the same thing.

**Alessio** [1:05:40]
Yeah.

**Jeremy Howard** [1:05:41]
Um, the problem is just we're really bad at it 'cause everybody's doing it dumb ways. So, you know, I... But it's, it's a good question, and it's not just new knowledge, but like new capabilities. Um, you know, I think like in, in my Hacker's Guide to LL, in to...

Hacker's Guide to LLMs, uh, talk, I show simple... I mean, it's funny that's a simple example 'cause it doesn't sound it, but like taking a pre-trained base model and getting it to generate SQL, and it took 15 minutes to train on a single GPU.

You know, I think that might surprise people that, that that capability is at your fingertips and, you know, 'cause it was already there, it's just latent in, in, in the base model. Really pushing the boundaries of what you can do with small models, um, I think is a really interesting question.

Like what can you do with a... Like I mean, there isn't much in the way of good small models. Um, a, a really underappreciated one is, uh, BTLM 3B, which is a like kind of 7B quality 3B model.

**Alessio** [1:06:43]
Mm-hmm.

**Jeremy Howard** [1:06:44]
Um, there's not much at the 1 to 2B range sadly. There are some code ones, but like the, the fact that there are some really good code ones in that 1 to 2B range shows you that that's a great size for doing complex tasks well.

**Swyx** [1:06:58]
There was Phi-1 recently which has had, has been the subject of a little bit of discussion about whether-

**Jeremy Howard** [1:07:04]
Yeah

**Swyx** [1:07:04]
... they train on benchmarks.

**Jeremy Howard** [1:07:05]
And Phi-1.5 as well. So that's not a good model yet. Um-

**Swyx** [1:07:11]
Why not?

**Jeremy Howard** [1:07:12]
It's, it's good at doing a... So Phi-1 in particular is good at doing a very specific thing, which is creating very small Python snippets. Um, the thing... Okay, so like Phi-1.5, uh, has never read Wikipedia, for example, so it doesn't know who Tom Cruise is, you know.

Um, it doesn't know who anybody is. It doesn't know about any movies. It doesn't, doesn't really know anything about anything, like, 'cause it was never, it's never read anything. You know, it was trained on a, um, nearly entirely synthetic data set, which was designed for it to learn reasoning.

And so it's a, it was a research project and a really good one, and it definitely shows us a powerful direction in terms of what can you do with synthetic data, and wow, gosh, even these tiny models can get pretty good reasoning skills, pretty good math skills, pretty good coding skills.

Um, but I don't know if it's a model you could necessarily build on. Some people have tried to do some fine tunes of it, and again, they're like surprisingly good in some ways for a 1.5B model, but not sure you'd find it useful for anything.

**Swyx** [1:08:26]
I think that's the struggle of pitching small models because, uh, small is great, you know. You, you don't have a lot, you don't need a lot of resources to run them, but the performance evaluation's always so iffy. It's always just like, yeah, it works on some things and-

**Jeremy Howard** [1:08:40]
Right

**Swyx** [1:08:40]
... we don't trust it for others.

**Jeremy Howard** [1:08:41]
Which is why I think... Yeah, so that's why it's back- we're back to fine-tuning. I would say a... So Microsoft did create a Phi-1.5 web, but they didn't release it unfortunately. Um, I would say a Phi-1.5 web with fine-tuning for your task, you know, might quite, you know, might solve a lot of tasks that people have in their kind of day-to-day lives.

Um, you know, particularly in kind of an enterprise setting, I think there's a lot of like repetitive kind of processing that has to be done. It's a useful thing for coders to know about, 'cause I think quite often you can like replace some thousands and thousands of lines of complex buggy code maybe with, with a fine tune, you know.

**Swyx** [1:09:27]
Got it.

**Alessio** [1:09:27]
Mm-hmm. Yep. Um, and Jeremy, before we let you go, I think one question on top of a lot of people's minds. So you've done Practical Deep Learning for Coders in 2018, '19, '21, '22. I feel like the more time goes by, the more the, the GPUs get concentrated.

Uh, if you're somebody who's interested in deep learning today and you don't wanna go join OpenAI, you don't wanna join Anthropic, what's like the best use of their time? Should they focus on, yeah, small model development? Should they focus on fine-tuning math and all of that?

Should they just like focus on making RAG not a hack and coming up with a better solution? Uh, yeah, what's a Practical Deep Learning for Coders 2024 gonna look like?

**Jeremy Howard** [1:10:11]
Yeah. I mean, good question. I'm trying to figure that out for myself. You know, like what should I teach? 'Cause I definitely feel like, uh, things have changed a bit, you know. Um, one of the ways in which things have changed is that coding is much more accessible now.

Um, so if you look at a lot of the folks in the kind of open source LLM community, they're folks who really hadn't coded before a year ago, and they're using these models to help them build stuff they couldn't build before, which is just fantastic, you know.

Um, so one thing I kind of think is like, okay, well, we need a lot more material to help these people use this newfound skill they have, 'cause they don't really know what they're doing, you know, and they don't claim to.

Um, but they're doing it anyway, and I think that's fantastic. You know, so like are there things we could do to help people, you know, bridge this gap? 'Cause previously, you know, I know folks who were, you know, doing menial jobs a year ago, and now they're training language models thanks to the help of Codex and Copilot and whatever.

So, you know, yeah, what does it look like to like really grab this opportunity? You know, maybe, maybe fastai's goals can be dramatically expanded now to being like, let's make coding more accessible, you know, or kind of AI-oriented coding more accessible.

Um, if so, our course should probably look very different, you know, and we'd have to throw away that like, oh, you have to have at least a year of full-time programming tech, you know, as a prerequisite. Um, yeah, what would happen if we got rid of that?

Uh, so that's kind of one thought that's in my head. Um, you know, as to what should other people do, honestly, uh, I don't think anybody has any idea, like the more I look at it, what's going on.

I know I don't. You know, like we don't really know how to do anything very well. Uh, clearly OpenAI do. Like they seem to be quite good at some things, although talking to folks at or who have recently left OpenAI, even there it's clear there's a lot of stuff they haven't really figured out and they're just kind of like using recipes that they've noticed have been okay.

So yeah, we don't really know how to train these models well. We don't know how to fine-tune them well. We don't know how to do RAG well. We don't know what they can do. We don't know what they can't do.

We don't know how big a model you need to solve different kinds of problems. We don't know what kind of problems they can't do. We don't know what good prompting strategies are for particular problems. You know, like somebody sent me a message the other day saying they've written something that is a prompting strategy for GPT-4.

They've written like 6,000 lines of Python code, um, and it's, uh, to, um, help it play chess. And then they've, uh, they said they've had it play against other chess engines, including the best Stockfish engines, and, uh, it's got an Elo of 3400-

**Alessio** [1:13:12]
Oh my God

**Jeremy Howard** [1:13:13]
... which would make it close to the best chess engine in existence. Um, and I think this is a good example of, like, people were saying, like, "GPT-4 can't play chess." I mean, I'm-- I was sure that was wrong.

I mean, obviously it can play chess. But the difference between, like, with no prompting strategy, it can't even make legal moves. With good prompting strategies, it might be just about the best chess engine in the world, far better than any human player.

**Alessio** [1:13:39]
Mm-hmm.

**Jeremy Howard** [1:13:40]
So yeah, I mean, we don't really know what the capabilities are yet. So I feel like it's all blue sky at this point. It feels like computer vision in 2013 to me, which was like in 2013 computer vision was like, "Okay, we've-"

**Alessio** [1:13:51]
We just had the AlexNet moment

**Jeremy Howard** [1:13:53]
... had AlexNet, we've had VGG Net. It's around the time Zaire and Fergus like... No, it's probably before that, so we hadn't yet had the Zaire and Fergus like, "Oh, this is actually what's going on inside the layers."

So, you know, we don't actually know what's happening inside these transformers. We don't know how to create good training dynamics.

**Alessio** [1:14:12]
Mm.

**Jeremy Howard** [1:14:13]
We don't really know anything much, and there's a reason for that, right? And the reason for that is language models, uh, suddenly got really useful. And so the kind of economically rational thing to do, like this is not a criticism, this is true.

The economic rational thing to do is to like, okay, like build that as fast as possible. You know, make something work, get it out there, and that's what, you know, OpenAI in particular did, Anthropic kind of did. And um, there's, there's a whole lot of technical debt everywhere.

You know, nobody's really figured this stuff out, um, because everybody's been so busy building what we know works as quickly as possible. So yeah, I think there's a huge amount of opportunity to... You know, I, I think we'll find things can be made to work a lot faster, a lot less memory.

I got a whole bunch of ideas I want to try. You know, um, every time I look at something closely, like really closely, I'm always like, "Oh, turns out this person actually had no idea what they were doing."

You know? Um, which is fine. Like, none of us know what we're doing.

**Alessio** [1:15:28]
Yeah.

**Jeremy Howard** [1:15:28]
Um, we should experiment with that.

**Alessio** [1:15:32]
Yeah. As we had, um, Tri Dao on the podcast who created Flash Attention.

**Jeremy Howard** [1:15:36]
Wow.

**Alessio** [1:15:36]
And I asked him, "Did nobody think of using SRAM before you? Like, were people just like, no..." And he was like, "Yeah, people just didn't, they didn't think of it. They didn't try. They didn't come from like a systems background."

And-

**Jeremy Howard** [1:15:49]
Yeah

**Alessio** [1:15:49]
... so.

**Jeremy Howard** [1:15:49]
I mean, the thing about Flash Attention is, I mean, lots of people absolutely had thought of that, and so had I, right? But I mean, the, the honest truth is, particularly before Triton, like everybody knew that tiling is the right way to solve anything, and everybody knew that attention, fused attention wasn't tiled, and that was stupid.

**Alessio** [1:16:12]
Mm-hmm.

**Jeremy Howard** [1:16:13]
But not everybody's got his ability to like, be like, "Oh, well, I, I, I'm confident enough in CUDA and or Triton to use that insight to write something better." You know, and this is where like I'm super excited about Mojo, right?

I always talk to Chris about Flash Attention 'cause I'm like, "You know, there is a thousand Flash Attentions out there, uh-"

**Alessio** [1:16:35]
Right

**Jeremy Howard** [1:16:35]
... "for us to build. Um, you just got to make it easy for us to build them." So like Triton definitely helps. Um, but it's still

not easy, you know.

**Alessio** [1:16:47]
Mm-hmm.

**Jeremy Howard** [1:16:47]
It still requires kind of really understanding the GPU architecture and writing it in that kind of very CUDA-ish way.

**Alessio** [1:16:56]
Right.

**Jeremy Howard** [1:16:56]
Um, so yeah, I, I think, I think, you know, if Mojo or something equivalent can really work well, we're gonna see a lot more Flash Attentions popping up.

### Lightning Round

**Alessio** [1:17:08]
Um, great, Jeremy. Before we wrap, we usually do a quick lightning round. Um, we kind of have three simple questions. So the first one is around acceleration, and you've been in this field a long time. What's something that it's already here today in AI that you thought would take much longer?

**Jeremy Howard** [1:17:24]
I don't think anything. So I've actually been slightly too bullish. So in my 2014 TED Talk, I had a graph and I said like, "This is like the slope of human capabilities and this is the slope of AI capabilities."

**Alessio** [1:17:40]
Mm-hmm.

**Jeremy Howard** [1:17:40]
And I said, "Oh, we're..." And I put a dot saying, "We are here." It was just before they passed. And I say-- I looked back at the transcript the other day and I said, "In five years, I think, we'll...

You know, we might have crossed that threshold in which computers will be better at most human tasks than most humans or most average humans." And so that might be almost true now for non-physical tasks. So I was like, took, you know, took about twice as long as I thought it might.

Um, yeah. No, I wouldn't say anything surprised me too much. It's still like definitely like... I got to admit, you know, I had a very visceral reaction using GPT-4 for the first time.

**Alessio** [1:18:27]
Mm.

**Jeremy Howard** [1:18:28]
Not because I found it surprising, but actually like actually doing it, like it's something I was pretty sure would exist by about now, um, maybe a bit earlier. But actually using it definitely is different to just feeling like it's probably on its way, you know.

And yeah, whatever GPT-5 looks like, um, I'm sure I imagine I'll have the same visceral reaction, you know.

**Alessio** [1:18:57]
It's, uh, it's really dif- uh, amazing to watch develop. Um, we also have an exploration question. So what do you think is the most interesting unsolved question in AI?

**Jeremy Howard** [1:19:09]
how do language models learn? You know, what are the training dynamics? Like, I wanna see... There was a great, um, paper about, uh, ResNets, uh, a few years ago, um, that showed how... that was able to, like, plot a kind of projected three-dimensional loss surface for a ConvNet, um, with and without, um, skip connections.

And, you know, you could very clearly see without the skip connections it was, like, super bumpy, and with the skip connections it was super smooth. Um, that's the kind of work we need. Like, so there was actually an interesting blog post that came out just today from the PyTorch team, uh, where some of them have created this, like, uh, 3D matrix product visualization thing.

**Alessio** [1:19:57]
Yeah. The, the MatMul Visualizer.

**Jeremy Howard** [1:19:58]
And-

**Alessio** [1:19:58]
Yeah

**Jeremy Howard** [1:19:59]
... yeah, and they actually showed some nice, nice examples of, like, a GPT-2 attention layer and, like, showed an animation and said, like, "If you look at this, we can actually see a bit about what it's doing." You know, so, you know, again, it reminds me a bit of the Zeiler and Fergus, you know, uh, ConvNet paper that was the first one to do these reverse convolutions to show what, what's actually being learned in each layer in a ConvNet.

Yeah, we need a lot more of this. Like, h- what is going on inside these models? How do they actually learn? And then how can we use those insights to help them to learn better? Um, so I think that would be one.

The other exploration I'd really like to see is a much more rigorous analysis of what kind of data do they need, at what level, and when do they need it, and how often. So that kind of, like, data set mixing, curation, so forth-

**Alessio** [1:20:52]
Right

**Jeremy Howard** [1:20:53]
... in order to get the best capabilities.

**Alessio** [1:20:56]
How much of it should be code? Yeah.

**Jeremy Howard** [1:20:56]
You know.

**Alessio** [1:20:56]
How much is Wikipedia? Um, yeah.

**Jeremy Howard** [1:20:58]
Yeah.

**Alessio** [1:20:59]
It's, it's very uncertain.

**Jeremy Howard** [1:21:00]
Yeah, and be able to fine-tune what, you know, what kind of mix do you need for it to keep its capabilities, and what are the kind of underlying capabilities that it most needs to keep, and if it loses those it would lose all these other ones, and what data do you need to keep those, and, you know, are there things we can do to change the loss function to help it to not forget to do things.

Stuff like that.

**Alessio** [1:21:21]
Awesome. Uh, and yeah, before wrapping, what's a one message, one idea you want everyone to remember and think about?

**Jeremy Howard** [1:21:28]
You know, I guess the main thing I want everybody to remember is that, you know, there's a lot of people in the world, and they have a lot of, you know, diverse experiences and capabilities. And, you know, they, they all matter.

And now that we have a, you know, newly powerful technology in our lives, we could think of that one of two ways. One would be, um, "Gee, that's really scary. What would happen if all of these people in the world had access to this technology?

Some of them might be bad people. Let's make sure they can't have it." Or one might be, um, "Wow, of all those people in the world, I bet a lot of them could, could really improve the lives of a lot of hu- humanity if they had this tool."

Um, this has always been the case, you know, from the invention of writing to the invention of the printing press to the, you know, development of education. And it's been a constant battle between people who think that the distributed power is unsafe and it should be, uh, held onto by an elite few, and, uh, people who think that, um, that humanity on net, you know, is, is a marvelous species, particularly when part of a, a society and a civilization, and we should do everything we can to enable more of them to contribute.

This is a really big conversation right now. And, you know, I, I, I wanna see, yeah, more and more people showing up and showing what, you know, what the, the, the great unwashed masses out there can actually achieve.

You know? That actually, you know, regular people are gonna do a lot of really valuable work, um, and, uh, and actually help us be, you know, m- more safe and also flourishing in our lives and providing a, a future for our children to flourish in.

Um, you know, um, if we, if we lock things down, um, to the people that we think... you know, the elites that we think can be trusted to run it for us, yeah, I think all bets are off about where that leaves us, uh, as a, as a, as a society, you know.

**Alessio** [1:24:02]
Mm-hmm. Yep. No, that's a important message. And yeah, that's why we- we've been promoting a lot of open source developers, open source communities. I think, uh-

**Jeremy Howard** [1:24:13]
Good on you

**Alessio** [1:24:13]
... letting the builders build, uh, and explore, it's always a good idea.

**Jeremy Howard** [1:24:17]
Yeah.

**Alessio** [1:24:18]
Um, thank you so much for coming on, Jeremy. This was great. Um-

**Jeremy Howard** [1:24:21]
Thank you for having me.

**Alessio** [1:24:23]
Yeah.

---

This library is powered by PodHood (https://podhood.com), the podcast website platform.
