Intro0:00
Let's go.
Um, I'll kill time while Eric, uh, figures out how to share his computer. Um, so why, why did we pick BERT today?
I, I'm, I'm kind of curious.
Eric picked it.
That's-- Yeah. Yeah, 'cause I'm working on a text classification problem at work and just wanna get the background.
Yeah. I wonder if there is a microservice that is easy to set up, I wonder if OpenPipe does this, that only does, uh, mirrors a structured output GPT-4o, um, call and just mirrors it until it has enough data for BERT, and then just switches you to BERT.
What do you mean by mirrors it?
Shadows it. Um, basically-
Basically you use it as start until you have enough data and then do else and then just swap it, so it's cheap.
Yep.
I think OpenPipe actually does that, but I don't think it's as automatic as and as seamless as how your b- your, your vision is, Six.
Yeah, I mean, like the amount of times people suggest using the BERT for classification, it seems like it's a path that is worth paving. Oh, Eric's rejoining. Okay.
BERT and T5 are good, but like, there's like a host. You need to-- Like one thing is just switch, but like for OpenAI or routing, you have endpoints and hosted model and all that. To, to do BERT, like you gotta go deploy it somewhere and, you know, handle that part.
Yeah. All right. Eric, we can see you.
Okay, perfect.
Paper Background1:48
Um, so we'll start by going through the paper and then, um, there's a ton of a-additional material out there on BERT. So, um, depending how much time we have, what we can look at some of the, the other things out there.
So first of all, BERT stands for Bidirectional Encoder Representations from Transformers. So this is one of the first, uh, transformer papers. Like I, I believe it was about a year after the, uh, Attention Is Everything You Need paper.
You can see it's from, uh, twenty nineteen, so ancient history in terms of, uh, deep learning and NLP.
I can put it in there.
Um, but still useful for a lot of use cases. The-- And, um, just for a little bit more context. So after, uh-- So this was a, a Google paper, and after they trained and released BERT, they started using it in, um, Google Search.
And so it provided context for search results so that, um, there are some examples maybe we can look at later, but it would-- If you gave like a, a query that could mean a couple different things, they use this model to discriminate between the two.
So let's, um,
let's, I guess, just walk down through the paper here. Uh, one thing they note is that there was this, uh, two different kinds of approaches early on: feature-based and fine-tuning. And so the-- like models like ELMo were feature-based, where, uh, they had task-specific architectures that were built into the model.
Whereas models like BERT and GPT were more, uh, used a fine-tuning approach where they just, uh, it's more general purpose. They just trained one model and then you could fine-tune it after the fact, um, for your particular, uh, tasks.
Uh, they talk a little bit here about the-- one of the limitations of standard language models is that they're unidirectional. So what that means is for like GPT and many other, um, models out there now, they only look forward, and so they take a sequence of words and, and try to predict the next word in the sequence.
Um, however, BERT also goes backwards, so it takes-- Like in the training data, it, it can also start at, say, the end of a piece of text and try to predict the previous word. And so we'll talk a little bit, um, how they avoid like contamination of, uh, the prediction because obviously if you're training from back to front, you get a peek at what the, the words are coming up.
So a, a lot of this section was, um, pre-decoder-only transformers, right? So a lot of what they reference is like RNNs, GRUs, LSTMs. So like pre-transformer stuff, you only get one pass and then they're like, "It's a crazy idea if you look at it from front to back."
Like, so a lot of the training tasks was like classification, right? When humans classify something, we don't do it token by token. We listen to the whole sentence, then we classify. So this was more so like, uh, LSTM, RNN era.
Yeah.
Right.
Yeah. It's a very early paper. Um-
So let's see. We talked about bidirectional. Um,
and then, yeah, so it's fine-tuned versus feature-based. And then they show the effectiveness on BERT on eleven different NLP tasks.
I guess one call-out on the, on the related work is just ELMo. So ELMo was another model. I believe it was also by Google. If anyone knows for sure, uh, feel free to correct. But they used different, um, representations for the same word in different contexts.
And so, uh, like the word stick, for example, you could say that means I'm going to chase that dog with a stick, or it could be like, "Hey, let's stick to the material that we're talking about," or maybe some other words or context.
And so those can mean different things, um, the same token, and because of that they wanna use, like, different representations, even though it's the same word in English. And so BERT leverages that feature as well.
So, um, slight correction. ELMo is from Allen Institute and University of Washington.
Mm. Okay.
And then further context. So one of the best BERT adaptations, I think in twenty nineteen was RoBERTa. RoBERTa is like BERT, but make it good and bigger. And it was also from the same team at University of Washington and Facebook, I think.
But, um, ELMo was just... Yeah, it wasn't from Google, but same thing. It went from like one-hot encoding bag of words to... ELMo was really popular for like a bunch of Kaggle competitions where you needed good embeddings and then BERT embeddings.
Yeah, it's not Google.
Mm-hmm. Great. Thanks for the additional info.
Architecture8:20
Uh, so let's take a stop here to look at, um, the-- this diagram. So this is the, the, uh, BERT like architecture. Essentially, you can see this pink row down here is a set of tokens. Um, let me see if I can zoom in here.
Hmm.
I'm not sure why I can't zoom in while I'm sharing. But anyway, this, this token is a classifier token, and then you see like tokens one through N. So those, uh, are the, are the first sentence. And then there's a separator token, and then there's another token one through M.
And the reason that this-- that BERT has this structure is because one of the tasks it does is, uh, sentence classifica-- or classification. Basically, it can take two sentences and determine if one of the sentences seems like it follows the other one.
And then you see... So this is a pre-trained model. So this is trained on, well, what was at that point a large amount of text. And then these are the different fine tunes of that model for different benchmarks, essentially.
Um, and FYI, I'm not looking at the chat, so if there's anything in there, I, I'm not seeing it. I can answer questions later or if someone just wants to unmute, feel free.
So here's those two tokens I mentioned that you might not have been able to see, the classifier token and then the separator token. The classifier token, uh, tells it like what task it's supposed to be performing.
So they talk about how there's two steps in the framework, retraining and fine-tuning. Uh, probably a lot of people are familiar with that, so I won't go into too much detail. Um, this is interesting. So this is two-- twenty nineteen numbers of what was at that point a large model.
You can see the, the base model was a hundred and ten million parameters, and the large model, which was, uh, I saw someone refer to as like gigantic or like unbelievably large or something like that, is three hundred and forty million parameters, which is, you know, like five percent the size of, uh, like a Llama 7B or something along those lines.
So what at the time seemed very large, uh, in retrospect is- isn't really. And that, that also goes for the training data. Um, let's sh...
Do they have it in here? Well, there's two, uh, training sets that they used on-- for it. One was a all of Wikipedia, the English version, which was like eight hundred million words, and then I think a set of books which was two point five billion words.
And again, those data sizes, uh, while they could have been large for the time, currently are relatively small. Um, typically, like at least for frontier models, you're talking about low trillions of tokens to train them
Uh, so here they talk about what I mentioned earlier about, uh, having a pair of sentences where you can have a question and answer in one token sequence with the separator token in the middle.
And so let's go down and talk about, uh, pre-training.
So-- And here we get to the, their answer to the left to right versus right to left dilemma. So because it's bi-directional, potentially, uh, each word could see itself, um, in the future. And so what they ended up doing was to mask fifteen percent of the words.
Pre-training12:44
So instead of, uh, having the actual word in the sequence, they have this mask token, and then they try to predict that word either going forward or backward, uh, a-and do the, you know, score the, the training on that.
Um, they do have an issue that, uh, let's see,
that since the mask token does not appear during fine-tuning, uh, they have to compensate for that in the pre-training step. So in order to do that, they don't always replace the masked words with this masked token. Uh, they do eighty percent of the time.
Um, ten percent of the time, they just stick in a random token from anything in the vocabulary, and then another ten percent of the time they just leave it unchanged. And so that helps, uh, during the fine-tuning stage so that fine-tuning doesn't expect these masked tokens, like, all the time.
And then the other task they give it during pre-training is this next sentence prediction. That is the, uh, you saw the two sentences, um, and whether they are related. So that is, uh,
you know, that's the sentence A separator sentence B.
And then they have, um, the fifty percent split of training data where half of it is labeled as is next and half of it is labeled as not next. So you can imagine a initial sentence of, "I walked down to the store."
Uh, o-one that might be marked as is next is something like, "I bought a burrito." And something that might be not next is, you know, "Salamanders have many spots," or something like that. And so part, this part of the pre-training figures out which of those sentences, uh, follow each other.
Here-- Oh, okay. So here's what I was telling you about earlier, the BookCorpus, eight hundred million words, and I guess I had them flipped. Wikipedia is the two point five billion words.
And th-this shows how the, uh, the input representation. So there's a couple of different, um, layers, and they just add these layers up to come up with the input. So the top layer here is the, the token, so it's the vector representation of that particular word.
Uh, the next is the segment. So this splits it up between sentence A and sentence B. Each one of those has a different, uh, factor that gets added into the input to help the model distinguish between them. And then finally is the positional embedding.
So each, uh, each place in the sequence gets its own value that when it's added to the, the other layers, helps the model to, um, distinguish, like, does my come before dog or vice versa?
Okay, and then they talk about a bunch of, uh, different benchmarks. We'll just maybe take a look at some of the results and then, uh, continue forward. I'll also look at the-- Let me take a break to look at the chat here.
Results17:06
There are no serious questions in chat. Um, I guess Sam had one. Uh, have there been revisitings of the weird now pre-training tasks that were used to convert, or has this all been done and it's somewhat of a shut book and next token prediction is all you need?
Any thoughts, comments?
So I, I didn't look at any of the following BERT papers like RoBERTa or DeBERTa for this. Uh, so I'm not sure. Uh, if there's anyone else on the call who knows, would be happy for a contribution.
Yeah, we had a slight follow-up in the comments there. Uh, basically for encoder-decoder stuff, it still makes sense. And a lot of my takeaway from this paper when I read it, like years ago was they had really interesting- Pre-training tasks, and if you look at it from a lens of what they're trying to do, it's actually pretty useless, right?
Like, next token prediction is still useful to generate words, but something like predicting sentence order is sentence one or sentence two. Should sentence one come before sentence two or whatever? It's not really useful, right? There's never a time where there's people trying to predict which sentence came before the other.
But it does really teach a model conceptually, like what word order is, right? So there's these words, and there's these words. You have to understand, should these set of words come before these set of words. So that helps in the classification task where you have to like group together words together.
So in some sense, for the tasks that BERT is trying to do, it's trying to be a small, efficient classification model. That's one of the tasks it's trained to do, right? It kind of makes sense to do these weird training objectives.
So, uh, next token prediction or like masking of words is still like there's not many use cases where you need to fill in the blank, right? We don't have like five hundred words and you have to fill in one word, but it does teach a model like over a billion words.
If you start to understand what words go in between other words, it's a good pre-training task, and it helps with classification in general. So it's like instead of having a very broad task of just predict next token and then, you know, extract like abstract that out to eventually like get this emergent capability to do classification, this is like here's twelve tasks that mimic understanding words very well.
You're a small model. Use all this to do classification. But, um, that concept is still applied for like current encoder, decoder models and small models that are not just next token prediction. So basically, if you dumb down the task to subsets of your like main goal, it's still very effective.
And that's why like if you take a step back and look at what BERT did, it makes no sense. Like Google doesn't need to spend millions of dollars to predict which sentence comes before or after another sentence, but it does help a small model better le- better learn embedding per se and abstract that to these tasks.
Yeah. Agreed. Uh, like the, the sentence to sentence prediction didn't seem like... You're right. There's not too many use cases where that would be helpful. Um, I wonder if it does help the model to like figure out the relationships between words and, and ideas, um, to some degree.
Yeah. That's, that's pretty much the whole point. It should help the model better understand, like breaking down the problem and understand word order. I think they also did something where they swapped words from different sentences or swapped sentences, right?
And like, that's even-
Mm-hmm
... more useless in real, in reality. Like in reality, there's not many times where you've got mixing of words from different chunks of sentences. But once again, it helps the model generalize towards understanding sentences. So stuff like that.
Um, it's just if you look at it in today's sense, if you have a niche topic that can benefit from encoder, decoder, like small, small, like on-device active model, uh, you wanna start to employ breaking down problems in this sort of methodology.
But there's, there's examples of papers that, that do this type of work.
Mm-hmm.
Yeah. For sure, um, would be great to follow on this presentation with m- some of the more recent work that kind of builds on it.
So-
Yeah. Not, not many other questions in chat, by the way.
Okay. Okay, cool. Um, you can see from this in the-- at least when it was released, BERT Large was state-of-the-art, even beating out, uh, GPT-1 and ELMO. So, uh, at least at the time it was released, it was a very capable model.
Um, and then, I don't know. I'm gonna kind of skip through these, except maybe just to show the, the results, um, of it being state-of-the-art in a lot of things. So, uh, the next section is ablation study, so removing different parts of the model to see, uh, like what the effects were.
And
let's see. So here's the... A couple different ablations they did was they, uh, they removed the next sentence prediction task. So I guess this is something we were just talking about, uh, but they still kept the, the mask LM.
And then the next thing they did is they only made it go left to right, and they also have the no sentence prediction. And so you can see the results from those attempts up here. Uh, the top is the kind of the, the standard model.
And then if you look at the no next sentence prediction, it does lose a little bit. Oh, actually in QNLI it looks like it has a significant loss. Maybe in other tasks, much less. But then as you also take away, um, the bidirectional, it becomes less capable.
Looks like it kind of varies between tasks, uh, like how much capability it loses. Uh, but this does show that there is some value to those components
Oh yeah, this is, um, maybe this is what I'm talking about, where they say, "We believe that this is the first work to demonstrate convincingly that scaling to extreme model sizes also leads to large improvements on a very small scale task."
This is kind of like the, the bitter lesson, um, but maybe a little bit exaggerated as far as the extreme model size at this point.
Resources25:05
Uh, and then we can just... They talk about feature, um, approach with BERT a little bit. So, uh, if there's questions, feel free to, um, unmute. Otherwise, let's go over to and look through. There's Jay Alammar has made some very helpful, uh, like illustrated BERT and ELMo articles that we can go through to just kind of, uh, cement our understanding.
So this is just kind of a comparison of some models that were out at the time. And then this is one thing that was a takeaway for me, is that, okay, this is the pre-training step that we talked about, but then on the supervised learning step, you basically stick a classifier after the BERT.
So BERT is in, in charge of essentially encoding the text into an embedding, and then you use that classifier to then classify, um, in this case, either spam or not spam. Let me see if I can find...
There's one diagram, uh, that I thought was especially helpful. Oh yeah, this one. So this one shows how BERT takes the entire sequence of tokens and then creates like for each input token, it has an output vector. However, for the purpose of classification, we only look at the first output vector.
Um, that one in-- contains essentially the entire, uh, sense of all of the input tokens, and then you can run that through. It can be a neural network, could be a logistic regression. And then from that, from the features here, and I think there's like something like seven hundred and sixty-eight, uh, dimensions in the embedding.
From the f- the features there, you can then predict, uh, you know, spam or not spam based on your training set.
And let's see. That shows the same thing.
And then here's like some illustrated, uh, illustration of the different encoder blocks. So as we mentioned earlier, BERT is encoder only. So the, the kind of classical transformer is an encoder and a decoder. Uh, many modern models are decoder only.
Um, and so encoder is like used mostly these days for text classification or text clustering. Um, to my knowledge, there's encoder only, uh, tran-transformers aren't really used for any kind of, uh, next like sequence generation or, or next token generation.
This talks about ELMo and the, the different context, uh, of words and how ELMo captures that.
Uh, it's GPT. I thought there was, um,
something here. Yeah. So this is just
like what we, uh, talked about, about, you know, if you have a BERT encoder, you can stick a- another model for training on the end of it and then, uh, go from there. So...
And then you can also use BERT for embedding. So if you have like a certain, uh, problem space with a lot of text that you want to embed, you can, uh, create... You can continue pre-training or do fine-tuning on BERT with your corpus in your industry specific, uh, like text corpus, and then create an encoder that's, uh, especially built for your,
uh, your needs.
So there's one more, uh, I'll pause any other questions in the chat.
So some context of, um, how they trained BERT. They had like those twelve paths, right? They had a BERT base model, math language model. They had the next sentence prediction, token classification, QA, uh, sequence classification. They had all these tasks, and basically what they did were they were BERT models with a layer added on top for a classification head.
Now, in the time of twenty nineteen when people started using these models, what was really common to do was you could either, if you had enough data, take the base model, add in a linear output head for classification, where you basically take all this.
There's no output head. It's just- Output is the last step of processing these tokens or sentences. Then you add just a linear head with a softmax for classification. Now, then you fine-tune it on a lot of your, your data itself.
If you had-- If you didn't have as much data, one thing that was popular was you take the, uh, sequence classification head, you just continue fine-tuning it on your data, and it's already somewhat good at sequence classification. But there is a whole, like, series of work that looked into based on how much data you have, where you should do your fine-tuning.
So if you have a lot, a lot of data, it was pretty common to not only add a classification head, but also peel back a few layers. So, uh, reset the weights of, like, the top three, the top two, the final layer, and then continue training those in as well for your task.
Because at some level, what people started to learn was these train-- these pre-training objective tasks of, like, mask word prediction and sentence ordering or QA, they were actually affecting the net output of sequence classification. And if you want it better, you could just train more of your whole model on that.
So there was a whole thing of, like, you should remove the top two layers, add a sequence classification head, train on tens of thousands of examples, and you'll get state-of-the-art, um, results. You could, if you have less data, freeze layers, you could unfreeze weights.
There was, like, a whole set of this, but it was pretty common to also just mess with the architecture and add classification heads.
Do you know if anyone, um, trained all the layers or, like, just used that as a starting point-
Yeah. Yeah
... to train all the layers? Okay.
So there's-- Today, there's, like, stuff that came out, like, a year or two ago where basically you could retrain BERT in twenty-four hours on, like, a budget of, like, sub five hundred dollars with regular A100s and how you can do this better.
So it was in the realm of, like, at the time, not as effective to y-you know, like, you don't have Google Compute to retrain BERT from scratch. But now there's stuff of, like, twenty-four hours and a couple hundred dollars could retrain your own better BERT.
There's, like, an academia paper that came out about this. If people down the rabbit hole of, like, encoder models and this stuff, it's a cool one to look into of how they can better objectify these twelve pre-training tasks to a few in a better curated data set and outperform it on a couple hundred dollars in twenty-four hours.
But then it was also common where, like, there was, um, uh, sentence classification and sentence extraction tasks that BERT was adapted towards. So, like, BERT for sequence classification or extraction for, like, abstractive summarization. And then companies that took it to production would do, like, significant, um, retrains or, like...
Yeah, they, they train a lot more of it. And then, um, this also just went into, like, at what part do you want to start training.
Mm-hmm. Yeah. I mean, that, that sounds, uh, like interesting stuff.
If you have any links, uh, please drop in the chat, and I'll check it out.
So maybe the last thing we can go through here is, um, the same author, Jay Alammar, has, uh, a notebook where he shows, like, hands-on how to, how to, um, do this movie review sentiment classification. He uses DistilBERT.
Notebook34:02
So DistilBERT is a Hugging Face, uh, like, recreation of BERT that, like, has very comparable performance on, uh, many fewer parameters. And then to do the classification, he just uses a basic logistic regression model from scikit-learn. Uh, and so then the features that go into this logistic regression model are just the, the vector of size seven sixty-eight that comes out of, uh, the, the DistilBERT embedding.
So a lot of this leans very heavily on the, uh, Hugging Face Transformers library. So let's see. That's just installing it, doing imports. Um,
he uses-- He must have mentioned it up above, but a...
Maybe it's... Anyway, there's a particular Hugging Face data set that he's using that has the, um, movie sentiment training data.
Maybe he just uploaded it to somewhere he had it. Um, so-
I think it's a IMDb data set.
Okay.
I think it's one of the Kaggle ones. Uh-
Okay.
It's just a Kaggle IMDb. You have movie reviews, you classify them.
Oh, and, and on that, actually reminds me, one of the big things that made BERT somewhat popular was there was another Kaggle competition on tweet classification of sentiment. So in tweets, like with previous embeddings, like a bag of words or GloVe or ELMo, if you have stuff like, um, you know, uh, "I hate this so much," that in some contexts in tweets could still be positive even though it's very negative.
And when you just look at, like, um, like lexical understanding of words, it's very negative. But then BERT embeddings were what really dominated that. And then for like a few years, they kept doing follow-ups on that. But IMDb and tweet classification were, um, versions that they used in a lot of these demos.
Mm-hmm.
So let's see. So here we're just, um, uploading or downloading BERT, DistilBERT from, uh, Hugging Face and, um, getting the model initialized. So you can see it's just a few lines of code there.
So we have to do a few things, um, like tokenize it and then add padding. So we're-- this is so the, all of the sequences can be run, run in parallel. Um, so we need to pad out so that they're all the same length.
And then we need to mark the padded sections as masked so that BERT doesn't get confused at thinking, like, what's empty space is, uh, actual sequence that we want it to, uh, to process. So then
you can see this diagram here, and again, apologies, I'll try to zoom in again. Oh, it worked this time. So this just takes the, the input text, runs it through DistilBERT, and comes out with the embeddings.
Um, so that's all of this thing does. And then the, the one tricky part about all of this is you need to pick out, uh, exactly which, uh, vec-- like, which values from this, I guess, this three-dimensional tensor you want to predict on.
So if you remember back, uh, from here, we just want the, the v-very first output. We wanna ignore all of these other ones.
So he draws out in detail, like, how exactly you, you pull just those vectors out of this, uh, three-dimensional tensor. And then it's pretty straightforward, uh, machine learning after that. Uh, you just turn those seven hundred and sixty-eight, um, dimensions into, uh, features and do a trust-- test train split, train your logistic regression model, and then, uh, run, you know, once you've got the model trained, you can run a score, and it gets eighty-two percent.
So, uh, assuming it's a fifty/fifty split, then the, the expected amount just from random chance would be fifty percent. So there is a significant, uh, increase using BERT to do classification, but obviously still plenty of room for improvement.
Uh, okay, and down here it says the highest accuracy score for this dataset is currently ninety-six point eight. So as you can see, things have come a long way since, uh, twenty nineteen, but, uh, still, you know, a useful model to, to start with for, uh, any classification tasks or clustering.
You know, if you just wanna see what text sequences are close to each other in embedding space, you can use it for that as well.
So that's, uh, about all I had for prepared stuff. I'll stop, um, sharing and then maybe go through the chat. If anyone wants to, uh, chime in, add any color commentary, feel free to do that.
Q&A40:34
Uh, someone linked the paper on the dynamic budget, the, uh, twenty-four hour BERT. And then I was also tr-- found the paper. Apparently MosaicML showed how you can do it for twenty dollars, pre-train BERT from scratch now. So, um, yeah, at twenty dollars they did like eight A100s for an hour, and they're able to match the GLUE score of basic BERT with their recipe.
Um, kinda interesting. So, uh, a note Eugene Chaw made is, uh, researchers felt 10K training is expensive. So I remember mathing this out. BERT... Th-- All these things that compare, like twenty-four hour or one-hour BERT, it-- they, they trained BERT for four days of TPU v3 equivalent, which is like, at the time, let's say eight to ten dollars an hour, which is like ten to fifteen K.
But then there wasn't just BERT base. There was BERT base, there was BERT large, there was BERT small. There was a bunch of experiments. The BERT large was trained for more than four days. Like the, the cost equivalent is fifty K on that, ten K on the regular BERT base, less on the little one.
And then you gotta add in like the time and the R&D. Oh, it was well more than a 10K project at Google. Uh, the, the BERT large itself was already a fifty K train run, plus ten to fifteen for BERT base, plus just experimentation.
So expensive, expensive.
I think it's more of-
But-
-the amount of labs that would love to hear those number for SOTA right now
.
True, true.
I'm just reading through the chat. Um-
Did we get a volunteer for next week or still waiting on that?
Anyone else next week? Any other questions on this by the way?
I do have a random one, uh, because this is regarding the embedding size, right? Um, e- even though I joke about this was, uh, the era before the GPU, uh, the gaming GPU, uh, folks came in and said, "Hey, you need to be divisible by sixty-four, thirty-two, and, or power two," right?
Um, does TPUs not have the, the divisible by sixty-four, uh, batch, I mean, like, uh, optimization when it comes to em- embedding size, uh, characteristic? That- that's why they have all these weird embedding size on TPUs related training.
I don't think it's TPU based. I have like old notes that I'm recalling where I dug through why they specifically did seven sixty-eight and five twelve, and also someone noted in chat that's a limitation. Um, there's other questio- there's other work that extends this out to like Sentence BERT that extends the embedding dimension and, um, they, they were all still pretty small.
But back to Eugene's point of is it hardware limitation? It's not. It was, well, the visibility between layers and adding layers and a bunch of stuff. I, I really can't remember the specifics of the reasons. I'll dig through some old notes, but someone broke down the per layer math and sending through, um, inputs and, and there was a, a decent ream- reason for why, why all this.
It's also like there's twelve layers divisible by seven sixty-eight. Um, it, it, it went down that, that path. But also it wasn't like someone from Google that worked on BERT. It was just mapping through the input through every layer and all this math working out.
And then the reason for like, oh, here's why this, why not this. And I was like, "Sounds, sounds good. Checks out to me." Um, I can probably find this if I look. It's, it's in some notes from a couple years ago.
Um, Eric, someone has asked, does BERT pre-training objective MLM, like vast language modeling, follow the same LLM scaling laws as GPTs?
Uh, that's a good question. I don't know if there's been enough, uh, research in that area to like come to any conclusion. Um, so I... like when I was researching this presentation, I went to, uh, I think it was paperswithcode.com or something like that, and looked at all the, the top papers for text classification.
And like a lot of them were pre... or twenty twenty-one or, or previous. And so it seems like, this, this direction of like encoder only or bi-directional has, well, I don't know about the bi-directional part, but at least encoder only research has been pretty, uh, sparse recently.
Um-
So for example, I don't, I don't know if anyone's spending millions of dollars to train a like a, you know, Super BERT or something like that.
Well, the Reka AI guys seems to.
Oh.
Uh, no, the last time I looked at this like leaderboard on Hugging Face, uh, I think it was all led by these, uh, transformed LLMs, uh, that now get the best performance, like the Mistral 7B turned into a embedding model.
Is that... Sorry, I missed. Is that for text classification?
Uh, well, yeah. I guess at the core it's all like turning text into an embedding. Um, so yeah.
Could you drop a link in the, in the chat?
Uh, yeah.
Yeah. It seems like that would lead to higher performance, uh, with classification, but I'd have to do a little bit more research on that.
So there's, there's two, two pieces to this as well, right? So for mass language modeling, uh, a lot of the scaling law papers directly showed why decoder-only token prediction is better scaling than mass language models. One is purely when you mask fifteen percent of tokens, you train on the mask, so you lose a lot of data.
You, you need to-- you need more quality data. You're just straight training on less, right? If you have a dataset of a trillion tokens, you can mask fifteen percent of them and train on learning the fifteen percent, or you can train on all trillion of them at the straight scaling.
Now, if you have fifteen trillion tokens versus one trillion, that's another question. But for embedding small, for embedding tasks on smaller models, there's a better trade-off scaling curve at the start for encode- the encoder learns better with less tokens at first.
But then extending this out in pure scaling laws, uh, yeah, you lose a lot of your training data, right? And then that was one of the big points of why do we do next token prediction, because it scales better than other tasks, right?
So scaling laws were made to show better objectives, so it's directly against it. But then, uh, at small scale and stuff, there's, um, there's benefit in this specifically for like, uh, edge models. Like, uh, you can deploy a BART as a guardrail live, and you can have it intercept every query because it can act in milliseconds versus LLMs will still take longer, right?
So at a smaller scale, they'll be better. Um- There was another part to this that I'm blanking on. Oh, Reka AI. They're doing encoder-decoder generation models where they're adding decoder heads to encoders, and they're scaling them up to billions of parameters.
Um, they're, they're a case study of spending money to train them up pretty big.
There's a question from Isaac in the chat about, um, my use case at work. So currently, where we're at in the project is we need to accumulate some, uh, good training data. So we don't have enough training data to, like, actually train a BERT or, um, that type of model.
So to start with, we're just using, uh, LLMs and prompts to,
like, do some logical, uh, classification to, like, kind of bootstrap until we get enough data, uh, and then also to create a feedback loop where we can get, uh, feedback from people so that we'll have enough, like, solid training data, so we can actually train a, a model.
Um, the main purpose of it being faster performance, as Vibhu mentioned. Uh, then you're-- you can respond in milliseconds versus multiple seconds, uh, or tens of seconds if you're u-using an LLM.
Cool. Thank you, Eric. Um-
Wrap-up50:53
Yeah.
Always appreciate the OG papers.
Yeah. It was good to-
Anyone else want to volunteer next week? Any paper. Doesn't have to be a look back thing if anyone's interested in.
Yeah. I don't know if there's any paper that's caught my eye recently. Um, I guess, like, we talked a little bit about embedding papers. People are interested in embeddings.
Did we ever do the Gina 2 embeddings paper?
Uh, no. Uh, there's also Nomic Embed. I, I thought, I thought, like, um, I didn't see the Gina one. Um, I thought the Nomic one was pretty detailed in terms of, uh, what, uh, their process was, so be interesting.
I might be mixing up-
I can-
... another paper plot, but I think I remember we went through one embedding paper.
It, it may be Nomic. Maybe it's Nomic. Uh, but yeah. Probably wasn't there.
Nomic compares directly to Gina. I have the exact same thing. I haven't seen the Nomic one as much. I just know Gina was the one open source 8K context, uh, very detailed, here's how to do embeddings from scratch and fine-tune them paper.
But, uh, I guess if Nomic's the same thing, 50/50 if anyone wants to take one or both. Would love both. Oh, there's Gina 3 now. Crazy.
Um, okay. Well, um, I will volunteer for, uh, Nomic or Gina. Um, if anyone else has, um, papers you wanna cover in the meantime, let's, uh, cover them. But otherwise I don't wanna drag this too long. Um, yeah.
Nice chat.
All right. And thanks once again, Eric.
Thanks, guys.
Yeah. Thank you.
Thank you, sir. Thanks, everyone.
Bye. See ya.






