# [Paper Club] Embeddings in 2024: OpenAI, Nomic Embed, Jina Embed, cde-small-v1 - with swyx

Latent Space · 2024-12-01

<https://addtry.com/4ca5d503-a264-475f-bb99-6e068b8de62b>

Shawn Wang (swyx) surveys four key embedding developments: OpenAI's text-embedding-3 models, Nomic Embed, Jina Embed v3, and CDE-small-v1. OpenAI's models introduce Matryoshka embeddings, compressing output from 1024 to 64 dimensions (94% storage reduction) with only an 8% performance drop. Nomic Embed is a fully open-source BERT-based model trained on English data, offering reproducible code. Jina Embed v3 supports 89 languages and uses task-specific LoRA adapters—e.g., a retrieval adapter yields a huge boost over the base model. CDE-small-v1 (143M parameters) employs a two-stage adaptation: first conditioning on the corpus, then embedding; it outperforms 7B-parameter models on many tasks, though its stateful API complicates deployment.

## Questions this episode answers

### How do contextual document embeddings (CDE) work and what makes them so efficient?

Shawn Wang presents the CDE-small-v1 paper's two-stage embedding method: first, the entire corpus is embedded to create a context, then queries are embedded conditioned on that context. Unlike fine-tuning, there are no gradient updates; it is more like context caching. This allows a 143M parameter model to outperform 7B models, offering huge efficiency gains.

[35:10](https://addtry.com/4ca5d503-a264-475f-bb99-6e068b8de62b?t=2110000)

### How much storage can Matryoshka embeddings save without a big performance drop?

Shawn Wang explains that OpenAI's text-embedding-3 models offer Matryoshka embeddings, allowing dimension reduction from 1024 to 64, which saves 94% storage space. This results in only an 8% drop in performance; accuracy at five drops from 75.3 to 69.4, making it very practical for production use.

[6:00](https://addtry.com/4ca5d503-a264-475f-bb99-6e068b8de62b?t=360000)

### What are task-specific LoRA adapters in Jina Embed v3?

Shawn Wang describes Jina Embed v3's task-specific LoRA adapters, which allow loading different fine-tuned adapters for tasks like retrieval, classification, and clustering. Eugene notes that this contrasts with typical use of the same embedding model for all RAG tasks, and the adapters are applied across model layers, not just the last, yielding significant performance boosts in benchmarks.

[19:11](https://addtry.com/4ca5d503-a264-475f-bb99-6e068b8de62b?t=1151000)

## Key moments

- **[0:00] MTEB Overview**
  - [0:36] "If you use embeddings and you don't know MTEB, you don't know embeddings at all."
  - [1:21] Stella, ranked #6 on MTEB, is an order of magnitude more efficient in model size than higher-ranked models.
- **[2:30] OpenAI Embeddings**
  - [6:00] OpenAI's Matryoshka embeddings cut dimension from 1024 to 64, saving 94% storage with only an 8% performance drop.
- **[7:23] Dimension Tradeoffs**
  - [7:42] Eugene: "1024 is just not usable in production; 64, one to eight, those are amazing for latency."
  - [9:23] "Embeddings are a fantastic form of lock-in for any API provider, 'cause once you've embedded something, you still have to get it back."
- **[9:37] Nomic Embed**
  - [10:15] Nomic Embed is built on a BERT architecture, updated with modern training techniques.
  - [14:54] All popular embedding models are general-purpose language models; no dedicated code embedding models exist.
- **[15:56] Jina Adapters**
  - [18:20] Jina Embed v3 supports 89 languages, making it one of the most multilingual embedding models.
  - [19:11] Jina Embed v3 introduces task-specific LoRA adapters for retrieval, clustering, classification, and text matching.
  - [22:26] Shawn Wang: "Blindly applying embedding models without reading how they were trained is asking for failure."
  - [23:20] Jina Embed v3's task adapter evaluation uses fewer than 10 examples, so performance gains should be interpreted cautiously.
- **[27:34] Multimodal CLIP**
  - [28:31] Shawn Wang: every embedding paper should include qualitative output comparisons, like Jina-CLIP2's visual QA examples.
- **[29:33] Medical Models**
  - [30:35] Q: Are there high-quality biomedical embedding models? A: John Snow Labs offers some, but fine-tuning a general model is recommended.
- **[35:10] Contextual Embeddings**
  - [35:10] Contextual Document Embedding (CDE) uses a two-stage process: first condition on the corpus, then embed queries.
  - [35:49] CDE-small-v1, a 143M-parameter model, outperforms 7B-parameter models through corpus conditioning.
  - [40:45] CDE's arbitrary corpus conditioning could replace fixed task adapters like Jina's five LoRAs.

## Speakers

- **Eugene** (host)
- **Swyx** (host)

## Topics

Models

## Mentioned

Alibaba (company), Codium (company), Jina (company), John Snow Labs (company), Morph Labs (company), Nomic (company), OpenAI (company), BERT (product), BioBERT (product), CLIP (product), Cursor (product), Jina CLIP2 (product), Jina Embed (product), MTEB (product), Nomic Atlas (product), Nomic Embed (product), OpenAI embeddings (product), Qwen (product), SigLIP (product), cde-small-v1 (product)

## Transcript

### MTEB Overview

**Swyx** [0:00]
There was a whole bunch of, uh, interesting embedding work piling up. Uh, and I figured, like, the- it'd be good to have, like, a state of embeddings, uh, overview. And so, uh, we have basically one blog post and three papers that we're, that we're...

that I've sort of defined in scope. Um, they're all listed in the, the meeting notes here. Um, and I f- I would consider this basically everything that is relevant for understanding embeddings as of today. Um, and so I think that the first thing is to understand MTEB, um, which is Massive Text Embedding Benchmark.

This is the sort of de facto benchmark. Um, there are criticisms of it. Um, but, um, you know, it's, it's a, it's a pretty... Like if you, if you don't know, if you do, if you use embeddings and you don't know MTEB, you don't know, uh, embeddings at all.

Um, this changes a lot. Uh, it used to be that, uh, the Chinese models, uh, were completely dominating the top 10. Uh, now we have American Chinese

models. So other Chinese models and, uh, I don't, I don't know some of these, these guys. Uh, so I wouldn't pay strict attention to the ranking of, of these things, but just to know the, the main, um, benchmarks that, that people care about, as well as, um, the, the trade-offs in model size and memory usage.

Um, this becomes extremely, extremely relevant when it comes to efficiency comments, right? Like, um, even though Stella is ranked as number six, um, they're, they're, uh, at least an order of magnitude more efficient in model size for the same amount of performance that you might get from a much larger model.

Um, so, uh, practically you might just use this instead of something that's higher ranked. Uh, the other thing I think is also relevant is, um, I think, I'm not sure, but, uh, I don't know if they actually... Yeah, so, so they have everything in here, including cde-small, which we're gonna cover.

Um, what is text embedding for? Oh, this is for text. Got it. Um, where is the... Uh, so I don't know where, uh, the, the OpenAI models land, but I, I think, uh, definitely people should understand, uh, that this is the latest update for o- the OpenAI offering.

Um, typically you want to at least be familiar with the OpenAI offerings just because that's tends to be the starting point. That's the, that's the API key that everyone already has, uh, rather than adding a new API key for other, other models.

Um, so they're not gonna be the best in the world, but they're gonna be very, very good, and usually that's good enough. Uh, I would say the other thing to be aware of is, um, for the first time, they're offering, like, two different sizes.

### OpenAI Embeddings

**Swyx** [2:43]
Uh, and also Matryoshka, uh, embeddings, which... Matryoshka. Oh, um,

um, do you know how to... Where, where do I find the, the document-

**Guest** [2:56]
I think they didn't mention it.

**Swyx** [2:58]
Uh, they did.

**Guest** [2:58]
I think they did.

**Swyx** [2:59]
They added the blog-

**Guest** [3:00]
I think.

**Swyx** [3:00]
They added it to the blog post at the end.

Uh, are you gonna see all my emails? Um, okay. Well, never mind. Um, uh, maybe so I, I'm just setting up this browser for the first time, so OpenAI texts embeddings. Um,

there we go. Here. No. That's 2022. Um,

does anyone have that, that, um, that link? Um, I'm, I'm-

**Guest** [3:34]
Are you looking for Matryoshka or the one where they ref- referenced Matryoshka?

**Swyx** [3:37]
Where they refreshed, where they refreshed it. Matryoshka. That'd be really awesome if they actually had it here. Uh, nope. Embeddings. Um, uh, huh. Okay. I can-

**Guest** [3:50]
Uh, okay. It's, it, it's actually the first pop-down. If you scroll the, if you scroll, uh, uh, if, if you scroll to this link I'm pasting in the chat here. Uh, give me a second. Shoot. Okay, I'm pasting this in the chat here.

If you open it. It's, it's the link that you had.

**Swyx** [4:10]
2024?

**Guest** [4:12]
No, no. It, it, uh, k- the, the link that you, you, you shared. Yeah, click on that, and if you scroll down, scroll down a little bit. Uh, scroll down a little bit more. Ah, reducing embedding dimensions.

**Swyx** [4:24]
There we go.

**Guest** [4:24]
Do you see that?

**Swyx** [4:25]
Uh-huh. Yes. Uh-

**Guest** [4:26]
It's actually hidden in there.

**Swyx** [4:27]
Is it-

**Guest** [4:28]
I don't know if actually the word Matryoshka shows up.

**Swyx** [4:30]
Let's see. Um, so, uh, whoever commented in the chat, uh, Kishore sent the ac- the, the blog post that they, that I was looking for. Uh-

**Guest** [4:40]
Oh, great

**Swyx** [4:41]
... and they did add it in the footnotes, um, because the author complained that they were not credited. Um, which is very, very shady of OpenAI. Um, so yeah, I mean, the- those, th- these guys were the first to offer it.

Uh, and it, uh, it is very good. Um, we'll see later in, in one of the Jina postings, um, how efficient it is. Um, I don't think they communicated very well here. Uh, but let me just skip ahead to the Jina posting, and then we'll, we'll, we'll show you.

So, um, uh, so the Matryoshka embeddings lets you reduce the amount of data that you store. Um,

this is so annoying. Wait, when I have the image in my head...

Wait, did they, did they get rid of it?

Wow. They got rid of it. Okay. So I guess I have to, uh, refer to my own blog post about it because they got rid of it. Oh. Oh, maybe it was in the paper. Ah, okay. Yeah, it was the paper.

Sorry. Um, sorry, sorry, sorry, sorry.

Um, I'm so sorry. Let me, let me refer to my own notes. Uh, here. Um, here.

There we go Um, so when you offer Matryoshka, um, you can do something like this, where you compress... Like, let's say the original output dimensions is 1024. You can compress it to 64, so you're reducing, uh, the amount of storage space by 94%, um, and that only results in an 8% drop.

So basically from here, 1024 down to 64, uh, goes... Your performance drops from 75 to, like, 69 or whatever, um, which is, which is pretty good. Um, maybe, maybe... So accuracy at one would be, like, 47.2 going down to 41.3, so that's like a 6% drop.

Uh, and then accuracy at, uh, five would be 75.3 going down to 69.4. Um, so, like, that's a huge amount of information, um, that is storage that is saved, um, as, as well as compute everything, um, uh, for, for, for, for a really good drop.

Um, and basically OpenAI pioneered, uh, this. Um, they were, they were, they were the first to, uh, acknowledge that this is relevant, and now basically every offering should do it. Um, and, uh, yeah, I, I think that's... Those are, those are state-of-the-art.

I, I don't know if anyone else has played with the OpenAI embedding models enough to, uh, offer any more notes before I move on to the open models. Uh, but I just want to start with OpenAI.

### Dimension Tradeoffs

**Guest** [7:23]
And the thing is, training these Matryoshka embeddings essentially comes for free. You just need to update the loss function.

**Swyx** [7:29]
Yes.

**Guest** [7:29]
Right? Uh, you can do it. And I've tried something like this, where I cut the embeddings by a quarter of the size, and it's, it's almost, uh, the fi- it's almost as good fidelity. And the thing is, okay, you might think that 1024 to 64, okay, that's not such a big drop.

But 1024 is just not usable in production. Depending on your production use cases, you may... will not be able to meet the latency. But 64, one to eight, those are amazing. Um, so it's, it's essentially the, the boundary between what's usable and what's not.

**Swyx** [7:56]
What exactly... Uh, so Dan says 1024 is enormous. I, I mean, I don't have a-

**Guest** [7:59]
Yeah, it-

**Swyx** [8:00]
Like, what do you mean en- It's just, it's just more numbers to store. Like, what, what's the-

**Guest** [8:04]
If, if you think about it this way, um, as your embedding size increases, your approximate nearest neighbors lookup will increase as well.

**Swyx** [8:11]
Okay, okay. So this is more about, like, the N squared explosion.

**Guest** [8:15]
It's like, yeah, looking up, doing dot products, et cetera. So it just costs more to compute.

**Swyx** [8:21]
Vibu, are you there? It's like, uh, I think, I think you have a lot more to, to add to the embedding search space segment.

Uh, I'm, I'm not sure. Vibu's, uh... I, I know he's in San Diego with family, so, um, I, I don't really, uh, know if he's, he's, uh, able to comment. Um-

**Guest** [8:38]
Oh, I think he dropped off.

**Swyx** [8:40]
Uh, yeah. Like, we-

**Guest** [8:41]
But-

**Swyx** [8:41]
... we've already done a session on, on MRL, um, so we can refer people to that MRL paper if, if they want to. Uh, I was just more, more just like, you know, what should you know with, like, state-of-the-art end-of-2024 embeddings?

Um, this would be it. That there- there's probably different sizes that you should, you should, uh, be aware of. You should know the models. You should know the costs. Um, uh, it, it, it's very cheap. Um, they, they...

I, I feel like they're basically embedding this, uh, this for you at cost. Um, mostly because embeddings are a fantastic form of lock-in for any API provider, um, 'cause once you've embedded something, you still have to, uh, get it back.

Um, okay. Um, all right. So, uh, let me just continue, unless Sam has... Yeah, Sam has other questions that people can answer. So, then we're going to move on to the papers. Uh, I think the first one I would highlight is Nomic.

Um, because, uh, the reason I picked this was because someone was asking, um, whether there's a good paper on the full, um, training process of a, of a, of a model. And, um, uh, Nomic is the closest that I can find, that I know of.

### Nomic Embed

**Swyx** [9:55]
Um, definitely, there's a US bias here because there's a whole bunch of, um, Chinese, uh, embedding papers that probably have some detail on, on their training process. Uh, but Nomic has open source, uh, code, open data, open training code, and, uh, full reproducibility, which, um, in my mind is, like, good enough, right?

Like, if you wanted to de- deep dive into, into that. Um, the main thing I, you know... The, the main thing I would, I would, uh, highlight is, uh, they actually use BERT, which is, um, uh, a, a good follow-up from, from, uh, last week.

Um, and, uh, basically, the, what, what we call the Nomic zero stack, um, is, uh, i- is pretty, is pretty standard. Like, these are all basically, like, state-of-the-art in terms of training processes, uh, and training tech, uh, as far as I understand from, like, every single model, model trainer that I've talked to.

Um, I'm not sure about the masking. Um, I, I actually did not understand. I thought that you just kind of mask individual tokens. I, I thought that was standard. I didn't know there was a hyperparameter here around mask rate of 30% versus 15%.

Um, and, uh, I... It's not something that I'm familiar with, and neither, neither am I familiar with, like, a lot of these other, uh, types of, uh, sort of, uh, BERT-based models. Um, but I'm c- I'm curious if anyone has, like, thoughts or questions around, um, what you would like to-

**Eugene** [11:20]
Dr. Charles already came and saw you, right?

**Guest** [11:21]
Yes.

**Eugene** [11:22]
Okay, so after this you're done.

**Swyx** [11:23]
Um, Daniel-

**Guest** [11:24]
All right.

**Swyx** [11:25]
You're unmuted. I don't know if that's on purpose.

**Eugene** [11:27]
Okay, sorry.

**Swyx** [11:29]
All right. Um, does... Yeah. Has, has anyone, has anyone, like, checked out Nomic? Um, are, are you interested in, like, going into any detail? Um, you know, I'm fairly friendly with that team. Uh, Kishor says, "Original BERT paper masked 15%."

Oh, I didn't know that. Um, yeah. So, um, so yeah. I mean, that, that was, that was a new finding for me. Um, the rest of it I think was, like, relatively, um, unsurprising, uh, you know, for its time.

Um, I would say, like, um, the, uh, the interesting thing... I mean, the, one of the reasons that Nomic is investing in this is because they sell, um, a cluster visualization tool. Which is Nomic Atlas. Um, and so they, they're interested in basically just building, um, tools for embeddings, uh, for you to explore your data sets.

Um, RJ says, "Is there a study of vector retrieval speed versus embedding size?" Uh, no, but I guess they're correlated. I don't know what, uh, specifically you would want. Um, yeah, more, more detail would be great. Um, yeah, but, but yeah.

I, so I would say, like, if you want the sort of s-state-of-the-art process or paper, uh, or like data or code, um, I would just come and grab it off of here. Um, they've, uh, they found this a lot.

Um,

RJ says, "We're discussing large embeddings equals bad," so you want to quantify. Um, yeah, I don't... How would you quantify it, um, Eugene? It sounds like you've, you've had some direction.

**Eugene** [13:00]
I got, I got you, RJ. Um, and I can address this, um, when you finish whatever you want to say and we turn off recordings.

**Swyx** [13:08]
Okay. All right, keep that in mind, uh, as, as we go. Um, but yeah. I mean, um, I like, uh, you know, a-and you can, you can go through the sort of no- the, the Nomic paper here. Um, I would say, like, pretty straightforward training stuff here.

Um, you know, like I, I, I just think it's nice to have a starting document where you just have all the, the tech choices, the hyperparameters, um, and, you know, and reproducible code. Um, I, I, I, I don't know.

Like, to me, this is, like, where you start. Um, I also think that, um, like the, uh, this is all... These, these prefixes and stuff, um, basically completely reflect BERT. Like, this is just updating BERT, uh, in, in every shape and form.

Um, which is kind of nice. Like, I, I, um, I never really thought about that. Um, these are all the Chinese models that I talked about. Um, so if you want the, uh, their papers, I'm sure they're all reflected here as well.

Uh, I have not read them. Um, but yeah.

**Eugene** [14:03]
I was a little surprised that it was just BERT updated and modified slightly. Um, but I wonder, I, I wonder if that's because the, like, true value out of this neural net would be in the data that's coming into it, meaning, like, it's more dependent on the data than the, than the architecture.

Uh, I say that, but it's likely, you know, both.

**Swyx** [14:29]
Yeah. Um, yeah, I, I, I've got nothing for you there. Um, one comment I'll make, I, I'll, uh, share as well about this that, uh, I have, I've had other founders tell me, uh, which is that it's very surprising that, um, all embedding models are effectively, um, like, general purpose, and there's no code embedding models.

Um, so if you look at the, the Nomic, um, data sets, um, code is like number 10 down the list, um, for less than 1% of the data set. Um, and then Stack Exchange is maybe down here. I don't even...

The Stack Exchange is not even like a code-specific thing. Um, and so, like, people have had, uh, like the, the, the Codiums and the Curses of the world, they've had to, and, and Morph Labs, they've had to create their own code embedding models that they don't release.

Um, uh, which is surprising. It's IP, but, like, it sounds like a, a high-potential thing for some, you know, PhD person to, to publish as their, as their sort of open research article. Um, 'cause as of right now, every single embedding model is just general purpose language.

Uh, and obviously, that is, uh, different. Uh, we'll, we'll cover a little bit of, like, how to change that with, uh, CDEs. Um, but, um, I think I'll move on to Jina unless, unless anyone has, has issues. Um, okay.

So, um, Nomic was very focused on, um, single language, um, English. Uh, I, I, I think they have some multilingual capability. Um, um, multi...

### Jina Adapters

**Swyx** [16:13]
Uh, I, I don't know what the... Spanish. Uh, okay. It, it wasn't, it wasn't covered, it wasn't covered in here, but I, I did talk to them about, about this. But anyway, so Jina specifically as a European company, uh, very, very focused on multilinguality.

Uh, so this would be kind of their update. Um, I would also say that I've been very impressed by their out-of-the-box offering. So when we talked about, um, clip embeddings, right? Um, so th-this is, this is one of the AI News, uh, articles from last week.

Um, you can see that, you can see the difference between, uh, a paper that is, uh, very focused on research technique or, and, and algorithms, and a very, a paper that is focused on being a technical specification for an API that they intend to offer.

Um, and so Jina-CLIP2, uh, is basically... I'm just gonna chuck that in here as part of the reading. Um, Jina-CLIP2 came out of the box. This is actually what I was looking for, by the way. Um, so, uh, Jina CLIP, CLIP2 came out of the box with like, "Here's how you deploy to AWS, Azure, Google Cloud."

You won't get this from, from an Apple paper. Um, and that's just because they're trying to make money off, off of, uh, their API calls, right? Um, but let's, let's rewind to embeddings. So basically, um, they, uh, they have, they, they've been running their own embeddings for a while.

Uh, they updated this, uh, in September. Um, and their focus has been multi- multilinguality. Um, so there's a variance of MTEB for multilinguality. I don't know if it's, uh, it's here. Yeah, it's just Chinese. Um, uh, French. Yeah.

Uh, there's, there's some Japanese somewhere as well. Like, I, I, I don't think it's like this specific leaderboard, but there's a, there's a different leaderboard as well. Um, and, uh, it, it's mostly a function of data set. Um, I don't think there's, there's anything particularly a call-out here apart from, like, um, they also, you know, have, have like really, really good thoughts on, um, scaling laws and- Uh, and, and the kind of data set that, like, works well for cross, cross-language transfer.

Um, so they support 89 languages, which is, uh, which is pretty massive. Uh, and I think they're also very practical around, around like, uh, the size of model. Um, you know, um, if you look at the size of the models here, some of them are like 7B models, uh, which is huge.

Um, and, and so like, yeah, I, I, I don't know if the people are actually interested in using these 7Bs models. Like, these... This is definitely sort of benchmark maxing, um, compared to the more practical, uh, oriented people who are like, "No, like you actually want to use like a Roberta."

Um, and, uh, and keep it, keep it to like sub-1B, uh, for, for actual sort of inferencing for, uh, for embeddings. Um, the other thing that, uh, I will call out here is, um, just like the, uh, LoRA adapters, um, they are s- they also introduce, um, these like...

this concept of like task-specific LoRA adapters. Um, so let me see if, if, if they cover it in... Uh,

here we go. Yeah. Um, so, uh, this is, this is start, this is where you start to see, like instead of the traditional, um, single model type of, um, of embedding model, they use, they use unomic, where it's like this is like a d- that...

Basically, what, everything that we had up till 2024 was just like single embedding model. Um, here we have task-specific adapters, and we'll see another form of adapters with, with the last paper today. Um, but, um, the... Uh, where, where am I looking at?

Uh,

where are the adapters? Okay. Yeah. So, so they have, um,

they have retrieval for documents, retrieval for queries, uh, separation of documents and, uh, and clustering them, class- classifying them, and then, uh, text fetching, which is I think the classic workload. Um, and I think like we tend to use, at least in traditional RAG in, in, in how I learned it and how I think most people use it, we tend to use the same embedding model and the same mode, uh, for all these things and maybe try to prompt differently or pre-process differently to, to get performance out of them.

Um, but, uh, training LoRAs for individual models for, for, for the different RAG tasks, uh, I think is very interesting, um, and probably a good, like a very good idea. 'Cause, 'cause they're basically doing different things. Um, they, they, they have different tasks, uh, over here, and yeah, I think there, there's like ablations for like, um, the, the accuracy and precision of, of, uh, each of these things.

Um, but like, you know, I, I would say the main contribution of this paper is, is just the idea that you should have task adapters. Uh, they also have MRLs. Um, so like w- I don't think we should just...

You know, I'm just gonna leave the MRL discussion aside. Like, we all know that it's good. Um, uh, I think the, the main idea to, to get from here is the sort of idea of task-specific adapters. Does anyone have questions?

I haven't been looking at the chat.

Okay, people are still talking.

**Guest** [21:23]
It's mostly the debate on the dimension size.

**Swyx** [21:27]
Uh, what?

**Guest** [21:29]
It's mostly the debate on dimension s- dim size.

**Swyx** [21:34]
Okay. Cool.

**Guest 2** [21:34]
I was hoping that they would have done an ablation of the adapters versus... They did ablation of one versus two adapters, but they didn't do one on no adapter versus adapter, which is, I think, in my opinion, more interesting.

**Swyx** [21:47]
Um, no adapter versus adapter. You mean on their own model?

**Guest** [21:53]
Isn't, isn't that table six?

**Swyx** [21:57]
Yeah. S- Oh, did they-

**Guest** [21:58]
Table six in the results? Um, yeah, it's like they did... If you look at it-

**Swyx** [22:05]
They did it in me too

**Guest** [22:05]
... the second row from the... Yeah, Jina is not, Jina v2 has no ad- adapters.

**Swyx** [22:09]
Mm-hmm.

**Guest** [22:09]
And then, you know, Jina v3 one star is like pair training. I think that's no adapter. And then they have a retrieval adapter, which is the la- last row, which you can see a huge boost.

**Swyx** [22:19]
Oh.

**Guest** [22:19]
Well, specifically for retrieval. Yeah.

**Swyx** [22:21]
I see.

**Guest** [22:22]
At least that's how I inter- at least that's how I interpreted it. Please let me know if I'm-

**Swyx** [22:25]
Yeah

**Guest** [22:25]
... uh, if I'm misreading that.

**Swyx** [22:26]
That makes sense. Yeah. Yeah. Okay. Got it. Thank you. It's intuitive that it's a fairly big lift. I, I mean, like, um, I think I, I ca- I, I can't remember the, the actual wording, but the most influential thing somebody said to me was like, you know, like, um, blindly, um, applying embedding models onto any arbitrary task without actually reading the paper and how it's trained, like it's, it's asking for failure.

Because, like, em- embeddings have, like embedding models have very specific ex- assumptions into it. Um, and so it makes sense that splitting out the assumptions into like the top five use cases and splitting them out would have very material impact on how the embedding, uh, works.

Um, and it, I mean, you can look at the, the numbers here. They're, they're pretty big lifts.

**Guest** [23:07]
Yeah. Also, that said, take the results here with a pinch of salt, though. If you scroll down a little bit more, swyx, in the second paragraph on the left, you can see that this, their evaluation set size is only 10, fewer than 10 examples.

Um, so they added synthetically generated data, et cetera, et cetera. Um, yeah. So we'll see, but it's a very good result, and I, I'm glad people, um, pushing on using adapters more.

**Swyx** [23:34]
Yeah. How about-

**Guest 2** [23:35]
Um, so I guess the other thing from the other paper, uh, the Nomic paper, they had used prefixes, which apparently is a thing that has been done in training for a little bit. And, and the, the prefix, um, is kind of, in my mind, similar to the adapter in that you're just training different-

**Swyx** [23:55]
Um-

**Guest 2** [23:56]
... um, tagging things. But the difference with the adapter is you have a different loss function, right? Per-

**Swyx** [24:02]
Yeah

**Guest 2** [24:02]
... per type. And so I would have been interested to see the, an ablation there as well.

**Swyx** [24:09]
I mean, um, I would say, I, I agree with you that prefixes are a standard part of the toolkit, therefore, um, that, that would be kind of covered in like the base v2 versus the v3s. No?

**Guest 2** [24:23]
So but when you have to, like, I, I was ac- this was another point that I was trying to understand is I couldn't find any evidence that there's any place where pe- they were, like, putting the prefix in the Nomic paper.

Like, it seems like you would be able to improve a s- task, task-specific embedding by putting the prefix into the query, right? So it, because it was trained on that prefix, so presumably it would be better if you also used it in the query.

**Swyx** [24:51]
Yeah. Yeah, no comment on that. Um, yeah.

**Guest 3** [24:56]
I, excuse me. I have a question about figure one. Does that... It shows two imp in the Jina, in the Jina paper. Does that mean that the input's both the query and the supplied classification, or is there, is there a classification model that is a part of this, uh, embeddings model that auto-classifies?

**Swyx** [25:20]
Uh, no, it's just loading in the classification LoRA.

**Guest** [25:23]
Exactly. There's no separate classification model. It's really just embedding the text, embedding the, um... Yeah, like right now, and then if there's a classification label, embedding a classification label and just doing the classification.

**Guest 3** [25:36]
So that means that the program or user supplies the, the class- the adapter task?

**Swyx** [25:46]
Uh, yes.

**Guest 3** [25:47]
Okay.

**Swyx** [25:47]
If you look at the, if you look at the API for this, um, you see this?

So it has-

**Guest 3** [25:55]
Oh, yeah

**Swyx** [25:56]
... that has, that's, that's the param that has to match with one of the five that they give you.

**Guest 3** [26:01]
Do they, does Jina supply a classification model, or is it up to the the-

**Swyx** [26:06]
Uh, no. You can only use what they, what they give you.

**Guest** [26:09]
They have a classification LoRA, but you have to provide your own labels.

**Guest 3** [26:13]
Okay. Cool.

**Swyx** [26:16]
Um, one thing, uh, uh, uh, for, for those who are more familiar with LORAs, isn't this wrong? Why, why, why is it side by side? Shouldn't the LORAs be the last layer?

**Guest** [26:28]
Mm, the LORAs are usually on the MLP layers. Um, so it's on, like, every MLP layer or every, like, query key attention value layer. So at least how I use it is I apply LORAs on all the, um, MLP key, query-key value layers.

So it's not just the last layer. Um, I think maybe what you're thinking of is maybe fine-tuning. By adding a special last layer to fine-tune, that you freeze all the weights and fine-tune that special last layer for the specific task.

**Swyx** [26:57]
Yeah.

**Guest** [26:57]
Yeah.

**Swyx** [26:57]
That's what I'm familiar with for LORAs, but I, I guess I might be, I, I might be very focused on, like, diffusion LORAs. Um...

**Guest** [27:04]
Yeah. Oh, I, I think that could be it. Yeah. That, that could be it. I think in LORAs-

**Swyx** [27:08]
Okay.

**Guest** [27:08]
It's mostly all the weights except for the embedding weights.

**Swyx** [27:12]
Okay. Got it. Um, that's, that's not very low rank to me, but okay.

**Guest** [27:18]
Well, it's low rank in the sense that the LoRA dimension is very small, in the sense you can compress it.

**Swyx** [27:23]
Yeah. Yeah. All right. Uh, I'll throw an honorable mention to CLIP, even though I, I didn't, I didn't, uh, mention this, just because, uh, th- this was also part of the reason why I chose this topic for this week.

Um, 'cause there's a whole bunch of embedding shit that just came out in the last two weeks. So, uh, I just wanted to, like, here, here is the state-of-the-art, here's everything I know, and then just, just kind of discuss it.

### Multimodal CLIP

**Swyx** [27:42]
So they took Embeddings V3 and then, uh, jammed it into CLIP. Um, so here's, here's the, the same weights for Embeddings V3, froze it, uh, and then they have, um, this other vision, um, sorry, this other vision adapter here.

Um, and that's it. Text embeddings, vision, vision embeddings, you get a CLIP model. Uh, we've covered CLIP in the past. Uh, but basically for, you know, for, for a refresher of those, those people who don't remember, uh... Where is the goddamn CLIP paper?

Ah, I hate it when they don't, when they don't show everything that's important. Okay, I have to go back to my, my new, my summarization again. Um, oh, okay. Um, it was a different paper, unfortunately. Okay. But this is the Apple one, but, like, I, I just, I really love this example.

Every example, every paper should have qualitative example of the output compared to competitors, right? 'Cause then, then you understand, uh, fundamentally what they're going for, because they are showing you how they want it to be used. And so, for example, visual QA stuff like this is really cool because, um, you, you, you could, you, you can definitely see yourself having a, having an image like this where, you know, there's a number on the screen and you say, "What is the weight of the luggage?"

OpenAI CLIP gets it wrong, SigLIP gets it wrong, and, uh, you know, your model gets it correct, right? Obviously, these are all gonna be cherry-picked, but at least it's, it gives you an idea of what's in the damn dataset.

Um,

uh, that I, that I find it hard to, to get. So, uh, Jina unfortunately did not do this. Um, at least that I can tell. Uh, but at least, you know, they, they publish a lot of, uh, really sort of technical, quantitative, uh, specs.

And it's based on Embeddings V3. Um, so this, this is how foundational embedding models are. Okay. I wanna move on to the last one, um, unless people have questions. I, I haven't been looking at the comments here. Uh, blah, blah, blah.

Blah, blah, blah. Ooh, uh, any- anyone have interesting comments? Um,

### Medical Models

**Swyx** [29:40]
yes. Yes, you have to, you have to click on my screen. Um, Zoom has made it, uh, easy to, to miss the, to miss the screen. Okay. K asks about-

**Guest 3** [29:50]
I have a... Oh.

**Swyx** [29:51]
Go ahead.

**Guest 3** [29:51]
Quick. So is Jina, like, a research lab or-

**Swyx** [29:55]
It's a startup

**Guest 3** [29:56]
... a company or-

**Swyx** [29:57]
No, it's a-

**Guest 3** [29:57]
Okay. Startup.

**Swyx** [29:58]
It's a, a Chinese founder, lives in Germany. I, I met him at, uh, in Vienna. Uh, a very nice guy, a big fan of Latent Space. Um, he- we'll have him on at some point. Um, it's, it's... For me, there's like 10 of these, uh, you know, and, uh, it's hard for me to, like, figure out who to talk to.

But Jina seems to do solid work-

**Guest 3** [30:16]
Mm-hmm

**Swyx** [30:16]
... and they're very, very serious, and I mean, look at the quality of their, their stuff. Like, it's obvious that they're-

**Guest 3** [30:20]
Yeah

**Swyx** [30:20]
... serious about it. Um, so yeah. They're, you know, they're a startup trying to, trying to make it. Um- Cool. Uh, has anyone ex- experimented with medical embedding models? Okay, um, I'm gonna go ahead and guess no, but, uh, can you, can you show...

Can you tell us your interest?

**Guest 4** [30:39]
Can you hear me?

**Swyx** [30:40]
Yeah. Hey, we can hear you.

**Guest 4** [30:41]
Uh, hi. Um, yeah, so I'm currently working on, like, um, with the QN, um, multimodal, um, model. I'm, um, working on, um, a retrieval system for medical papers and, um, I'm currently, yeah, trying different models, and, uh, there's this BioBERT.

That's a Q-Qwen from, from Alibaba.

**Swyx** [31:09]
Oh, okay.

**Guest 4** [31:09]
It's this new Qwen, yeah. Um, and BioBERT was one, but I'm not 100% co- uh, yeah. Uh, it's, it's not really, um, good for my use case, so... And there is not, not a ton. So there's John Snow Labs, they are quite active in this area.

So John Snow, like, from "Game of Thrones." Um, yeah. But I, I was wondering if any of you have, have experience with, uh, models that were trained on, like, biomedical data and have high embedding quality.

**Swyx** [31:47]
Um-

**Guest 4** [31:47]
Also for, like, the, uh, especially, like, chemical or biochemical, uh, like, protein pathways, um, like this.

**Swyx** [31:57]
So I, I don't think anyone here does medical stuff, but Tanishq in our Discord does. Um, so, um, Tanishq, uh, I science lover, I think. Do you have any recs on biomedical embedding? I think he got you. Um, if he, if it doesn't exist, he'll train it for you.

Uh, for-

**Guest 4** [32:20]
Yeah, I was also maybe starting on the weekend, like, a, um, my own embedding model just-

**Swyx** [32:26]
Ooh

**Guest 4** [32:26]
... with a budget of a few. Yep.

**Swyx** [32:29]
Yeah.

**Guest 4** [32:30]
Maybe let's see what come, comes out.

**Swyx** [32:32]
Uh, I'd be, I'd be very interested to see if, uh, the, uh, Nomic code works for you. Um, because-

**Guest 4** [32:38]
Yep

**Swyx** [32:38]
... uh, th- th- this is supposed to be the, the re- Like, you just swap out the dataset and you, you, you know, just run the same code again. You know? Um, I don't know.

**Guest 4** [32:47]
Yeah, but that's, that's the theory. In practice, you know, there's, it will not work in the, in the first go. But, uh, yeah.

**Swyx** [32:54]
Uh, uh, someone also, Navs, Navs also says you can just fine-tune a gener- generic one. I definitely agree with that. Um, uh, yeah, anyone from the sort of fast AI community would be horrified that you s- you should start from random weights.

Uh, you should just start from something that's a decent weight. Um, okay. Um, oh, uh, Khalid says, "Med-Gemini AI Healthcare." What is that? Is that, is that a, is that a thing?

Oh, okay. Uh, ooh. Yeah. You know, Google keeps doing this stuff, and then I, I just ignore it 'cause I don't do m- any medic- medical stuff, but yeah, this sounds really awesome. Oh, Sam. Sam, Sam has a med model.

I f- I forgot.

**Guest 5** [33:35]
Yeah, but Sam is for segmentation. Uh, it's not a foundation model.

**Swyx** [33:42]
Got it.

**Guest 5** [33:42]
They, they have a, a, a Med, uh, SAM, which is good for, uh, segmentation.

**Swyx** [33:50]
Oh, no, no, no, no. Uh, when I say SAM, I mean Sam Julian, uh, who's in the, in the chat.

**Guest 5** [33:54]
Oh, okay.

**Swyx** [33:56]
Not, not, not segment anything model, no.

**Guest 5** [33:59]
There is a SAM that is a segment anything, uh, model from, uh-

**Swyx** [34:04]
Yeah, yeah

**Guest 5** [34:04]
... Medcom.

**Swyx** [34:05]
We've, uh, we've interviewed them. Twice, actually. Um, so we have-

**Guest 5** [34:09]
Yeah. Okay.

**Swyx** [34:10]
Uh, I think it's, I think it's Sam too. Uh, yeah, these guys. Uh, Nick- Nikki is a friend. Um, and Roboflow is a, a very def- friend of ours. Um, but they all-

**Guest 5** [34:19]
Yeah, and the good, the good thing is that they also have a specific, uh, model for, uh, medical imaging, which is-

**Swyx** [34:26]
Yeah

**Guest 5** [34:26]
... uh, what I work on. Uh, but, but since, uh, it's a coincidence that you mentioned SAM, uh, as we were talking about, uh, Med-Gemini and, uh...

**Swyx** [34:37]
Yes. People have used it for medical applications. I, I believe, um, uh, Joseph in this podcast actually mentions it. Um, but it, it's... Yeah, I, I don't have the domain expertise to, to go beyond that. But yes, people have fine-tuned, uh, SAM to do, uh, medical imaging.

**Guest 5** [34:52]
Y- no, you can just, uh, write MedSAM, and you will get the GitHub part.

**Swyx** [34:57]
Okay. Yeah, yeah. Sorry. All right. Um, cool. All right. Um, let me, let me round out the other stuff, and then, and then we can sort of, uh, jump to Q&A, 'cause I'm also keen to hear, uh, Eugene's, um, take on, on embeddings.

Um, so the last thing I'll, I'll highlight for you guys is contextual embeddings. So I, I'm basically trying to organizing it in terms of progression of what I've seen in embeddings this year. So there was Nomic, which has started the year, Jina, which introduced Task Loras, and contextual embed- embeddings, um, now introduced this idea of a two-stage adaptation, where, um, uh, they, they, they specifically help you...

### Contextual Embeddings

**Swyx** [35:35]
Uh, they specifically train the model to be conditioned on the corpus first to then, uh, be used on, in an embedding context. Which is, which is a little bit weird, um, but it also helps them be very, very OP in one specific aspect, which is efficiency.

Um, so they are, um... I think if we, we go to em... Uh, they're still up here somewhere. It's very hard to, like, keep up. Um, so, uh, they are 143 million parameter model, uh, competing with 7 billion parameter models, um, because of this adaptation.

Uh, so it's a little bit cheating to put them, uh, as an apple- on an apples-to-apples comparison with these guys, because their deployment model is a little bit different. Um, what they're doing is basically, uh, first you consume a context.

Uh, let me s- let me see if I can, uh, show the code.

Um,

uh, where is it? I don't know if I saw it here inside the... Um, I think it's, uh, there might be the GitHub. Um, where's the GitHub? Uh, let's see. GitHub, GitHub, GitHub. Um,

sorry, I d- I, I don't, I don't think I put it in my notes here. Contextual embed- embeddings. There we go. Okay. Uh, yeah, here. Um,

okay. So yeah, th- this is what I wanted to show you. So instead of just, like, a single shot, uh, "Here's a bunch of text embedded. Embed this, please," um, that's basically what all the other models did, uh, in, in, in the Jina model, you maybe specify, like, a task, right?

So to load the LoRA. Here, you actually kind of construct the LoRA as you go, right? So you feed in the, the corpus first, feed all of it, and then you get dataset embeddings for the first stage on the whole thing.

Then the second stage, you, you, you, you use it, um, to actually do your prompt query. Uh, which is kind of slow, uh, for loading, but then, uh, you can understand why this w- uh, domain adapted so much better than basically every other method out there.

Um, and it's such a simple idea that you just, um, train your model in a sort of two-stage process. Um, so these guys worked it out. Um, and, and, uh, the technique is, you know, above my pay grade, but it's a whole bunch of math, whatever.

Uh, but, like, the, the conditional, that conditional aspect, I think, uh, makes a ton of sense to me. And, uh, this... I- in my mind, like, if this method proves popular enough, basically everyone is gonna do it because it's such a cheap win, uh, especially for the efficiency.

So, uh, I'll pause there. Um-

**Guest** [38:14]
Is that the contextual embedding paper that you mentioned?

**Swyx** [38:17]
Yeah, yeah. Uh, s- yeah, I, I just talked.

**Guest** [38:18]
Yeah. I, I, I think most, I think even Jina and Nomic, they actually adopt that methodology. I, I think there's, there's two things. One is updates to the architecture. Another one is updates to the training methodology. Essentially, they say that they do some clustering, and then they feed in the batches from the same cluster.

I think Jina and Nomic, um, also do that, where they say that they feed in the batch to make sure that the data comes from the same, uh, dataset. They don't actually mix data sources across different datasets.

**Swyx** [38:47]
Oh, but-

**Guest** [38:47]
Um, but I think what's unique-

**Swyx** [38:48]
... inference-

**Guest** [38:48]
Oh, sorry

**Swyx** [38:49]
... inference is only one, one run. Like, you know what I mean? Like, they, they don't let you domain adapt this thing.

**Guest** [38:55]
Yeah, that's true. That's true. Exactly. So what's unique is their architecture, whereby the, the inference, they actually allow you to provide some priors on your existing domain. That's quite interesting to me, and that was new to me as well.

**Swyx** [39:08]
Yeah. Um, so I, I think it's a very good idea. Um, I would love for other people to adopt it. Uh, this might be one of those things where, um, it just takes one of the big labs to read the paper and figure out that this, this makes sense.

I think, uh, the, the, the other deployment issue is that it's, it's a basically a stateful API. Um, so you cannot... Like, all these are stateless, which is great. Um, sorry, this is stateless. So you just call an endpoint, right?

All the model labs love this kind of model. But here you're gonna have to, like, call an endpoint to, um, to embed this model first and then return a, a new endpoint that you can, that you can actually do the embeddings on.

Um, so it might be a little bit annoying to, for, for these guys to figure out. But if it's a, if it's a big enough deal, they'll figure it out. Um, but, like, the, the lifts are, the lifts are very, very great.

Uh, if you look at, um,

um, some of the, the, the data that they have, um,

like, yeah, just um, just across all these models, uh, n- keep in mind that these, the- they're at least an order of magnitude smaller than all these guys. They actually perform better on basically every task. Um, it, it, it's, it's pretty crazy.

Um, so, uh, it will be interesting, like, there's no reason for it to be small if you can just make it big, but you just keep the, the same, the technique the same. Um, this was trained on a grad student budget.

Like, if you just keep, if you just scale this up, I, I, like, I think it, I think it would work. I think people would use it. It basically is a more generalized version of, uh, this, this, uh, you know, task adapter API, right?

So instead of having only five, uh, task adapters, what if you could just come up with your own task adapters just arbitrarily by feeding in the corpus that you're trying to embed?

That, that's, to me, that's a big idea anyway. Should I pause there? I don't know if there's any other questions. Uh, you can see the first set. Isn't that just fine-tuning? Um, no, because there's no gradient updates here.

Um,

yeah.

It, it's more like, uh, context caching maybe. I'll, I'll kind of liken it to that, where, um, you, you know, like, the, the, the, the, the initial context is pre-processed and s- as a KV cache, and you just kind of keep the KV cache around.

Um, that's effectively what you do for context caching. Okay.

**Guest** [41:39]
I think we can pause the recording then let Eugene do his hot takes.

**Swyx** [41:42]
Eugene hot takes, let's go. All right.

---

This library is powered by PodHood (https://podhood.com), the podcast website platform.
