LALatent SpaceNov 2, 2024· 52:22

[Paper Club] Intro to Diffusion Models and OpenAI sCM: Simple, Stable, Scalable Consistency Models

RJ Honicky presents OpenAI's sCM paper, which introduces techniques to simplify, stabilize, and scale continuous-time consistency models for faster image generation. The paper addresses instability issues by modifying skip connections, using cosine/sine schedules, adjusting the noise embedding scale to 0.02, applying adaptive double normalization, and implementing tangent warm-up over the first 10K iterations. These innovations enable one-step generation with FID scores rivaling GANs and multi-step diffusion models, while scaling experiments demonstrate performance up to 1.5B parameters. Distillation from scratch outperforms distillation from a teacher, and the episode contrasts these continuous-time methods with discrete-time approaches, explaining the intuition behind the mathematical choices.

  1. 0:00Intro
  2. 1:33Diffusion Recap
  3. 12:51SDEs & PF-ODE
  4. 18:27Consistency Models
  5. 27:21The Math
  6. 30:41Making It Stable
  7. 43:59Evaluation
  8. 49:24Q&A

Powered by PodHood

Transcript

Intro0:00

Host0:01

It should just record to my, my account so that I can upload it to YouTube.

RJ Honicky0:08

Um-

Host0:08

Is-

RJ Honicky0:09

Done. I did it.

Host0:10

Done. All right. Great.

RJ Honicky0:11

Okay, great. Oops.

Host0:12

All right, we're off.

RJ Honicky0:13

Great. So we're off. Uh, so, um, hi, everybody. I'm RJ. Um, and we're gonna talk about this paper. I hope people, uh, had a chance to at least pick through it. I know this is one of the hardest papers I've read in a long time.

Um, so I understand if you di- you didn't understand everything that's going on. I certainly didn't. And, um, I had to dig through a bunch of other papers to have the context also to understand it, so... And then there's a lot of proofs that are...

Like, take too long to read and are not as important to understanding the paper. So, like, there's a lot, a lot there. Um, but without... So without further ado, um, this, th- this sCM, there's a simple, stable, and scalable, and they actually...

This, this, um, the authors do a great job of actually going through all three of those things, um, in the paper. Um, the consistency model, we'll talk about what that means in a little bit. Um, but that sCM.

And then, like, why are they doing this? Well, because these continuous time consistency models show promise, but they're, um, but, but they're hard to train and hard to scale. So the paper kinda presents some techniques for getting over those challenges.

Um, okay, so I'm gonna... You know, as me and 6 were talking about, uh, in the beginning, um, there's... I think it is good to have just a recap of diffusion in general and, um, this is m- more complex version of diffusion, but I thought a lot of people have seen this diagram, so it'd be useful to talk about what we're talking about here.

Diffusion Recap1:33

RJ Honicky1:56

And, um, and in, in reference to, you know, sort of stable diffusion, um, and latent diffusion. Um, so the, the, you know, sort of the left, uh, space is the v- the variational autoencoder, and that lifts the pixels into latent space.

Um, and, uh, so we're not gonna talk about that. And then the right-hand side is the conditioning for, on images and text. So that's how you actually prompt the model instead of just gen- generating random images. And so we're also not gonna really talk about that.

Um, but suffice, these mechanisms are, um, they, they, they're compatible with the existing technologies for these things. So you can just view it kind of as a drop-in replacement for this green part. Um, and actually, uh, w- keep this diagram in mind.

Um, the U-net in the sort of, in the middle, the big part in the middle with the attention blocks there are, um, is, is basically identical even. So a lot of, a lot of this diagram is the same.

Um, and it's really about the training process and, you know, some tweaks to make the, the main network work well with that training process. Um, and, and you know, specifically, uh, with, with the regular diffusion models that we all know and love, they have to iterate multiple times generally to, um, generate a, a good image.

Whereas, um, the, the, these consistency models are designed so they can generate a good image with only one pass through the network, and there's a mechanism by which you can iterate, um, to refine if you want to. But the, you...

The idea is that you don't need to e- even one shot through the network is enough to get a really good image. Um, and I tried to put all the links that I took images and other things from.

So I'll share this, and everyone can, um, everyone can, uh, uh, click on those if they want. Um, so ple- and also, please, if you guys have questions, I know they go in the chat. I'm not watching the chat, so let me know if, um, you wanna interrupt and ask a question.

Um, so, uh, you know, this is just regular diffusion, and the idea is quite simple. You're just adding step by step some l- amount of noise. Um, and there's... This is actually scheduled. Um, I learned this, that it's not actually...

It's, um, y- uh, you know, like, it's not adding a constant amount of noise at a time. It does work that way, and so the top image is... Like that, they found, um, in this paper that it's actually inefficient, that you end up, um, with a lot of wasted, uh, a wa- wasted backwards proce- backwards diffusion, um, steps that you don't need, and you can cut out a whole bunch of it just by using a cosine schedule.

So the intensity of the noise is modulated by this cosine. Um, yeah. Um, and then I think this is where it gets interesting, the, the reverse diffusion process. And I actually learned, um, a little bit about this, uh, during...

wh- when I was studying this. Because, um, my impression was that the... You know, you sort of, like, crank in the image, um, and then it just sort of, like, progressively refines and refines and refines. But that's not actually quite what's happening.

Um, and this, this, um, diagram here kind of shows this. And so I wanna... We're gonna have to look at the next one to really understand what's going on here maybe fully. But the, the, um, if you look on the right, maybe starting on the right, um, you see the, like, the predicted noise at every step is actually, uh, like, you can't really make any sense of it, right?

It's all just noise. Um, and then in the middle column, you s- you have, like, slightly more, um- More and more refined picture. And then on the left you have this like sort of full noise removed, um, column.

And the weird thing is that, like if I look at, um, for example, uh, on the T equals forty row, um, and then I look at the T equals thirty, the full noise removed, uh, that doesn't look like...

Uh, like the one in the full noise removed column looks a lot more refined than the one in the input at T, right? And so like in the thirty... You would expect thirty, the output at thirty would be more than, like nicer than the input at forty.

And it's actually not the case, and the reason is because that's not exactly what's happening, right? It's not the output goes into the input, but rather the, um, the input is only being used to improve the noise estimate that's in the right-hand column.

So, so, um, that noise estimate is, is, uh, added to the noise in the, the... Like from the, um, the, the like sort of generated noise is added... The, the right-hand column is added to that generated noise to get the left-hand column.

So in the in... The middle column is only input to the model in order to refine how to generate that. Is that... I, I hope that... I want that thought to make sense. Um, and I'll switch this page, but if anyone wants to ask a question, now would be a good time 'cause I, I think...

And I tried to i-illuminate a little bit what's happening here. Um, uh, but I wanna... Okay. So I'll, I'll... If... Please interrupt if you wanna ask about this.

Host 27:54

It's good so far, RJ. Go for it.

RJ Honicky7:56

Okay, great. Um, so, uh, so now what's happening, and this is sort of like a bad illustration of the same concept, but in a different way. So the... If you look on the T equals ten, um, right-hand column, that's like just random noise that was generated as input into the model, right?

And then on the right-hand column you have, um, the noised versions of that noise. And so if you look at... Like, let's, let's look at the bottom, T equals eight, right? So time goes kinda backwards here. So left and...

Left-hand is T equals zero. Right hand is T equals ten. Um, and then T equals eight is two to the left of T equals ten. Um, and, and so the... And because the backwards diffusion process goes backwards in time.

So that T equals eight, what's happening here is that, that, um, the, the, the line... The thing at the top of the line is actually the input to the model. It goes through, um, you know, that, that, that thing with the attention block, and then it outputs something to add to T equals ten.

And then when you add those two things together, you get the thing at T equals zero down at the very left, lower left. Right? So then as time goes on, once we get to T five and we have a better estimate of that noise because we've gone through several steps, and then we add that to the T equals ten and we get that airplane that's kinda blurry to the...

In the top, top left. And then T equals three, and then finally we get to something that's really close to the original data. So, so and I'm making this distinction because it's very different. I, I hope this is clear because it's different than...

This is sort of the motivation for how to understand what's happening with these trajectories in the paper, is that you can see that these, these are separate trajectories and I, I've intentionally drawn that. It's not just like me being inaccurate.

I've intentionally drawn it that way because the, the locations that you're learning are not basically on the same trajec- in, in the, in this latent space. They're, they're not in the same trajectory as y- uh, uh, as each other, right?

So they like, um... And this causes a lot of inefficiency, and that's sort of the whole point to this, the, these consistency models and other full matching and other things that use this technique, is, is to sort of like be more efficient about the places you're sampling.

And use-

Host 210:37

RJ, and you're saying that, you're saying that the trajectories are not the same... Like T equals eight, T equals five, T equals three, they're not the same trajectories on each, each other-

RJ Honicky10:44

Yeah

Host 210:44

... in the sense that with T equals three you could not have gotten what you have gotten at T equals five?

RJ Honicky10:51

That, that's right.

Host 210:52

Just-

RJ Honicky10:52

That's right.

Host 210:52

No, don't restart your computer.

RJ Honicky10:57

Yeah.

Host 210:57

Yeah, okay. I, I understand the intuition. Thank you. And why is that?

RJ Honicky11:02

Um, I... So I, I... This is slightly vague in my mind, but basically because the... You know, all that the, y- like the, um... Maybe I've drawn them like further than they might be in reality, right? They'll probably be pretty close.

But the point is that, um, when they're, uh... When the input to the model is like the T equals ten and the T equals eight noise at the top, right? And that doesn't necessarily like... And I'm just trying to generate a, something that I can add to T equals ten to do a little better than that T equals eight one, right?

So I'm trying to like, I'm trying to ge- I didn't draw this, the lines around the square, but I'm trying to generate this thing. And so there, there's nothing constraining it in the latent space to be on that same trajectory.

So it doesn't necessarily pick something.

Host 211:55

I see. So each trajectory at T eight, T five, T three, they each try to optimize as much as they can, and therefore they arrive at-

RJ Honicky12:02

Yeah

Host 212:02

... the different tra- That's the intuition. Thank you.

RJ Honicky12:05

Yeah. They're alway- always... Like this T equals three is just trying to get this, this like little fuzzy airplane noise- Um, back or something like that. Not... Actually, not the fuzzy airplane. Sorry, I misspoke. Just some noise that produces something like that corresponds to that fuzzy airplane thing.

Does that, does that make sense?

A little bit?

Host 212:30

Yes. Yes, it does. Thank you.

RJ Honicky12:32

Okay, great. Okay. So, okay, so this fancy plot is sort of the alternative that I just alluded to. And, um, we're not... Th-these, um, uh, these stochastic, uh, differential equations, we're not gonna really talk about this, and this is, like, a thing that you don't really need to understand.

SDEs & PF-ODE12:51

RJ Honicky12:51

So I just kept... got it from this diagram. So you can just kind of ignore the SDE part. It looks almost the same. It's slightly different. Doesn't really matter, but it's actually a cool diagram because you can see this-- there's these stochastic differential equations that are, like, sort of modeling this random process by which, um, you know, sort of, you have, um, you have noise, and it has a continuous instead of a discrete, um, uh...

It's a continuous instead of a discrete process. So, like, adding noise to... Like, on your left-hand side, very, very left is the data, and these, these two n-like normal-looking distributions. So this, like, bimodal normal distribution. And then, um, and then what this differential...

So this is sort of like heat difu- think heat diffusion, right? And as ti-- like, time is going to the right, and as, um, as the, you know, sort of like heat that is, um, con-sort of, uh, condensed in those two, the modes of these two Gaussians, like, uh, over time it kinda spreads out and, and ends up, you know, sort of, like, in the middle with this prior here.

And that's not exactly what, what happened 'cause it's not quite this-- it's not the heat equation, but, like, it's a useful intuition for what's happening here. So that, like, there's some differential equation, and it's describing this, like, sort of diffusion process by which everything ends up in this sort of, like, literally Gaussian, uh, prior for the distribution.

And so th-this is where we generate our noise, and then we use the backwards process to get back to these, uh, you know, these distributions, th-these bimodal distribution. And there's like a... Each one of these, uh, squiggly lines is a trajectory from...

that, that some random process could have taken based on the sort of, uh, based on sort of likelihood, um, that is imposed on the field by the, um, by the diffusion process it... in the, the colored diffusion process in the background.

Right? So, so... And then, and then this probability flow ODE is sort of a deterministic version that looks at what is the maximum likelihood path if I started at that trajectory, right? So in this-- So, like, looking at...

What's... The easiest one to see is this bottom white one. So if you look, what this is saying is, "I started here. What is the maximum likelihood path for me to get to somewhere in th-in this middle area?"

Right? So, so instead of, like, samp... So you can imagine if I sampled, like, a huge number of these SDEs, and then I took the, uh, each one of these columns, I took the, you know, sort of the means, then I, I would get...

If I started here, I would sort of follow this path, right, to get to, to this point. Uh, I'm, I'm not speaking super accurately, but that's sort of the intuition. So, um, and... Okay. So, and I... This looks, uh, quite different from what this previous diagram, but it's actually the same kind of space.

So I wanna... And that's... So I wanna, um... So in, in this diagram, you would have, like, y-like, you know, maybe a chunk right here, which is, uh, you know, like I did a, like a discrete, um, Gaussian, uh, blurring of my data.

And then I did another one and another one and another one. So this is discrete. And so what happen-- this is what happens if you just make the number of those discrete, um, uh, d-discrete Gaussian blurings infinite and infinitely small.

Okay. So, and then the other thing just to look at, this is a score function. We don't have to really understand it right now and/or not too much ever in this talk, but the score function is, uh, more or less, um, uh...

This is sort of how the... This is the mathematical mechanism by which, uh, the, the reverse diffusion can happen. And inside of this score function is our, uh, U-Net model, right? So our U-Net is there. Is... And, and so I'm using that U-Net to predict what's called the marginal distribution on that probability at-- for that data, and then that's what helps it to sort of figure out how to do this process in reverse.

Okay? Um, I, I know this is not super clear. Uh, um, hopefully it will become slightly more clear. Um, and then just to be clear, this is like a U-Net. This is a... I don't think it really matters, but this is like the name of the U-- particular U-Net model.

Um, and then I... This is sort of what I just said. It estimates a gradient of the log probability des-density with respect to data at a given time step. Um, and but, uh, most importantly, that's where your U-Net lives.

And then just... This is just, this works in latent space too. So, um, this is like all the other stuff- That is in that, um, in that first diagram. That's all here in the left-hand side.

Um, okay, good. So now let's talk about consistency model. So you can see this, this looks a lot like this, but it's slightly different, right? And the same author, so not surprising. But, um, so this is k- sort of the left-hand side, right?

Consistency Models18:27

RJ Honicky18:40

And then we got rid of all the squiggly SDE stuff. So, uh, so what's happening here, and this is... this goes back to the discussion about trajectories. These are trajectories. So you start, you know, in the forward direction.

You start at one of these green dots, and then you follow this trajectory, and you end up somewhere in this noise, um, distribution here so that, like, this, like, structured data turns into a Gaussian distribution. And what are dif- So we have a...

Then we have this, um, uh, uh, a particle function, ordinary differential equation. I think it's particle. It could be probability function. I think it's probability. And it, um, and it sort of is, um, is a o- a differential equation that can map from a given point here back to a given point here.

So again, b- we can do that because since it's deterministic, then there's exactly one for any one place here, there's exactly one place here, right? So, and so once... i- because you have that, what a consistency model does is it says that everything should be on the same trajectory, right?

So I'm gonna... If I estimate it, I, I can, um, I'm gonna learn how to map from any point on this trajectory to the, um, to this point in the data. So including these very left-hand things. Um, and you'll see that it's useful to be able to map to the middle in a minute.

But, um, but the, the point is that the, the differential equation solver tries to find the parameters of the differential equation such that, uh, if I know this, then I know this. Okay? Um, and the link, if you wanna learn about that, that's the sort of some...

the first link about that. And then so just to tie it, I wanna just tie this back to this is the same information, even the same paper. Um, but I just wanna tie it back to our, you know, sort of noising and denoising, um, idea, is that we're, you know, sort of like mapping.

So this is one trajectory. In this case, it's straight, but, like, this is one of those curvy trajectories from the previous diagram. And it's mapping, uh, back, um, to this data, this image, um, at X zero. This is the data that this particular noise, uh, best corresponds to, right?

So, um, uh, and, and, you know, so the idea is just that I'm mapping all the way, um, from any point on that trajectory. And then this is why it's useful to, um, useful to be able to map from any point in the trajectory, is that now I can, if I want to, and most of the time it's two times, but if I want to, um, uh, if I want to do...

like refine my image by throwing more compute at the problem, what I can do is I can, uh, I can add, uh, some... this, the, uh, you know, sort of these, these time steps. I can divide my, uh, space up into time steps.

Um, and then I can a- I can sample noise. I can re-add it to the trajectory at that timestamp, and then I can, uh, refine based on that noise, and that seems to help the model sort of zone, um, hone in on the...

Like, just having another opportunity to refine. Um, it- it'll put it on a slightly different trajectory and help the model to learn from a closer space, uh, what the original data was. So, like, for example, uh, let's go back here.

You know, I might, you know, f- step one, I start at T, I have this noise. I go all the way to X zero, but then I add back enough noise to get back to this step here, and then I follow the trajectory that that lands me on.

So I will be, like, slightly off this trajectory. So maybe this is a better model. Like, you think multiple trajectories, maybe I end up on a different trajectory and that's not what this diagram is for, but I may end up on a different trajectory so that my, my data is slightly different.

The data that I'm, uh, that, that I'm outputting is slightly different. So I wanna... Just any questions in the comments? Um, should I keep just plowing ahead, or do people wanna comment or ask questions?

Host 223:12

I'm not seeing any questions in the comments, um-

RJ Honicky23:15

Okay. Yeah

Host 223:15

... other than what are some good open source models. I don't, I don't know if you wanna chime in on that, RJ.

RJ Honicky23:19

Yeah, uh, yeah. Um, there, there's actually a whole-

Host 223:22

Especially that can generate images with good text.

RJ Honicky23:26

Oh, uh, I... Well, I, I, I'm actually not the best person to answer that. I think that, uh, like the SDXL 3 and Flux, but those are not this type of model, I don't think. Um-

Host 223:41

Perfect. That's what Alexander suggested as well.

RJ Honicky23:44

Yeah.

Host 223:44

And then you would add your own text Loras.

RJ Honicky23:47

Yeah. That's it.

Host 323:47

Yeah, exactly. Flux, Flux is, is great for text as well as, as the, uh, 3.5.

RJ Honicky23:54

Yeah.

Host 323:55

Um, additionally, you can alw- always add some Loras. Uh, there's plenty of Loras on Civitai, uh, for making better text. Uh, yeah, I can, I can do this right now. Actually, I have, uh, ComfyUI open. So if anyone wants to test, I can just throw and, uh, from the text and, uh, generate some, uh, images quickly.

RJ Honicky24:18

Awesome. So, so, like, but I think there's a... Like, maybe this is actually a good, slightly good talking point. So these, these types of models are not... Like, these consistency models, and there's, like, related models. They're... I, I...

With the exception of maybe some of the closed sourced models from- I think that, well, a, a lot of this, these papers come from OpenAI, so I suspect they're maybe using some of these techniques. Um, maybe. So, uh, but...

And I haven't really looked at the sort of big models that are out there to see if any of them are using any of these techniques. But, like, th- these, this is pretty, um, cutting-edge research, so the, the...

none of this technology that we're discussing today is in any of, really, in any of the big models that we know and love. With the exce- maybe of with- maybe with the exception of Flux, um, I don't... But I...

That's actually a great question. I, I'm gonna look, l- look into that as soon as this is over. Um, okay, so let's, let's keep going then. So this is, um... Now, so, uh, this, um, these ODE solvers, they kind of...

they're discrete, um, time solvers. So they, they sort of chunk up the time into little chunks, and then they, um, and then they solve the ODE as best they can given the discretization, the quantization. Um, but that can cause errors.

And so th- that's what this diagram is ch- trying to present, is that if you have, like, ver... this, if this delta T here is very big, you see it goes, like, far from XT to X to minus delta T, then the error that c- it can have is very big, and that can put it on a different trajectory.

So you get the wrong trajectory. Um, and then, um, if the discretization is maybe a bit smaller, then the error might be a little bit smaller, and then you'll get a, um, a closer, but not the same trajectory, and that's sort of the point.

And then, like, in continuous time, in theory, then because your discretization is infinitely small, then you, uh, you, you... it's, like, quote-unquote, imp- impossible to get it on the wrong trajectory although, um, there are still obviously ways that you can.

Um, but that's the theory anyway. Um, so and the, th- this, um, I was a little scratching my head a little bit. I didn't see the, the ODE solver. Uh, sorry, the continuous models don't use an ODE solver.

They do actually in the paper reference using ODE solvers in certain steps, but they're not essential to the process here, is, I think, the point that they're making. Um, and this unbiased estimator, um, they don't give any information on how to do this.

So I think that what they just mean is an empirical estimation of the marginal distribution, but they think they left that at, as an exercise to the reader or somewhere in the appendix that I didn't see. Um, okay.

The Math27:21

RJ Honicky27:21

So I'm gonna get very mathy for a minute. Um, I know, uh, I, like, I understand that, like, first of all, even, you know, the math geniuses among us, um, might have trouble following because there's just too little context, and unless you really read the paper carefully, you're not gonna...

So i- you're not gonna be able to follow super well. But so I, I wanna just call out a few things, and then there-- Like, these two, one and two, equations one and two, we're gonna reference a little bit, um, to explain the, how they accomplished what they did.

And, and also these, um... There's this, this is sort of the, this ca- canonical equation that everybody has been using. And when I say everybody, I, I think I mean mostly this same author in previous papers, but also some other people.

Um, and so the... like, the, this, this is sort of like the canonical setup, and then there's this, these constraints that are on these, the C skip and the C out. And the interesting th- thing to note and the import- the reason those are there is it...

So if this is the time variable, this C out, um, and C skip, and if time is zero, that means that, uh, you're just... you have the data, right? And, and C out, you have none of the input from your dif- from your, uh, from your, your neural network, right?

So this is just saying at time zero, this is a way to guarantee that at time zero you're getting back your data. Um, and then, then the skip and C out, they just trade off in some way between each other, and as time goes, gets later, then you have more, uh, um...

then you, you are giving more weight to the, uh, um, to the neural network. Okay. And then, so then... And then this is just... I don't think we need to go through this in any detail. This is just the training objective.

Um, and then if you look at, um, this, equation two is just that, um, you know, sort of instantiated whe- with this, um, certain kind of distance function. Um, so you'll see that a distance function appears here. This, you know, L2 nor...

th- this L2 loss is, uh, being used, um, which is just squared error loss. Um, that's being used. Um, and then you, you know, sort of plug everything in. You, uh, you know, you sort of, uh, take the limit as T goes to zero or delta T goes to zero, and you get this thing out, and there's a proof of it elsewhere.

Um, and this, uh, this tangent function is kind of the, the, the main actor here. So, like, it... you don't really have to care again what this is. Just if you see this thing or if you see, like, this delta maybe, then think tangent function mostly.

Um, okay. So I know th- this is, like, probably, like, pretty opaque. It's... I would be surprised if it wasn't, unless you read the pa- paper carefully. Um, and then... So then they... So, um, you know, again, reminding you this, this is great, but it has a problem in that it's unstable and, and hard to understand.

Making It Stable30:41

RJ Honicky30:41

So, um, what they did was they took... Th-th-this is just a repeat of what we saw. There's nothing different here, um, just previously. So they took, and they... This is, like, the old way of doing things where you have the...

And this is for, um, you know, for these discrete mostly being used for discrete, but also for continuous. And they have these, you know, sort of they did some derivations and got these values for C skip and for C out and for C in.

And, um, and then they have all these stability problems. And so, um, they came up with this other idea, and we'll motivate this a little bit in a minute, um, uh, to use cosine and sine, um, and this one over, uh, sigma D, uh, for the, for the, um, C skip, C out, and C in.

And then when you plug that in, then you end up with this instead, right? So that... Nothing, nothing magical happened here. I'm just plugging my CF... Uh, my F theta. I'm just plugging in these values and I get this.

Okay, so now here's... This is the thing, and this is I-in my opinion, the meat of the paper. So you have this, uh, you have this, um, part of the, the, the... You have this tangent function that I had called out, and this is, like, if you plug stuff in again into that tangent function, there's nothing magical here.

I'm just plugging that cosine and sine and everything in there. And you get this big long thing here that's really hard to read. And in the paper, they talk about, okay, th-this thing, if you look at just this thing, like the stability is instant.

Like, th-this is not causing instability. This thing here is not causing stability. Oh, this part, like this sine times this, this expression is causing instability, so let's look there. Okay, so then they'd say, "Okay, well, when I... Like, this part is also not causing instability.

It's this, this part." Right? So we have the sine, um, uh, with a differential F. And then so I take that and, um, and, and so now we're going to address all the sources of instability in these three...

Or not necessarily these three classes, but in these, this expression, right? So let's go through piece by piece and decide and figure out what is causing all of the instability. You'll notice this is the chain rule.

Um, I hope you guys are following. I'm trying to keep it very as understandable as possible. So a-any questions?

Host 233:24

No. I was able to follow the last slide. Thank you for breaking it up and reminding us of chain rule, RJ.

RJ Honicky33:29

Yeah.

Host 233:30

I think I'm gonna get... Start getting lost here, but- You're doing a great, you're doing a great job.

RJ Honicky33:34

Okay, great. Good. So, uh, this... And then the noise, um, this, this was just from the previous paper, papers again. Um, I... They just, you know, sort of give it to you and say, "Go read the bu- paper if you wanna know."

And they... So they have this, this C noise, which you'll recall appears, uh, here, right? The C noise. Um, and, uh, and this is basically just a, a way to, um, a coefficient function that can, like, alter the time, the, like, the rate at which time changes basically.

And they're saying this actually causes a lot of... Uh, this cau- like, definitely will cause stability at I equals two because this is zero, and so therefore, this goes to infinity. So definitely can't have that. So they just say, "Ah, let's just set it to T."

And they don't really motivate this except for this is just makes time pass cont- con, uh, at a constant rate, which makes, I think, intuitive sense. Um, okay. And then so now... So that's like this C noise from this page, um, that, uh, this...

It, it appears here, it appears here, appears here, right? Uh, and then, um, okay, so now what about this embedding thing? Um, so what they say is, um, this is sort of basically, you know, similar to the Fourier.

This is a Fourier embedding similar to what you have, um, you know, in the, in the positional embeddings in transformers. Um, and there's a reason for that because this is used in the attention block. Um, but what... So they point out that, um, you know, uh, the, the scale that they wo- they had set here was sixteen.

It's very high and it causes lots of instability, including right, right when you get to pi over two it goes to infinity. So, um, that's obviously bad. Um, so they, um, they played with this and they, I think, empirically determined that, oh, if we just make this really small to this value, I don't know where they got point two, but it just basically corresponds to the same positional embeddings that you see in the transformer, uh, and attention is all you me- need.

Um,

uh, hopefully that's also not so hard. I think they were very hand-wavy, so I don't think, um-

Host 235:54

So what they say is just set S-

RJ Honicky35:57

Yeah

Host 235:57

... to zero point zero two?

RJ Honicky35:59

Yep, that's right.

Host 236:00

Man, I wish I had the math to understand how they got to their intuition, or maybe it's empirical, but okay.

RJ Honicky36:06

I, I think, I think they, um... Yeah, I think that basically what it scales the, um, coefficients to the same scale that the ones in attention all is all you need are at. I think that's... And I don't, I don't think it's...

I think it's all algebra to do that. Um, if I understood, if I recall and understood correctly.

Host 236:29

Thank you.

RJ Honicky36:29

Um, yeah. And then this adaptive double normalization, I don't think it's that important. They didn't really talk... This is the whole... This is everything they say about it here. Um, but they basically, they just say, "Okay, there's this adaptive normalization.

We think it works, but it doesn't work in our case, so we just do it twice and seems to work." Um, so no, I don't think we need to talk about it very much. Um, so you can see, like, there's a bunch of stuff they're stacking on top of each other.

Um, okay, and then there's this, like, tangent normalization. They tried this, and then they also tried just, um, clipping to between one... negative one and one. Um, and they, they, they see, like, oh yeah, like with our, um, with our models, normalization has a lower FID score.

This is First... Fre- Fréchet Inception Distance. It's a measure of quality of, uh, or, or closeness of... perceptual closeness of images. So, uh, um, and so you can see, yeah, with either two steps or one step in our consistency model, uh, um, then we do better, uh, you know, if we have, uh, normalization.

And maybe if clipping is... Clipping might be good enough, uh, because it's obviously cheaper than doing this normalization here.

Um, and then finally, they, they have this adaptive weighting. Basically, I, I think all you need to understand is they, uh, throw this, uh, um, weight term. They th- throw the weight value into the loss function. So this is a loss function for the optimizer when they're, you know, sort of doing the gradient descent.

Then they, uh, they throw this, um, weight term in there. Um, and that is... comes from right, uh, somewhere. Right here, right? So this is the, our gradient, right? And this weight term is in here. So they just in...

You know, when they derive the loss function, they just throw in that weight term, and that seems to help a little bit, right? Um, this, uh, yellow, uh... Oh, no. Actually, are they saying... No, it looks like it's...

Oh, it's, it's better in some cases, but as you go, it looks like it's worse. Or no, I guess, uh, if you have, uh... No, sorry. If you have... If you do two steps, then it's worse. If you do one step, it's slightly better.

So it doesn't look like it matters that much.

And then tangent warm up. I don't, again, probably don't need to understand the, the sine t. They just put r in front of it, and that r just linearly increases from zero to one over the first 10K iterations, and they just do this because it's, uh, it's instable in the beginning.

Um, no need to discuss a lot. Okay. And so, like, when you stack all of these things together, then, um, you're able to train much more effectively, and continuous time does much better, uh, than, um, these, uh, discretes.

The... So this N is the number of discrete steps that your model is taking and, you know, maybe one interesting thing about this plot is that, you know, sort of the best you can do is at 1,024, and then it gets worse, right?

So it goes... So, like, gets better from here to green and then green to purple. Gets way better, and then it gets worse again. So, like, at some point, um, it, it sort... like it, the issues that they brought up start to, uh, start to matter.

Host 240:11

Why is continuous so much better than discrete? Is it because it's just continuous and therefore it's easy to learn? But I thought the continuous is unstable as well.

RJ Honicky40:20

No. So like all-

Host 240:20

But you made it stable

RJ Honicky40:21

... all... Yeah. All of these techniques that we just talked about in the last few slides are making continuous stable, so then they are able to, uh, train more effectively. Um, yeah, I think that in previous, like, in the previous attempts, um, they...

actually other authors did, um, and they did. People fo- I think they found that, um, they, they had to be very conservative in the, you know, sort of like the way that they trained the model in order to avoid all the instabilities.

And so because of that, they were not able to get good results. Whereas now, they're able to c- sort of like this intuition that they s- right here, um, the, like, this quant- It's sort of like because I...

if I have only a few steps, you know, this, the time delta is very big between these, then I'm gonna have a lot of error in my tangent calculation, right? So this, this tangent is that tangent that we talked about.

And there, if there's a error, like if there's a big, uh, quantization here or, like, only a few time steps, then the ODE solver has to, you know, sort of has only a limited amount of data to work with, and it ends up making mistakes due to that, um, discretization.

And so then it ends up having big errors, and you get on the wrong trajectory. Does that make sense?

Host 241:41

That makes sense. That makes sense. Yeah. And we also have a question from the chat. What does the FID metric evaluate?

RJ Honicky41:47

Uh, yeah. It's a, that's a great question. I had to look it up myself 'cause I've seen it before, and I forgot it completely. It's, um, so it is just... It is, like, a, it is a numerical, um, measure of the difference between two images.

But, and the reason why it's popular is because it seemed... it was explicitly designed to match the, it matched closely to human perceptual, uh, difference, right? So they, there's a paper that talks about it, and one of the things that they talk about is...

And I think, uh, it's, it's... Well, it's in the, definitely in the bibliography, but I can dig it if, if, um, if anyone's interested. Um, it basically describes, uh, s- the, the, the, the previous methods that were being used, which were just like, uh, there was one of the, um, divergence, uh, divergence measures, or not measures, divergences.

Uh, the, they didn't match to what- People visually... W-what humans actually were visually using to distinguish between images. So that w-this one was designed to do a better job of that. So it's, it's supposed to be like a human perceptual, um, uh, distance metric between the, um, the, the input, uh, like an input and output im- uh, image.

Um, so I think that meaning, uh, um... Actually, you... Yeah, I think there's a, there's like a... And I don't know exactly. That's a, that's a great question. I don't know exactly how they do the experiment. Um, I th- I think that they go do a forward pass with the image, and then they, um, add some noise, and then they do the backwards pass, and they see the diff- difference between the images.

But I'm not 100% sure of that. Is- does anyone know this, how this works exactly?

No. Okay. Well, so yeah, that's actually a... I-- that's something I als- that I wanna follow up on.

Evaluation43:59

RJ Honicky43:59

Okay, uh, and then I talked about the tangent warm up, continuous versus discrete. Okay, so here's the part that everybody should be mo- a little more comfortable with, hopefully, is just l- like we're evaluating models. Uh, the, you know, y- you have this NFE stands for, um, um, uh, number of, uh, function evaluations.

So this is how many times did I... did my, um, sort of core... How many iterations did I have to do to generate my images? And so, um, uh, and you see that they, all, all their numbers in this section of the paper, they have, they, they compare to two evaluations in one.

Um, and they do slightly better whenever they do two. Um, so and then, so there's, uh, several things to note about this. One thing that they said in the paper, it's not here, but that, uh, that it, they...

It takes about two X to compute to train the, this consistency model from as a, um, as a distillation of whatever that it's distilled from. So ap- approximately twice the compute. So, um, i- if you spend a lot of compute to train a model and then you wanna distill it with this mechanism, you're gonna pay about twice as much.

So there could be-- that could pres- present an operational problem, or it might not. Um, another thing to note, these joint training sections, these are basically GANs, but, uh, uh, like you'll see some of the more, um, common GANs here.

But... And I don't know exactly the difference between these guys, but this CTM, like, was the sort of overall winner for this, um, this CFAR, uh, data s- uh, benchmark. And then, um, and then, you know, for a conditional class, uh, ImageNet sixty-four by sixty-four, it's a different one, but it's also in this joint training.

So the GANs tend to be, tend to be winning on these benchmarks. Um, and so I-- my belief is that these are... The reason why people don't use them is because they're hard to train and they're, they, um, have, they're very, they have mode-seeking behavior, meaning it's hard to get any diversity and hard to control them.

But for these benchmarks, they do the best. Um, and you see down here, uh, you know, sort of their, um, you know, their consistency model is actually quite close, um, for what it's worth. And then, uh, so this is from distillation.

And then this is if you train from scratch, then they do better. So, uh, that's maybe another interesting thing. So the distillation doesn't work quite as well as the training from scratch does. Does anyone... This is a place where I'd-- someone might wanna, like, comment or ask questions.

So I wanna pause and make sure.

Okay. So and then, um, so they also compared this variational score distillation, which is kind of in the same ballpark, um, in terms of effectiveness. And they found that, um, they, they have, uh, or VSD has higher precision and lower recall, meaning lower diversity.

Um, and like as the guidance scale gets higher, it does worse. And um, uh, you know, this because if you-- you'll notice the, the diffusion teacher model, um, and, and the, um, the, the, uh, uh, the models that they built are very close to each other in all, in these cases, both for precision and recall.

So you end up having a sim-- very similar FIE score as well.

Um, and then, like, uh, another interesting thing here, uh, a, a couple of things. So their, you know, their model, uh, does quite well, um, when it's, uh, trained and not distilled. Um, so... Or sorry, uh, let's see.

Uh, the, the, yeah, the distillation doesn't do quite as well as the training. Um, and but neither of them do as well as the diffusion teacher, um, in- including the one that was trained from scratch. Um, and that, like, you'll see also, interestingly, one s- uh, so the two-step, um, does actually, um, the two-step trained model does worse with, um, you know, sort of like, uh, um, in the smaller models, but better in the bigger models.

Um, and- Okay, and then this is sort of like their scaling study. We're kinda out of time, so I won't talk too much. But like, um, yeah, you can see they, they did well. They have up to 1.5 B model.

Um, uh, let's see. The... Yeah, those are... That's sort of the main takeawa- Okay, yeah. So that's, that's all I have. Um, I can take any questions if people wanna stick around for a few minutes. Or...

Q&A49:24

RJ Honicky49:24

I know this is not at all an easy paper.

Host 249:32

I actually think it's one of the hardest papers. Um-

RJ Honicky49:34

Yeah.

Host 249:35

I think the other one was probably PPO or DPO.

RJ Honicky49:39

I, yeah. I, I missed that paper, so

Host 249:42

Yeah.

RJ Honicky49:42

But yeah. And I, I mean, to me this, this was challenging for several reasons, right? Like, one is just the topic. Like, diffusion by itself is hard, and then you have this like really obscured diffusion that is like even more complicated, and you have to understand a little bit about differential equations and whatever.

And then on top of that, there's just a ton of literature to read in order to read the paper. So all those things like kinda... It's like a triple whammy.

Um, but I, I've, I must say, I really, really enjoyed learning, uh, by reading. Like, 'cause I, you know, I dug into all these papers, which I normally don't have time to do, but I, you know, because I'm presenting, I took the time to really look at all the references and try to understand things well enough to hopefully explain them right to other people.

And that was super valuable experience, and I'm glad I did that on such a hard paper.

I hope that like, uh, you guys felt that you got some of the intuition behind this. Um, I, I know it's not easy. I tried to focus on the intuitions for the paper, um, and not so much on the math.

Uh, so I'll... Let me stop sharing.

Um, okay, guys. If there's nothing else, um, you know, I thoroughly enjoyed, um, doing this. I hope, hope, uh, somebody new will, will, um... Someone who's never presented before will, will take, take up the mantle for the next session or the next open session.

That'd be awesome.

Host 451:31

Yep. If I, if anyone has any paper that they wanna cover, it's always great to, for y'all to voice out and say, because we always welcome new, new paper presenters.

RJ Honicky51:41

Yeah.

Host51:41

Yeah, I think we, I think we had a volunteer for next week, but, uh, it's on the Luma. I, I, uh, don't remember which one it was, but Deepu was signing someone up.

RJ Honicky51:50

Awesome.

Host 451:50

Yay.

Host51:52

More need, more, more still needed.

RJ Honicky51:56

Yeah. Okay, guys. Well, enjoy your Wednesday then. I will, uh, I'll see you guys on Discord.

Host 452:04

Yep. Thank you very much once again for presenting.

Host 252:07

Thanks, RJ.

RJ Honicky52:08

No problem. My, my pleasure. Yep, you're welcome. Bye-bye.