Beginnings0:00
Hey everyone, welcome to the Latent Space Podcast. This is Alessio, partner and CTO of Decibel Partners, and I'm joined by my co-host, swyx, founder of Smol AI.
Hey everyone. We are in the Chroma studio again, but with our first ever anonymous guest. Comfy Anonymous, welcome.
Yeah. Well, hello.
I feel like that's your full name. You, you just go by Comfy, right?
Yeah, well, a lot of people just call me Comfy even though even if- even when they know my real name. They say, "Hey, hey Comfy."
Yeah. swyx is the same. Like, you know, not a lot of people call you Sean.
Yeah, it's-- Yeah, you have a professional name, right-
Yeah
... that people know you by, and then, then you have a legal name. Yeah, it's fine. How do I phrase this? Like I think a lot of people who are in the know know that Comfy is like the tool for image generation and now other multimodality stuff.
I would say that when I first got started with Stable Diffusion, the star of the show was Automatic 11 11, right? And I, I actually looked back at my notes from 2022-ish, like Comfy was already getting started back then, but it was kind of like the up-and-comer and like your main feature was a flowchart.
Can you just kind of rewind to that moment, so that, that year and like, you know, how, how you looked at the landscape there and decided to start Comfy?
Yeah. I discovered Stable Diffusion in, uh, 2022, in October 2022, and, uh, well, I kind of started playing around with it. Yes, I-- And back then I was using Automatic, which was what everyone was using back then, and I-- So I started with that.
'Cause I, I had the-- 'Cause when I started, I had no idea like how diffusion models work, how any of this works. So
Oh, yeah. What was your prior background as an engineer?
Uh, just a software engineer. Yeah, boring software engineer.
But like any, any image stuff, any orchestration-
No
... distributed systems, GPUs?
No, I was doing, uh, basically nothing interesting.
Crud web development.
Yeah. Well, not web development. Just, yeah, some basic-- maybe some basic like automation stuff and, uh-
Okay
... just, yeah. But, uh, no, like, uh, no big companies or anything.
Yeah, but like already some interest in automations, probably a lot of Python.
Yeah, yeah, of course, Python. But, uh, like I wasn't actually used to like, uh, the Node graph interface-
Mm
... before I started ComfyUI. It was just, uh, I just thought it was like, "Oh, like what's the best way to represent the diffusion process in the user interface?" And then like, oh, well, like naturally, this is the best way I found.
And this was, uh, like with the Node interface.
Mm-hmm.
So how I got started was, uh, yeah, so basic October 2022, just, uh, like I hadn't written a line of PyTorch before that.
Mm-hmm.
So it's, uh, completely new. What happened was I kind of got addicted to generating images.
As we all did.
Yeah. And then I started experimenting with, uh, like the high-res fix in Auto, which was, uh, for those that don't know, the high-res fix is just to gen-- Since the diffusion models back then could only generate at low resolution.
Mm-hmm.
So what you would do, you would generate low-resolution image, then upscale, then kind of pa- refine it again, and that was kind of the hack to generate high-resolution images. I really liked generating like higher resolution images, so I was experimenting with that.
And so I modified the code a bit. Okay, what happens if I, if I use different samplers on the second pass?
Mm-hmm.
So I must edited the code of Auto. So what happens if I use a different sampler? What happens if I use a different, uh, like a different settings and a different number of steps? Uh, and because back then the high-res fix was very basic.
Mm-hmm.
Just, uh, so.
Yeah. Now there's a whole library of just, uh, the upsamplers that's very popular.
Yeah, yeah, right. I think, yeah, I think they added a bunch of, uh, of, uh, options to the high-res fix since, uh, since, since then. But before that it was just so basic, so I wanted to go further.
I wanted to try, okay, what happens if I use a different model for the second, the second pass? And then, well, then the Auto code base was, wasn't good enough for it. Like it would have been, uh, harder to implement that in the Auto interface than to create my own interface.
So that's when I decided to create my own.
And you were doing that mostly on your own when you started, or did you already have kind of like a subgroup of people?
No, I was, uh, on my own. 'Cause it was just me experimenting with stuff. So yeah, that was, uh, bas-- So I started writing the code January 1, 2023.
Mm-hmm.
And then I released the first version on GitHub January 16, 2023. That's how things got started.
And was, was the name ComfyUI right away or-
Yeah, yeah. ComfyUI. The reason the name, my name is Comfy is people thought my pictures were comfy. So I just, uh, just name it, uh, uh, it's my ComfyUI. So yeah, that's, uh-
Is there a particular segment of the community that you targeted as users, like more intensive workflow artists, you know, compared to the Automatic crowd or, you know?
This was my way of like experimenting with, uh, with new things, like the high-res fix thing I mentioned, which was like in Comfy, the first thing you could easily do was just chain different models together. And then one of the first things-- I think the first times it got a bit of popularity was when I started experimenting with the different, like applying prompts to different areas of the image.
Mm-hmm.
Yeah, I called it Area Conditioning. Posted it, it on Reddit, and it got a, a bunch of upvotes. So I think that's when, like, when people first, uh, learned of ComfyUI.
Mm.
Is that mostly, like, fixing hands?
Oh, no, that was just, uh, like, let's say... Well, it was very-- Well, it still is kind of difficult to, like... Let's say you want a, a mountain. You have an image, and then you, "Okay, I want a mountain here, and I want, uh, like, a, a fox here."
Mm.
Like-
Yeah, so compositing the image.
Yeah. By the way, it was very easy. It was just like, oh, when you run the diffusion process, you kind of generate, okay, you do pass, one pass through the diffusion model. Every step you do one pass. Okay, this place of the image with this prompt, this place, place of the image with the other prompt, and then the entire image with another prompt, and then just average everything together every step, and that was, uh, area composition, which I call it.
And then, then a, a month later, there was a paper that came out called Multi-Diffusion, which was the same thing. But, uh, yeah. That's, uh-
Could you do area composition with different models or, or because you're averaging out, you kinda need the same model?
Well, you could do it with-- But yeah, I hadn't implemented it for different models. But, uh, you, you can do it with, uh, with different models if you want, as long as the models share the same latent space.
Mm-hmm.
Like you-
We h- we're supposed to ring a bell every time someone says latent space.
Yeah, it's from-
Yeah. Like, for example, you couldn't use, like, XL and SD 1.5 'cause those have a different latent space. But, like, uh, yeah, like SD 1.5 models, different ones, you could, you could do that, uh-
There's some models that try to work in pixel space, right?
Yeah. They're very slow.
Of course.
That's the problem. That, that's the, the reason why Stable Diffusion actually became, like, popular, like, 'cause-- was because of the latent space.
Yeah, small in... Yeah.
Yeah.
Because it used to be latent diffusion models, and then they trained it up.
Yeah, 'cause the pixel, pixel diffusion models are just, uh, too slow, so.
Yeah. Have you ever tried to talk to, like, uh, like Stability, the latent diffusion guys, like, you know, Robin Rombach, that, that crew?
The-- Yeah. Well, I used to work at Stability.
Oh, I actually didn't know.
Yeah. I used to work at Stability. I got, uh, I got hired, uh, in June 2023.
Ah. That's the part of the story I didn't know about. Okay.
Yeah. So the, the reason I was hired is because they were doing, uh, SD XL at the time.
Mm.
And they were-- Basically, SD XL, I don't know if you remember, it was a base model and then a refiner model.
Mm-hmm.
Basically, they wanted to experiment, like, chaining them together, and then, uh, they saw, "Oh."
Right.
"Can-- Oh, this, we can use this to do that. Well, let's hire that guy."
But they didn't, they didn't pursue it for, like, SD3.
What do you mean?
Like, the SD XL approach.
Yeah. The reason for, for that approach was because basically they had two models, and then they wanted to publish both of them. So they, they trained one on lower time steps, which was the refiner model, and then they-- the first one, well, was trained normally.
And then they w-- During their test, they realized, "Oh, like, if we string these models together, our, like, quality increases. So let's publish that."
It worked.
Yeah. But, uh, like, right now, I don't think many people actually use the refiner anymore, even though it is actually a full diffusion model. Like, you can use it on its own, and it's gonna generate images. I don't think anyone-- People have, uh, mostly forgotten about it.
But, uh-
Can we talk about models a little bit? So Stable Diffusion obviously is the most known.
Model Wars10:06
Yeah.
I know Flux has gotten a lot of-
Yeah
... traction. Are there any underrated models that people should use more, or what's the state of the union?
Well, the, the latest, uh, state-of-the-art, at least, yeah, for images, there's, uh, yeah, there's Flux. Uh, there's also SD 3.5. SD 3.5 is two models. There's a, there's a small one, 2.5B, and there's the bigger one, 8B. So it's, it's smaller than Flux, so...
And it's more, uh, creative in a way.
Mm-hmm.
But Flux, yeah, Flux is the best. People should give SD 3.5 a try 'cause it's, uh, it's different. I won't say it's better. Well, it's better for some, like, specific use cases. Like, if you want some-- to make something more, like, creative, maybe SD 3.5.
If you want to make something more consistent, then Flux is probably better.
Do you ever consider supporting the closed source model APIs?
Uh, well, they-- we do support them with the custom nodes. We actually have some, uh, official custom nodes from, uh, different-
Ideogram.
Yeah. Have-
I guess DALL-E would have one.
Yeah. That's, uh, it's just not I'm not the person that handles that.
Mm-hmm. Sure, sure. Quick question on, on SD. There's a lot of community discussion about the transition from SD 1.5 to SD 2, and then SD 2 to SD 3. People still, like, you know, very loyal to the previous generations of SDs?
Uh, yeah. SD 1.5 then still has a lot of, uh- ... a lot of users.
The last based model.
Yeah.
Yeah. Then SD 2 was mostly ignored as it wasn't, uh, it wasn't a big enough improvement over the previous one.
Okay, so SD 1.5, SD3, Flux, and whatever else
Yeah, SDXL
SDXL
SDXL, that's the main one
Stable Cascade, you know?
Stable Cascade, that was a good model, but, uh, that's-- uh, the, the problem with that one is, uh, it got, uh, like SD3 was announced one week after
Yeah, it was like a weird release, uh... What was it like inside of Stability, actually?
I don't know.
I mean, statute of limitations expired, you know, management has, has moved somewhere. It's easier to talk about now.
Yeah. I think the inside Stability, actually, that model was ready, uh, like three months before, but it got, uh, stuck in, uh, red teaming. So basically, the problem that if that model had released or was supposed to be released- ...by the authors, then it would probably have gotten very popular since it's a, it's a step up from SDXL.
But it got all of its momentum stolen by the SD3 announcement, so people kind of-
Yeah
... didn't develop anything on top of it, even though it's, uh, yeah. It was a good model, at least, uh, completely mostly ignored for some reason. Like-
It seem-- I think the naming as well matters. It seemed like a branch off of the main ra- main-
Yeah
... tree of development.
Yeah. Well, it was different researchers-
Yeah
... that did it. Like-
Different branch
... yeah.
Yeah.
V- very, like a good model. Like, it's the Wurster-Schun authors. I'm not sure if I'm pronouncing it correct.
Wurschin, yeah.
Wurschin, yeah. Yeah, they-
I actually met them in, uh, Vienna.
Yeah, they worked at Stability for a bit, and they left right after the Cascade release.
This is Dustin, right?
No.
Uh, Dustin's SD3. Yeah.
No, uh, Dustin is, uh, SD3, SDXL. That's, uh, Pablo and-
Yeah. Yeah, yeah
... uh, Dome. Besides, I think I'm pronouncing his name correctly.
Ah.
Yeah, that's, uh-
Yeah, very cool
... you know, very, very good.
It seems like the community is very-- they move very quickly.
Yeah.
Like, when there's a new model out, they just drop whatever the current one is, and they just all move wholesale over. Like, they don't really stay to explore the full capabilities.
Mm.
Like if, if the Stable Cascade was that good, they would have A/B tested a bit more. Instead, they're like, "Okay, SD3 is out. Make this go." You know?
Well, I find the opposite actually.
Okay.
The community doesn't-- like they only jump on a new model when there's a significant improvement.
I see.
Like, if there's, uh, only like a incremental improvement, which is what, uh, most of these models are, are going to have, especially if you-- 'cause, uh, stay the same parameter count, uh-
Yeah
... like you're not gonna get a massive improvement, uh, into like... Unless there's something big that, that changes, so, uh, yeah.
And how are they evaluating these improvements? Like, um, because this-- it's a, it's a whole chain of, you know, Comfy workflows.
Yeah.
How does, how does one part of the chain actually affect the whole process?
Are you talking on the model side specific?
Model specific, right. But like once you have your own whole workflow based on a model, it's very hard to move.
Uh, not... Well-
No?
... not really.
Maybe not.
Depends on your, uh, depends on the specific kind of the workflow.
Yeah. So I do a lot of like text and image.
Yeah. When you do change, uh, like most workflows are kind of gonna be compatible between different models. It's just like you might have to completely change your prompt, completely change.
Okay. Well, I mean, then maybe the question is really about evals. Like what does, uh, the Comfy community do for evals? Just, you know-
Well, they, they don't really do. They-- It's more like, oh-
Vibe check
... I think this image is nice.
Yeah.
So that's, uh
They just subscribe to FOOFER AI and just see like, you know, what FOOFER is doing.
Yeah. They just, they, they just generate like, like I don't see anyone really doing like, uh, at least on the Comfy side, Comfy users, they-- it's more like, oh, generate images and see, oh, this one's nice.
Yeah, yeah.
This is nice.
Yeah, vibes.
Yeah. Yeah. It's not, uh, like the, the more, uh, like, uh, scientific, uh, like, uh, like checking, that's more on specifically on like model side if, uh... Yeah. But there is a lot of, uh, vibes also 'cause it is a, like- ...
a artistic, uh... You can create a very good model that doesn't generate nice images. 'Cause mo- most images on the internet are ugly. So if you-
LoRAs & Prompts16:17
Mm
... if that's like, if you just, "Oh, I, I have the best model that can, like, it's super smart. I trained it on all the, like... I've trained it on just all the images on the internet." The images are not gonna look good.
So-
Yeah. Yeah.
They're gonna be very consistent. But yeah, peop- like it's not gonna be like the, the look that people are gonna be expecting from, uh, from a model, so yeah.
Can we talk about LoRAs? 'Cause we talked, we talked about models, then like the next step is probably LoRAs. Before-- Uh, actually, I'm kinda curious how LoRAs entered the tool set of the image community, because the LoRA paper was 2021, and then like there was like other methods like textual inversion that was popular at, at the early SD stage.
Yeah, I can explain the difference between that. Like textual inversions, that's, uh, basically what you're doing is you're, you're training a... 'cause, well, yeah, Stable Diffusion, you have the diffusion model, you have the text encoder. So basically, what you're doing is training a vector that you're gonna pass to the text decoder.
It's basically you're training a new word.
Yeah. It's a little bit like representation engineering now.
Yeah. Basically, yeah, you're just, uh... So yeah, if you know how, uh, like, uh, the text encoder works, basically you have, uh, you, you take your, your words of your prompt, you convert those into tokens with the tokenizer, and those are converted into vectors.
Basically, yeah, each token represents a different vector, so each word represents a vector, and those Depending on your words, that's the list of vectors that get passed to the text encoder, which is just, uh, yeah, just a stack of, uh, of attention.
Like, basically, it's, uh, very close to LLM architecture. Yeah. Yeah, so basically what you're doing is just training a new vector. We're saying, "Well, I have all these images, and I want to know which word does that represent."
And it's gonna get-- Like, you train this vector, and then, and then when you use this vector, it hopefully generates, uh, like something similar to your images. Yeah.
Yeah. I would say it's, like, surprisingly sample efficient in picking up the concept that you're trying to train it on.
Yeah. Well, people have kind of stopped doing that.
Yeah.
Even though, uh, back at, like, when I was at Stability, we, we actually did train internally some, like, textual versions on, like, T5-XXL.
Mm.
Actually worked pretty well. But, uh, for some reason, yeah, people don't use them. And also, they might also work, uh, like, like... Yeah, that's just something you'd probably have to test, but maybe if you train a textual version, like on T5-XXL, it might also work with all the other models that use T5-XXL.
'Cause same thing with, uh, like, uh, like the textual versions that, uh, that were trained for SD 1.5, they also kind of work on SDXL, because SDXL has the, has two text encoders, and one of them is the same as the, as the SD 1.5 Clip L.
So those, they actually w- they don't work as strongly 'cause they're only applied to one of the text encoders. But, uh, and the same thing for SD3. Three-- SD3 has three text encoders, so-
Mm.
It works. It's still w-- You can still use your text conversion SD 1.5 on SD3, but it's just a lot weaker because now there's three text encoders, so it's gets even more diluted. Yeah.
Yeah. Do people experiment a lot on, on... Just on the Clip side, uh, there's like SigLIP, there's BLIP. Like, do people experiment a lot on, on those?
Uh, well, you, you can't really replace-
Yeah, 'cause they're, they're trained together, right?
Yeah, they're trained together, so you, you can't, uh... Like, well, what I've seen people experimenting with is, uh, LongCLIP. So basically someone, uh, fine-tuned the CLIP model to accept longer prompts.
Oh.
Yeah.
It's kind of like long context fine-tuning.
Yeah. So, so like it's, it's actually supported in core Comfy.
How long is long?
Regular CLIP is 77 tokens.
Yeah.
LongCLIP is 256.
Okay.
So, but the hack that, uh, like you've-- if you've used Stable Diffusion 1.5, you've probably noticed, "Oh, it still works if I, if I use long prompts," prompts longer than 77 words. Well, that's because the hack is to just, uh, well, you split, you split it up in chunks of seven.
Your whole-
Mm.
Your big prompt. Let's say you, you give it like the massive text, like the Bible or something.
Oh, my.
And it would split it up in chunks of 77-
Yeah. It's like-
And then just pass each one through the, the CLIP.
Yeah.
And then just concatenate-
Revive
... everything together at the end. It's not-
Ooh
... ideal, but it actually-
That's messy
... works.
Like the positioning of the words really, really matters then, right?
Yeah.
Like, this is why order matters in prompts.
Yeah. Yeah, like it, it works, but it's, it's not ideal, but it's what people expect. Like if, if someone gives-
Yeah
... a huge prompt, they expect at least some of the concepts at the end-
Yeah
... to be like present in the image. But usually when they give long prompts, they, they don't, they like, they don't expect, uh, like detail, I think. So that's why it works very well.
And while we're on this topic, uh, prompt weighting, neg- negative prompting, all, all sort of similar part of this-
Yeah
... layer of the stack.
Yeah. The, the hack for that which works on CLIP, like it w- basically it's just, uh, for SD 1 point f-- well, for SD 1.5, the prompt weighting works well because, uh, Clip L is a, is not a very deep model.
So you have a very high correlation between you have the input token, the index of the input token vector, and the output token. They're very-- the concepts are very close, closely linked. So that means if you interpolate the vector from what...
Well, the, the way ComfyUI does it is it has, okay, you have the vector. You have a empty prompt. So you have a, a child, like a CLIP output for the empty prompt, and then you have the one for your prompt, and then it interpolates from that depending on your prompt weight, the weight of your, of your tokens.
So, so if you... Yeah. So that's how it, how it does, uh, prompt weighting, but this stops working the deeper your text encoder is.
Mm.
So on T5-XXL, it doesn't work at all, so.
Wow. Is that a problem for people? I mean, 'cause I, I'm used to just move, moving up numbers.
Probably not-
Yeah
... 'cause, uh, well-
So you just use words to describe, right? 'Cause it's a bigger language model.
Yeah. Yeah. So honestly, it might be good, but I haven't seen many complaints on Flex that, uh, it's not working. So-
Yeah
... 'cause I guess people can sort of, uh, get around it with, uh, with language. So, uh, yeah.
Yeah. And then coming back to LoRAs, now the, the popular way to, to customize models is LoRAs. And I s- I saw you also support LoCon and LoHa, which I've never heard of before.
There's a bunch of, uh... 'Cause, uh, what, what the LoRA is essentially is, uh, instead of, uh, like, okay, you have your, your model, and then you want to fine-tune it. So instead of, uh, like what you could do is you could fine-tune the entire thing.
But that's-
Yeah, full fine-tune. Yeah
... but that's a bit, uh- heavy. So to speed things up and make things less heavy, what you can do is just fine-tune some smaller weights. Like basically two, two matrices that when you multiply, like two low-rank matrices, that when you multiply them together gives a-- represents a difference between trained weights and your base weights.
Mm-hmm.
Design Philosophy24:39
So by training those two smaller matrices, that's a lot less heavy.
Yeah. And they're portable, so you can share them.
Yeah.
It's like easier.
And also smaller. Yeah. That's the-- how LoRAs work. So basic-- So when, when inferencing, you can inference with them pretty efficiently, like how ComfyUI does it. It just-- When you use a LoRA, it just applies it straight on the weights so that there's only a small delay at the be- like before the sampling to-- when it applies the weights and then it just same speed as, uh, as before.
So for, for inference, it's, it's not that bad. But, uh... And then you have, uh... So basically, all the LoRA types like LoHA, LoVA, everything, that's just different ways of, uh, representing that, uh. Like, basically, you can call it kind of like compression even though it's not really compression.
It's just different ways of represent, like just, okay, I want to train a different on the-- difference on the weights. What's the best way to represent that difference? There's the basic LoRA, which is just, "Oh, let's multiply these two matrices together," and then there's all the other ones, which are all different, uh, algorithms.
So yeah.
Yep. So let's talk about what ComfyUI actually is. I think most people have heard of it. Some people might have seen screenshots. I, I think fewer people have built very complex workflows.
Yeah.
So when you started, Automatic was like the super simple way. What were some of the choices that you made? So the node workflow, is there anything else that stands out as like this was like a unique take on how to do image generation workflows?
Well, I feel like, yeah, back then everyone was trying to make like easy-to-use interface. Then I'm like, "Well, everyone's trying to make an easy-to-use interface."
Let's make a hard-to-use interface.
Like so, like, I like, I don't need to do that.
Yeah.
Everyone else doing it, so let me try something, uh, like let me try to make a powerful interface-
Yeah, yeah
... that's, uh, not easy to use, so.
So like, yeah, there's a sort of node execution engine. Your README actually lists, uh, this really good list of features of things you prioritize, right? Like, um, let me see, like, uh, sort of re-executing from a, from any parts of-
Engine Room27:05
Yeah
... this workflow that was changed, asynchronous queue system, smart memory management. Like, all, all this seems like a lot of engineering that-
Yeah, there's a lot of engineering in the, in the back end to make, uh, things, uh... 'Cause I was always, uh, focused on making things work locally very well 'cause that's-- 'cause I was using it, uh, locally, so everything, uh...
So there's a lot of, uh, a lot of thought and work in, like, getting everything to run as well as possible. So yeah, ComfyUI is actually more of a back end-
Mm-hmm
... or at least, uh, well, now, now the front end's getting a, a lot more development . But, but before, before it was-- I was pretty much only focused on the back end.
Yeah. So v0.1 was only August this year-
Yeah, before there was-
With the new front end
... no versioning, so.
Yeah, yeah.
Yeah.
And so what was the big rewrite for the 0.1 and then the 1.0?
Uh, well, that's more, uh, on the front end side. That's-
Okay
... 'cause b-before that, it was just, uh, like the UI what-- 'Cause when I first wrote it, I just, uh, I said, "Okay, how can I make..." Like I can do web development, but I don't like doing it.
Mm-hmm.
Like what's the easiest way I can slap a node interface on this? And then I found this library, Light Graph, like JavaScript library.
Life Graph?
Light Graph.
Usually people will go for like React Flow for like a flow builder.
Yeah, but that seems like too complicated
'Cause of React.
So I didn't really want to spend time at like developing the, the front end. So I'm like, "Well, oh, Light Graph. This has the whole node interface." So okay, let me just plug that into, to my back end then.
I feel like if Streamlit or Gradio offered something, you would have used Streamlit or Gradio 'cause it's Python.
Yeah, Streamlit and Gradio-- Like Gradio, I don't like Gradio. It's-
Why?
... it's ba- Like the-- That's, that's one of the reasons why like Automatic was very bad. It's Gra- 'cause, uh, the problem with Gradio, it forces you to... Well, not forces you, but it kind of, uh, makes your, your interface logic and your back end logic and what, just sticks them together.
It, it is supposed to be easy for you guys for-- if you're a Python main. You know, I'm a JS main, right?
Yeah.
If you're a Python main, it's supposed to be easy.
Yeah. It's e- well, it's easy, but it makes your whole software a huge mess.
I see, I see. So you're mixing concerns instead of separating concerns?
Well, it's 'cause-
Like front end and back end
... the front end and back end should be well separated with a-
Yeah, yeah
... defined API. Like that's, that's how you're supposed to do it. Gradio just like-
People, smart people disagree, but yeah
... it ju- it just stick e- sticks everything together. It makes, uh-
Yeah
... makes it easy to like make a huge mess. And also it's, uh... And there, there's a lot of issues with, uh, with Gradio. Like it, it's very good if all you want to do is just get, like slap a quick interface on your, uh, like to, to show off your, uh, like your ML project.
Like that's what it's made for.
Yeah, yeah.
Like, uh, like there-there's no problem using it like, oh, I have my I have my code. I just want a quick interface on it. That's perfect.
Mm-hmm.
Like, use Gradio. But if you want to make something that's, like, a real, like, real software that will last a long time and will be easy to maintain, then I would avoid it.
Yeah.
So.
So your criticism is Streamlit and Gradio are the same-- I mean, those are the same criticisms, uh-
Yeah.
Okay.
Streamlit I haven't-
Haven't used as much.
Yeah, I've s- just, uh, looked a bit, uh-
Similar philosophy.
Yeah, it's similar. It's just-- it, it just seems to me like, okay, for quick, like, AI demos, it's perfect.
Yeah. Going back to, like, the, the, the core tech, like asynchronous queues, slow re-execution, smart memory management, you know, anything that you, you, you're very proud of or was very hard to figure out?
Yeah. The, the thing that's the biggest, uh, pain in the ass is probably the memory management.
Mm-hmm.
Yeah. Were you just paging models in and out, or...?
Yeah. Be-before it was just, okay, load the model, completely unload it, load the new model, completely unload it. Then okay, that, that works well when your, your model are small. But if your models are big and it takes sort of like-- Let's say if someone has a, like a, A forty ninety and the model size is ten gigabytes, that can take a few seconds to, like, load, unload, load, unload.
So you want to try to keep things, like, in memory, in the GPU memory as much as possible. What ComfyUI does, uh, right now is it, uh, it tries to, like, estimate, okay, like, okay, you're gonna sample this model.
It's gonna take probably this amount of memory. Let's remove the models, like, this amount of memory that you load-
Mm-hmm
... that's been loaded on the GPU and then just execute it. But, uh, so there's a fine line between just-- 'cause try to remove the least amount of models that are already loaded 'cause as for apps like Windows driver, the-- And another problem is, uh, the NVIDIA driver on Windows by default.
Because there's a way to-- There's an option to disable that feature. But by default, it, uh, like, if you start loading, you can overflow your GPU memory, and then it's-- the driver's gonna automatically start paging to RAM. But the problem with that is it's, it makes everything extremely slow.
Mm.
So when you see people complaining, "Oh, this model, it works, but oh, shit, it starts slowing down a lot," that's probably what's happening. So y- it's basically you have to just try to get-- use as much memory as possible, but not too much, or else things start slowing down, or people get out of memory.
And then just find-- try to find that line where, oh, like, the drive around window starts paging and stuff.
Yeah.
And yeah. And the problem with PyTorch is it's, uh, it's high levels don't have that much fine grain control over, like, uh, specific, uh, memory stuff.
Mm-hmm.
So, uh, kinda have to leave, like, the memory freeing to, to Python and PyTorch, which is-- can be annoying sometimes.
So, you know, I, I think one thing as a, as a maintainer of this project, like, you're designing for a very wide-
Yeah
... surface area of compute. Like, you even support CPUs.
Yeah. Well, that's-
It's-
... that's just, uh-
For fun.
For PyTorch. PyTorch CPUs.
Yeah.
So yeah, it's just-- that's not, that's not hard to support.
First of all, is there a market share estimate? Like, is it, like, seventy percent NVIDIA and, like, thirty percent AMD and then, like, miscellaneous on-
Uh-
... Apple Silicon or whatever?
Uh, for Comfy?
Yeah.
Yeah, I'm-- Well, yeah, I don't know the market share.
Can you guess?
Uh, I think it's mostly NVIDIA.
Right.
'Cause-
Yeah
... 'cause AM-- the problem, like, AMD works horribly on Windows.
Mm.
Like, on, on Linux, it, it works fine. It's, it's slower than the price equivalent, uh, NVIDIA GPU. But, uh, it works. Like, you can use it, generate images, everything works on Linux. On Windows, you might have a hard time.
So that's the problem. And most people-- I think most people who, uh, bought, uh, AMD probably use Windows.
Mm-hmm.
Yeah, yeah.
They probably aren't, uh, gonna switch to Linux, so So u-until AMD actually, like, uh, ports their, like, RAWCM to, to Windows properly, and then there's actually PyTorch. I think they're, they're doing that.
Mm-hmm.
Uh, they're in the process of doing that. But, uh, until they get a, they get a good, like, PyTorch RAWCM build that works on Windows, it's, uh, like, they're gonna have a hard time.
Yeah.
We gotta get George on it.
Node Playground35:08
Yeah. Well, he's trying to get Lisa Su to do one, but- Let's talk a bit about, like, the node design. So unlike all the other text to image, you have a very, like, deep-- So you have, like, a separate node for, like, clip and code.
You have a separate node for, like, the K sampler. You have, like, all these nodes. Going back to, like, the making it easy versus making it hard, but, like, how much do people actually play with all the settings, you know?
Kinda like how do you guide people to like, "Hey, this is actually gonna be very impactful," versus, "This is maybe, like, less impactful, but we still wanna expose it to you"?
Uh, well, I try to expose, uh, like, uh, I try to expose everything or-- But, uh, yeah. At least for the-- But for things like, for example, for the samplers, like, there's, uh, like yeah, four different sampler nodes-
Mm
... which go in easiest to most advanced. So yeah, if you go, like, the easy node, the regular sampler node, that's-- you have just, uh, basic settings. But if you use, like, the sampler advanced-- custom advanced node, that, that one you can actually-- you'll see you have, uh- Like different nodes
I'm looking it up now.
Yeah.
What are, like, the most impap- impactful parameters-
Yeah
... that you use? So it's like, you know, you can have more, but, like, w- which ones, like, really make a difference?
CFG.
Yeah, they all do. They all have- ... their own, like, uh... They all, like, for example, yeah, steps. Usually you want, uh, steps, you want them to be as low as possible, but y- you want-- If you're optimizing your, your workflow, you want to...
You lower the steps until, like, the images start deteriorating too much.
Mm.
'Cause that, um... Yeah, that's the number of steps you're, you're running the diffusion process, so if you want things to be fast, that's, uh, lower is better. But, uh, yeah, CFG, that's more... You can kind of see that as the contrast of the image.
Like, if your image looks too burnt out-
Mm
... then you lower the CFG. So yeah, CFG, that's how... Yeah. That's how strongly the, like, the negative versus positive prompts. 'Cause when you sample a diffusion model, it's, uh, it's basically a negative prompt. It's just, yeah, positive prediction minus negative prediction.
Mm-hmm.
Contrastive loss.
Yeah. So it's positive minus negative, and the CFG, that's the multiplier.
Yeah.
Yeah.
That's just... Yeah, so.
What are, like, good resources to understand what the parameters do? I think most people start with Automatic, and then they move over, and it's like steps, CFG, sampler name, scheduler, denoise.
Reddit.
But, uh, honestly, well, it's more-- it's something you should, like, try out yourself.
Mm-hmm.
And I don't think you, you don't necessarily need to know how it works to-
Right.
Right.
... like what it does, 'cause even if you know, like CFG, oh, it's like positive minus negative prompt.
Yeah.
So the only thing you know it's CFG is if it's one point zero, then that means the negative prompt is applied. It also means sampling is two times faster. But, uh-
Yeah.
Yeah
... yeah. But other than that, it's more like you should really just see what it does to the images yourself, and you'll, you'll probably get the more intuitive understanding of what these things do.
Mm-hmm.
Mm-hmm. Any other nodes or things you wanna shout out? Like, I know the AnimateDiff, IPAdapter, those are like some of the most popular ones. Um, yeah, what else comes to mind?
Uh, not nodes, but there's, uh, like I... What I like is when, when some people, sometimes they make, uh, they make things that use ComfyUI as their backend. Like there's a, a plugin, uh, c- for Krita that, uh, uses ComfyUI as its backend.
So you can use, uh, like all the models that work in Comfy in Krita, and, uh, I think I've tried it once, uh, but, uh, I know a lot of people use it-
Mm-hmm
... and find it really nice, so.
What's the craziest node that people have built? Like the most complicated?
Craziest node, like I, I like... Yeah, I know some people have, uh, made like video games and- ... in Comfy, uh, with like stuff like that. So- ... like someone... Like I remember like, yeah, last... Think it was la- last year, someone made a, like a, like Wolfenstein 3D in Comfy.
Of course.
And then one of the inputs was, oh, you can generate a texture, and then it-
Oh
... changes the texture in the game.
Nice.
So you could plug it to like-
Yeah
... your workflow, and there's a lot of... If you look, there- there's a lot of crazy things people do, so.
Yeah.
And now there's like a node register that people can use to like download nodes and-
Yeah, like well, there's always been the, like the ComfyUI manager.
Yeah.
But we're trying to make this more like, I don't know, official, like, uh, with, uh, yeah, with the, the node registry. 'Cause b- before, before the node registry, the... It like, okay, how did your custom node get in the ComfyUI manager?
Right.
That's the guy running it who like, uh, every day, he search GitHub for new custom nodes and added them manually to his, uh, to his custom node manager. So we're trying to make it, uh, a less effort for him basically.
Yeah.
Yeah, but I was looking. I mean, there's like a YouTube download node. There's like th- this is almost like, you know, a data pipeline more than like an image generation thing at this point. It's like you can get data in, you can like apply filters to it, you can generate data out.
Yeah. You can do a, a lot of, uh, different things. Yeah. Something I think, uh, what I did is, uh, I made it, uh, easy to make custom nodes, so I think that, that helped a lot for-
Mm
... like the, the ecosystem, 'cause it is very easy. Just make a node. So-
Mm-hmm
... yeah, a bit too easy sometimes. Then y- then we have the issue where there's a lot of custom node packs which share similar nodes.
Mm.
So, but, uh, well, that's, uh, yeah, something we're trying to solve by maybe bringing some of the functionality into core.
Yeah.
Video Models41:35
Yeah.
But, uh, yeah.
Yeah.
And then there's like video. People can do video generation.
Yeah. Video, that's, uh... Well, the, the first video model was, uh, like Stable Video Diffusion, which was last, yeah, exactly last year, I think. Like one year ago. But that wasn't a true video model.
Mm-hmm.
So it was, uh-
As in like it was like moving images?
Yeah. It generated video. What, what I mean by that is it's, uh, like, uh, it's still 2D latents. It's basically what they did is they took SD2, and then they added some temporal attention to it, and then, uh, trained it on videos.
And, uh, so it's, it's kind of like animate- Diff, like-
Yeah
... those same, same idea basically. Why I say it's not a true video model is that you still have, like, the 2D latents. Like, a true video model like, uh, Mochi, for example, would have 3D latents.
Mm-hmm.
So imagine-
Which means you can, like, move through the space basically is the, the difference. You're not just kind of like-
Well-
... reorienting
... so, yeah, and it's also... Well, it's also 'cause you have a temporal VAE-
Mm-hmm
... that also. Like, Mochi has a temporal VAE that compresses, um, like, the temporal direction also. So that's something you don't have with, like, yeah, AnimateDiff and, uh, Stable Video Diffusion. They only, like, compress it spatially, not temporally.
Mm-hmm.
So yeah. So these models, uh, that's why I call them, like, true video models. There, there's a-- Yeah, there's actually a few of them, but, uh, the, the one I've implemented, uh, in Comfy is Mochi 'cause that, that seems to be the best one so far.
Yeah. We had, uh, AJ come and speak at the State of Diffusion meetup. Other open one I think I've seen is CogVideo.
Yeah, CogVideo. Yeah. That one's in... Yeah, it also seems decent, but, uh, yeah.
Uh, Chinese, so we don't use it.
No, it's fine. It's just, yeah, I could... Yeah, it's just it, uh, there's a-- It's not the only one. There's also a few others which, uh-
The rest are, like, cr-closed source, right? Like Kling and-
Yeah. The closed source, there's a bunch of them, but I mean open. I've seen a few of them. Like, uh, yeah, I can't remember their names, but there's Cog-CogVideo is the big, uh, the big one. Then there's also a few of them that, uh, released at the same time.
Uh, there's one that released at the, at the same time as SD 3.5, same day, which is why I don't remember the name.
Oh. We should have a release schedule, so we don't conflict on each of these things.
Yeah. Yeah, I think SD 3.5 and Mochi released on the same day. So-
Uh-huh
... everything else was kind of drowned-
Mm
... completely drowned out. So, uh, for some reason, lots of people picked that day to release their stuff.
Yeah. Which is, uh, well, shame for those, uh, I think, yes. I think, uh, Omnigen also released the same day-
Mm-hmm
... which also seems interesting, but, uh-
Yeah
... well.
Yeah. What's Comfy... So you are Comfy, and then there's, like, comfy.org. Um, I know we do a lot of things for, like, news research, and those guys also have kind of like a more open source and on, uh, thing going on.
Comfy's Arc44:31
How do you work? Like, you mentioned you mostly work on, like, the, the core piece of it, and then what-
Maybe I should fill in because I, yeah, I g- Yeah, I feel like maybe, uh, yeah, I only explained part of the story.
Right. Yeah, yeah.
Yeah. Maybe I should explain the rest. So yeah. So yeah, basically, uh, January, that's when the first-- January 2023. January 16, 2023, that's when Comfy was, uh, first, uh, released, uh, to the public. Then, yeah, did a Reddit post about the area composition thing somewhere in, uh, I don't remember exactly.
Maybe end of January, beginning of February. And then someone, a YouTuber made a video about it. Uh, like, uh, Olivio, he made a video about Comfy in, uh, March 2023. I think that's when it got its real burst of attention, and by, yeah, that time I was continued to developing it, and it was getting, uh...
People were starting to use it, uh, more, which unfortunately meant that I, I, my-- 'cause, uh, I had first written it to do, like, experiments, but then, well, my time to do experiments-
Mm-hmm
... well, started going down 'cause, yeah. 'Cause, uh, yeah, people were actually starting to use it then.
Right.
Like, I had to-- And I say, "Well, yeah," had to add all, all these features and stuff. Uh, yeah, and then I got hired by Stability June 2023. Then I made the basically... Yeah, they hired me 'cause they wanted, uh, SDXL.
So I got SDXL working very well in ComfyUI because, yeah, they were experimenting with it. Actually, the SDX-- how the SDXL release worked is they released, uh, for some reason, like, they released the code first.
Mm-hmm.
But they didn't release the model checkpoint. Uh, yeah.
Oh.
So they released the code, and then, well, if since the research was released the code, I, I released the code in Comfy 2. And then the checkpoints were basically early access. People had to sign up, and they only allowed, uh, a lot of people from edu emails.
Like- ... if you had a edu email, like, uh, like, they gave you access basically to the zero-- SDXL 0.9, and well, that, uh, leaked.
Right.
Of course, because of course it's gonna leak if, if you do that. Uh, well, the only way people could easily use it was with Comfy. So yeah, people started using it, and then I fixed a few of the issue that people had.
So then the big 1.0 release happened, and well, ComfyUI was the only way a lot of people could actually run it on their computers. 'Cause it just, like, Automatic was so, like, inefficient and bad that, uh, most people couldn't act...
Like, it just wouldn't, wouldn't work. Like, 'cause he, he did a quick implementation. So people were forced to use ComfyUI, and that's how it became popular because people had no choice.
The growth hack.
Yeah.
Yeah.
Yeah. Like, everywhere, like, peop-people who didn't have the 4090, they had, like, who had just regular GPUs.
Yeah, yeah.
They ju- they didn't have a choice, so.
Yeah. I got a 4070, so think of me. And so today, what's-- Is there, like, a core Comfy team or...?
Uh, yeah. Well, right now, um, yeah, we are hiring actually. So right now, core co- like, on the core core itself, it's, it's me Uh, but because, uh, reason we're foc- like all the focus has been mostly on the front end right now-
Mm-hmm
... 'cause that's the thing that's been neglected for a long time.
Right.
So, uh, so most of the focus right now is, uh, all on the front end, but we are, uh, yeah, we will, uh, soon get, uh, more people to, like, help me with the actual backend stuff. Because that's...
O- once the f- once we have our V1 release, which is gonna be the package ComfyUI with the, the nice interface and easy to install on Windows, and hopefully Mac. Yeah.
Yeah.
Once we have that, uh, we're gonna have to l- lots of stuff to do on the backend side and also the front end side. But, uh-
Yeah.
What's the release da- I'm on the wait list. What's the timing?
Uh, soo- soon. Yeah. Like, I don't want to promise a release date.
Yeah, yeah, yeah.
Because, yeah, we, we do have a, a release date we're targeting, but yeah, I'm not sure if, uh, if it's public.
Yeah.
Yeah. And how we're gonna... Like, we're, we're still gonna continue, like, for f- doing the open source, like making ComfyUI the best way to run, uh, like Stable Diffusion models. Like, at least the, the open source side then, like, it's gonna be best way to run, uh, models locally.
Mm-hmm.
But we will have a few, uh, like a few things to, to make money from it, like, uh, cloud inference or, like, that type of, that type of thing. So-
Right
... and maybe some, like, some things for some enterprises.
I mean, a few questions on that. How do you feel about the other Comfy startups?
I mean, I, I think it's great.
They're using your name, you know?
Yeah. Well, it's better they use Comfy than they use something else.
Yeah. That's true.
Yeah. Like, yeah, it's fine. I don't... Like, we're... Like, yeah, I'm-- We're gonna try not to... We, we don't want to... Like, we want them to-- people to use Comfy, 'cause, uh, like I said, it's better that people use Comfy than use something else.
So as long as they use Comfy, it's, uh... I think it helps, it helps the ecosystem and stuff. 'Cause more peop- even if they don't, they, they, like... Even if they don't contribute directly, the fact that they are using Comfy means that, like, people are more likely to, like, join the ecosystem.
So yeah.
What's Next50:58
And then would you ever do text?
Yeah. Well, the... You can already do text with some custom nodes. So yeah, it's something we, we like. Yeah. It's something I've wanted to eventually add to core, but it's more, like, not a, not a very high priority.
But, uh, 'cause a lot of people use text for, like, prompt enhancement and, like, other things like that. So it's, uh, yeah, it's just that my focus has always been, like, um, diffusion models. Yeah, unless some text diffusion model comes out.
Yeah. Uh, David Holtz is investing a lot in text diffusion.
Yeah. Yeah. Well, if a, if a good one comes out, then well, he'll probably implant it since it fits with the whole-
Yeah. I mean, I, I imagine it's gonna be closed source to Midjourney, so.
Yeah. Well, if a, yeah, if an open one comes out.
Yeah.
Yeah. Then, uh, yeah, I'll probably, yeah, yeah, I'll probably implement it. It's, uh-
Cool, Comfy. Thanks so much for, for coming on. This was fun.






