Intros0:00
Hey, everyone. Welcome to the Latent Space Podcast. This is Alessio, founder of Kernel Labs, and I'm joined by Swyx, founder of Smol.ai.
Hello, hello. Uh, today we're so excited to be in the studio with Gorkem and Batuhan of Fal. Fal, uh, welcome.
Yeah, thanks for having us.
Thank you for having us.
Long time listening, first time calling.
Gorkem, you and I actually go back a long way, uh, to when it was still features and labels-
Yeah
... and you were just coming out of Amazon.
That's correct, yeah.
Um, and, uh, and you know, it was, it was so-- It was-- I don't even, I don't even remember the pitch. I honestly, I should look at my own notes, but you were optimizing runtimes.
Yeah. It was first, like, we, we were building a feature store, and then we took a step back, and then we decided to build a Python runtime in the cloud, and that evolved into, uh, an inference system. That evolved into what Fal is today, which is, uh, a generative media platform.
So we, we optimize inference for image and video models and audio models, but we, we, we do a lot more. We try to own this whole generative media space for, for developers, basically.
Yeah, amazing. And we can talk about that-
Yeah
... that journey. I wanted to also introduce Batuhan. Uh, I'm new-- we're newer to each other, but you've, you've come to some of my meetups before. You're head of engineering. Uh, yeah.
Yeah, I, I lead engineering here at Fal. You know, glad, glad to be here.
And what, what's your journey, uh-
Uh, I met Burkay in 2021, uh, when they were just starting the company, and, like, just before the seed round, you know, Burkay and Gorkem, we, we met online. We are all both Turkish-
Yeah
... uh, so I think that's, that was the connection. We just met, and then they said, "Oh, why don't you join us?" And, like, I, I'm coming from a-- Uh, I was one of the core developers of Python language, so I, I had, like, really, you know, really good experience with developer tools around the Python language.
So I started coming here to build the Python cloud, which evolved into this, like, inference engine and the generative media cloud that we are building today.
Yeah. And now you spend time-- less time with Python and more time with, I don't know, uh-
CUDA.
CUDA. Custom kernels and-
Yes, exactly.
Yeah, yeah.
Um, yeah, I remember the dbt Fal-
Yeah. Yeah, yeah
... uh, when the modern data stack was hot. Can you guys maybe just give a quick sense of the scale of Fal? So you just raised $125 million Series C.
That's correct. Yes.
Uh, we can talk about how I passed on one of your early rounds. Uh, we can go through, through that. How many of the developers, how many models do you serve, and maybe any other cool numbers. Yeah. We have around two million developers on the platform, um, and, like, we-- for, for the longest time, we required GitHub login.
It recently changed, but... So I'm assuming everyone who has a GitHub account is a, is a developer. Um, we have around 350 models in the platform. These are mostly image, video, and audio models. It used to be only image, and then we added audio and, uh, the, the space evolved into video as well.
Um, and yeah, that's, that's pretty much the scale. We just closed-- announced our Series C round, and w- we've been growing a lot in the past year, and, and it still continues.
Yeah, you had a very nice, uh, Series C party. Um-
Yeah. Thank you
... and you guys are over 100 million of revenue, right? Just this is not-
Yeah
... you know, just developers kinda kicking the tires.
That's correct. Yeah.
Um, that's great. When you say 350 models-
I think-
... uh, what percentage of all the models that you could serve is that? Because, uh, you know, especially in-
There is an infinite amount of-
Right
... fine-tunes, post-training versions of these models. We are trying to serve the models that fit-- that fix a gap, uh, you know, that fill a gap in, in, in the stack. So we don't add a model that's, like, significantly worse in any aspect compared to other models that we have.
We're pr- We are trying to bring unique models that solve a customer's needs, so that's like these are 350 models. You know, there's like 20, 30 text image models, but, like, one of them excels in logo generation. Another one excels in human face generation.
So, like, th-- every model has a unique personality, but if a model is, like, significantly worse in all aspects, we don't add that to the platform. So there's, like, infinite amount of models that we cannot.
A-and do you re-rely on your own evals or just like-
Yeah
... what the community does?
We, we, we, we mainly, uh, rely on our, our, our own evals as well as, you know, like we a- We are, like, very-- We are in the community, so we also, like, follow the community very well to see, like, what the-- what is gonna be the thing that's gonna be in the next generation of apps.
So if we think something-- like, we have a good intuition, if we think something is gonna pop up, we just add it.
Yeah.
Yeah. Uh, to my knowledge, you haven't published your own evals, right?
No, we don't publish evals.
It's, it's internal.
It's-
And the community is Reddit and Twitter.
Twitter, Reddit, uh-
Yeah
... you know, Hugging Face, seeing how popular the models are on Hugging Face, uh-
Yeah
... and other demos.
Okay. I just want to give people a sense of where to get this info.
The, the best part of the job is the day of a model release, the adrenaline rush that comes with it- ... the whole team trying to scramble something together and release it, and it, it happens every week.
Right.
Uh, every week is, is, is exciting. Yeah.
Model Wave4:46
Can we do maybe a brief history of, like, the models that were, like, the biggest spikes maybe in usage?
Mm-hmm.
You know, you kinda-- I, I think everybody knows Stable Diffusion.
Yeah.
You know, and then you have maybe, like, the Flux models, and then you have Blackboard Labs.
Yeah.
You have, like, all these different models.
History-wise, I think the biggest, like, the, the initial hit was Stable Diffusion 1.5, which is when we actually pivoted into this new paradigm of Fal Generative Media Cloud. We started hosting it. We noticed, like, we had the serverless runtime, and everyone was running the Stable Diffusion 1.5 by themselves, and we noticed it's terrible for utilization, and they are not optimizing it.
So let's just offer an optimized version of this that's ready for API to be scaled and doesn't require people to deploy Python code because we want product engineers to start using it. We want mobile engineers to start using it.
So we start offering one-- the Stable Diffusion 1.5. It was very popular. The fine-tunes around it was very popular. Uh, Stable Diffusion 2.1 came. It was a bit of a flop, so it didn't, like, you know, you know, got that much attention.
And then SDXL came, which was, like, the first major model that brought, like, our first million in revenue, uh, if, if you con-consider that. Uh, and with SDXL, obviously, like the po-- uh, small fine-tuning ecosystem also, like, tried-- exploded.
People started fine-tuning their faces, their objects, whatever, and generations with this, like LoRAs started becoming very popular. And then after Stable Diffusion XL, there was, like, a bit of a Quietness around it, you know, the SD3, uh, there was, like, some drama around it.
And, uh, the t- the team at Stability left to start Black Forest Labs, which released Flux models. And that was the first model to, you know, breach the barrier of commercially usable, you know, enterprise-ready great models, where in the first month of Flux models we reached from, like, 2 million to 10 million in revenue.
Uh, that was, like, a big jump. Next month we were at 20, like, it, it just started going from there. And then Video Model started ca- came around. You know, we partnered with Luma Labs, we partnered with, uh, o- other video model companies in China.
We partnered with Kling, Kuaishou, uh, MiniMax. And the- with, with these models, like, you know, it created another market segment. That was a big jump. And with, with... I think the, the s- the final biggest thing was Veo 3, where it actually created this, like, usable text-to-video component, where before text-to-video was, like, a very boring, uh, soundless video that you would like, you would not get enjoyment out, whereas now it's, like, a, such a great experience.
You can create all the good- all these, like, memes that we're seeing online, all these ads. So that was, like, another big jump for us, like, partnering with Google DeepMind for Veo 3.
Yeah. Uh, well, actually that's a really good history of generative media-
Yeah.
... that, that soundbite. Uh, uh, so I wanted to double-click on that because obviously we can s- we can, um, dive, I, w- I, I think everyone's interested in video, but there's a whole history of the image side that I wanted to cover first.
Just definitely wanted to start with was just the decision to pivot. I think I just wanna double-click on that. You know, it, it's not a trivial decision, but obviously the right one.
Mm-hmm.
Pivot7:29
At the time, uh, I would say, like, a lot of people were hosting Stable Diffusion, right? So it wasn't obvious that you can just build an entire company around effectively just specializing in diffusion in inference. What gave you the confidence?
What were the, the debates back and forth?
Yeah, I mean, uh, I, I think, so a couple decisions we had to make there. What, w- we could have evolved the company into more towards GP orchestration and, like, essentially we had this Python runtime, we were running it on top of GPUs.
Like, that could have been the company. But we saw every single person, every single company who are using what we had, like, a little SDK to, to run Python code on, on GPUs, they were doing the same thing.
They were, um, you know, deploying a Stable Diffusion application, maybe using some LoRAs on top of it, different versions of it, inpainting, outpainting, things like that. I mean, it was very wasteful. We, we decided, okay, this needs to be an API where we actually optimize the inference process and everyone benefits from, from it.
And, like, you can run it multi-tenant, you know, the utilization m- is much higher than. So that was the decision number one. And then obviously after Stable Diffusion, I think, like, four or five months later, Llama 2 came out and, um-
Right. So th- then-
There was a, there was a decision point again
... you could do language models.
Yes. Exactly.
You know?
Uh, and a lot of the inference providers at the time, there were maybe a couple of them, and they all went all in on language models, and we decided, you know, language models, hosting language models is, is not a good business.
At the time we thought, okay, we are gonna be competing against OpenAI and Anthropic and all these labs. Turn- turned out that it was even worse because the killer application of language models is search, and you, you are competing against Google at the end, and Google can basically give this for free if they can- ...
because it, it's s- so important for them and, you know, it threatens their business right away. And with image and video models, it was a net new market. Uh, we were, we weren't going against any incumbent. We weren't, like, trying to get market share from someone much bigger than us, and w- we liked that aspect of it.
We, we, we thought we could be a leader here. It was a niche market, but it was very fast-growing. So we, we chose to be a leader or, or play a-- to be a leader in th- in this, in this fast-growing niche market rather than trying to go against Google or OpenAI or Anthropic.
So that was the decision we made, and turns out it's a good one because we are able to define the market we are in and educate the people and grow with it. And so far it's been growing fast enough that we are able to build a whole company around it.
Yeah. Uh-
Yeah.
And I think you noted at, at, uh, AIE that, uh, you know, now there's a generative media track-
Yeah.
... and a gen- generative media specialist-
Thank you for calling it generative media, by the way.
Um, yeah. I mean, obviously it's, it's a, it's a thing and people care about it-
Mm-hmm
... and I, I do think it's gonna change the economy. A- and as a creative person, I think, like, I also wonder-
Yeah
... uh, what it's gonna do for, for us. Um, a- and but, like, I think just so I wanna keep it technical-
Yeah
... and keep it, keep thinking about the pivot because I think it's, uh, still, like, one of the most interesting pivots I've, I've seen in the AI era. Um, you were not, like, CUDA kernel specialist at the time, right?
Performance10:46
I come from a compilers background-
Yeah
... so my, my job was optimizing, you know, like, Python bytecode interpreter to make stuff faster, which is performance engineering. And, like, yes, like, I don't think at the time there was that many CUDA kernel specialists either.
Yeah.
So it's like we were at, like, the right time, you know, the very... It was like, uh, the, the-- actually, like, the, the, the space was actually so, so much worse than what we have today, where, like, the running basic, like, Stable Diffusion 1.5 was, like, a UNet with convolutions, and the convolution performance on A1s was like you're getting, like, 30% of the GPU power if you just use raw touch because no one cared about it.
So there was, like, so many low-hanging fruits that we s-started to pick up and started optimizing, and it kinda evolved, evolved, evolved. Right now it's, like, much more competitive space with, like, NVIDIA has, like, a 50 person, 100 person kernel team that's writing kernels.
You're competing against that. At the time, no one really cared about it. So it was, like, a good, uh, new field for us to go thrive.
And there's no community effort like a VLM-
Not, not-
Exactly. When these models were first released, like, no one in the world has ran them in production. Like, it just w- didn't exist.
It's like a research output of-
Exactly. Yeah
... of, uh-
It was-
... Stability.
Yeah. You had your maybe local GPU, maybe you had, like, a single GPU that you rented from the cloud, and basically this was a research interest rather than, uh, a product interest. And no one at Meta, no one at Google had run this in production.
So we also thought this is a good, good time to start a company around this and actually spend time optimizing it as much as we can because, you know, if, if we can get millions of people to use this, uh, there is a lot of economical value to be created there.
Can you talk a bit about
... how much of a performance boost you get because I know when I met you guys, you were about a million in revenue. You were like, "Well, we're writing all these custom kernels."
Yeah.
And, uh, maybe part of it is like, okay, how many kernels can you actually write?
Sure.
You know, as you-
Yeah, yeah, yeah
... as you support all these different models. Like, what's kinda like the breadth of them? Like, are you writing kernels that you can reuse across models? Like, how much work do you have to do on a per model basis?
It really evolved in the past three years. You know, when we first started, there was a single model, Stable Diffusion 1.5. So every-- all of our kernel efforts were, how do we make Stable Diffusion 1.5 as fast as possible?
You know, you go from, like, 10 seconds with PyTorch. At the time, there was no, not even, like, a Torch compile, Torch inductor, whatever. So you were going from 10 seconds to maybe, like, two seconds on, on the same GPU.
Uh, and, like, we, we started with that. The next thing, like, you know, uh, we, we-- the, uh, like with, with adding more models, you know, like Stable Diffusion XL was a different architecture. PixArt was a different architecture.
All these, like, different architectures started coming around. We said, "Let's build an inference engine," which is what we call a collection of kernels, parallelization utilities, diffusion caching methods, quantization, all that stuff combined into one package. And so we built this inference engine.
Uh, the same time, PyTorch 2.0 was released with Torch Inductor and Torch Dynamo to do, like, Torch compile, which is, like, essentially a way to trace the execution of, of, of, uh, your neural net and generate Triton kernels.
There are a few as, like, that are more efficient. Uh, and I'm a big sucker for just-in-time compilers. I used to work at PyPy, uh, like a just-in-time compiler for Python. And we said, "This is a great idea.
Let's apply this, but a more specialized, more vertical way for diffusion models." At the time, it was UNets. Now it's diffusion transformers, which are significantly different than your autoregressive transformers in terms of, like, the profiles of, like, how compute-bounded it is, what sort of the kernels are taking the majority of the time, you know, if they are doing bi-directional attention or causal attention.
So we started doing that, and now it's like what we have today is a inference engine that's, like, applicable, that, that gets you, like, 70, 80% of majority of the models on diffusion transformers, and we still have, like, a lot of custom kernels for a lot of models to squeeze out, like, because they're still, still small.
Every model wants to make an architectural difference. You guys see this on like, you know, even for stuff like, you know, Qwen, DeepSeek, whatever. People want-- Like, even if we know an architecture is the best, they want to tweak it a little bit just to make sure, "Oh, we are releasing something cool."
So we, we saw this and then, like, you know, for that we have to write, like, custom kernels for custom RMS norms that people are doing or whatever, like, stuff like that. So we, we have, we have a decent amount of kernels, like over 100 of custom kernels.
Uh, uh, this doesn't include the auto-generated ones. You know, we have templates of kernels that generates, like, you know, for thousands of, uh, different shapes, problem spaces, whatever. But, like, if you consider those, you know, like, we have tens of thousands of kernels obviously at runtime that we are running and dispatching, but that's pretty much, like, the depth and-
Yeah
... breadth of it.
And on average, a model on Fal runs 10x faster than I would self-host it. Like, if I just take Stable Diffusion, right, and upload it-
It, it... So this, this is, like... Do we-- Like I, the, I, I know, like, this might be a bigger discussion point. Do we consider speed as a mode? Uh, this, this comes to that.
Right.
It e- Like the existing open source industry evolves so fast where, you know, like, if you go to, like-- This might have been true three years ago. Now PyTorch is, like, already, like, very, very good for H100s, right?
But what about P200s? Uh, when you use PyTorch with P200s Blackwell chips, you're not getting the best performance. So our main objective and our main goal is for the, for whatever, whatever GP type you're using, for dif- these diffusion models, we're gonna extract the best performance.
At any point in time, it could be 1.5x, it could be 3x, it could be 5x. For certain models, it could be 10x. So I, like it would, like, it would be a bit of a unfair thing to say, "Oh, we're gonna make everything magically 10 f- 10x faster."
No one in the world can do that.
We are lucky that this is a moving target and open source community, everyone catches up, but at the same time, new chips come out, new architectures are released. So we are always ahead of, like, what's possible, but then they catch up.
But we have to, we have to stay ahead of it, and that's how we can create differentiation because it's a moving target, because there's so much going on. We are-- When, whenever something new comes out, we are the first one to optimize it, first one to adapt our inference engine to it.
So we are, at that time, the fastest place to run it. That, that helps with margins, things like that. But eventually, people do catch up. I, I think it's very hard to create, uh, this differentiation over long term if there is no new architectures, if there's no new chips.
But luckily, there is all the time.
Right.
Yeah.
Yeah, and I think with image specifically, um, there's kind of like you cannot stream a response, so to speak. So when you have a language model, it's like you're kind of bound by, like, how quickly you can read.
How quickly you can read.
So even with, like, Groq, it's like it's impressive to show-
It's a good demo. Yeah.
... 1,000 token a second, but it's like I'm not reading that fast, right? So it can go slower.
Yeah.
Versus with images, it's like you just need to see it. That's why Midjourney now has, like, the draft mode, for example.
Mm-hmm.
It just gives you this, like, very low quality thing.
Low resolution. Yeah.
Yeah, but at least you can see whether or not it's going in the right direction. How much of that is actually true for, like, your customers? Like, what do, what do they care about the most? Like, is latency that important?
Yeah, latency-
Like, what, what's the range of latency that matters?
Yeah. Ra- latency is really important. One of our customers actually did a very extensive A/B test of, like, they on purposely slowed down latency on Fal to see how it impacts, you know, their metrics, and it had a huge part in it, and it's, it's almost like page load time.
Right.
When, when the page loads slower, you know, you make less money. I think Amazon famously did a, did a, uh, a b- very big A/B test on this. It's, it's very similar. Like, when you, when the user asks for an image and, you know, iterating on it, if it's slower to create, then they, they are less engaged.
They create fewer number of images and, and things like that. Yeah.
It's the same learning that Amazon has, like every, you know, 10% inc-
Improvement
... improvement in speed.
Yeah, exactly.
And, uh, the, the elasticity is high. Um, yeah, so, so and then the other thing I wanted to also d- dive into, you know, like, um, uh, I-- putting my, a, a little bit of the investor hat on, one of the reasons for Fal's success is kind of with- not within your control, which is h- when and how people release, release open, uh, models for dif- for diffusion.
Open Source18:18
Mm-hmm.
Um, which, like, at the time, it was just stability and, like, there wasn't, there was no Chinese-
Yeah
... uh, you know, output. I mean, what we... We did have other image models, but they were not great.
Yeah, yeah.
And, like, and so, so you, you made a bet on w- when it was just, like, wasn't super obvious, I think. But then the other thing is, is what you're touching on, the diffusion workload is very different from the language workload, and the language w- workload is being super optimized whereas diffusion is not.
So you, you just, like, had kinda no competition for a while, which is fantastic for you.
100%. And, like, the open source, we benefit a lot from it obviously, but, m- like, in the, in the past six months, a year, we started working with some of the closed source model developers as well, like behind the scenes helping them with inference.
But they're not sending you their weights.
They do.
They are?
They are.
They do.
Yeah, yeah.
Wow.
Yeah.
Like how-- what, what do you have to give secu- like- ... guarantee them?
I mean, we are any cloud pro-
Yes
... like, what, what, what do they think in like AWS or Google Cloud or these neo clouds? There's like 50 neo clouds, right?
Uh-huh.
Like we, we are not that different from any other cloud provider, and this is why we packed the inference engine in a way that, you know, they can self-service and get 80%, 90% of the performance. So they don't even have to show us like their code.
They deploy to our-- we, we have our own cloud platform where your-- our inference engine is a-available only in that platform. So they can tap into that when they deploy their code and their model weights to us, and we don't really have to look at it.
And if they want to collaborate with us, which some companies did, uh, in the past, where we would just essentially we have performance engineers acting as forward deployed engineers on their behalf and writing custom kernels for them.
Okay.
Yeah.
Wow. So, so, like, uh, a-and you're doing this-- have you disclosed who, who you're doing this for? Like-
We, I-- we disclosed Play.ht, Play.ai.
Okay.
That was one of those. We have like four different companies, uh, four, four major video companies that we are doing this with-
Yeah
... and one image company, uh, that I don't think we disclosed.
Yeah.
As you can imagine, it's a little sensitive for them, so it's-
Yeah.
Yeah.
Yeah.
Um, yeah.
I, I would say like, so like Replicate started serving v03 models and I was-- and we were like, "Okay, are you just wrapping their APIs or something?" And, and I think so it would... I-it's not obvious-
Mm-hmm
... like how much integration there is going on and how much that's on your infra or like your tech.
Just to be honest, like some of that is happening too.
Yeah.
v03, uh, like I think everyone is-
It's just API wrapper.
Yes. Yes.
It, it is, yeah. You, you have a dedicated pool that you can serve to your customers with like different speed, SLA guarantees, whatever.
Okay.
That's like how it, how it would work for something like v03.
So, so, uh, but your objective is to be one-stop shop, but then also you can do inference better than some of these other-
We, we-
Like Play.ht
... we think like the, the, the-- like with Google it's obviously hard, but like with other vendors, like our goal is like helping them run inference because these are research labs that doesn't necessarily invest heavily on inference optimizations.
Yeah.
Scaling up infrastructure, that's like another challenge that we can talk about. Like all-- at launch days, like some of these models, like in their website, they just like explode and Fal API is working fine because they deploy them.
We can scale up to like thousands of GPUs instantly. So the, the, the, there, there is that aspect too. When we pitch this value prop as well as the distribution that we bring to them, it's a no-brainer for them to just like deploy their model to Fal and use it for both the Fal marketplace as well as their own distribution channels.
Yeah. Uh, a couple follow questions. Uh, just on Play.ht, just 'cause you mentioned it.
Mm-hmm.
Uh, m-music, audio, uh, is that a different workload than normal diffusion, or is that the same?
Some of-- like I, I can't really comment on their architecture, uh, but like some of them are autoregressive models in the o- like open source world, some of them are au-autoregressive, some of them are diffusion based. So, you know, there's like notorious ones that's known for diffusion, uh, as you guys can, uh, guess, like one of the biggest companies.
So it, it, it's, it's similar workloads but at the end of the day, you know, uh, our performance, uh, our inference engine is like very versatile and our performance team is very versatile. With Play.ht, we had like very deeper, very deep collaborations where we had like three engineers at some point, you know, helping them, uh, optimize their inference process as well as infrastructures to get them like ADMS end-to-end time to first audio chunk, which is like a very impressive thing for real time text-to-speech workloads.
And then the other hard, uh, known hard problem is serverless GPUs-
Infrastructure22:09
Yeah
... which is a, a thing that, uh, everyone has chased a lot and many people have failed. Um, what can you say about like what you've done there to, to make it happen? Like, uh, so for example, Model has been talking a lot about their, uh, G-GPU snapshotting, uh, but like I, I imagine it's like a stack of technologies in order to-
It is a stack-
... achieve the, the scaling
... stack of technologies. Es- the, the biggest problem with serverless GPUs is are you just build- are you just wrapping another like if you have like a Kubernetes deployment, are you just wrapping it and giving people access? Or are you actually like multi-cloud?
Do you manage your own orchestration chain? Do you manage your own container runtime? Do you manage like all this like stack? And in, in our case, like we, we started with a Kubernetes version when we were just doing it for ourselves, and we s- like, uh, and Kubernetes version at Google Cloud was fine in 2022 when we wanted to get eight A100s.
But when we wanted to go to like thousands of H100s, it's not, it's not gonna work. It's a, it's a terrible position to be bound by a single cloud. So right now we work with six cloud providers, and we have 24 different data centers in four different countries.
And we, we now do like long-term data center leases as well to manage like the, some of the, the, uh, hardware chain ourselves. And in, in, in this world, like we had to build our own orchestration layer, we had to build our own distributed file system, we had to build our own container runtimes, all, all the stack to make sure that the cold starts are extremely, extremely fast, which is one of the things when you're scaling up, as well as actual...
handle actual scale where, you know, like we are managing over like 10,000 plus H100 equivalent today.
Yeah, and a CDN for caching.
CDN, that's like outside of the serverless infrastructure-
Yeah
... but like CDN, content moderation systems, like all of these like all, uh, consists the platform. Like there's like so, so much-
You also do moderation. Do you, do you-
We also offer content moderation services to the foundational model companies for them to like moderate their inputs and outputs.
Yeah, I see. I see. As a separate, uh, product.
Yes.
Oh, interesting.
Um, from a GPU perspective, do you always need to be on the latest? You keep mentioning H100s. Um-
Majority of our workloads are in H100s because price perf-wise, it makes sense. But like Blackwell is, is obvious, like we, we have a-- like we have five people dedicated to writing Blackwell kernels right now to make sure we can like...
Because theoretically it looks good, right? Like, uh, FLOPS dollar-wise it makes sense, but can you reach the actual FLOPS? No. So we have a dedicated team that's like working with Nvidia directly to write custom kernels for Blackwell for diffusion transformers to get to the, get to the point where it makes perf-- dollar makes sense, and then, then we would start with our own workloads, as well as some of our foundational model companies.
We would ask them, "Oh, if you wanna migrate to Blackwell, here's an inference stack that already works."
We are at that point where we should be the ones pushing the boundaries on Blackwells because no one else is doing this work, and, uh, maybe it doesn't make sense economically right now price perf-wise, but we know it can.
So we are like working towards maybe like couple months away from that point, and then when-whenever it does, we'll probably switch as many workloads to Blackwells as possible.
Just to be super crazy-
Hardware & Chips24:56
Mm-hmm
... when does it make sense to just work on a ASIC?
I don't think it does. Uh, that, that's like honest opinion. You know, like the peop- this is like one of the most controversial topics, right? Like is, is, is all these like ASICs a great idea? If you're like SRAM, if you're a memory bandwidth bound and like you can put all of SRAM, is it like even economically viable at that point?
I don't know. But like there's, there's, uh, you know, there's... If you look at like, you know, the, the, the summits ar- around like th- these chip designs, you see like, okay, what is the overhead of an N- Nvidia GEMM instruction, right?
It's like 16%. So like you're essentially buying a matrix multiplication machine. So at like, it doesn't really make sense to specialize it that much. Uh, and like some of the like, like V300s are gonna have like a better softmax instruction that gets like 1.5x, whatever.
And like that might be one way where NV- like Nvidia gets like, you know, better performance out of the majority of like, for the majority of workloads, which is like attention-heavy stuff. So I think it might make sense for like Nvidia to like add more specialized stuff, but like for us, I don't think it will ever make sense to build ASICs.
Yeah, uh, just, just thinking about, uh, like from first principles that the diffusion workload is very different, but al- also obviously there's still a lot of changes in the architecture that, uh, you need to just-
Correct
... be general purpose for.
We don't have a single model where we are trying-
Correct
... to optimize. We are trying to do it for the newest, the best, like always. The flexibility is therefore really important.
I was gonna show, I'm gonna pull up the Qwen MMDIT where there's-
Yeah
... like this dual streaming thing-
Yes
... which I last, I think, uh, SD3 had it.
Yes, SD3 Flux.
Uh, yeah. Is that the standard, standard model now instead of when-
MMDiT, so the, it's also an controversial topic. You know, uh, Scaling Rectified Flow of Transformers paper, the SD3 paper came up with this architecture and then, uh, o- one of our, uh, research team actually, like Simo Raio, he's like our head of research, he found out that like just using MMDITs are inefficient.
You need to mix them, and now there's like controversial opinions. You know, like Movie Jam paper were saying, "Oh, MMDIT is complete unnecessary. You can just use a single stream DIT," whatever. So like there, there's like controversial opinions happening in terms of architecture changes, which I understand because everyone wants to do a different architecture.
No one wants to do the same architecture-
Yeah
... because it's lame. Like otherwise, like it's just a matter of compute and data, and these researchers don't feel proud that their model is an output of data and compute. They wanna make a novel research change. So I, I think the architectures is gonna ke- keep changing until this like paradigm of like, you know, researchers changing stuff for the sake of changing, you know, finishes.
Yeah. Uh, I'll talk about a couple other architectural things just to keep it bounded within this topic. Uh, the d- uh, distillation was a thing for a while. SDXL Lightning, you guys did-
Yeah. Yeah
... fantastic demos of, uh, tail draw, which we've also had it on podcasts. Fantastic episode. What happened to those things? How come they're not popular anymore?
I think it makes for a good demo. Uh, you know, you could build real-time applications. You could build these drawing applications, things like that. But I don't think people could, uh, build applications that like have user retention long term.
Like, people couldn't really build useful things with it.
So let me-
Yeah
... let me play out what I thought was going to happen-
Mm-hmm
... and, and then you tell me why it didn't happen.
Sure.
Which is consistency models for drafting or like quick... Like it's, it's like you, you use your, uh, uh, uh, a hand pen to draw things-
Mm-hmm
... and it, it creates the, the, the draft, then you, you upscale, right, with, with a, with a real model. But like that's it. Like why, why, why can't it be a two-stage process instead of one stage?
Yeah. Uh, and I think one thing that happened is Flux gen- like that generational models were not good at image to image when it first came out. So like you need a good image to image model-
Yeah
... to be able to draw. Um, maybe it needs to be revisited around this time with some of the editing models maybe.
Like i- image to image and ControlNets. Flux like-
Yeah
... SDXL or ControlNets were very popular where like people were used to do this stuff like sketch to image, whatever.
Mm-hmm.
And like with, with Flux, I think people cared less about it. Uh, one, one, one, one thing that I, I, I keep thinking about this is like is this true for LLMs too? You know, like do you, like would you, like I, I always default to Claude 4.1 Opus, right?
Even if it's slower than Sonnet, it's just like I know I'm gonna get the best quality and like-
Exactly. That's what's happening here.
Yeah.
Yeah. That is, it seems like what, what is happening here as well.
But as a... Okay, anyway. As a creator- ... I'm like I want fast, quick drafts and then, and then I can refine, right? Like, so I don't know, I don't know why it didn't happen.
More true for video models, right? It used to be like five minutes, four minutes for a single five-second generation. Now it's mostly under a minute, but you want like 10 second, five second generation and then because the, the, the workflows of like creatives when they're working with it, they generate a ton of videos and then like pick one and then like m- create a story around it.
So when, when you watch these people actually generate videos, they generate like hundreds at a time, and they have to like kind of sit, sit around and wait and then like- ... iterate on it. Like the, the faster speeds mean a lot for, for creators.
Yeah. It, it, it does.
Yeah.
Uh, the, the, the other thing I wanted to-
Mm-hmm
... also briefly touch on before we, we go back to, uh, the, the, the main topics are, is the autoregressive models which you mentioned, right?
Video & Worlds29:56
Yeah.
Like, uh, obviously Gem- I, honestly I still think Gemini is underrated because they were first.
Mm-hmm.
And then, but then obviously, uh, OpenAI did the 4.0 Image Gen and that was a huge thing. I actually even wonder if there was a panic for you guys because obviously it like this is Sora Image Gen and like no one else has it.
It's not in, it's not open source.
Have you passed through those eras so many times- ... when, when, you know, like when, when-
Stop worrying about it.
Yeah. You know, Gorkem has like good stories around Dolly.
Yeah, I mean, so I, I talk about this how like when DALL-E 2 first came out, I was like, "Okay, OpenAI is so far ahead of- ... anyone else, like it's impossible for-"
And Midjourney.
Uh, like and then people caught up within months, and then Stable Diffusion was even maybe better or just as good as DALL-E like couple months later, and it was open source. So like a year later, same thing happened with Sora.
Like they, they put-
Yeah
... out those videos, and that time around, like I think we were excited because now that people see that it's possible, like this is like actually doable, researchers get motivated, and we- they, they see the hype. They see that this is possible, so they work on it, and within couple months we had maybe not Sora level, but much better video models.
Now we have video models that are much better than Sora. So whenever we see someone actually pushing the frontier, it's, it's a reason for excitement because now that's possible, other people are just gonna do it within a couple months.
So we don't panic anymore.
Is the fact that Anthropic doesn't have a image generation model tell you anything about what the larger labs care about?
Or do you feel like, "Hey, you know-
It, it, it tells more about Anthropic's own, like, personality-
Priorities
... than-
Yeah
... like, the in general, like what other, like ... Because every, like, you- if you look at xAI, if you look at Meta, if you look at, uh, OpenAI, if you look at Google, they all have really good image models.
Yeah. Like Google, in their last, you know, uh, announcement, they, they used the word generative media, by the way, which was a proud moment-
That's a win
... for us. Like-
That's a win
... and, and a lot. And, you know, they focused on generative media as much as th- their new LLM models. So some labs definitely care about it, and some labs it's not a priority for them.
Ex- look at xAI.
Yeah.
They keep pushing, like, image and videos.
You like AI slob.
Yeah, I know. I know, it's crazy.
And, and waifus.
And levels of interactivity. You have images. You have video. Uh, now you have Genie, this kind of like more a world model. You have kind of like gaming applications of that. Um, how far are we from, like, Fal getting a lot of traffic on, on those models?
Like, is it mostly experimental today-
Yeah
... in open source? Obviously Genie is impressive, but, like, it's a Google model, you know?
I have a very optimistic take on this, and that may be like, you know, uh, a normal outcome. I think at, at, at worst we are gonna have very capable video models that come out of world models, right?
It's gonna be a very controllable video model, and the use cases will be similar to what video models are. You're gonna create content, but you're able to control the camera angles. You're able to control the, the video model a lot better than what you can do today.
Like at worst, we are gonna get that from world, world models. And at best, I think it's very hard for anyone to predict what's gonna happen. Yeah, movies and games, like, it's gonna be something in the middle where you can, uh, be, be part of, uh, the, a whole, whole movie universe, uh, that c- that's gonna be playable.
So it's boundless possibilities what's, what's gonna happen at best. And how affordable is it gonna be? Like, is this-
Right
... ever gonna reach, uh, you know, mainstream adoption? We'll, we'll see all that, but it's definitely technically incredibly exciting and impressive what's, what's coming out of these labs.
Yeah. I need to find the paper again, but, uh, there was this study on, like, um, um, video models and, like, image generation understanding physics.
Yeah.
And, like, it could, like, predict the orbit of a planet, but then when it actually had it draw out the, uh, gravitational forces, it was, like, completely wrong.
Yeah, yeah.
You know? And so I, I think, like, that's my thing with world models, like, I understand the creator application, which is how you can create consistent world.
Uh-huh.
But I don't know if, like, the other side of people that are like, "Hey, these are, like, the best way to, like, simulate the world and, like, get intelligence and things like that." I don't know if that's-
There's so much optimism around it, too, because whenever you talk to someone who's working on robotics, they're bottlenecked by the amount of data they have. And from all these, you know, past three years of AI innovation, we've seen that, you know, whe- whenever there's an abundance of data, that type of models actually, like, improve a lot and you see it.
So, like, robotics, we expect something similar. Whenever they figure out this data problem, uh, those models are gonna get better as well. So that's why people are so optimistic, "Okay, maybe this solves the robotics data problem." And it's, yeah, boundless opportunities there.
And regarding, like, the, the, the example you mentioned about gravitational forces- ... like, I think this is still the same problem as, oh, LLMs can't do 9.9 plus, like n- uh, 9.0. Like, no, it, it ... Yes, it can.
It- like you just need to train it with more data. You need to have a better tokenizer. It, it is the reason whatever. It's just a matter of, like, data scale and, like, the underlying fundamental architectures. But, like, I, I don't think it's gonna change that much.
We just, like, we're all gonna put 1000x more data, 1000x more compute, and we'll get, like, the best physics simulators. Uh, and I think this, this should be possible with the existing signals coming from the data.
Just to double-click on video stuff as well.
Yeah.
Uh, you had a great slide in the, in, at AIE where you're like, "Currently 18% of Fal's revenue comes from video models," and it might be a m-
Business & Ads35:27
That was February. So- ... now, now it's probably over 80.
No, 50%.
50? Okay.
Yeah. It's like-
Over 50.
Yeah, yeah. 100%.
Wow. Okay.
I, I guess editing models brought some life into-
Yeah.
Yes
... the image as well. So, like, both of them grew. Uh, but yeah-
Yeah
... video grew faster for sure.
Vi- vi- video is, like, pretty, pretty, uh, significant, and the, one of the main drivers is open source models, where, you know, in February there was Hunyuan video. I think that was, like, pretty good. There was Mochi from Genmo, uh, but, like, the quality still wasn't there.
And one from Alibaba was, like, insanely good model, and they released a newer version of this at, like, a month ago-
Yeah
... I think, or, like, a couple of weeks ago. And now it's getting so, so popular and, like, we can run this model, like, for, like, 480p, like the draft mode version, we can run it, like, f- five seconds, under five seconds.
So people can have, like, instant feedback loop, and then when they wanna go to 720p, like, full resolution's just, like, 20 seconds. And we're, we're planning to bring it down to 10 seconds.
Yeah. That's amazing. Uh, and I wanna double-click on that. Uh, like, for, for a while I was kind of bearish on Alibaba because they kept releasing papers with very cherry-picked-
Yeah, yeah, yeah
... th- and, like, it was like a, it was like, "Okay, we're on GitHub," and then you go to GitHub, it's a README. Like
Yeah.
Yeah.
There's no
I mean, you, you can see something really changed. They've been-
What happened? Yeah
... releasing new image models-
Do, do you talk to them?
... new video models.
No.
Yeah.
No. We, we, we haven't talked to them, but, like, it seems like they, they actually like the ... And now there's, like, competing teams inside Alibaba. You know, one is a really good image model, but they released Qwen as a, as a competing Qwen's image model.
Like, you know, we, we, we all, we think, like, one image model is actually very, very good. If you'll do one with single frame instead of, like, 81 frames or whatever, you get a really good text image model out of it, and this is just because of the pure amount of data that you put from the videos.
So now there's like, you know, Alibaba has, like, two, two of the really good models, like, fro- from their, from their lab, and then there's, like, smaller labs in China that, like, you might not hear about but, like, StepFun.
They, they released, like, a image editing model. Uh, Hydream-
Yeah, yeah
... V1.0. Like, they're, like, there's, like, all these, like, small labs they're releasing. Because I don't think training these image or editing models are that expensive, and video models might be, like, slightly more expensive. My guess is, like, training these costs, like, a couple million dollars, which is not that much, especially, you know, like, you know, uh, it- they're- Probably backed by some sort of entity, you know, like, like o- other than Alibaba, you know, like, the, there's, there's, like, Stefan, whatever, they probably raise, like, a, a really good amount of money.
So training these models are-- will bring you a lot of attention, and it's more attention than you would get releasing a subpar LLM because, like, LLM space s- has so much more competition. So just training this for, like, a million dollars a video model and then releasing it, I think that brings you a lot of attention.
It is a hack. When you look at Hugging Face, let's look right now, I'm sure the top models are image models. Like, it's-- I-- Kohan image editing probably is, is-
Probably up there
... probably up there.
Yeah, number one.
Number one.
DeepSeek.
Yeah.
Hunyuan, Gamecraft.
Is number three.
Number, number four.
Number four.
Then you get Gemma 270 million.
ByteDance had, had some image stuff.
ByteDance hasn't open sourced, uh, but they have a really good team, Seed, th- that's their new lab. They're, they're working on SeaDream, SeaDance, OmniHuman, like, stuff like that. We, we have a good, a good, good partnership going with them to hopefully, you know, have their models hosted in US as well.
Yeah.
And the, the idea is, you know, like, the-- I think that the team that they, they were able to assemble is very good, and it's coming from, you know, their previous researchers, whatever. Like, the, the-- like, they, they were-- ByteDance was doing really good open source stuff on, like, SDXL Lightning.
They, they released, like, SDXL Lightning paper, Animative Lightning, so I'm, I'm pretty, pretty hopeful about them.
Yeah. Uh, so I'm just gonna, like, you know-- First of all, uh, I, I-- you know, hopefully they d- they don't reach out to you when they, when they launch. They just drop it and you, you have to rush.
Oh, we, we-
Okay
... at this point, like, people reach out to us because, like, we, we, we, uh, we, we are the market leader, so, like, they, they just like-
Yeah
... reach out to us for, like, m- the getting distribution.
Yeah. There, there's a Chinese platform that they always launch on first, uh, which I for- I forget the name of it, but, like, you have to immediately go reach out.
We, we, we all also get, like, day zero, uh-
Yeah
... launch with the majority of these models.
But, like, so, so basically I think, like, the question is always, like, you know, you're the ones making money. Stability did not make money from Stable Diffusion.
I, I-
You know?
... I, I think the, the, the, the thing that Black Forest Labs did was-
Yeah
... very, very interesting in this aspect. They released three different models, uh, and Apache 2 licensed extremely distilled model, which is good for-
Dev
... uh, Schnell. This is the Schnell version.
Schnell.
This is like the for four-step generations for, like, lower quality stuff. They released a dev model with a non-commercial license where their inference partners are, you know, like, you're paying a revenue share with, like-- And this is, like, very good way.
And then there is a pro version where you can, like, collaborate for hosting it in a-- Like, the revenue share is obviously different for that as well.
Yeah.
And, like, th- this is, I think, a very smart choice for labs whose whole premise is releasing models. But if you're a lab-- if, if, if you're a company that is doing a product in the side, you don't necessarily need to make money from the open source models.
You're doing it for getting researchers, get-- you know, hiring people, getting distribution, whatever. So it's-- it, it really depends on, like, the company's goals. Like, for, for Alibaba's case, like, they don't care if one model is, like, hosted in their API.
Right? Like, Alibaba makes a-- It doesn't, it doesn't touch the Alibaba's top line revenue, whatever one makes. So it's like, for them, it's like it's a no-brainer to release it and get attention and maybe, like, get some leads to their Alibaba Cloud offerings.
But, like, in, in, in general, for Black Forest Labs or companies like that, I think it's, like, a smart move to release, like, a distilled version as, like, fully open source and then like a less distilled or, like, the actual model as non-commercial.
Yeah.
And then p- partner with inference companies and stuff like that.
What's the distribution of usage? So, like, is eighty percent of your revenue, like, five models? Like-
Mm
... or are people really using, like, the, the long tail of all these open models outside of, like, the initial launch?
I think, like, there is some power law, but not as much as you would think. But-- And it keeps changing. That's the other part. Like, it's not, like, only a single model that's being used a lot. Like, month to month, it, it changes a lot.
Um, this summer has been crazy. There have been, like- ... like-
Yeah
... just countless amount of new video models, new image editing models. Like, the leader kept changing week over week even. Uh, but, like, if you, if you, like, take a step back and look which models are being used, people wanna use either the, the best, most expensive video model, or they, they wanna use, like, a cost-efficient, like, good but cheap enough video model.
So those two models are usually, like, used a lot. And whatever those models are, it changes week over week and yeah.
Yeah. Uh, like, the one good example is, like, Flux Context was released, you know, like, on late May. And Kohan image edit was released, like, two weeks ago, and now it's, like, topping out, uh, context dev. You know, it's, it's, it's in- insane how quickly these stuff transition just because it's-- there's a better quality model, and the-- that's the va- value prop.
Like, you don't have to set up the infrastructure to manage Flux Context dev. As soon as, like, Kohan image edit is available, you can just switch to that with Fal.
I mean, it seems to me that if some models are open and some models you have to pay a revenue share-
Yeah
... you ideally wanna move the people off the revenue share models into the open models, right? Like, what, what's that dynamic?
They're, they're priced-- Like, the, the pri-- Like, it's all priced in.
I'm, I, I'm also thinking, like, okay, we'll do whatever our customers are gonna be successful with. Uh, we are still, like, early enough that these small calculations-
Right. It doesn't-
... I don't think it matters. I rather people actually go to production and build products with it and be successful rather than, like, okay, twenty percent here, ten percent there. It doesn't matter.
I, I mean, you're doing a hundred million in revenue.
Cool.
I'll just ask a, a few more questions we had around, uh, just, like, the-- how people really use this. Um, okay, I'll ask the super obvious question.
Yeah.
How much is not safe for work?
Almost none.
Negligible, yeah.
You don't moderate everything, and moderation is optional, right? It's just-
Moderation is, is optional to a level where, like, illegal content is moderated.
Okay.
And we, we also track, like, the, the non-illegal content, NSFW moderation, and, like, we, we haven't seen-- like, we haven't seen more than one percent.
Like, the models themselves are actually not generating that type of content. Like-
So this, this is the, like-
Yeah
... some model providers especially, like, if you look at Black Forest Labs models-
Yeah
... the models are not for, uh-- Like, it's, it's incapable of generating because it's like-
It's not in the dataset
... it's in the-- Or, like, it's enabled in a way that it's, it's prevented. So-- And, like, ma-- like, we, we-- Majority of our customer base, like, if you look revenue-wise, it's, like, enterprises or, like, a bit more on, like, the higher level of stuff where some of them might be, like, user space mobile applications.
But, you know, like, but for the last six months, nine months, we've been transitioning more and more to enterprise where it's, like, less of a need for them.
Yeah.
Like, or less of a-
So, so what are those enterprises doing? You know, apart from like-
Yeah
... building a general purpose chatbot that can generate images. Like, I'm trying to... You know, I think like maybe Canva, you know-
Mm-hmm
... would, would be like a, a, a good use case, but my imagination is a bit limited beyond that.
Advertising seems to be absolutely, like, growing, and if you think about it, it fits very well, and let's talk about video advertising. So like, I keep repeating this, but, uh, some companies talk about, "Oh, we are gonna change Hollywood.
Uh, filmmaking is gonna be revolutionized." Like, I don't think it's that interesting. Like, how many movies do you watch a year? Like maybe 20, 25 movies. How many movies in the theater you watch? Three, four at most. So if, if there are like thousands of movies that a year, like people won't be able to watch all of these movies.
Like there's just-
It's, it's-
... not enough time
... you still, it's a max quality game.
Exactly. Exactly.
Yeah, yeah. Uh-
And with advertising, it's the exact opposite. The more content there is, the, the different like, you know, ways you can create ads, there is always economical value attached to it. So you can create unlimited number of ads, unlimited different versions of it, and more personalized it is, more, you know, economical value there is behind it.
So with ads, it fits really well to this type of technology because there is no limit what you can create.
As... I- I tell you a side comment the- about a ch- Silicon Valley trend I'm seeing which I cannot explain. Which is that all these YC startups and all these, uh, they're spending between 10 to $70,000 per launch video-
Yeah
... in the age of generative video. Like they're hiring, you know, actual creative directors, hiring a studio, hiring actors. Uh, I was in one of them. And like do you need all that when you have generative video?
Like I, I think clearly Roy started talking about generative video as well and-
I don't know if you guys know PJ Ace.
Yes.
Yes.
I think he's like the absolute killer for this stuff.
We talked about this.
Uh, he, uh, launched like a... Is it a Super Bowl ad or some- he, he-
Like a play- basketball playoff ad.
Yeah, yeah, yeah.
Yeah, yeah, yeah.
NBA Finals.
Yeah, NBA Finals.
He also did our like Series B announcement video, like we're, we're, like pretty close with him and like that it's, it's insane that like what he's able to come up with and how viral it goes over like, you know, these videos where you spend like hundreds of thousands of dollars, right?
You know, like it's just like you just need to create viral content and these generative media models are the best way to do it. And z- we are still at the infancy of this, right? Like obviously it might not be professional quality.
I'd still think like, you know, human-in-the-loop like mixed like, you know, content is, is, is, is the way for today. But in six months who, who would know? Like 12 months, I think like 80% of it's gonna be generated.
Like we, like we were wa- watching Super Bowl and we were like saying, "Oh, how much of this video is like AI generated? It looks like AI generated." Like it could be, right?
It could be.
You, you can't tell. So like I, I, I think at some point, you know, we're gonna have like 80, 90% AI generated- ... all ads.
Uh, it, it reminds me of, uh, I think, um, who's, who's that guy? Uh, Phofer from Replicate.
Yeah.
Obviously he is the best, uh, ins- inspiration for all these workflows. He like lay- overlaid some kind of NBA realistic, uh, sort of LoRA on top of game footage. So you could play like NBA 2K, but it looks like a real, uh, video.
I saw that. Yeah, yeah.
Yeah. I was like, "What the hell?"
Uh, it's pretty cool. Yeah, so, so maybe and that's the, the other part of my question and I wanted to get into ComfyUI, which is, um, how much LoRA serving is going on, right? How much custom-
A lot. A lot
... you know? Okay. Is it like majority?
That's one of the reasons-
Not majority, but like, like if it's like 30%, is it like the major- you know, it's like-
And everyone trains their own LoRAs or you pick it off of like a LoRA marketplace?
Both. Both. There is like a lot of-
That's why open source works very well with image and video models because you tap into this big LoRA ecosystem, everyone. Like it only-- Like I've never seen a closed source model ca- that can create a good LoRA ecosystem.
It just basically doesn't exist. Like maybe there is Midjourney SREFs, but I like I don't know if you can consider them LoRAs. Maybe it's like-
SREFs are just seeds, right?
Yeah.
Uh, conditioning, let's call it like another conditioning-
Okay
... like a prompt. Yeah.
Yeah. And then like only the open source models have these rich LoRA ecosystems and it's extremely, extremely popular.
It, it... But even for like the oldest models, it, it brings a new life, you know, like when you see these cool LoRAs, like you, like there's-- we have like still a lot of people using SDXL with their own LoRAs because they're happy with the quality.
It's fast enough.
Yeah.
It's cheap enough, you know, like that. It's, it's, it's, it's, it's amazing. And these models are not single shottable like the language models. You can like even the editing models, you know, GPT Image One or like Flux Context, whatever, Qwen Image.
If you put your face or like if you put like multiple people, whatever, you can't get the quality. Like it's gonna be like 90% there. But if you train this like if you train for like 1,000 steps with like 6 to 20 images, you're getting like 99% accuracy with like, like the, like we, we worked a lot on like fine-tuning the right hyperparameters, like writing like distributed trainers, uh, distributed optimizer stuff.
And with those, like people can train like their LoRAs under 30 seconds now on the platform, run an inference with them in the same, in the same job, and get like 99% accuracy for the same face, character consistency, which is one of the har- biggest challenges that maybe like more on the enterprise side they're facing and less on the consumer side.
You know, if you're creating ASL, you don't really care who it looks like.
Right.
But if, if, if you, if you are actually like doing a product ad, you want it to like exactly look like the product. Every single pixel on the product's, you know, banner, whatever, you want it to look like that.
So you train like a LaCroix LoRA with like 20 images and then like, you know, after that you're, you're, you have almost a perfect, uh, pixel perfect model.
All right. We have to train a LoRA for every guest.
No.
Then we can make thumbnails-
Yeah
... with the guest doing.
Yeah.
Yeah, yeah, yeah. No, I, I actually think that's a very good application because it's a nice way to in- inject brand, but not in a strict style.
And we are just entering like post-training on video models and what's that, that's gonna mean because we didn't have a good base video model that it made sense. But now we have like companies really investing into like post-training on 12.2 or Hunyuan and like creating lip sync models o- on top of it, creating like different, uh, video effects, came- camera angles.
Seems like there's a lot of possibilities with, with like creative, you know, data sets that people can do. I think in the next six months to a year, we are gonna have a lot of like companies that are just built on post-training of, of open source video models.
Wow.
Let's talk about pipelines. So we are Comfy anonymous-
Mm-hmm
... on the podcast. Um, Comfy UI is kinda like this community that, like, if you're into it, you love it. If you, if you don't know about it, you kinda underestimate in a way, but people create all kind of crazy, uh, workflows.
One, have you thought about doing pipelines? Obviously, you host all the models. There's kinda discussion-
We do have a pipeline product called File Workflows, and you can chain models together, but it's obviously less flexible than Comfy where, like, the, the-- you can only change, like, chain, chain different models' outputs and not the intermediary stuff.
In Comfy UI, you can, like, use, like, access the latents from, like, one model and then pass it to, like, a latent upscaler, whatever. In our case, it's, like, more limited, but we have a Workflows product, and we have a serverless Comfy product where people can bring their own Comfy UI workflow and run it as an API with just, like, posting the workflow and inputs.
And have the models be served by you.
Yes. So-
So i- is, is that a-
The-
... a bullish thing? Is that-
Yeah, like-
... gonna be commoditized by big models?
The, the thing that we saw is, as the models get better, the, like, Comfy UI was a much bigger thing or, like, you know, relatively much bigger thing in two years ago or, like, a year ago-
Yeah
... when the models were like... You know, one of the biggest Comfy UI use case was you were, like, generating an SD, SD 1.5 or SDXL image, and you were fixing six-fingers situation. You know, like, you were fixing, like, the resolution, you were, whatever, upscaling.
Now that the models are actually so, so good, the Comfy UI workflows are actually getting simpler for image side. For video, it's still very crazy. Like, if you look at some video workflows, it's like there's, like, 50 nodes, whatever, that you're, you're processing.
So I think this is still a matter of, like, how good the models are and how much extra stuff you need to do around it for majority of use cases. For artistic use cases, you still are doing, like, a lot of stuff, and that's, like, something we wanna support.
But, like, that-- we don't see that happening at super scale, you know, like super scale. You're, you're not gonna... Like, there's not companies that are spending, like, $10 million-plus on running this as an API, so that doesn't seem to be happening yet, uh, just because it's, like, a bit inefficient and the more existing...
Like, it's more reliable to use an existing model than patching together 50 different things because you don't know when it's gonna not work.
Yeah. But it feels like for things like ads-
Mm-hmm
... you wanna do, you know, one step, which is, like, maybe generate the backdrop. One step is, like, adding copy. Like, are you saying-
Oh, there's-
... the busy-- the models are so good?
Yeah, chaining of models happening for sure, but I think Comfy UI did very well is you can also, you know, play with the pieces of the model inside-
Yeah, yeah. You're basically saying that's all-
Yeah, yeah.
Exactly.
So, like, chaining of the models is like that's what the File Workflow product does. It basically calls many different APIs back to back or in parallel and then creates a result at the end, I think.
Yeah.
And it is very popular.
Yeah.
We, we have, like, enterprise adoption from it, like, from very big names.
Hiring & Future52:30
Amazing. Uh, cool. Uh, you know, I was just gonna go into the sort of broader topics. Uh, the, the first thing that comes to mind was request for startups. If you're not working on Fal, but you see a lot of things in the ecosystem, right?
Mm-hmm.
What's the most obvious thing that people should be working on?
More model companies. Go raise more money- ... and train models obviously.
That's obviously good for Fal.
And host it on Fal.
If you're not interested in training models, but, like, if other people are trained, that's amazing. Go raise more money. There's so much money.
Yes. Or, like, uh, scale AI for i-image and video models, like data, data collection, like more, more prepared data sets for, for video models, like effects, different camera angles.
Yeah.
Everyone seems to be reinventing the wheel when it comes to collecting that data.
Yeah.
I, I think it's a great opportunity for someone to come in and do this at scale. So it, it's really interesting 'cause, like, I think this is what Together AI did with Red Pajama is they, they actually built a data set for language models to help people create more open language models-
Open model. Yeah
... that they can serve. So at some point, like, actually it might make sense for you guys to do that, the-
Yeah
... the, the offer.
The, the image data is a bit more finicky situation. You-- like, the-- in terms of, like, copyright stuff, whatever, but it's an interesting area.
Uh, do it in Japan. Then you're protected.
I think, I think it requires focus. It require-- like, this needs to be like-
Yeah, yeah.
Like, one connected thing to what Gorkem said is image/video RL. That's an unknown unknown for us.
Ooh, say more.
I can't.
It, it, it's like, what does it look like? Can you, like, can you RL a video model to, uh, be a world model? You can, right? Like, if you consider it, like, it-it's essentially the world models are an RL, uh, video models where you condition it for, like, you know, moving around.
So what are the use cases for RL-ing image and video models? I don't know. But that's, like, uh, if I wasn't working at Fal, that would be something that, like, that might be fun to explore.
Yeah. And, and is this is specifically for editing? Because the RL-
No
... is for-- the reward is the edit or-
Uh, no. Anything.
Moving around like that.
Well, well, well, what, what, uh, that's, that's the thing. That's, that's what you should look for, right? Like, what, what is the reward function? Like, what is the interesting reward functions that you can apply on top of these base models?
I see. Interesting. Okay, got it. Uh, a-actually, I was really asking about, like, uh, if I were to build a Fal wrapper startup, like on top of Fal.
Oh, oh. I see. Yeah.
'Cause, uh, you, you guys are very low level, which is fantastic. Uh- You know, but I also wanna give our, our listeners some ideas, you know, if they're not gonna work at that level.
I think I'm gonna say it again, advertising.
Advertising.
Like, there's so much opportunity there, and everyone's, like, still trying to create these horizontal applications where, like, any creative can come in and do something, but, like, a lot more targeted to specific industries, a lot more targeted to, like, different kinds of, like, ad networks.
Uh, there's, there's a lot of potential there.
Good. Cool. And then requests for models. Uh, you-- obviously you want more open models. That's good for you. But, like, any, like, specialization in the models, like, um, you know, uh-
Yeah
... I think image, image editing was a huge unlock, which I didn't foresee until this year.
Yeah.
Right? Where the-- And where we're like, "Ob-obviously we're gonna use it-
No one saw it
... edit image-
Like, even OpenAI didn't like guess it was gonna go this big. It was like, it's insane, like, how, how popular it became, and then, like, everyone started, like, catching up.
Uh, I was part of a group at, at NeurIPS, uh, that we meet at NeurIPS every year, and, like, they were talking a-about this at the, at the last NeurIPS. And so, like, it, it like, it was-- it's in the air, but you have to be at the researcher level.
And, like, everyone moved to video, left image behind a little bit, so there was, like, a little vacuum of research on, on image. But luckily, people, people saw it, that it works very well, and then they went back to it.
It's so much cheaper to train image models. Like, right now- ... like, if you, if you look-- if you want to train a state image model, I don't think it's gonna cost more than a million dollars. It's extremely cheap.
It's like a matter of data engineering effort, cleaning. Look, it's-- I think it's a function of dataset. Like, image models are really, really, really affected by the dataset that you use.
Yeah.
I think one obvious thing that there's a gap in the market is, like, Veo 3 is very expensive-
Sound
... and, and the way-- the, the reason why people like it is conversation, right? If you can create maybe a smaller, cheaper video model that is less capable but can do conversation and sound very well, uh, I think there is definitely-
Yeah
... demand for it.
One open source, one that we saw was MultiTalk. It was a post-trained version of one. Uh, and it's, like, really, really good for conversations, but it lost the ability to generalize. You know, it's sort of like you-- it's only, like, talking faces at some point versus, like, Veo 3, it, it could generalize, and it could do scenes and whatever.
Yeah.
So I, I think that there, there needs to be, like, some middle ground between these two, between talking faces and, like, you know, uh, extremely generalized video models where it's, like, much cheaper to run. But at the same time, you know, uh, you, you get, like, this conversation because it's very memetic.
You know, people-- there's infinite amount of memes that you can post with this, infinite amount of ads that you can do with this. So it's like-
But you don't see a world in which you have a video model, and then you have a separate maybe audio-only model that can generate the audio for that video.
This is the workflows question, right?
Yeah.
Like, do you stitch together a whole bunch of things or-
Yeah
... do you do better than-
People did that before Veo 3, but what Veo 3 gets very well is, like, the timing. It, it-
Right.
Almost like-
Totally
... you ask for a joke, and you know the delivery and the timing and the laugh and, like, you know-
It, it even-
... waiting right before, like, the joke drops. Like, all of that is so perfectly timed. I don't think when you do it separately, you, you get it all-
It also matches, like, the so- like, the s- human accent sound to the face that is talking, right? It's like, it's, it's an unknown challenge for, like, other, other te-text-to-text-to-speech models. Like, you-- like, it feels very natural.
Is Veo 3 the best text-to-speech model?
It is also one of the best text-to-speech models.
What? Oh my God.
It is so good. Like, I don't think any model can do what it, what it is doing emotionally.
But, but I would say the counterargument is that we dub movies, so there's already-
The, the only-
... you know, obviously you can-
It is also the best lip-sync model. Like, Veo 3 has the best, most accurate lip sync because it's generating very natively, and I feel like you-- I, I can tell if a mo- Like, even, like-- So there's really good lip-sync models.
I think they're, like, ninety-five percent there, but Veo 3 is, like, hundred percent there. Like, ninety-nine percent.
To me, this is, like, the single most bearish thing about workflows, right? And, and ComfyUI and all this stuff, like, because just wait for a bigger-
As a model
... model. Like, it's just pure bitter lesson. Like-
Yeah, we love ComfyUI.
Uh, but, but i- I mean, obviously, like, when the technology doesn't exist yet, you, you have to stitch together things.
Yeah.
So, uh, request for engineers?
Yeah. I mean, what-- I, I'm sure you're hiring, right? You just raised a hundred and twenty-five million.
We are very selective on hiring. Like, we, we just recently crossed 40 people, but, like, for, like, the-
Wow
... for, like, three months ago, we were, like, you know, 20.
20, yeah.
So, like, for the last three months, we have been actually hi- a-accelerating, you know, best kernel engineers, best infrastructure engineers, best product engineers, best ML engineers. Like, best of-- If you're the best at what you do, just come join us.
I think it doesn't really matter what you do. Just, like, we're, we're just hiring the best talent right now.
Even on the go-to-market side, we are hiring-
Right
... account executives, customer success managers, uh, because we, we work with very large enterprises. We gotta grow that size of the-- that side of the company as well.
Yeah. On the engineering side specifically, how do you think about how many people you need? There's, like, this whole question of, like, lean AI. It's like, you know, coding agent-
Our, our performance team is, like, seven people. I think seven people, like, focusing on performance. Always, like, some of-- There's some overlap with our applied ML team, which is taking these models, productionizing them, exposing new capabilities, building fine tuners.
So it's like a-- And then helping customers adapt these models. So that team, I think we can scale to, like, double, triple the amount because, like, there's infinite amount of models and, like, you know, it's better because, like, we're, we're gonna have more customers with more proprietary models, so just, like, helping them optimize it.
That-- It's just, like, a really good function that we have.
That team scales very well because there's always, like, independent work that can be done. Oh, okay, so these two people are working on this new model, trying to optimize that, and it's completely independent from trying to optimize this, this other model.
So we've been hiring a lot for, for that applied ML team. Yeah.
Our infra team, you know, we-we're probably gonna keep it lean as well. Uh, l- keep it lean, uh, uh, in, in contrast to the applied ML team and the product team. Maybe, like, we wanna build more, uh, higher-level components where people can directly integrate to their applications because, like, that's even, like, now there's this, like-
Like, just SDK or-
SDKs, but think about with component. Like, if you-- Imagine you're an e-commerce website designer, and you're not really, like, the best, like, component designer. So here's an, uh, virtual try-on component that you can put to your app, stuff like that, more higher-level components.
Uh, and this is, like, also coming from the fact that, you know, vibe coding has been very, very insane. We a- we, we see, like, significant-- Like, revenue-wise, like, we're small, but, like, we see a significant amount of user adoption just coming from people who are vi- like, fr- just from looking at su- our support tickets.
Maybe they're, like, more-- they need more support, but, like, there's, like, a lot of people who are, like, you know, coding these applications without, like, that much expertise in, in, in the product building. So we wanna provide you-- provide them more guardrail experiences where they can integrate much easier without, like, messing with, like, a-all the other, like, lower-level components.
That's really nice care of, um, developer experience.
So crack low-level engineers-
Yeah, crack high-level
... that's probably fair. Well, yeah, high le-
Crack engineers in general
... yeah, crack engineer, crack go-to-market people. Crack whatever. Just your power.
Yeah. Uh, you know, I'm, I'm always trying to ref- refine the definition of crack, you know. Like, both of you, like, you lead the technical side of Fal. Uh, so, like, if, if someone has, like-- Like, what's a really hard technical problem that if someone has the solution, they should talk to you immediately, you know?
That-- Maybe that's the way to frame it.
Write a sparse attention kernel with FP8 on Blackwell and tell- ... you know, like, if, if you can do that, come join us. We'll do it for, like-- We, we already have, like, a good base.
Hired on the spot.
Hired on the spot, you know.
Yeah.
Like, stuff like that. There, there's like-- The-- I, I really like picking, like, all these, like-- Some of these applied ML people, like, we just pick them from Discords who are working on these sort of generative media, who are, like, already interested.
We really have a high culture bar too, where everyone in the team loves generative media. Like, they're obsessed with it. They would have done this if this wasn't their job. So it's, it's, it's like we, we have this great composition.
We are... Always it's not like a prerequisite, but it, it's just naturally happened where we hire these people from, like, discourse, Twitters, Hugging Face. Like, one, one of our, uh... One, one of our applied ML engineers is, like, had the number one top Hu- Hugging Face space with, like, you know, creative workflows, whatever.
So we, we hired a person who was training Loras on Fal, like, just because they were training and posting cool Loras, you know? Like, just do cool stuff and we'll find you, or, like, you can reach-
Yeah
... out to us.
That's the master builder-
Yeah
... is what I've been calling this, this person.
But why not make it more explicit? So if I go on your careers website, right? It's like applied ML engineer. It just kind of looks like any job description.
Ooh.
I, I feel like there's, like, this question of like-
It, it does. Uh, yeah.
That's why we have to do a podcast
Well, yeah, but no-
Yeah.
But I think it's not, it's not just about Fal. I think in general-
It is, it is more like if you know, you know, which I know is not the best way . You know, people know about Fal already, so it's like we, we haven't really cared that much, but you're absolutely right.
Like, we, we should make it more explicit.
I- if I look at, like, George Hotz, like on Tinygrad you have this balance if like, "Hey, we'll just... If you can solve this, you should probably work here." Like, do you see-
I'm adding the bounty-
Right?
... today.
It's like this seems like, "Hey, look, if you can write this kernel," it's like-
Yeah
... yeah, you'll just get hired.
It, it, it is-
So-
... also about, like, one thing that we saw even with, like... There's a lot of people who are just, like, wipe coding stuff, and reviewing those code, like, there's a limited amount of people who can review those, right?
Like, so, like, how, like, how can you tell it's, like, not a shitty kernel versus like a good kernel?
Well, but then you're-
You add
... spending the time interviewing too, right?
Yeah, yeah.
It's like-
So, like, we, we, we have, like, first line of defense with our recruiters, whatever. So like-
Yeah, yeah, yeah
... we only get like the, like... So there, there is, like, trade-offs, but I, I, I absolutely agree. Like, maybe we should have, like, a kernel bench, uh, version that you can upload your kernels-
Yes
... automatically evaluate the-
Right
... stability, performance, whatever, and then if you do it, you get our email unlocked, whatever. A special email for your- for you. But yeah, great ideas. Come join us. Build this.
Um, awesome, guys. Anything else? Parting thoughts?
Yeah. I love your rants.
This was great. Yeah.
I'm happy to rant.
Batuhan is the podcast star.
Yeah.
No, con- congrats on all your success. Um, I should also say, uh, it's, it's fun to do karaoke with you guys.
Yes.
Like you, you, you're-
Let's do it again.
... both extremely technical, but also, like, a fun crew that, like... And I think that's pretty hard to, and rare to, to see. So-
Thank you
... it's nice to see the good guys win.
Cool.
Awesome, guys.
Awesome.






