LALatent SpaceJan 2, 2024· 1:09:23

The AI-First Graphics Editor - with Suhail Doshi of Playground AI

Suhail Doshi, co-founder of Mixpanel and founder of Playground AI, argues that image generation is still in a GPT-2 moment and that training open-source foundation models from scratch is necessary to unlock real utility, exemplified by Playground v2's 2.5x preference over Stable Diffusion XL on a 1K prompt benchmark. He explains the pivot from Mighty (cloud-streamed browser) to AI, driven by the belief that shifting compute elsewhere aligns with AI's parallel computation. The episode covers Playground's unique UI—not just a prompt box but a full canvas with preview rendering, seed control, and style filters—and the difficulty of balancing safety with artistic expression, especially around NSFW content. Suhail discusses the under-investment in graphics AI compared to language, the open-source community's role, and his choice to release pre-trained weights for academic research. He also shares lessons from running GPU infrastructure (harder for training than inference) and advises founders to follow curiosity and build projects rather than relying on books.

  1. 0:00Introduction
  2. 9:00AI Awakening
  3. 17:58Training From Scratch
  4. 27:52Aesthetic Benchmarks
  5. 30:59Emergent Styles
  6. 43:18Safety and Limits
  7. 49:47AI-First UX
  8. 57:13LoRA Ecosystem
  9. 1:02:07GPU Engineering
  10. 1:04:44Lightning Round

Powered by PodHood

Transcript

Introduction0:00

Alessio0:01

Hey everyone. Welcome to the Latent Space Podcast. This is Alessio, partner and CTO in residence at Decibel Partners, and I'm joined by my co-host Swyx, founder of Smol.ai.

Swyx0:10

Hey, and today in the studio we have Suhail Doshi. Welcome.

Suhail Doshi0:13

Yeah, thanks. Thanks for having me.

Swyx0:14

Among many things, you're a CEO and co-founder of Mixpanel.

Suhail Doshi0:17

Yep.

Swyx0:17

Uh, and, uh, I think about three years ago you s- you left to start Might- Mighty.

Suhail Doshi0:22

Mm-hmm.

Swyx0:22

Um, and more recently, I think about a year ago, uh, it transitioned into Playground. Uh, and you've just announced your, your, your new round. Uh, I'd just like to start touch on Mixpanel a little bit 'cause it's obviously like one of the more, uh, sort of successful, uh, analytics companies.

Uh, we previously, uh, had Amplitude on, um, and I'm curious if like you had any sort of, uh, reflections on like just your... the, that overall, the interaction of like that am- that am- that amount of data, um, that people would want to use for AI.

Like, uh, I don't know if, like, there's, there's still a part of you that stays in touch with that world.

Suhail Doshi0:57

Yeah. I mean, it's, I mean, you know, the short version is, is that, um, maybe back in like 2015 or '16, I don't really remember exactly 'cause it was a while ago, we had a m- ML team at Mixpanel and, um, I think this is like when maybe deep learning or something like really just started getting kind of exciting and we were thinking that maybe we, you know, given that we had such vast amounts of data, perhaps we could predict things.

So we built, you know, two or three different features. I think we built a feature where we could predict whether users would churn from your product. Uh, we made a feature that could predict whether users would convert. We tried to...

We built a feature that could do anomaly detection, like if something occurred in your product that was just very surprising, maybe a spike in traffic in a particular region, could we tell you in advan- could we tell you that that happened?

'Cause it's really hard to like know everything that's going on with your data. Could we tell you something surprising about your data? And we tried all of these various features. Most of it boiled down to just like, you know, using, uh, logistic regression, and it never quite seemed very groundbreaking in the end.

And so I think, um, you know, we had a four or five person ML team and, um, yeah, I think we never expanded it from there. And I did all these fast.ai courses trying to learn about ML and that was the, that's the-

Swyx2:15

That was the first time you did fast.ai.

Suhail Doshi2:17

Yeah, that was the first time I did fast.ai.

Swyx2:18

Oh.

Suhail Doshi2:18

Yeah, I think I've done it now three times maybe.

Swyx2:21

Oh, okay. I didn't-

Suhail Doshi2:21

Yeah

Swyx2:21

... realize the third. Okay. Um-

Suhail Doshi2:23

No, no, I just mean reviewing it-

Swyx2:25

Right

Suhail Doshi2:25

... is maybe three times, but yeah.

Swyx2:26

Yeah, yeah. Yeah. I mean, uh, um, I think you, you mentioned prediction, but honestly, like it's, um, also just about the feedback, right? The, the quality of feedback from, uh, from, from users, I think it's, uh, it's useful for anyone building AI applications.

Suhail Doshi2:39

Yeah.

Swyx2:39

Yeah. Self-evident.

Suhail Doshi2:41

Yeah, I think, I think I haven't spent a lot of time thinking about Mixpanel 'cause it's been a long time, but yeah, I wonder, I wonder now, given everything that's happened, like upon, you know, sometimes I, sometimes I'm like, "Oh, I wonder what we could do now," and then I kind of like move on to whatever I'm working on.

Swyx2:54

Yeah. No more-

Suhail Doshi2:55

But things have changed significantly since, so, uh, yeah.

Swyx2:58

Yeah. Awesome. Um, and then maybe we'll touch on Mighty a little bit. Uh, Mighty was very, very bold. Uh, it was basically... Well, m- my framing of it was, uh, you will run our browsers for us because-

Suhail Doshi3:09

Yep

Swyx3:09

... um, everyone has too many tabs open. I have too many tabs open and it's slowing down your machines, like you can do it better for us, uh, i- in a centralized data center.

Suhail Doshi3:16

Yeah. We were first trying to make a b- a browser, uh, that we would stream from a data center to your computer at extremely low latency. Um, and I... But the, the real objective wasn't trying to make a browser or anything like that.

The real objective was to try to make a new kind of computer, and the thought was just that like, you know, we have these computers in front of us today and we upgrade them or they run out of RAM or they don't have enough RAM or not enough disk or, you know, there's some limitation with our computers.

Uh, per- perhaps like data locality is a problem. Um, could we... You know, why, why do I need to think about upgrading my computer ever? And so, you know, we just had to kind of observe that like, well, actually it seems like a lot of applications are just now in the browser.

You know, it's like how many real desktop applications do we use relative to the number of applications we use in the browser? So there's just this realization that actually like, you know, the browser was effectively becoming more or less our operating system over time, and so then that's why we kind, kind of decided to go, "Hmm, maybe we can stream the browser."

Unfortunately, the idea did not work for, for a couple d- different reasons, but, uh, yeah, but the objective was try to make a new, new r- true new computer.

Swyx4:21

Yeah. Very, very bold. Very bold.

Alessio4:23

Yeah. And, um, I was there at, at YC Demo Day when you first announced it.

Suhail Doshi4:26

Oh, okay.

Alessio4:27

At, uh... It was, I think the last or one of the last in-person ones, so like the Pier 34 in-

Suhail Doshi4:31

Yes

Alessio4:32

... in Mission Bay.

Suhail Doshi4:32

Yeah.

Alessio4:33

Um-

Suhail Doshi4:33

Yeah, before COVID

Alessio4:35

... how, how do you think about that now when everybody wants to put some of these models in people's machines and some of them want to stream them in? Do you think there's maybe another wave of the same problem?

Before it was like browser apps too slow, now it's like model's too slow to run on device.

Suhail Doshi4:49

Yeah. I think, you know, we, you, we obviously pivoted away from Mighty, but a lot of what I somewhat believe- believed at Mighty is like still somewhat very true, what maybe why I'm so excited about AI and what's happening.

A lot of what Mighty was about was like moving compute somewhere else.

Swyx5:06

Mm.

Suhail Doshi5:06

Right? Right now, applications, they get limited quantities of memory, disk, uh, networking, whatever your home network has, uh, et cetera. You know, what, what if these applications could somehow, if we could shift compute and then these applications could have vastly more compute than they do today.

Uh, right now it's just like client backend services, but, um, you know, what if we could change the shape of how, how applications, uh, could interact with things? And it's changed my thinking. It... In some ways, AI is like a bit of a, a continuation of my belief that like perhaps we can really shift compute somewhere else.

One of the problems with, with Mighty was that, um, JavaScript is single-threaded in the browser, and what we learned, you know, the reason, reason why we kind of abandoned Mighty was because I didn't believe we could make a new kind of computer.

We could have made some kind of enterprise business, probably could made, could have made maybe a lot of money, but it wasn't going to be what I hoped it was going to be. And so once, once I realized that most Of a web app, it is just going to be single-threaded JavaScript, then the only thing you could do largely, uh, withstanding changing JavaScript, which, uh, is a fool's errand most likely, uh, is make a better pro- CPU.

Right? And there's like three CPU manufacturers, two of which sell the, you know, c- you know, big ones, you know, AMD, Intel, and then of course like Apple made the M1. And it's not like single-threaded CPQ- CPU core performance.

Single core performance was like very- increasing very fast. It's plateauing rapidly. And even these different like companies were not doing as good of a job, you know, sort of with the continuation of Moore's law. But what happened in AI was that you got like, like if you think of the AI model as like a computer program, like just like a compiled computer program, it is literally built and designed to do massive parallel computations.

And so, uh, if you could take like the universal approximation theorem to its like kind of logical, complete point, you know, you're like, "Wow, I can get- make computation happen really rapidly and parallel somewhere else."

Swyx7:06

Mm.

Suhail Doshi7:07

Um, you know, so you end up with these like really amazing models that can like do anything. It just turned out like perhaps, perhaps the new kind of computer, uh, would just simply be shifted, um, you know, into these like really amazing AI models in reality.

Swyx7:23

Yeah. Like, I think, uh, Andrej Karpathy has always been, has been making a lot of analogies with the LLM OS.

Suhail Doshi7:29

Yeah, I saw his, yeah, I saw his video and I, I watched that, you know, maybe two weeks ago or something like that, and I was like, "Oh man, this-

Swyx7:35

Yeah.

Suhail Doshi7:35

... I very much resonate with this like idea."

Swyx7:37

Why didn't I see this three years ago?

Suhail Doshi7:39

Yeah. I, I think, I think there still will be, you know, local models and then there'll be these very large models that have to be run in data centers. Um, yeah, I think it just depends on kind of like the right tool for the job like any, like any engineer, uh, would probably care about.

But I think that, uh, you know, by and large, like if, if the models continue to kind of keep getting bigger, you're just gonna, it's... You're always going to be wondering whether you, whether you should use the big thing or the small, you know, the tiny little model.

Um, and it, it might just depend on like, you know, do you need 30 FPS or 60 FPS? Um, maybe that would be hard, uh-

Swyx8:10

Yeah

Suhail Doshi8:10

... to do, you know, over, over, over a network.

Swyx8:12

Yeah. You, you tackled a much, uh, harder problem latency-wise, um, you know, than, than the AI models actually require.

Suhail Doshi8:19

Yeah.

Swyx8:19

Uh, so

Suhail Doshi8:20

Yeah, you can do quite well. You can do quite well. Uh, you know, we definitely did, um, 30 FPS video streaming.

Swyx8:26

Yeah.

Suhail Doshi8:26

Did, did very crazy things to make that work. So I, I'm actually quite bullish on the kinds of things you can do with networking.

Swyx8:33

Yeah. All right. Maybe someday you'll, uh, come back to that at some point. Um, but so for those that, for those that don't know, you're very transparent on Twitter. Uh, very good to follow you just to, just to learn your insights, and you actually published a postmortem on Mighty that people can read up-

Suhail Doshi8:47

Yep

Swyx8:47

... if they're willing to. Um, and so there was a bit of an overlap. Uh, you started exploring, um, the AI stuff in June 2022.

Suhail Doshi8:56

Mm-hmm.

Swyx8:57

Which is when you started saying like, "I'm taking fast AI again."

Suhail Doshi8:59

Mm-hmm.

Swyx8:59

Um, maybe, was there more context around that? Um...

AI Awakening9:00

Suhail Doshi9:03

Yeah. Uh, I think, I think I was kind of like waiting for the team at Mighty to finish up- ... you know, of something, and I was like, "Okay, well, what can I do? Uh, I guess I will, uh, make some kind of like address bar predictor in the browser."

So we had, we know we had forked Chrome and Chromium and, um, I was like, you know, one thing that's kind of lame is that like this browser should be like a lot better at predicting what I might do, where I might wanna go.

You know, it, it's, it struck me as really odd that, you know, Chrome had very little AI actually or ML inside this browser. And then for a company like Google, you'd think there's a lot. But it's actually like, it's actually just like the code is actually just very, you know, it's just a bunch of if/then statements is more or less the address bar.

So it seemed like a pretty big opportunity, and that's also where a lot of people interact with the browser. So, you know, long story short, I was like, "Hmm, I wonder what I could build here." So I started to, yeah, take some AI courses and try to get, take some- review the material again and get back to figuring it out.

But I think that was somewhat serendipitous because, um, right around April was, I think, a very big watershed moment in AI 'cause that's when DALL-E 2 came out, and I think that was the first like truly big viral moment, uh, for generative AI.

Swyx10:18

Because of the avocado chair.

Suhail Doshi10:20

Because of the avocado chair and, uh, yeah, exactly. Yeah. I mean, it was just so novel.

Swyx10:26

It wasn't as big, it wasn't as big for me as Stable Diffusion. Like-

Suhail Doshi10:28

Really?

Swyx10:29

Yeah. I don't know. DALL-E was like, all right, that's, that's cool. I don't know.

Suhail Doshi10:33

Yeah.

Swyx10:33

I mean, they, they had some flashy videos, but like I never really... Like, it, it didn't really register to me as like big-

Suhail Doshi10:38

Well, just, just that moment of images was just such a viral, novel moment. I think it just blew people's mind.

Swyx10:44

Yeah.

Suhail Doshi10:44

Um, and-

Swyx10:46

Yeah, I mean, that's the first time I like encountered Sam Altman because like they had this like DALL-E 2 hackathon, and open- they opened up the OpenAI office for developers to walk in, like back when, you know, it wasn't as, uh, I guess, uh, much of a security issue as it is-

Suhail Doshi10:59

I see

Swyx10:59

... today.

Suhail Doshi11:00

Yeah.

Swyx11:00

Maybe take us through like the, the journey to, to decide to pivot into, into this. And but, and, and also like choosing images. Obviously, you, you, you were inspired by DALL-E.

Suhail Doshi11:08

Yeah.

Swyx11:09

But there could be any number of, um, AI co- companies and businesses that you could start, like and why this one, right?

Suhail Doshi11:17

Yeah. Um, well-

Swyx11:18

So there must be an idea made from June to September.

Suhail Doshi11:21

Yeah, yeah. Yeah, there, there definitely was. So I think at that time, Mighty, we, we were, Mighty and, and OpenAI was, you know, not quite as popular as it is all of a sudden now these days. But back then, I think they were, they were like more than hap- They, it, they had a lot more bandwidth to like kind of help anybody.

And so, you know, we had, we had been talking with, um, the team there around like trying to see if we could do like really fast low latency, uh, address bar prediction with like GPT-3 and 3, 3.5, and that kind of thing.

And so, um, you know, we were sort of figuring out how, how could we make that low latency. Um, I think that just being able to talk to them and kind of being involved gave me a bird's eye view into a bunch of things that started to happen.

Um- You know, obviously first was the DALL-E team- DALL-E 2 moment, but then Stable Diffusion came out, and that, i- that was a big moment for me as well. And I remember just kind of like sitting up one night thinking...

I was like, "You know what, what are the kinds of companies one could build? Like, what matters right now?" I- one thing that I observed is that I find a lot of great-- I find a lot of inspiration when I'm working in a field in something, and then I can identify a bunch of problems.

Like for Mixpanel, I was an intern at a company, and I just noticed that they were doing all this data analysis, and so I thought, "Hmm, I wonder if I could make a product, and then maybe they would use it."

And in this case, you know, the same thing kind of occurred. It was like, okay, there are a bunch of like infrastructure companies that are, you know, uh, doing y- g- y- they put a model up, and then you can use their API, like Replicate is a really good example of that.

Um, there are a bunch of companies that are like helping you with training, model optimization, um, Mosaic at the time, uh, I mean, and probably still, you know, was doing stuff like that. So I just started listing out like every category of everything, of every company that was doing something interesting.

Obviously, Weights & Biases. Um, I was like, "Oh man, Weights & Biases is like this great company. If, uh, do I wanna compete with that company? I might be really good at competing with that company because of Mixpanel, because it's so much of like analysis."

Um, but I was like, "No, I don't wanna do anything related to that. That would-- I think that would be too boring now at this point." Um, but, uh, um, so I started to list out all these ideas, and one thing I observed was that at, at OpenAI, they had like a playground for GPT-3.

Swyx13:35

Mm-hmm.

Suhail Doshi13:35

Right? And all it was is just like a text box more or less, and then there were some settings on the right, like temperature and whatever.

Swyx13:41

Top K, top N.

Suhail Doshi13:42

Yeah, top K. Uh, you know, what's your end stop sequence. I mean, that was like their product before ChatGPT. You know, really difficult to use but fun if you're like an engineer. And I just noticed that their product kind of was evolving a little bit, where the interface kind of was getting more and more, a little bit more complex.

They had like a way where you could like generate something in the middle of a sentence and all those-

Swyx14:03

Mm

Suhail Doshi14:03

... kinds of things. And I just thought to myself, I was like, you know, there's not-- everything is just like this text box, and you generate something, and that's about it. And Stable Diffusion had kinda come out, and it was all like hugging face and code.

Nobody was really building any UI. And so I had this kind of thing where I wrote prompt dash like question mark in my notes, and I didn't know what was like the product for that at the time.

Swyx14:25

Mm-hmm.

Suhail Doshi14:25

I mean, it seems kind of trite now. Um, but yeah, I just like wrote "Prompt, what's the thing for that?"

Swyx14:30

Manager prompt.

Suhail Doshi14:31

Prompt manager.

Swyx14:32

Tool.

Suhail Doshi14:32

Do you organize them? Like, do you like have a UI that can like-

Swyx14:36

Library

Suhail Doshi14:37

... play with them? Yeah, like a library. What, what would you make?

Swyx14:40

Yeah.

Suhail Doshi14:40

Uh, and so then of course then you thought about what would the modalities be given that, how would you build a UI for each kind of modality. Uh, and so there are a couple people working on some pretty cool things.

Um, and, uh, and I, and I ch- and I basically chose graphics because it seemed like the most obvious place where you could build a really powerful, complex UI that's not just only typing in a box, that y- y- it would very much evolve beyond that.

Like what would be the best thing for something that's visual? Probably something visual. Um, so yeah, I, I think that just that progression kind of happened, and it just seemed like there was a lot of effort going into language but not a lot of effort going into graphics.

And then the l- and maybe the very last thing was I, I think, um, I was talking to Aditya Ramesh, who was the co-f- co-creator of DALL-E 2 and Sam, and I just kind of w- went to these guys, and I was just like, "Hey, are you gonna make like a UI for this thing?

Like a true UI. Are you gonna go for this? Are you gonna make a product?"

Swyx15:42

For DALL-E? Yeah.

Suhail Doshi15:43

For DALL-E, yeah.

Swyx15:44

Yeah.

Suhail Doshi15:44

Uh, are you gonna do anything here? Um, 'cause if you're not gonna do it... If you are gonna do it, just let me know, and I will stop, and I will go-

Swyx15:51

Yeah

Suhail Doshi15:51

... do something else.

Swyx15:51

Yeah.

Suhail Doshi15:52

But if you're not gonna do anything, I'll just do it.

Swyx15:54

Mm-hmm.

Suhail Doshi15:54

Um, and so we had a couple conversations around what, what, what, what that would look like. Um, and then I think ultimately they decided that they were gonna focus on language primarily. Um, and, uh, yeah, I just felt like it was gonna be very under-invested in.

Swyx16:08

Yes, there, there is, there's that sort of, um, under-investment from OpenAI-

Suhail Doshi16:12

Mm-hmm

Swyx16:12

... which I, I, I can see that.

Suhail Doshi16:14

Mm-hmm.

Swyx16:14

Um, but also it's a different type of customer-

Suhail Doshi16:18

Mm-hmm

Swyx16:18

... than you're used to. Uh, presumably, you know, and Mix- Mixpanel are very sel- very good at selling to B2B-

Suhail Doshi16:23

Right

Swyx16:23

... developers. Uh, with Figure, and you're not.

Suhail Doshi16:26

Yeah.

Swyx16:27

Was, was that, was that not a concern?

Suhail Doshi16:29

Well, not, not so much because I think that, um, you know, right now I would say graphics is in this very nascent phase. Like most of the customers are just like hobbyists, right?

Swyx16:37

Yeah.

Suhail Doshi16:38

Like they-- It's a little bit of like a novel toy-

Swyx16:40

Yeah

Suhail Doshi16:40

... as opposed to being this like very high utility thing. But I think ultimately if you believe that you could make it very high utility, then probably the next customers will end up being B2B. It'll, it'll probably not be like consumer.

Like there are-- There will certainly be a variation of this idea that's in consumer. If you- if your quest is to kind of make like a, a super, uh, something that surpasses, um, human ability for graphics, like ultimately it will end up being used for business.

Swyx17:06

Yeah.

Suhail Doshi17:07

So I, I think it's maybe more of a progression. In fact, for me, it's maybe more like Mixpanel started out as SMB, and then very much like ended up starting to grow up towards enterprise. So for, for me, it's a little...

It's-- I think it will be a very similar progression.

Swyx17:18

Yeah. Yeah.

Suhail Doshi17:19

Yeah. But yeah, I mean, the, the reason why I was excited about it is 'cause it was a creative tool. I make music, and, uh, it's AI. It's like something that I could-- I know I could stay up till three o'clock in the morning doing.

Those are kind of like very simple bars for me.

Swyx17:33

Yeah.

Suhail Doshi17:33

Yeah.

Swyx17:34

It's good decision criteria.

Alessio17:36

Um, so you mentioned DALL-E, Stable Diffusion. You just had Playground v2-

Suhail Doshi17:41

Yep

Alessio17:41

... come out two days ago?

Suhail Doshi17:42

Yeah, two days ago. Yeah.

Alessio17:43

Two days ago. So this is a model you trained completely from scratch, so it's not a cheap fine-tune on, on something. You open source everything, including the weights. Um, why did you decide to do it? I know you supported Stable Diffusion XL in Playground before, right?

Training From Scratch17:58

Suhail Doshi17:58

Yep.

Alessio17:59

Um, yeah. What, what made you want to come up with v2 and maybe some of the interesting, you know, technical research work you've done?

Suhail Doshi18:07

Yeah, so I think, I think that we, we continue to feel like graphics and these foundation models for, uh, anything really related to pixels, but also definitely images continues to be very under-invested. It feels a little like graphics is in like this GPT-2 moment, right?

Like even GPT-3, even when GPT-3 came out, it was exciting, but it was like, what are you gonna use this for? You know, yeah, we'll do some text classification and some semantic analysis, and maybe it'll sometimes like make a summary of something and it'll hallucinate.

But no one really had like a very significant like business application for GPT-3. Um, and in images we're kind of stuck in the same place. We're kind of like, okay, I write this thing in a box and I get some cool piece of artwork and the hands are kind of messed up and sometimes the eyes are a little weird.

Uh, may- maybe I'll use it for a blog post, you know, that kind of thing. The utility feels so limited and so, you know, and then we... You sort of look at Stable Diffusion and, and we definitely use that model in our product and our users like it and use it and love it and, and enjoy it, but it hasn't gone nearly far enough.

So we, we were kind of faced with the choice of, you know, do we wait for progress to occur or do we make that progress happen? Uh, so yeah, we, we kind of embarked on a plan to just decide to go train these things from scratch, and I think the community has given us so much.

The, the community for Stable Diffusion, I think is one of the most vibrant communities on the internet. It's like amazing. It feels like if the-- I hope this is what like Homebrew Club felt like when computers like showed up because it's like amazing what that community will do and it moves so fast.

I've never seen anything in my life where so far, and heard other people's stories around this, where a research- an academic research paper comes out and then like two days later someone has sample code for it, and then two days later there's a model, and then two days later it's like in nine products.

Alessio19:58

Yeah.

Suhail Doshi19:58

You know, they're all competing with each other.

Alessio20:00

Yeah, yeah.

Suhail Doshi20:00

It's incredible to see like math symbols on a academic paper go to-

Alessio20:04

Yeah

Suhail Doshi20:04

... features, well-designed features in a product. So, um, I think the community has done so much. So I think we wanted to give back to, to the community kind of on our way. We knew, we knew it wasn't gonna be...

It-- We knew it was not ever gonna be-- Certainly we would train a better model than, than what we, what we gave out on Tuesday, but we definitely felt like, uh, there needs to be some kind of progress in these open source models.

The last kind of milestone was in July when Stable Diffusion XL came out, but there hasn't been anything really since, right?

Alessio20:35

And there's, uh, XL Turbo now.

Suhail Doshi20:36

Well, XL Turbo is like this distilled model, right?

Alessio20:39

Yeah.

Suhail Doshi20:39

So it's like lower quality but fast, so you have to decide, you know, what your trade-off is there.

Alessio20:44

Uh, and, and it's also a consistency model?

Suhail Doshi20:46

It's not-

Alessio20:47

I'm not sure.

Suhail Doshi20:47

I don't think it's a consistency model.

Alessio20:48

Sure. Okay.

Suhail Doshi20:48

It's like it's they, they did like a different thing.

Alessio20:51

Yeah.

Suhail Doshi20:51

Yeah. I think it's like... I don't, I, I don't wanna get quoted for this, but it's like some- something called ad, like adversarial something or another.

Alessio20:56

That's exactly right. Yeah.

Suhail Doshi20:58

Um-

Alessio20:58

Yeah

Suhail Doshi20:59

... yeah, I think it's, uh, it's-- I've read something about the, maybe it's like closer to GANs or something, but I didn't really read the p- the full paper. But, but yeah, there hasn't been quite enough progress in terms of, you know, there's no multitask image model.

You know, the closest thing would be something called like EmuEdit, but there's no model for that.

Alessio21:14

Mm-hmm.

Suhail Doshi21:14

It's just a paper that's within Meta. So, um, we did that and we also gave out, uh, pre-trained weights, which is very rare. Um, usually you just get the aligned model and then you have to like see if you can do anything with it.

We actually gave out, um... There's like a 256 p- uh, pixel pre-trained stage and a 512, and we did that for academic research 'cause there's a whole bunch of... You know, we come across people all the time in academia and they have like, they have access to like one A100 or eight at best.

Uh, and so if we can give them kind of like a 512 pre-trained model, it might m- our, our hope is that there'll be interesting novel research that occurs from that. Um-

Alessio21:52

What research do you want to happen?

Suhail Doshi21:54

I would love to see more research around, uh, you know, things that users care about tend to be things like character consistency.

Alessio22:01

Uh, between frames?

Suhail Doshi22:02

Um-

Alessio22:02

For video

Suhail Doshi22:02

... more like if you had like a face.

Alessio22:04

Yeah.

Suhail Doshi22:04

Yeah, yeah. Basically between frames, but more just like, you know, you have your face and it's in, you know, one image and then you want it to be like in another.

Alessio22:11

And-

Suhail Doshi22:11

And users are very particular and sensitive to faces changing 'cause we know, we know what, you know-

Alessio22:17

The faces

Suhail Doshi22:17

... we're, we're trained on faces-

Alessio22:18

Yeah.

Suhail Doshi22:18

... as humans. Um, and you know, that's something, um, I don't- I'm not seeing a lot of innovation, enough innovation around multitask editing. You know, there are two things like Instruct Pix2Pix and then the EmuEdit paper that are maybe very interesting, um, but, uh, we, we certainly are not pushing the fold on that in that regard.

Um, yeah, yeah just all kinds of things, uh, like around that. Rotation, um, you know, being able to keep coherence across images, style transfer is still very limited. Um, just even reasoning around images, you know, what's going on in an image-

Alessio22:53

Mm-hmm

Suhail Doshi22:54

... that kind of thing. Uh, things are still very, very, uh, underpowered, very nascent, so therefore the, the l- utility is very, very limited.

Alessio23:01

Mm-hmm. On the 1K prompt benchmark, you are 2.5x prefer-

Suhail Doshi23:05

Yep

Alessio23:05

... to Stable Diffusion XL. How do you get there? Is it better images in the training corpus? Is it... Yeah, can you, can you maybe talk through the improvements in the model?

Suhail Doshi23:16

I think we're still very early on in the recipe, um, but I think it's a lot of like little things and, you know, every now and then there are some big important things. Like certainly, uh, your data quality is really, really important, so we spend a lot of time, uh, thinking about that.

Um, but it-- I would say it's a lot of, lot of things that you kind of clean up along the way as you train your model, everything from captions to the data that you align with, um, after pre-train, to how you're picking your data sets, um, how you filter your data sets.

Uh, there's a lo- I feel like there's a lot of work in AI that's like doesn't really feel like AI, it just really feels like-

Alessio23:54

Mm

Suhail Doshi23:54

... just data set filtering and systems engineering and just like, you know, and the recipe is all there, but it's like a lot of extra work to do that. Um, so I think these, these models, I think, I think whatever version, I think we, we plan to do a Playground v-, uh, 2.1 maybe either by the end of the year or early next year, and we're just like watching what the community does with the model.

And then we're just gonna take a, a lot of the things that they're unhappy about and just like fix them. Um, you know, so for example, like maybe the eyes of people in an image don't feel right. They feel like they're a little misshapen or they're kind of blurry feeling.

That's something that we already know we wanna fix, so I think in that case it's gonna be about data quality. Um, or maybe we wanna improve the kind of the dynamic range of color. You know, we wanna make sure that that's like got a good range in any image, so what technique can we use there?

There's different things like offset noise, pyramid noise, uh, terminal zero SNR. Like there are all these various interesting things that you can do. So I think it's like a lot of just like tricks. Some are tricks, some are data, and some is just like cleaning.

Yeah.

Swyx24:57

If specifically for faces, it's very common to use a pipeline rather than just train, train the base model more.

Suhail Doshi25:05

Mm-hmm.

Swyx25:06

Uh, do you have a strong belief either way on like, oh, they should be separated out to different stages for like improving the eyes, improving the face or-

Suhail Doshi25:13

Mm-hmm

Swyx25:13

... even the hands or whatever, or do you think like it can all be done in one model?

Suhail Doshi25:18

I think we will make a unified model.

Swyx25:19

Okay.

Suhail Doshi25:20

Yeah, I think it will, I think we'll certainly in the end ultimately make a unified model. Um, you know, i- there, there, there's not enough research about this. There's pro- maybe there is something out there that we haven't read.

There are some bottlenecks, like for example, in the VAE. Um, like the VAEs are ultimately like compressing these things, and so you don't know, and then you might have like a big informational, information bottleneck. Or so maybe you would use a pixel-based model perhaps.

Um, you know, there's a lot of belief, I think we've talked from, to people, everyone from like Rombach to various people. Uh, Rombach trained Stable Diffusion. You know, th- there's, I think there's like a big question around the architecture of these things.

It's still kind of unknown, right? Like we've got transformers and we've got g- like a GPT architecture model, but then there's this like weird thing that's also seemingly working with diffusion. And so we, you know, are we gonna use vision transformers?

Are we gonna move to pixel-based models? Is there a different kind of architecture? We don't really-- I don't think there have been enough experiments in this area.

Swyx26:19

Still? Oh my God.

Suhail Doshi26:21

Yeah.

Swyx26:22

That's surprising.

Suhail Doshi26:23

Yeah. I think it's very computationally expensive to do a pipeline model where you're like fixing the eyes and you're fixing the mouth and you're fixing the hands.

Swyx26:29

That's what everyone does, as far as I understand.

Suhail Doshi26:31

Well, I, I'm not sure, I'm not exactly sure what you mean, but if you mean like you get an image and then you will like make another model specifically to fix a face, yeah, I don- I think that's very computationally-- that's fairly computationally expensive, and I think it's like not probably not the right thi- right way.

Swyx26:45

Yeah.

Suhail Doshi26:46

Yeah, and it, it doesn't generalize very well.

Swyx26:48

It doesn't.

Suhail Doshi26:48

Now you have to pick all these different things.

Swyx26:49

It's, yeah, you're just kind of glomming things on together.

Suhail Doshi26:51

Yeah.

Swyx26:51

But like when I look at AI artists, like that's what they do, so.

Suhail Doshi26:54

Ah, yeah, yeah, yeah. They'll do things like, you know, um, I think a lot of ARs will do, you know, control net tiling to do kind of generative upscaling-

Swyx27:02

Yeah. Yeah

Suhail Doshi27:02

... of all these different pieces of the image. Yeah, I mean, to me these are all just like, they're all hacks. Yeah, ultimately in the end. I mean, it just, to me it's like let's go back to where we were just three years, four years ago with where deep learning was at and where language was at.

You know, it's the same thing. It's like we were like, "Okay, well, I'll just train these very narrow models to try to do these things and kind of ensemble them or pipeline them to try to get to a best-in-class result" and, uh, and, and here we are with like where the, the models are gigantic and like very capable of solving huge amounts of tasks, uh, when given like lots of great data.

Swyx27:36

Yeah.

Suhail Doshi27:36

So yeah.

Swyx27:37

Makes sense.

Alessio27:38

Um, you also released a new benchmark called MJHQ 30K for automatic evaluation of a model's aesthetic quality.

Suhail Doshi27:46

Mm-hmm.

Alessio27:46

Um, I have one question. Um, the data set that you used for the benchmark is from Midjourney.

Aesthetic Benchmarks27:52

Suhail Doshi27:52

Yes.

Alessio27:52

You have-

Suhail Doshi27:53

Yeah

Alessio27:53

... ten categories. Um, how do you think about the Playground model, Midjourney?

Suhail Doshi27:59

You know, there are a lot of people-- a lot of people in research like to come up with, uh, they like to compare themselves to something they know they can, uh, beat.

Alessio28:06

Mm.

Suhail Doshi28:06

Right? But, um, a- and maybe this is the best reason why, uh, it's, it can be helpful to not be a researcher also sometimes. Like I'm not, I'm not like trained as a researcher. I don't have a PhD in anything AI related, for example.

Um, but I, I think if you care about products and you care about your users, then the, the most important thing that you wanna figure out is like we, you know, we have-- everyone has to acknowledge that Midjourney is very good.

You know? They're, they are the best at this thing. We would, I would happily, I'm happy to admit that. I have no, no problem admitting that. Uh, it's just, it's just easy. It's very visual, um, to tell. So, you know, I think it's incumbent on us to try to compare ourselves to the thing that's best, even if we lose, even if we're not the best, right?

And, um, you know, at some point, if we are able to surpass Midjourney, then we, you know, we only have ourselves to compare ourselves to. But on first blush, you know, I think it's worth comparing yourself to maybe the best thing and try to find like a really fair way of, um, of doing that.

So I think, I think more people should try to do that. I definitely don't think you should be kind of comparing yourself on like some Google model or some old SD, you know, Stable Diffusion model and be like, "Look, we beat, you know, Stable Diffusion one point five."

I think, I think users uh, you know, ultimately want care, you know, how close are you getting to the thing that like I, I also, you mostly, you- people mostly agree with. So we put out that benchmark not because, and, and for no other reason to say like this seems like a worthy thing for us to at least try, you know, for people to try to get to.

Uh, and then if we surpass it, great, we'll come up with another one.

Alessio29:40

Yeah. No, that's awesome. And, um, you killed Stable Diffusion XL and everything. Um, in the benchmark chart, it says Playground V2 1024 pixel dash aesthetic.

Suhail Doshi29:52

Yes.

Alessio29:52

Do you have, um, kind of like, yeah, style fine tunes or like what's the dash aesthetic for?

Suhail Doshi29:57

Yeah. We, we debated this. Maybe we named it wrong or something, but we were like how do we help people realize, uh, you know, the model that's, that's aligned versus the models that weren't? So because, because we gave out pre-trained models, we didn't want people to like use those.

Um, so we-- that's why they're called base. And then the aesthetic model, yeah, we wanted people to pick up the thing that we thought would be like the thing that makes things pretty. Um, who wouldn't want the thing that's aesthetic?

Alessio30:23

Mm-hmm.

Suhail Doshi30:23

But if, if there's a better name Yeah, we're, we're-- we definitely are open to feedback

Swyx30:27

No, no, that's cool. Um, I was using the product. You also have the style filter-

Suhail Doshi30:31

Uh-huh

Swyx30:32

... and you have all these different style.

Suhail Doshi30:33

Yeah.

Swyx30:33

And it seems like the styles are tied to the model, so there's some, like, SDXL styles-

Suhail Doshi30:39

Yeah

Swyx30:39

... there's some Playground v2 styles. Um, can you maybe give listeners a overview of how that works? Because in, in language, there's not this idea of, like, style, right?

Suhail Doshi30:49

Right.

Swyx30:50

Versus like in, in vision model there is, and you cannot get certain styles in different models. Um-

Suhail Doshi30:55

Mm-hmm

Swyx30:56

... how do styles emerge and how do you categorize them and find them?

Emergent Styles30:59

Suhail Doshi30:59

Yeah, I mean it's, it's so fun having a community where they-- people are just trying a model. Like we- we-- it's only been two days for Playground v2, and we actually don't know what the model's capable of and not capable of.

You know, we certainly see problems with it, but we have yet to see, uh, what emergent behavior is. And we, we've just sort of discovered that it takes about like a week before you start to see like new things.

But uh, I think like a lot of that style kind of emerges after that week, where you start to see, you know... You know, there's some styles that are very like well known to us, like maybe like pixel art is a well-known style.

But then there's some style-- you know, photorealism is like another one that's like well known to us. But there are some styles that cannot be easily named. Uh, you know, it's not as simple as like, okay, that's an anime style.

Swyx31:46

Yeah.

Suhail Doshi31:46

Uh, it's very visual. And, and in, and in the end you end up making up the name, uh, for what that style represents. And so the community kind of shapes itself around these different things. And so if you-- if anyone that's in- into stable diffusion and into any- building anything with graphics and stuff with these models, you know, you might have heard of like Protovision or DreamShaper- ...

some of these weird names. Uh, but they're just, you know, invented by these authors, but they have a sort of je ne sais quoi that, you know, appeals to users. Um, yeah, there's this-

Swyx32:16

Because it like roughly embeds to what you, what you want.

Suhail Doshi32:19

I, it- I guess so. I mean it's like, you know, there's this, uh, one of my favorite ones that's fine-tuned, it's not made by us, it's called like Starlight XL, um, it's just this beautiful model. It's got really great color contrast and, and visual elements, and the users love it.

I love it. And yeah, it's, it's so hard. I think that's like a very big open question with graphics that I'm not totally sure how we'll solve. Um, yeah, I think a lot of styles are sort of... I don't know, it's, it's like an evolving situation too, 'cause styles get boring-

Swyx32:51

Mm

Suhail Doshi32:51

... right? They get fatigue. Like it's like listening to the same style of pop song. I kind of-- I try to relate to graphics a little bit like with music because I think it gives you a little bit of a different shape to things.

Like in music it's not just-- it's not as if we just have pop music and-

Swyx33:07

Mm

Suhail Doshi33:07

... you know, rap music and country music. Like there-- all of these-- Like the EDM genre alone has like subgenres. And I think that's very true in, in graphics and painting and art and anything that we're doing. There's just these subgenres even if we can't quite always name them.

Swyx33:22

Yeah.

Suhail Doshi33:23

Uh, but I think they are emergent from the community which is why we're so always happy to work with the community.

Swyx33:27

Yeah. That is a struggle, you know, coming back to this like B2B versus B2C thing. Uh, B to- B2C you're gonna have a huge amount of diversity and then it's gonna reduce as you get towards more sort of B2B type use cases.

I, I, I'm making this up here. I'm-

Suhail Doshi33:40

Yeah, yeah

Swyx33:40

... tell me if you disagree. Um, so like you might be optimizing for a thing that you m- may eventually not need.

Suhail Doshi33:46

Yeah, possibly. Yeah, possibly. Um, yeah, I try not to share-- I think like a simple thing with startups is that I worry sometimes by like, like by trying to be, uh, overly ambitious and like really scrutinizing like what something is in its most nascent phase that you miss the most ambitious thing you could have done.

Like just having like very basic curiosity-

Swyx34:07

Okay

Suhail Doshi34:07

... um, with something very small, um, can like ki- kind of lead you to something amazing. Like Einstein definitely did that and then when-- and then he like, you know, he basically won all the prizes and- ... got everything he wanted and then basically did like kind-- didn't really-

Swyx34:21

Nothing Yeah. He was done.

Suhail Doshi34:22

He kind of dismissed quantum-

Swyx34:24

Yeah.

Suhail Doshi34:24

... uh, and then just kind of was still searching, you know, for the unifying theory and he like had this quest, and I think that happens a lot with like Nobel Prize people. I think there's like a, a term for it that I forget.

Um, I actually wanted to go after a toy almost intentionally.

Swyx34:39

Huh.

Suhail Doshi34:40

Um, so long as that I could see, I could imagine that it would lead to something, uh, very, very large later. And so yeah, it's a very-- like I said, it's very hobbyist but you need to start somewhere.

You need to start with something that there ha- has a m- big gravitational pull, um, even if these hobbyists are, aren't likely to be the people that you know ha- have a way to monetize it or whatever. Even if they're-- But they're doing it for fun so there's something, something there that I think is really important.

But I agree with you that, you know, in time, uh, we're gonna have to foc-- we, we will absolutely focus on, um, more u- utilitarian things, like things that are more related to editing feats that are much harder but...

And so I think like a very simple use case is just, you know, I'm not a graphics designer. Um, I don't know if, I don't know if you guys are.

Swyx35:27

Mm-mm.

Suhail Doshi35:29

Um, but it su- you know, it, it seems like very simple that like you-- if we could give you the ability to do really complex graphics w- without skill, wouldn't you want that? You know, like my wife the other day was said, you know, said, "Ah, I wish Playground was better because I wish that...

You know, don't you-- D- d-- When are you guys ha- gonna have a feature where like we could make my son, his name's Devin, smile when he was not smiling in the picture for the holiday card?" Right? You know, just being able to highlight his, his mouth and just say like, "Make him smile."

Like why can't we do that with like high fidelity and coherence?

Swyx36:00

Mm-hmm.

Suhail Doshi36:00

Uh, little things like that all the way to, um, you know, uh, putting you in completely different scenarios.

Swyx36:06

Is that true? Can we not do that in painting?

Suhail Doshi36:09

You can do in painting but it's-- the quality is just so bad. Yeah.

Swyx36:14

Oh.

Suhail Doshi36:15

It's just really terrible quality. You know, it's, it's like you'll do it five times and it'll still like kinda look like crooked or just the artifact.

Swyx36:22

Mm.

Suhail Doshi36:22

Part of it's like, you know, the lips on the face are so-- there's such, it gi- there's such little information there. It's so small that the models really struggle with it. Yeah.

Swyx36:31

Make the picture smaller and you won't see it.

Suhail Doshi36:32

Well, I think, I think one-

Swyx36:33

That's my trick. I don't know.

Suhail Doshi36:35

Well, uh, yeah, yeah, that's true. Or y- you know, you could take that region and make it re- get really big and then like say it's a mouth and then like shrink it. It-

Swyx36:41

Yeah

Suhail Doshi36:41

... it feels like you're wrestling with it, um, more than it's doing something that kind of-

Swyx36:46

Yeah

Suhail Doshi36:46

... uh, surprises you. Yeah.

Swyx36:48

It, it feels like you are very much the internal tastemaker, like you carry in your head this vision for what a, a good art model should look like.

Suhail Doshi36:56

Mm-hmm.

Swyx36:57

Um, is it-- do you find it hard to like communicate it to like your team and, and, you know, other, other people just because it's obviously it's, it's hard to put into words like we just said.

Suhail Doshi37:06

Yeah. It's, uh, it's very hard to explain, uh, like images have such, like such high bit rate compared to just words.

Swyx37:15

Yeah.

Suhail Doshi37:16

And word-- we don't have enough words to describe-

Swyx37:18

Yep

Suhail Doshi37:18

... um, these, these things. Difficult. I think everyone on the team, if, if they don't have good kind of like judgment taste or like an eye for some of these things, they're like subtly building it because they have no choice.

Right? So in that realm, I don't worry too much-

Swyx37:33

Sure

Suhail Doshi37:33

... actually. Like everyone is kind of like l- learning, uh, to, to get the eye, is what I would call it. But I also have, you know, my own narrow taste, like I'm at my, you know, I'm not-- I don't represent the whole population either.

Swyx37:45

True, true.

Suhail Doshi37:46

So...

Swyx37:46

Um, y- when, when you benchmark models, you know, like this benchmark we're talking about, we use FID, uh-

Suhail Doshi37:52

Yeah

Swyx37:53

... Fisher Input Distance. Um, okay, that's one, one measure, but like doesn't capture anything you just said about smiles.

Suhail Doshi38:00

Yeah. FID, FID is, FID is generally a bad metric. Um-

Swyx38:03

So-

Suhail Doshi38:03

You know, it's good up, up to a point, and then it kind of like is irrelevant.

Swyx38:07

Yeah.

Suhail Doshi38:07

Yeah.

Swyx38:07

And then so w- are there any other metrics that you like, um, apart from vibes? I'm, I'm always looking for alternatives to vibes.

Suhail Doshi38:14

Apart from vibes.

Swyx38:14

Because vibes don't scale, you know?

Suhail Doshi38:16

You know, it might be fun to kind of talk about this, um, because it's actually kind of fresh. So up till now, we haven't needed to do a ton of like benchmarking because it's-- we hadn't trained our own model, and now we have.

So now what? What does that mean? How do we evaluate it? You know, we're kind of like living with the last forty-eight, seventy-two hours of going, "Did the way that we benchmark actually succeed? Did it deliver?"

Swyx38:38

Yeah.

Suhail Doshi38:38

Right? You know, like I think Gemini just came out. They just put out a bunch of benchmarks, but all these benchmarks are, are just an approximation of how you think it's going to end up with real world performance, and I think that's like very fascinating to me.

Um, so if you fake that benchmark, you'll, you'll still end up in a really bad scenario at the, at the end of the day. And so, you know, what, one of the benchmarks we did was we did a-- we kind of curated like a thousand prompts.

That's what, that's what we published in our blog post, you know. Of all these tasks that we-- a lot of the-- some of them are curated by our team where we know the models all suck at it. Like my favorite prompt that no model is really capable of is a, a horse riding an astronaut.

Swyx39:16

Yep.

Suhail Doshi39:16

The inverse one, and it's really, really hard-

Swyx39:19

Yep

Suhail Doshi39:19

... to do. Um-

Swyx39:20

Not in data.

Suhail Doshi39:21

You know, another one is like a giraffe underneath a microwave. How does that work? Right. There's so many of these little funny ones. We do-- we have prompts that are just like misspellings of things.

Swyx39:32

Yeah.

Suhail Doshi39:33

Right? Just to see if the models will figure it out. Uh, so sp-

Swyx39:35

That's easy. That's, that should, uh, embed to the same space.

Suhail Doshi39:39

Yeah.

Swyx39:39

Yeah.

Suhail Doshi39:39

And, and, and, and just like all these very interesting, weird, weirdo things. And so we have so many of these, and then we kind of like evaluate whether the models are any good at it, and the reality is that they're all bad at it, and so then you're just picking the most aesthetic image.

But, uh, but I think, you know, we're just-- we're still at the beginning of building like our, uh, like the best benchmark we can that aligns most with just user happiness-

Swyx40:01

Mm-hmm

Suhail Doshi40:01

... I think. 'Cause we're not, we're not like putting these in papers and trying to like win, you know, I don't know, awards at ICCV or something, if they have awards. Sorry if they don't. Um, and um-

Swyx40:11

You could. Well, that's absolutely a valid strategy.

Suhail Doshi40:13

Yeah, y- you could. I don't think it would correlate necessarily with the impact we want to have on humanity. I think we're still evolving whatever our benchmarks are. So the first benchmark was just like very difficult tasks that we know the models are bad at.

Can we come up with a thousand of these? Um, whether they're hand-rated and some of them are generated, uh, and then can we ask the users like, "How do we do?" Um, and then we wanted to use a benchmark like party prompts so that people in academi- we mostly did that so people in academia could measure their models against ours versus others.

Um, and, uh, but yeah, I mean, FID, FID is pretty bad, and I think, yeah, you-- in terms of vibes, it's like you put out the model, and then you try to see like what users make. And I think my sense is that we're gonna take all the things that we notice that the users kind of were failing at, um, and try to find like new ways to measure that, whether that's like a smile or, you know, color contrast or lighting.

Um, one benefit of Playground is that we have users making millions of images, um, every single day, and so we can just ask them. Um, and that-

Swyx41:18

Like go for like a post-generation feedback.

Suhail Doshi41:20

Yeah. We can just ask them. We can just say like, "How, how good was the lighting here? How was, um, how was the subject? How was the background?"

Swyx41:27

Yeah.

Suhail Doshi41:27

Uh-

Swyx41:28

Oh, like a for- like a proper form of like-

Suhail Doshi41:31

It's just like you make it-

Swyx41:32

... some like six-

Suhail Doshi41:32

You come to our site, you make an image, and then we say-

Swyx41:34

Yeah

Suhail Doshi41:34

... and then maybe randomly you just say, "Hey, you know, like how was, how was the color and contrast of this image?" And you say, "It was, it was not very good," and then you just tell us. So I think, I think we can get like tens, uh, tens of thousands of these, uh, evaluations every single day to, to truly measure real world performance-

Swyx41:52

Yeah

Suhail Doshi41:52

... as opposed to just like benchmark performance. Hopefully next year, I think we will try to publish kind of like a, like a, a benchmark that anyone could use, that we evaluate ourselves on-

Swyx42:03

Yep

Suhail Doshi42:03

... and that other people can-

Swyx42:04

That's an ideal goal

Suhail Doshi42:05

... that we think does a good job of approximating real world performance because we've tried it and done it and noticed that it did. Yeah. I think, I think we will do that.

Swyx42:13

Yeah. Um, we're-- I think we're gonna ask a few more like sort of product-y questions. Um, I, I, and I, I personally have a few like categories that I consider special-

Suhail Doshi42:22

Mm-hmm

Swyx42:22

... among, you know, you know, you have like animals, art, fashion, food. Um- There are some s- categories which I consider like a different tier of image. Uh, so f- the top among them is text in images.

Suhail Doshi42:33

Mm. Mm-hmm.

Swyx42:34

Um, how do you think about that? Um, so one of the big wow moments for me, or something I've been looking out for the entire year is just the progress of text in images. Like do you, can you write in an image?

Suhail Doshi42:45

Yeah.

Swyx42:45

Or, um, and Ideogram-

Suhail Doshi42:47

Mm-hmm

Swyx42:47

... I think came out recently-

Suhail Doshi42:48

Mm-hmm

Swyx42:48

... which had decent, uh, but not perfect text in images. Um, DALL-E 3 had improves, uh, some, and all they said in their, their, uh, paper was that they just included more text in the dataset and it just worked.

I was like, "That's just, that's just lazy." All right. But anyway, uh, do you care about that? Uh, because I don't see any, any of that in like your samples.

Suhail Doshi43:08

Yeah, yeah. We're, our, yeah, the, the V2 model is, um, was, was mostly focused on, um, image quality versus like the feature of, uh, y- text synthesis.

Swyx43:18

Yeah. 'Cause I, well, as a business user, I care a lot about that.

Safety and Limits43:18

Suhail Doshi43:21

Yeah.

Swyx43:21

Right.

Suhail Doshi43:22

Yeah. I'm very excited about text synthesis, and yeah, I think Ideogram has done a good job of may- maybe the best job. Uh, DALL-E kind of ha- it's like a, it has like a hit rate. You know, you don't want just text effects.

I think where this has to go is h- it has to be like you could like write little tiny pieces of text like on like a milk carton.

Swyx43:41

Yeah.

Suhail Doshi43:41

That's maybe not even the focal point of a scene.

Swyx43:44

Yeah.

Suhail Doshi43:44

I think that's like a very hard task that, um, you know, if you could do something like that, then there's a lot of other possibilities.

Swyx43:50

Well, you don't have to zero shot it. You can just be like, "Here, f- focus on this."

Suhail Doshi43:54

Sure, yeah, yeah. Definitely. Yeah, yeah. So I think text synthesis would be very exciting.

Swyx43:58

Yeah. Uh, and then also, I'll also flag that, um, Max Wolf, Minimaxir, which you must have come across his work, um, he's done a lot of stuff about w- using like logo masks-

Suhail Doshi44:08

Mm-hmm

Swyx44:09

... that then map onto like a, like food or-

Suhail Doshi44:13

Mm-hmm

Swyx44:13

... vegetables and it, and it looks, looks like text, uh, which, which can be pretty funny.

Suhail Doshi44:17

Yeah. Yeah, I mean, you, you... It's very interesting to-- That, that's the wonderful thing about like the open source community is that you get things like ControlNet-

Swyx44:25

Yeah

Suhail Doshi44:25

... and then you see all these people do these just amazing things with ControlNet, and then you wonder, uh, I think from our point of view, we, we sort of go that, that's really wonderful, but how, how do we end up with like a unified model that can do that?

What are the bottlenecks? What are the issues? Um, because the community ultimately has very limited resources.

Swyx44:42

Yeah.

Suhail Doshi44:42

And so they, they need these kinds of like workaround, um, workaround research ideas to get there. Um, but yeah.

Swyx44:50

Yeah. Are, are techniques like ControlNet portable to your architecture?

Suhail Doshi44:54

Definitely.

Swyx44:55

Okay.

Suhail Doshi44:55

Yeah. It, we kept the Playground v2 arc exactly the same as SDXL, not because, not out of laziness, but just because we wanted-- we knew that the community already had tools.

Swyx45:04

Yeah.

Suhail Doshi45:05

It's, you know, all you have to do is maybe change a string in your code and then, you know, retrain a ControlNet for it, so that it was very intentional to do that. We didn't wanna fragment the community with different architectures.

Swyx45:15

Yeah. Yeah.

Suhail Doshi45:15

Yeah.

Swyx45:15

Uh, I, I have more questions about that. I, I don't know. I don't, I don't wanna DDoS you with, uh- ... with topics. But, but okay, I was basically gonna, gonna go over three more categories.

Suhail Doshi45:24

All right.

Swyx45:24

One is, uh, UIs, like, um, app UIs, like mock UIs. Uh, th- third is, uh, not safe for work.

Suhail Doshi45:31

Mm-hmm.

Swyx45:31

Um, obviously. Uh, and then copyrighted stuff.

Suhail Doshi45:34

Mm-hmm.

Swyx45:34

Um, I don't know if you care to comment on any, any of those.

Suhail Doshi45:37

The NSFW kind of like safety stuff is really important. Um, part, part of-- I, I kind of think that one of the biggest risks kind of going into maybe the US election year will probably be inter, very interrelated with like graphics, audio, um, video.

I think it's gonna be very hard to explain, you know, to a family relative who's not kind of in our world, and our w- our world is like sometimes very, you know, we think it's very big, but it's very tiny-

Swyx46:05

Yeah

Suhail Doshi46:05

... compared to the rest of the world-

Swyx46:05

Absolutely

Suhail Doshi46:06

... sometimes like there's still lots of humanity have no idea what ChatGPT is. And I think it's gonna be very hard to explain, you know, to your uncle, aunt, whoever, you know, "Hey, I saw, you know, I saw President Biden say this thing on a video," you know, "I, I can't believe, you know, he said that."

I think that's gonna be a very troubling thing going into, um, going into the world next year or the year after.

Swyx46:30

Oh, uh, uh, I didn't, that, that's more of like a risk thing-

Suhail Doshi46:32

Yeah

Swyx46:32

... or like deepfakes, uh, well, faking, political faking. But, uh, there's just, there's a lot of, um, studies on how, um, yeah, for most businesses you don't wanna train on not safe for work images-

Suhail Doshi46:44

Mm-hmm

Swyx46:44

... except that it makes you v- really good at bodies.

Suhail Doshi46:48

Yeah, I mean- ... uh, yeah, I mean, we, we personally, we filter out, um, NSFW type of, uh, images in our dataset so that it's, you know, so our safety filter stuff doesn't have to work as hard.

Swyx47:00

But you've, you've heard this argument that it get, it makes you worse at, uh, because obviously not safe for work images are very good at, um, human anatomy-

Suhail Doshi47:08

Mm

Swyx47:08

... which you do wanna be good at.

Suhail Doshi47:10

Yeah, it's, it's not about like, it's not like necessarily a bad thing to train on that data. It's more about like how you go and use it. That's why I was kind of talking about safety, um-

Swyx47:18

Yeah, yeah, I see

Suhail Doshi47:19

... you know, in part because there are very terrible things that can happen in the world. If you have a sufficiently, you know, extremely powerful graphics model, you know, suddenly like you can kind of imagine, you know, now if you can like generate nudes and then there's like you could do very character consistent things with faces, like what does that lead to?

Swyx47:34

Yeah.

Suhail Doshi47:34

Yeah. I think it's like more what occurs after that, right? Even if you train on, let's say, you know, new data, if it does something to kind of help, there's nothing wrong with the human anatomy, um, it's very valid for a model to learn that, uh, but then it's kind of like how does that get used?

And, uh, you know, I, I, I won't bring up all of the very, very unsavory terrible things that we see, uh, on, on a daily basis-

Swyx47:57

Oh, God

Suhail Doshi47:57

... on the site. I think it's more about what, what occurs. And so we, you know, we just recently did like a big sprint on safety internally around... And, and it's very, it's very difficult with graphics and art, right?

Because there is tasteful art that has nudity.

Swyx48:12

Yeah.

Suhail Doshi48:13

Right? They're all over in museums, like, you know, it, it's very, very valid situations for that, and then there's, you know, there's the things that are the gray line of that. You know, what I might not find tasteful, someone might fi- be like, "That is completely tasteful," right?

And then, and then there are things that are way over the line.

Swyx48:29

Yeah.

Suhail Doshi48:30

Um, and then there are things that are, you know, maybe, maybe you or, you know, maybe I would, you know, be okay with, but society isn't.

Swyx48:37

Yeah.

Suhail Doshi48:38

I think it's really hard with art. Think it's really, really hard. Sometimes if even if you have like even if you have, um, things that are not new, if, if a child goes to, to your site, scrolls down some images, you know, classrooms of kids, you know, using our product, it's a really difficult problem.

And, um, and it, and it stretches mostly culture, society, politics, everything. Yeah.

Alessio49:00

Okay. Um, another favorite topic of our listeners is, um, UX in AI, and I think you're probably one of the best all-inclusive editors for these things.

Suhail Doshi49:12

Mm.

Alessio49:12

So you don't just have the, you know, prompt images come out, you pray and if no, you do it again. Uh, first you let people, um, pick a seed so they can kind of have semi-repeatable generation. Um-

Suhail Doshi49:26

Yeah.

Alessio49:26

You also have-

Suhail Doshi49:27

Absolutely.

Alessio49:27

Yeah, you can pick how many images and then you leave all of them in the canvas, and then you have kind of like this box, the generation box, and you can even cross between them and outpaint.

Suhail Doshi49:38

Yeah.

Alessio49:38

There's all these things.

How did you get here? You know?

Suhail Doshi49:42

Yeah.

Alessio49:43

Most people, most people are kind of like, "Give me text, I give you image," you know?

Suhail Doshi49:46

Yeah.

Alessio49:46

And you're like, "These are all the tools for you."

AI-First UX49:47

Suhail Doshi49:48

Even though we are trying to make, um, a, a graphics foundation model, I think we think that we're also trying to f- like reimagine like what a graphics editor might look like given the change in technology. So you know, we-- I don't think we're trying to build Photoshop, but it's the only thing that we could say that people are fam- you know, largely familiar with.

"Oh, okay. There's Photoshop." Uh, I think... You know, I don't think you would think of Photoshop without like the com- you know, you don't, you wouldn't think what would Photoshop compare itself to pre, pre-computer? I don't know, right?

Alessio50:22

Mm-hmm. Mm-hmm.

Suhail Doshi50:22

It's like, oh, we're kind of like a, a canvas.

Alessio50:25

Mm-hmm.

Suhail Doshi50:25

But you know, there's these menu options, and you can use your mouse. What's a mouse? Um, so I, I think that we're trying to make like-- we're trying to reimagine what a graphics editor might look like, not, not just for the fun of it, but because we kind of have no choice.

Like, there's this idea in, in image generation where you can gen-generate images. That's like a super weird thing. What is that in Photoshop, right? You have to wait right now for the time being, um, but the wait is worth it often for a lot of people because they can't make that with their own skills.

So I, I think it goes back to, you know, how we started the company, which was kind of looking at GPT-3's playground. The, that the reason why we're named Playground is, is a homage to that actually. Um, and you know, it's like shouldn't these products be more visual?

Shouldn't, you know, shouldn't they... These prompt boxes are like, like a terminal window, right?

Alessio51:14

Mm-hmm.

Suhail Doshi51:15

We're kind of at this weird point where it's just like CLI. It's like MS-DOS. I remember my mom using MS-DOS, and I memorized the keywords like dir, ls, all those things. Right? It feels a little like we're there, right?

Prompt engineering is this like-

Alessio51:28

The, the shirt I'm wearing, you know, it's, it's-

Suhail Doshi51:29

Yeah

Alessio51:29

... it's a bug, not a feature.

Suhail Doshi51:30

Yeah. Exactly. Parentheses to say beautiful or whatever, which weights the word token more in the model or whatever. Um, yeah, it's that- that's like super strange. I think that's not... I think everybody, I think a large portion of humanity would agree that that's not user-friendly, right?

So how do we think about the products to be more user-friendly? Well, sure. You know, sure would be nice if I could like, you know, uh, if I wanted to get rid of like the headphones on my, my head, you know, it'd be nice to mask it and then say, you know, "Can you remove the headphones?"

Um, you know, if I want to grow the, expand the image, it should... You know, how can we make that feel easier without typing lots of words and being really confused? And by no, by no m- stretch of the imagination, I don't even think we've nailed the UI/UX yet.

Um, part of that is because we don't-- we're still experimenting, and part of that is because the model and the technology's gonna get better. And whatever felt like the right UX six months ago is gonna feel very broken now.

Um, and, uh, so that, that's a little bit of how we got there is kind of saying, "Does everything have to be like a prompt in a box, or can we do, can we do things that make it very intuitive for users?"

Alessio52:42

How do you s- decide what to give access to? So you have things like, um, expand prompt-

Suhail Doshi52:47

Mm-hmm

Alessio52:48

... uh, which DALL-E 3 just does.

Suhail Doshi52:50

Mm-hmm.

Alessio52:50

It doesn't let you decide-

Suhail Doshi52:51

Yeah

Alessio52:51

... whether you should or not. Um, yeah.

Swyx52:54

It as in like, uh, rewrites your prompts for you.

Alessio52:56

Yeah.

Suhail Doshi52:56

Mm-hmm.

Swyx52:56

Yeah.

Suhail Doshi52:57

Yeah, for that feature, I, I think we'll probably... I, I think once we get it to be, uh, cheaper, we'll probably just give it out. We'll probably just give it away. But we also decided something that we-- that might be a little bit different.

We noticed that most of image generation is just like kind of casual. You know, it's in WhatsApp, it's, you know, it's in a Discord bot somewhere with Midjourney. It's in ChatGPT. One of the differentiators I think we provide is at the expense of just lots of users necessarily, mainstream consumers, is that we provide as much like power and tweakability and configurability as possible.

So the only reason why it's a tog- it's a toggle because we know that users might wanna use it and might not wanna use it, right? There are some us- there are some really powerful power user hobbyists that know what they're doing, and then there's a lot of people that, um, uh, you know, just want something that looks cool, but they don't know how to prompt.

And so I think a lot of Playground is more about, um, going after that core user base that like knows-- has a little bit more savviness, uh, in how to use these tools. Yeah. So they might not use like these users probably...

You know, the average DALL-E user is probably not gonna use ControlNet. They probably don't even know what that is.

Alessio54:07

Hmm.

Suhail Doshi54:08

Um, and so I think that like as the models get more powerful, as the, the-- there's more tooling, um, yeah, I think you could imagine it. Hopefully, you'll imagine a new sort of AI first graphics editor that's just as like powerful and configurable as Photoshop.

Uh, and you might have to master a new kind of tool.

Swyx54:27

Yeah.

Suhail Doshi54:28

Yeah.

Swyx54:28

Well, um, th- uh, there are so many things I could, I could go bounce off of that. Um, one, one you, you mentioned about waiting.

Suhail Doshi54:36

Mm-hmm.

Swyx54:37

Um, w- we have to kind of somewhat address the elephant in the room. Uh, uh, consistency models have been blowing up, uh, uh, the past month. Um, is that... Like how do you think about integrating that? Um, obviously there's, there's a lot of other companies also trying to, uh, beat you to that space as well.

Suhail Doshi54:53

I think we were the first company to integrate it. Well, we integrated it in a different way. There are like 10 companies right now that have kind of tried to do like interactive editing where-

Swyx55:01

Yeah

Suhail Doshi55:02

... you can like draw on the left side and then you get an image on the right side. We decided to kind of like wait and see whether there's like true utility on that. Um, we have a different feature that's like unique, uh, in our product that, um, that, that's called preview rendering.

And so you go to the product and you, and you say... You know, we, we're like, "What is the most common use case?" The most common use case is you write a prompt and then you get an image.

But what's the most annoying thing about that? The most annoying thing is like it's like feels like a slot machine, right?

Swyx55:28

Mm-hmm.

Suhail Doshi55:28

You're like, "Okay, I'm gonna put it in and I'm gonna... Maybe I'll get something cool." So we did something that seemed a lot simpler but a lot more relevant to how users already use these products, which is preview rendering.

You toggle it on and it will show you a render of the image, and then it's just like a pr- graphics tools already have this. Like if you use Cinema 4D or After Effects or something, it's called viewport rendering.

Swyx55:49

Mm-hmm.

Suhail Doshi55:50

And so we try to take, take something that exists in the real world that has familiarity and say, "Okay, you're gonna get a rough sense of an early preview of this thing, and then when you're ready to generate, it's...

We're gonna try to be as coherent about that image that you saw." That way you're not spending so much time just like, you know, uh, pulling down the slot machine lever. So I, I... We were, we were actually the first company, I think we were the first company to actually ship-

Swyx56:13

My bad

Suhail Doshi56:13

... that quick LCM-

Swyx56:14

Yeah

Suhail Doshi56:15

... uh, thing. Yeah.

Swyx56:16

Okay.

Suhail Doshi56:18

Yeah. We were very excited about it, so we shipped it very quick. Yeah.

Swyx56:21

Yeah, yeah. I, I think m- like the other, um, the... Well, the demos I've been seeing, it, it's, it's also, I guess, it's not like a preview necessarily. They're almost using it to animate theirs, their, their generations. Like to, you can- because you can kind of move shapes over-

Suhail Doshi56:36

Yeah, yeah. They're, they're like doing it. They're like animating it, but they're sort of showing like if I move a moon, you know, can I... Yeah.

Swyx56:43

Yeah. I don't know. It, it, it to me unlock, un-unlocks video in a way.

Suhail Doshi56:47

Yeah.

Swyx56:47

Um, that, uh

Suhail Doshi56:48

But the video-

Swyx56:49

I've seen

Suhail Doshi56:49

... the video models are already so much better than that.

Swyx56:51

Yeah.

Suhail Doshi56:51

Yeah. So.

Swyx56:52

So . Uh, there's a- there's another one which I, I think is, uh, um, um, like h- how about like the just g- general ecosystem of LoRAs?

Suhail Doshi57:01

Uh-huh.

Swyx57:02

Right? That, um, Civit is obviously the most popular, um, repository of LoRAs. Um, how do you t- think about sort of interacting with that ecosystem?

Suhail Doshi57:12

Yeah, I mean, uh, the guy that, that did LoRA, not the guy that invented LoRAs, but the person that brought LoRAs to Stable Diffusion, uh, actually works with us, um, on, on some projects. Uh, his name is Simu.

LoRA Ecosystem57:13

Suhail Doshi57:24

Um, shout out to Simu. Um, and I, I think LoRAs are, are wonderful. Um, obviously fine-tuning all these DreamBooth models and such are j- is just so heavy and giving... And I- it's obvious in our conversation around styles and vibes and, you know, it's very hard to evaluate the artistry of these things.

LoRAs give people, uh, this wonderful like opportunity, uh, to create like subgenres of art, and I think they're amazing. And so any graphics tool, any kind of thing that's expressing art has to provide some level of customization to its, its user base that goes beyond, you know, just typing like Greg Rutkowski in a prompt.

Right? We have to give more than that. Um, y- it's not like users wanna type these, you know, art- real artist names. It's that they don't know how else to get an image that looks interesting. They, they truly want like originality and uniqueness, and I think LoRAs provide that, and they provide it in a very nice scalable way.

Um, I hope that we find something even better than LoRAs in the, in the long term, in the long term. Um, 'cause there, there are still weaknesses to LoRAs, um, but I think they do a good job for now.

Swyx58:33

Yeah. And so you don't want to be the... Like you don't, you would never compete with Civit. You would just kind of let people in.

Suhail Doshi58:38

Civit's a site where like all these things get kind of hosted by the community, right?

Swyx58:41

Yeah.

Suhail Doshi58:41

Um, and so yeah, we'll often pull down the thing, like some of the best things there. Um, I think, I think when we have a significantly better model, uh, we will certainly build something-

Swyx58:53

I see

Suhail Doshi58:54

... that gets closer to that. I still, again, I go back to saying just I still think this is like very nascent. Things are very underpowered, right? We, you know, LoRAs are not easy for people to train. You know, they're easy for an engineer.

Swyx59:07

Okay.

Suhail Doshi59:08

But they're not e- they're not easy, you know... It sure would be nicer if I could just pick, you know, five or six reference images-

Swyx59:14

Yeah

Suhail Doshi59:14

... right? And, and then say, "Hey, you know, this is, this is..." And it, and there might even be five or six different reference images that are not... They're just very different actually. Like they're, they're, they communicate a style, but they're actually like, it's like a mood board, right?

And it takes, you have to be kind of an engineer almost to train these LoRAs or go to some site and be technically savvy at least. Um, it seems like it'd be much better if I could say, "I love this style.

I love, I love this, this style. Here are five images" And you tell the model like, "This is what I want" and the model gives you, gives you something that's very aligned with what your style is, what you're talking about.

And it's a style you couldn't even communicate, right? There's no word. You know, this is, you know, if you have a Tron image, it's not just Tron. It's like Tron plus like s- four or five different weird things.

Swyx59:59

And cyberpunk. Yeah.

Suhail Doshi1:00:00

Yeah. Um, even cyberpunk can have its like subgenre, right? But I just think training LoRAs and doing that is very heavy, so I hope we can do better than that.

Swyx1:00:09

Cool.

Suhail Doshi1:00:09

Yeah.

Swyx1:00:10

Yeah. Um, we had Sharif from Lexica on the podcast before.

Suhail Doshi1:00:14

Oh, nice.

Swyx1:00:15

And both of you have like a landing page with just a bunch of images where you can like explore things.

Suhail Doshi1:00:21

Yeah.

Swyx1:00:21

Um-

Suhail Doshi1:00:22

Yeah, we have a feed.

Swyx1:00:23

Yeah, yeah. Is that something you see more and more of in terms of like coming up with these styles? Is that why you, you have that as the starting point versus a lot of other products you just go in, you have the generation prompt, you don't see a lot of examples.

Suhail Doshi1:00:36

Right. Our feed is a little different than, than their feed. Our feed is more about community, so we have kind of like a Reddit thing going on where it's a kind of a competition like every day, loose competition, mo- mostly fun competition of like making things, and there's just this wonderful community of people where they're liking each other's images and just showing their like, their genuine interest in each other, and I think we definitely learn about styles that way.

One of the funniest polls, uh, i- if you go to the Midjourney c- c- polls, they'll sometimes put these polls out and they'll say, "You know, what do you wish you could like learn more from?" And like one of the, one of the things that people vote the most for is like learning how to prompt.

Right? And so I think like, you know, if you, if you put away your research hat for a minute and you just put on like your product hat for, for a second, you're kind of like, "Well, why do people wanna learn how to prompt?"

Right? It's because they wanna get higher quality images. Well, what's higher quality? Composition, lighting, aesthetics, so on and so forth. And I think that the community on our feed, I think, I think we have-- I think we might have the biggest community and, uh, and it gives all of the users a way to learn how to prompt.

Because they're just seeing this huge rising tide of all these images that are super cool and interesting, and they can kinda like take each other's prompts and like kind of learn how to do that. Um, I, I think that'll be short-lived because I think the complexity of these things is gonna get higher.

Um, but, um, but that, that's more about why we have that feed, is to help each other. To help teach us-users, and then also just, you know, celebrate people's art.

GPU Engineering1:02:07

Swyx1:02:09

You run your own infra.

Suhail Doshi1:02:10

We do.

Swyx1:02:11

Yeah. That's unusual.

Suhail Doshi1:02:14

Uh, it's necessary.

Swyx1:02:16

It's necessary.

Suhail Doshi1:02:16

Yeah.

Swyx1:02:17

Uh, what have you learned running DevOps for GPUs? Uh, you-- I, I... You had a tweet about like how many A100s you have, but I feel like it's out of date probably.

Suhail Doshi1:02:26

Uh, yeah, we, uh... I think, I mean, it just comes down to cost. These things are very expensive, so, uh, we just wanna make it as affordable for everybody as possible. Um, I don't find... I find the DevOps for inference to be rel-relatively easy.

Swyx1:02:40

Okay.

Suhail Doshi1:02:40

It doesn't feel that different than, you know... I think we had thousands and thousands of servers at Mixpanel, uh, just for dealing with the, the API. It had such huge quantities of volume that I didn't find it, I don't find it particularly very different.

Um, I do find, uh, GP, model optimization performance is very new to me, so I think that, I find that very difficult at the moment. So that's very interesting. But, uh, scaling inference is not, not, not terrible. Tr- scaling a training cluster is very, much, much harder, um, than I perhaps anticipated.

Swyx1:03:12

Why is that?

Suhail Doshi1:03:13

Well, you have, you know, you have to, it's just like a very large distributed system with, um, you know, if you have like a, a node that goes down, then your tr- you know, training run crashes, and then you have to somehow be resilient to that.

And I would say training infra software is very early. It feels very broken. Feel-

Swyx1:03:30

Like a, like a-

Suhail Doshi1:03:30

I can tell in 10 years it would be a lot better

Swyx1:03:32

... like a Mosaic or whatever.

Suhail Doshi1:03:34

We don't, yeah, we don't even... No, we don't. We, we think we use very basic tools like, you know, Slurm for scheduling and just normal PyTorch, PyTorch Lightning, that kind of thing. I think our tooling is nascent. I think I talked to a friend that's over at xAI.

They just, they like built their own scheduler, you know, and doing things with Kubernetes. Like, when people are building out tools because the existing open source stuff doesn't work and everyone's doing their own bespoke thing, you know there's a ma- there's a valuable company to be formed.

Swyx1:03:58

Yeah. Uh, I think it's Mosaic. I don't know.

Suhail Doshi1:04:01

Well, Mos- with Mosaic, yeah. It's t- it's tough with Mosaic 'cause, um... Anyway, I won't, I won't go into the details why, but yeah, we, we found it, uh, difficult to do. You, you... It might be worth like wondering like why, why, why not everyone is going to Mosaic.

Swyx1:04:13

Yeah.

Suhail Doshi1:04:13

And perhaps it's still, it's, I, I just think it's nascent.

Swyx1:04:16

Cool.

Suhail Doshi1:04:16

And perhaps Mosaic will come through.

Swyx1:04:18

Cool. Anything for you?

Alessio1:04:19

Um, no, no. This was great. And just to wrap, we, we talked about some of the pivotal moments in your mind with like DALL-E and, and whatnot. If you were not doing this, what's the most interesting unsolved question in AI that you would try and build then?

Suhail Doshi1:04:36

Oh, man. Coming up with startup ideas is very hard on the spot. Uh... And-

Swyx1:04:41

You sh- you have to have them. I mean, you're a founder. You're a repeat founder, like...

Lightning Round1:04:44

Suhail Doshi1:04:46

I, I'm very picky about my startup ideas. Um, so I don't, I, you know, don't have any great ones. Uh, the only thing that I... I, I don't have an idea per se, as much as a curiosity. Uh, and I'll po- I suppose I'll pose it to you guys.

Right now, we sort of think that a lot of the modalities just kinda feel like they're, you know, vision, language, audio, like that's roughly it. And somehow all this will like turn into something. It'll be multimodal, and then we'll end up with AGI, um, perhaps.

And I just think that there are probably far more modalities than maybe we, than meets the eye. And it just seems hard for us to see it right now because it's sort, sort of like we have tunnel vision on the moment.

Swyx1:05:32

We're, we're just like code, image, audio, video.

Suhail Doshi1:05:35

Yeah. I think I-

Swyx1:05:36

Very, very broad categories

Suhail Doshi1:05:37

... I think we are lacking imagination as a species in this regard.

Swyx1:05:41

Yeah.

Alessio1:05:41

I see it. I see it.

Suhail Doshi1:05:42

And, and I think like, you know, just like, uh, you know, it's not, I don't know what company would, would form as a result of this, but you know, like there's some, some very difficult problems like just tr- like a true actual, like not a meta world model, but an actual world model that truly maps everything that's going in terms of like physics and fluids and all these various kinds of interactions.

And y- what does that kind of model, like a true physics foundation model of sorts that represents Earth. And that in of itself seems very difficult, you know? But we just think of, but we're, but we're kinda stuck on like thinking that we can approximate everything with like, you know, a word or a token, if you will.

And I went, you know, I had a dinner last night where we were kind of debating this philosophically, and I think someone, you know, said something that I also believe in, which is like, at the end of the day, it doesn't really matter that it's like a token or a byte.

At the end of the day, it's just like some, you know, unit of information that it emits. But you know, you do, I do wonder if there are more, far more modalities than, um, than meets the eye. And if, if you could create that, then what would that, what would, what would that company become?

What problems could you solve? So I, I don't, I don't know yet, so I don't have a great company for it.

Swyx1:06:53

I don't know.

Suhail Doshi1:06:53

But-

Swyx1:06:54

Maybe you just inspire somebody to, to try, so.

Suhail Doshi1:06:56

Yeah. Hopefully.

Swyx1:06:57

Yeah. Uh, my personal response to that is I'm, I'm less interested in physics and more interested in people. Uh-

Suhail Doshi1:07:02

Mm-hmm

Swyx1:07:02

... like, like how do I, how do I mind upload? Because that is-

Suhail Doshi1:07:05

Right. Exactly

Swyx1:07:06

... teleportation, that is immortality, that is everything.

Suhail Doshi1:07:09

Yeah. Yeah. Can we, can we model our own... Rather than trying to create consciousness, could we model our own- Um-

Swyx1:07:16

Yeah

Suhail Doshi1:07:16

... even if it was lossy-

Swyx1:07:17

Yeah

Suhail Doshi1:07:17

... to some extent. Yeah.

Swyx1:07:19

Yeah. Um, well, we won't solve that here.

Suhail Doshi1:07:21

Yeah.

Swyx1:07:22

Um, if I were to take a Bill Gates book trip- ... uh, and had a week, uh, what should I take with me to learn AI?

Suhail Doshi1:07:30

Oh, man. Oh, gosh. You shouldn't take a book. You should just go to- ... uh, YouTube and visit Karpathy's, uh- ... class-

Swyx1:07:38

Zero to hero

Suhail Doshi1:07:39

... and just do it, do it. Grind through it.

Swyx1:07:42

Is, was that actually the most useful thing for you?

Suhail Doshi1:07:44

I wish it came out when I started-

Swyx1:07:45

Wow

Suhail Doshi1:07:45

... back last year. I, I'm, I'm as, as bummed that I didn't get to take it at the beginning. Um, but I did, I did do a few of his classes regardless. I, I don't think books-- Every time I buy a programming book, I never read it.

I always find that just writing code helps cement my internal understanding.

Swyx1:08:02

Yeah. So, so, so more generally, advice for founders who are not PhDs and are effectively self-taught, like, like you are. Like, what should they do? What should they avoid?

Suhail Doshi1:08:11

Same thing as if, as that I would advise if you were programming. Pick a project that seems very exciting to you, but don't, you know, it doesn't have to be too serious, and build it and learn every detail of it while you do it.

Swyx1:08:22

And it must be, uh, like, should you train, or can you, can you go f-far enough not training, just-

Suhail Doshi1:08:29

It-

Swyx1:08:29

... fine-tuning?

Suhail Doshi1:08:29

It depends. I would s- I would just follow your curiosity. If, like, you want, if what you wanna do is something that requires fundamental understanding of training models, then you should learn it. You don't have to be a P- you don't have to get to become a five, you know, five-year whatever PhD, but if that's necessary, I would do it.

If it's not necessary, then go as far as you need to go. But I would learn, you know, pick something that motivates. I think most people tap out on motivation, but they're deeply curious.

Swyx1:08:53

Yeah. Cool.

Suhail Doshi1:08:54

Yeah.

Swyx1:08:54

Cool. Excellent. Thank you so much for coming out, man.

Suhail Doshi1:08:56

Thank you. Thank you for having me.

Swyx1:08:57

This was fun.

Suhail Doshi1:08:57

Appreciate it.