# The AI-First Graphics Editor - with Suhail Doshi of Playground AI

Latent Space · 2024-01-02

<https://addtry.com/bc5c22b8-9e47-48fc-a66f-6f9b761cb0e7>

Suhail Doshi, co-founder of Mixpanel and founder of Playground AI, argues that image generation is still in a GPT-2 moment and that training open-source foundation models from scratch is necessary to unlock real utility, exemplified by Playground v2's 2.5x preference over Stable Diffusion XL on a 1K prompt benchmark. He explains the pivot from Mighty (cloud-streamed browser) to AI, driven by the belief that shifting compute elsewhere aligns with AI's parallel computation. The episode covers Playground's unique UI—not just a prompt box but a full canvas with preview rendering, seed control, and style filters—and the difficulty of balancing safety with artistic expression, especially around NSFW content. Suhail discusses the under-investment in graphics AI compared to language, the open-source community's role, and his choice to release pre-trained weights for academic research. He also shares lessons from running GPU infrastructure (harder for training than inference) and advises founders to follow curiosity and build projects rather than relying on books.

## Questions this episode answers

### How does Playground v2 perform compared to Stable Diffusion XL?

Suhail Doshi explained that Playground v2 achieved a 2.5x preference over Stable Diffusion XL on a 1,000-prompt benchmark they curated. They trained the model from scratch, open-sourced the weights including pre-trained checkpoints, and kept the architecture identical to SDXL to maintain community compatibility. The improvement came from meticulous data filtering, captioning, and alignment, with plans for a v2.1 update addressing issues like eye quality.

[0:23](https://addtry.com/bc5c22b8-9e47-48fc-a66f-6f9b761cb0e7?t=23000)

### What is the MJHQ 30K benchmark and why use Midjourney images for evaluation?

Suhail Doshi introduced the MJHQ 30K benchmark to automatically evaluate a model's aesthetic quality by using 30,000 Midjourney images across 10 categories. He believes in comparing to the best—Midjourney—rather than weaker models, to fairly push progress. The benchmark includes diverse prompts; the community can use it to measure their own models. Doshi emphasized that if they surpass Midjourney, they'll create a new benchmark.

[0:27](https://addtry.com/bc5c22b8-9e47-48fc-a66f-6f9b761cb0e7?t=27000)

### How does Suhail Doshi approach UX design for AI image generation tools?

Suhail Doshi sees current AI image tools as akin to MS-DOS command lines; Playground aims to be a visual, intuitive graphics editor. Features like preview rendering (using consistency models), canvas-based generation, and outpainting move beyond the text box. Inspired by the GPT-3 playground, they prioritize power and configurability for enthusiast users while gradually making the tool more accessible. The design evolves as model capabilities improve.

[0:49](https://addtry.com/bc5c22b8-9e47-48fc-a66f-6f9b761cb0e7?t=49000)

## Key moments

- **[0:00] Introduction**
  - [0:56] Suhail Doshi built an ML team at Mixpanel in 2015 that used logistic regression for churn prediction and anomaly detection
  - [3:09] Suhail Doshi's Mighty aimed to stream a browser from a data center to create a new kind of computer
  - [6:18] Mighty failed because most web apps are single-threaded JavaScript, limiting the ability to move compute
- **[9:00] AI Awakening**
  - [10:03] DALL-E 2's avocado chair in April 2022 was the first viral moment for generative AI, says Suhail Doshi
  - [10:24] Swyx argues Stable Diffusion had a bigger impact, while Suhail Doshi credits DALL-E 2's avocado chair as the first viral generative AI moment
  - [14:02] Stable Diffusion had only HuggingFace code and no UI, which Suhail Doshi saw as an opportunity for a graphics editor
  - [15:39] Suhail Doshi asked Aditya Ramesh and Sam Altman if OpenAI would build a DALL-E UI; they said no, focusing on language
  - [16:14] Suhail Doshi expects Playground's customer base to start with hobbyists and later move to B2B, similar to Mixpanel's progression
- **[17:58] Training From Scratch**
  - [17:58] Playground v2 was trained from scratch because Suhail Doshi says graphics is in a GPT-2 moment with limited utility
  - [19:27] Suhail Doshi praises the Stable Diffusion community as one of the most vibrant, comparing it to the Homebrew Computer Club
  - [20:15] Playground released pre-trained weights for Playground v2 to give back to the open-source community and help academia
  - [22:01] Q: What research does Suhail Doshi hope to see from the Playground v2 release? A: Character consistency, multitask editing, rotation coherence, style transfer, and reasoning about images
  - [23:01] Playground v2 is 2.5x preferred over Stable Diffusion XL on a 1K prompt benchmark, says Suhail Doshi
  - [23:47] Suhail Doshi attributes Playground v2's improvement over SDXL to data quality, caption cleaning, and filtering
- **[27:52] Aesthetic Benchmarks**
- **[30:59] Emergent Styles**
  - [31:22] Q: How do styles emerge in AI image models? A: The community tests the model and discovers styles, some are well-known like pixel art, others get invented names by the community, says Suhail Doshi
  - [32:35] Suhail Doshi's favorite community fine-tune is Starlight XL for its beautiful color contrast
  - [33:46] Suhail Doshi intentionally targeted a hobbyist toy to discover the most ambitious direction later
  - [36:00] Suhail Doshi says current inpainting quality is terrible, failing to convincingly change facial expressions like a smile
  - [38:03] Suhail Doshi says FID is a bad metric beyond a certain point and becomes irrelevant
  - [39:01] Suhail Doshi's favorite hard prompt is 'a horse riding an astronaut' which no model can do well
- **[43:18] Safety and Limits**
  - [43:18] Suhail Doshi says the ultimate text synthesis challenge is generating tiny text on a milk carton, not just large effects
  - [45:24] Suhail Doshi sees NSFW content as a major risk ahead of the US election, as deepfakes become harder to explain to non-technical people
  - [48:30] Suhail Doshi says the boundary between tasteful art with nudity and NSFW is subjective and hard to moderate
  - [49:12] Playground offers a canvas with seed selection, image expansion, and outpainting to move beyond simple text-to-image
- **[49:47] AI-First UX**
  - [51:30] Suhail Doshi says prompt engineering, like weighting words with parentheses, is a bug that shouldn't exist
  - [54:28] Playground introduced preview rendering using LCM to show an early image preview, reducing the slot-machine feel, says Suhail Doshi
  - [57:01] Suhail Doshi says Simu, who brought LoRAs to Stable Diffusion, collaborates with Playground
- **[57:13] LoRA Ecosystem**
  - [59:14] Suhail Doshi envisions a future where users pick a mood board of images to define a style without training LoRAs
- **[1:02:07] GPU Engineering**
  - [1:02:09] Suhail Doshi says running their own GPU infrastructure is necessary for cost control
  - [1:03:13] Suhail Doshi says training infrastructure software is very broken, requiring custom schedulers like those at xAI
  - [1:04:36] Suhail Doshi believes there are more modalities than vision, language, and audio, and a true physics foundation model is an unsolved question
- **[1:04:44] Lightning Round**
  - [1:07:22] Suhail Doshi recommends Karpathy's Zero to Hero YouTube series over any book for learning AI
  - [1:08:11] Suhail Doshi advises non-PhD founders to pick an exciting project and learn every detail while building, not just reading

## Speakers

- **Alessio** (host)
- **Swyx** (host)
- **Suhail Doshi** (guest)

## Topics

Image Generation, Diffusion Models, Compute

## Mentioned

HuggingFace (company), Midjourney (company), Mighty (company), Mixpanel (company), Mosaic (company), OpenAI (company), Playground AI (company), Replicate (company), Weights & Biases (company), ControlNet (product), DALL-E (product), MJHQ 30K (product), Playground (product), Stable Diffusion (product), fast.ai (product)

## Transcript

### Introduction

**Alessio** [0:01]
Hey everyone. Welcome to the Latent Space Podcast. This is Alessio, partner and CTO in residence at Decibel Partners, and I'm joined by my co-host Swyx, founder of Smol.ai.

**Swyx** [0:10]
Hey, and today in the studio we have Suhail Doshi. Welcome.

**Suhail Doshi** [0:13]
Yeah, thanks. Thanks for having me.

**Swyx** [0:14]
Among many things, you're a CEO and co-founder of Mixpanel.

**Suhail Doshi** [0:17]
Yep.

**Swyx** [0:17]
Uh, and, uh, I think about three years ago you s- you left to start Might- Mighty.

**Suhail Doshi** [0:22]
Mm-hmm.

**Swyx** [0:22]
Um, and more recently, I think about a year ago, uh, it transitioned into Playground. Uh, and you've just announced your, your, your new round. Uh, I'd just like to start touch on Mixpanel a little bit 'cause it's obviously like one of the more, uh, sort of successful, uh, analytics companies.

Uh, we previously, uh, had Amplitude on, um, and I'm curious if like you had any sort of, uh, reflections on like just your... the, that overall, the interaction of like that am- that am- that amount of data, um, that people would want to use for AI.

Like, uh, I don't know if, like, there's, there's still a part of you that stays in touch with that world.

**Suhail Doshi** [0:57]
Yeah. I mean, it's, I mean, you know, the short version is, is that, um, maybe back in like 2015 or '16, I don't really remember exactly 'cause it was a while ago, we had a m- ML team at Mixpanel and, um, I think this is like when maybe deep learning or something like really just started getting kind of exciting and we were thinking that maybe we, you know, given that we had such vast amounts of data, perhaps we could predict things.

So we built, you know, two or three different features. I think we built a feature where we could predict whether users would churn from your product. Uh, we made a feature that could predict whether users would convert. We tried to...

We built a feature that could do anomaly detection, like if something occurred in your product that was just very surprising, maybe a spike in traffic in a particular region, could we tell you in advan- could we tell you that that happened?

'Cause it's really hard to like know everything that's going on with your data. Could we tell you something surprising about your data? And we tried all of these various features. Most of it boiled down to just like, you know, using, uh, logistic regression, and it never quite seemed very groundbreaking in the end.

And so I think, um, you know, we had a four or five person ML team and, um, yeah, I think we never expanded it from there. And I did all these fast.ai courses trying to learn about ML and that was the, that's the-

**Swyx** [2:15]
That was the first time you did fast.ai.

**Suhail Doshi** [2:17]
Yeah, that was the first time I did fast.ai.

**Swyx** [2:18]
Oh.

**Suhail Doshi** [2:18]
Yeah, I think I've done it now three times maybe.

**Swyx** [2:21]
Oh, okay. I didn't-

**Suhail Doshi** [2:21]
Yeah

**Swyx** [2:21]
... realize the third. Okay. Um-

**Suhail Doshi** [2:23]
No, no, I just mean reviewing it-

**Swyx** [2:25]
Right

**Suhail Doshi** [2:25]
... is maybe three times, but yeah.

**Swyx** [2:26]
Yeah, yeah. Yeah. I mean, uh, um, I think you, you mentioned prediction, but honestly, like it's, um, also just about the feedback, right? The, the quality of feedback from, uh, from, from users, I think it's, uh, it's useful for anyone building AI applications.

**Suhail Doshi** [2:39]
Yeah.

**Swyx** [2:39]
Yeah. Self-evident.

**Suhail Doshi** [2:41]
Yeah, I think, I think I haven't spent a lot of time thinking about Mixpanel 'cause it's been a long time, but yeah, I wonder, I wonder now, given everything that's happened, like upon, you know, sometimes I, sometimes I'm like, "Oh, I wonder what we could do now," and then I kind of like move on to whatever I'm working on.

**Swyx** [2:54]
Yeah. No more-

**Suhail Doshi** [2:55]
But things have changed significantly since, so, uh, yeah.

**Swyx** [2:58]
Yeah. Awesome. Um, and then maybe we'll touch on Mighty a little bit. Uh, Mighty was very, very bold. Uh, it was basically... Well, m- my framing of it was, uh, you will run our browsers for us because-

**Suhail Doshi** [3:09]
Yep

**Swyx** [3:09]
... um, everyone has too many tabs open. I have too many tabs open and it's slowing down your machines, like you can do it better for us, uh, i- in a centralized data center.

**Suhail Doshi** [3:16]
Yeah. We were first trying to make a b- a browser, uh, that we would stream from a data center to your computer at extremely low latency. Um, and I... But the, the real objective wasn't trying to make a browser or anything like that.

The real objective was to try to make a new kind of computer, and the thought was just that like, you know, we have these computers in front of us today and we upgrade them or they run out of RAM or they don't have enough RAM or not enough disk or, you know, there's some limitation with our computers.

Uh, per- perhaps like data locality is a problem. Um, could we... You know, why, why do I need to think about upgrading my computer ever? And so, you know, we just had to kind of observe that like, well, actually it seems like a lot of applications are just now in the browser.

You know, it's like how many real desktop applications do we use relative to the number of applications we use in the browser? So there's just this realization that actually like, you know, the browser was effectively becoming more or less our operating system over time, and so then that's why we kind, kind of decided to go, "Hmm, maybe we can stream the browser."

Unfortunately, the idea did not work for, for a couple d- different reasons, but, uh, yeah, but the objective was try to make a new, new r- true new computer.

**Swyx** [4:21]
Yeah. Very, very bold. Very bold.

**Alessio** [4:23]
Yeah. And, um, I was there at, at YC Demo Day when you first announced it.

**Suhail Doshi** [4:26]
Oh, okay.

**Alessio** [4:27]
At, uh... It was, I think the last or one of the last in-person ones, so like the Pier 34 in-

**Suhail Doshi** [4:31]
Yes

**Alessio** [4:32]
... in Mission Bay.

**Suhail Doshi** [4:32]
Yeah.

**Alessio** [4:33]
Um-

**Suhail Doshi** [4:33]
Yeah, before COVID

**Alessio** [4:35]
... how, how do you think about that now when everybody wants to put some of these models in people's machines and some of them want to stream them in? Do you think there's maybe another wave of the same problem?

Before it was like browser apps too slow, now it's like model's too slow to run on device.

**Suhail Doshi** [4:49]
Yeah. I think, you know, we, you, we obviously pivoted away from Mighty, but a lot of what I somewhat believe- believed at Mighty is like still somewhat very true, what maybe why I'm so excited about AI and what's happening.

A lot of what Mighty was about was like moving compute somewhere else.

**Swyx** [5:06]
Mm.

**Suhail Doshi** [5:06]
Right? Right now, applications, they get limited quantities of memory, disk, uh, networking, whatever your home network has, uh, et cetera. You know, what, what if these applications could somehow, if we could shift compute and then these applications could have vastly more compute than they do today.

Uh, right now it's just like client backend services, but, um, you know, what if we could change the shape of how, how applications, uh, could interact with things? And it's changed my thinking. It... In some ways, AI is like a bit of a, a continuation of my belief that like perhaps we can really shift compute somewhere else.

One of the problems with, with Mighty was that, um, JavaScript is single-threaded in the browser, and what we learned, you know, the reason, reason why we kind of abandoned Mighty was because I didn't believe we could make a new kind of computer.

We could have made some kind of enterprise business, probably could made, could have made maybe a lot of money, but it wasn't going to be what I hoped it was going to be. And so once, once I realized that most Of a web app, it is just going to be single-threaded JavaScript, then the only thing you could do largely, uh, withstanding changing JavaScript, which, uh, is a fool's errand most likely, uh, is make a better pro- CPU.

Right? And there's like three CPU manufacturers, two of which sell the, you know, c- you know, big ones, you know, AMD, Intel, and then of course like Apple made the M1. And it's not like single-threaded CPQ- CPU core performance.

Single core performance was like very- increasing very fast. It's plateauing rapidly. And even these different like companies were not doing as good of a job, you know, sort of with the continuation of Moore's law. But what happened in AI was that you got like, like if you think of the AI model as like a computer program, like just like a compiled computer program, it is literally built and designed to do massive parallel computations.

And so, uh, if you could take like the universal approximation theorem to its like kind of logical, complete point, you know, you're like, "Wow, I can get- make computation happen really rapidly and parallel somewhere else."

**Swyx** [7:06]
Mm.

**Suhail Doshi** [7:07]
Um, you know, so you end up with these like really amazing models that can like do anything. It just turned out like perhaps, perhaps the new kind of computer, uh, would just simply be shifted, um, you know, into these like really amazing AI models in reality.

**Swyx** [7:23]
Yeah. Like, I think, uh, Andrej Karpathy has always been, has been making a lot of analogies with the LLM OS.

**Suhail Doshi** [7:29]
Yeah, I saw his, yeah, I saw his video and I, I watched that, you know, maybe two weeks ago or something like that, and I was like, "Oh man, this-

**Swyx** [7:35]
Yeah.

**Suhail Doshi** [7:35]
... I very much resonate with this like idea."

**Swyx** [7:37]
Why didn't I see this three years ago?

**Suhail Doshi** [7:39]
Yeah. I, I think, I think there still will be, you know, local models and then there'll be these very large models that have to be run in data centers. Um, yeah, I think it just depends on kind of like the right tool for the job like any, like any engineer, uh, would probably care about.

But I think that, uh, you know, by and large, like if, if the models continue to kind of keep getting bigger, you're just gonna, it's... You're always going to be wondering whether you, whether you should use the big thing or the small, you know, the tiny little model.

Um, and it, it might just depend on like, you know, do you need 30 FPS or 60 FPS? Um, maybe that would be hard, uh-

**Swyx** [8:10]
Yeah

**Suhail Doshi** [8:10]
... to do, you know, over, over, over a network.

**Swyx** [8:12]
Yeah. You, you tackled a much, uh, harder problem latency-wise, um, you know, than, than the AI models actually require.

**Suhail Doshi** [8:19]
Yeah.

**Swyx** [8:19]
Uh, so

**Suhail Doshi** [8:20]
Yeah, you can do quite well. You can do quite well. Uh, you know, we definitely did, um, 30 FPS video streaming.

**Swyx** [8:26]
Yeah.

**Suhail Doshi** [8:26]
Did, did very crazy things to make that work. So I, I'm actually quite bullish on the kinds of things you can do with networking.

**Swyx** [8:33]
Yeah. All right. Maybe someday you'll, uh, come back to that at some point. Um, but so for those that, for those that don't know, you're very transparent on Twitter. Uh, very good to follow you just to, just to learn your insights, and you actually published a postmortem on Mighty that people can read up-

**Suhail Doshi** [8:47]
Yep

**Swyx** [8:47]
... if they're willing to. Um, and so there was a bit of an overlap. Uh, you started exploring, um, the AI stuff in June 2022.

**Suhail Doshi** [8:56]
Mm-hmm.

**Swyx** [8:57]
Which is when you started saying like, "I'm taking fast AI again."

**Suhail Doshi** [8:59]
Mm-hmm.

**Swyx** [8:59]
Um, maybe, was there more context around that? Um...

### AI Awakening

**Suhail Doshi** [9:03]
Yeah. Uh, I think, I think I was kind of like waiting for the team at Mighty to finish up- ... you know, of something, and I was like, "Okay, well, what can I do? Uh, I guess I will, uh, make some kind of like address bar predictor in the browser."

So we had, we know we had forked Chrome and Chromium and, um, I was like, you know, one thing that's kind of lame is that like this browser should be like a lot better at predicting what I might do, where I might wanna go.

You know, it, it's, it struck me as really odd that, you know, Chrome had very little AI actually or ML inside this browser. And then for a company like Google, you'd think there's a lot. But it's actually like, it's actually just like the code is actually just very, you know, it's just a bunch of if/then statements is more or less the address bar.

So it seemed like a pretty big opportunity, and that's also where a lot of people interact with the browser. So, you know, long story short, I was like, "Hmm, I wonder what I could build here." So I started to, yeah, take some AI courses and try to get, take some- review the material again and get back to figuring it out.

But I think that was somewhat serendipitous because, um, right around April was, I think, a very big watershed moment in AI 'cause that's when DALL-E 2 came out, and I think that was the first like truly big viral moment, uh, for generative AI.

**Swyx** [10:18]
Because of the avocado chair.

**Suhail Doshi** [10:20]
Because of the avocado chair and, uh, yeah, exactly. Yeah. I mean, it was just so novel.

**Swyx** [10:26]
It wasn't as big, it wasn't as big for me as Stable Diffusion. Like-

**Suhail Doshi** [10:28]
Really?

**Swyx** [10:29]
Yeah. I don't know. DALL-E was like, all right, that's, that's cool. I don't know.

**Suhail Doshi** [10:33]
Yeah.

**Swyx** [10:33]
I mean, they, they had some flashy videos, but like I never really... Like, it, it didn't really register to me as like big-

**Suhail Doshi** [10:38]
Well, just, just that moment of images was just such a viral, novel moment. I think it just blew people's mind.

**Swyx** [10:44]
Yeah.

**Suhail Doshi** [10:44]
Um, and-

**Swyx** [10:46]
Yeah, I mean, that's the first time I like encountered Sam Altman because like they had this like DALL-E 2 hackathon, and open- they opened up the OpenAI office for developers to walk in, like back when, you know, it wasn't as, uh, I guess, uh, much of a security issue as it is-

**Suhail Doshi** [10:59]
I see

**Swyx** [10:59]
... today.

**Suhail Doshi** [11:00]
Yeah.

**Swyx** [11:00]
Maybe take us through like the, the journey to, to decide to pivot into, into this. And but, and, and also like choosing images. Obviously, you, you, you were inspired by DALL-E.

**Suhail Doshi** [11:08]
Yeah.

**Swyx** [11:09]
But there could be any number of, um, AI co- companies and businesses that you could start, like and why this one, right?

**Suhail Doshi** [11:17]
Yeah. Um, well-

**Swyx** [11:18]
So there must be an idea made from June to September.

**Suhail Doshi** [11:21]
Yeah, yeah. Yeah, there, there definitely was. So I think at that time, Mighty, we, we were, Mighty and, and OpenAI was, you know, not quite as popular as it is all of a sudden now these days. But back then, I think they were, they were like more than hap- They, it, they had a lot more bandwidth to like kind of help anybody.

And so, you know, we had, we had been talking with, um, the team there around like trying to see if we could do like really fast low latency, uh, address bar prediction with like GPT-3 and 3, 3.5, and that kind of thing.

And so, um, you know, we were sort of figuring out how, how could we make that low latency. Um, I think that just being able to talk to them and kind of being involved gave me a bird's eye view into a bunch of things that started to happen.

Um- You know, obviously first was the DALL-E team- DALL-E 2 moment, but then Stable Diffusion came out, and that, i- that was a big moment for me as well. And I remember just kind of like sitting up one night thinking...

I was like, "You know what, what are the kinds of companies one could build? Like, what matters right now?" I- one thing that I observed is that I find a lot of great-- I find a lot of inspiration when I'm working in a field in something, and then I can identify a bunch of problems.

Like for Mixpanel, I was an intern at a company, and I just noticed that they were doing all this data analysis, and so I thought, "Hmm, I wonder if I could make a product, and then maybe they would use it."

And in this case, you know, the same thing kind of occurred. It was like, okay, there are a bunch of like infrastructure companies that are, you know, uh, doing y- g- y- they put a model up, and then you can use their API, like Replicate is a really good example of that.

Um, there are a bunch of companies that are like helping you with training, model optimization, um, Mosaic at the time, uh, I mean, and probably still, you know, was doing stuff like that. So I just started listing out like every category of everything, of every company that was doing something interesting.

Obviously, Weights & Biases. Um, I was like, "Oh man, Weights & Biases is like this great company. If, uh, do I wanna compete with that company? I might be really good at competing with that company because of Mixpanel, because it's so much of like analysis."

Um, but I was like, "No, I don't wanna do anything related to that. That would-- I think that would be too boring now at this point." Um, but, uh, um, so I started to list out all these ideas, and one thing I observed was that at, at OpenAI, they had like a playground for GPT-3.

**Swyx** [13:35]
Mm-hmm.

**Suhail Doshi** [13:35]
Right? And all it was is just like a text box more or less, and then there were some settings on the right, like temperature and whatever.

**Swyx** [13:41]
Top K, top N.

**Suhail Doshi** [13:42]
Yeah, top K. Uh, you know, what's your end stop sequence. I mean, that was like their product before ChatGPT. You know, really difficult to use but fun if you're like an engineer. And I just noticed that their product kind of was evolving a little bit, where the interface kind of was getting more and more, a little bit more complex.

They had like a way where you could like generate something in the middle of a sentence and all those-

**Swyx** [14:03]
Mm

**Suhail Doshi** [14:03]
... kinds of things. And I just thought to myself, I was like, you know, there's not-- everything is just like this text box, and you generate something, and that's about it. And Stable Diffusion had kinda come out, and it was all like hugging face and code.

Nobody was really building any UI. And so I had this kind of thing where I wrote prompt dash like question mark in my notes, and I didn't know what was like the product for that at the time.

**Swyx** [14:25]
Mm-hmm.

**Suhail Doshi** [14:25]
I mean, it seems kind of trite now. Um, but yeah, I just like wrote "Prompt, what's the thing for that?"

**Swyx** [14:30]
Manager prompt.

**Suhail Doshi** [14:31]
Prompt manager.

**Swyx** [14:32]
Tool.

**Suhail Doshi** [14:32]
Do you organize them? Like, do you like have a UI that can like-

**Swyx** [14:36]
Library

**Suhail Doshi** [14:37]
... play with them? Yeah, like a library. What, what would you make?

**Swyx** [14:40]
Yeah.

**Suhail Doshi** [14:40]
Uh, and so then of course then you thought about what would the modalities be given that, how would you build a UI for each kind of modality. Uh, and so there are a couple people working on some pretty cool things.

Um, and, uh, and I, and I ch- and I basically chose graphics because it seemed like the most obvious place where you could build a really powerful, complex UI that's not just only typing in a box, that y- y- it would very much evolve beyond that.

Like what would be the best thing for something that's visual? Probably something visual. Um, so yeah, I, I think that just that progression kind of happened, and it just seemed like there was a lot of effort going into language but not a lot of effort going into graphics.

And then the l- and maybe the very last thing was I, I think, um, I was talking to Aditya Ramesh, who was the co-f- co-creator of DALL-E 2 and Sam, and I just kind of w- went to these guys, and I was just like, "Hey, are you gonna make like a UI for this thing?

Like a true UI. Are you gonna go for this? Are you gonna make a product?"

**Swyx** [15:42]
For DALL-E? Yeah.

**Suhail Doshi** [15:43]
For DALL-E, yeah.

**Swyx** [15:44]
Yeah.

**Suhail Doshi** [15:44]
Uh, are you gonna do anything here? Um, 'cause if you're not gonna do it... If you are gonna do it, just let me know, and I will stop, and I will go-

**Swyx** [15:51]
Yeah

**Suhail Doshi** [15:51]
... do something else.

**Swyx** [15:51]
Yeah.

**Suhail Doshi** [15:52]
But if you're not gonna do anything, I'll just do it.

**Swyx** [15:54]
Mm-hmm.

**Suhail Doshi** [15:54]
Um, and so we had a couple conversations around what, what, what, what that would look like. Um, and then I think ultimately they decided that they were gonna focus on language primarily. Um, and, uh, yeah, I just felt like it was gonna be very under-invested in.

**Swyx** [16:08]
Yes, there, there is, there's that sort of, um, under-investment from OpenAI-

**Suhail Doshi** [16:12]
Mm-hmm

**Swyx** [16:12]
... which I, I, I can see that.

**Suhail Doshi** [16:14]
Mm-hmm.

**Swyx** [16:14]
Um, but also it's a different type of customer-

**Suhail Doshi** [16:18]
Mm-hmm

**Swyx** [16:18]
... than you're used to. Uh, presumably, you know, and Mix- Mixpanel are very sel- very good at selling to B2B-

**Suhail Doshi** [16:23]
Right

**Swyx** [16:23]
... developers. Uh, with Figure, and you're not.

**Suhail Doshi** [16:26]
Yeah.

**Swyx** [16:27]
Was, was that, was that not a concern?

**Suhail Doshi** [16:29]
Well, not, not so much because I think that, um, you know, right now I would say graphics is in this very nascent phase. Like most of the customers are just like hobbyists, right?

**Swyx** [16:37]
Yeah.

**Suhail Doshi** [16:38]
Like they-- It's a little bit of like a novel toy-

**Swyx** [16:40]
Yeah

**Suhail Doshi** [16:40]
... as opposed to being this like very high utility thing. But I think ultimately if you believe that you could make it very high utility, then probably the next customers will end up being B2B. It'll, it'll probably not be like consumer.

Like there are-- There will certainly be a variation of this idea that's in consumer. If you- if your quest is to kind of make like a, a super, uh, something that surpasses, um, human ability for graphics, like ultimately it will end up being used for business.

**Swyx** [17:06]
Yeah.

**Suhail Doshi** [17:07]
So I, I think it's maybe more of a progression. In fact, for me, it's maybe more like Mixpanel started out as SMB, and then very much like ended up starting to grow up towards enterprise. So for, for me, it's a little...

It's-- I think it will be a very similar progression.

**Swyx** [17:18]
Yeah. Yeah.

**Suhail Doshi** [17:19]
Yeah. But yeah, I mean, the, the reason why I was excited about it is 'cause it was a creative tool. I make music, and, uh, it's AI. It's like something that I could-- I know I could stay up till three o'clock in the morning doing.

Those are kind of like very simple bars for me.

**Swyx** [17:33]
Yeah.

**Suhail Doshi** [17:33]
Yeah.

**Swyx** [17:34]
It's good decision criteria.

**Alessio** [17:36]
Um, so you mentioned DALL-E, Stable Diffusion. You just had Playground v2-

**Suhail Doshi** [17:41]
Yep

**Alessio** [17:41]
... come out two days ago?

**Suhail Doshi** [17:42]
Yeah, two days ago. Yeah.

**Alessio** [17:43]
Two days ago. So this is a model you trained completely from scratch, so it's not a cheap fine-tune on, on something. You open source everything, including the weights. Um, why did you decide to do it? I know you supported Stable Diffusion XL in Playground before, right?

### Training From Scratch

**Suhail Doshi** [17:58]
Yep.

**Alessio** [17:59]
Um, yeah. What, what made you want to come up with v2 and maybe some of the interesting, you know, technical research work you've done?

**Suhail Doshi** [18:07]
Yeah, so I think, I think that we, we continue to feel like graphics and these foundation models for, uh, anything really related to pixels, but also definitely images continues to be very under-invested. It feels a little like graphics is in like this GPT-2 moment, right?

Like even GPT-3, even when GPT-3 came out, it was exciting, but it was like, what are you gonna use this for? You know, yeah, we'll do some text classification and some semantic analysis, and maybe it'll sometimes like make a summary of something and it'll hallucinate.

But no one really had like a very significant like business application for GPT-3. Um, and in images we're kind of stuck in the same place. We're kind of like, okay, I write this thing in a box and I get some cool piece of artwork and the hands are kind of messed up and sometimes the eyes are a little weird.

Uh, may- maybe I'll use it for a blog post, you know, that kind of thing. The utility feels so limited and so, you know, and then we... You sort of look at Stable Diffusion and, and we definitely use that model in our product and our users like it and use it and love it and, and enjoy it, but it hasn't gone nearly far enough.

So we, we were kind of faced with the choice of, you know, do we wait for progress to occur or do we make that progress happen? Uh, so yeah, we, we kind of embarked on a plan to just decide to go train these things from scratch, and I think the community has given us so much.

The, the community for Stable Diffusion, I think is one of the most vibrant communities on the internet. It's like amazing. It feels like if the-- I hope this is what like Homebrew Club felt like when computers like showed up because it's like amazing what that community will do and it moves so fast.

I've never seen anything in my life where so far, and heard other people's stories around this, where a research- an academic research paper comes out and then like two days later someone has sample code for it, and then two days later there's a model, and then two days later it's like in nine products.

**Alessio** [19:58]
Yeah.

**Suhail Doshi** [19:58]
You know, they're all competing with each other.

**Alessio** [20:00]
Yeah, yeah.

**Suhail Doshi** [20:00]
It's incredible to see like math symbols on a academic paper go to-

**Alessio** [20:04]
Yeah

**Suhail Doshi** [20:04]
... features, well-designed features in a product. So, um, I think the community has done so much. So I think we wanted to give back to, to the community kind of on our way. We knew, we knew it wasn't gonna be...

It-- We knew it was not ever gonna be-- Certainly we would train a better model than, than what we, what we gave out on Tuesday, but we definitely felt like, uh, there needs to be some kind of progress in these open source models.

The last kind of milestone was in July when Stable Diffusion XL came out, but there hasn't been anything really since, right?

**Alessio** [20:35]
And there's, uh, XL Turbo now.

**Suhail Doshi** [20:36]
Well, XL Turbo is like this distilled model, right?

**Alessio** [20:39]
Yeah.

**Suhail Doshi** [20:39]
So it's like lower quality but fast, so you have to decide, you know, what your trade-off is there.

**Alessio** [20:44]
Uh, and, and it's also a consistency model?

**Suhail Doshi** [20:46]
It's not-

**Alessio** [20:47]
I'm not sure.

**Suhail Doshi** [20:47]
I don't think it's a consistency model.

**Alessio** [20:48]
Sure. Okay.

**Suhail Doshi** [20:48]
It's like it's they, they did like a different thing.

**Alessio** [20:51]
Yeah.

**Suhail Doshi** [20:51]
Yeah. I think it's like... I don't, I, I don't wanna get quoted for this, but it's like some- something called ad, like adversarial something or another.

**Alessio** [20:56]
That's exactly right. Yeah.

**Suhail Doshi** [20:58]
Um-

**Alessio** [20:58]
Yeah

**Suhail Doshi** [20:59]
... yeah, I think it's, uh, it's-- I've read something about the, maybe it's like closer to GANs or something, but I didn't really read the p- the full paper. But, but yeah, there hasn't been quite enough progress in terms of, you know, there's no multitask image model.

You know, the closest thing would be something called like EmuEdit, but there's no model for that.

**Alessio** [21:14]
Mm-hmm.

**Suhail Doshi** [21:14]
It's just a paper that's within Meta. So, um, we did that and we also gave out, uh, pre-trained weights, which is very rare. Um, usually you just get the aligned model and then you have to like see if you can do anything with it.

We actually gave out, um... There's like a 256 p- uh, pixel pre-trained stage and a 512, and we did that for academic research 'cause there's a whole bunch of... You know, we come across people all the time in academia and they have like, they have access to like one A100 or eight at best.

Uh, and so if we can give them kind of like a 512 pre-trained model, it might m- our, our hope is that there'll be interesting novel research that occurs from that. Um-

**Alessio** [21:52]
What research do you want to happen?

**Suhail Doshi** [21:54]
I would love to see more research around, uh, you know, things that users care about tend to be things like character consistency.

**Alessio** [22:01]
Uh, between frames?

**Suhail Doshi** [22:02]
Um-

**Alessio** [22:02]
For video

**Suhail Doshi** [22:02]
... more like if you had like a face.

**Alessio** [22:04]
Yeah.

**Suhail Doshi** [22:04]
Yeah, yeah. Basically between frames, but more just like, you know, you have your face and it's in, you know, one image and then you want it to be like in another.

**Alessio** [22:11]
And-

**Suhail Doshi** [22:11]
And users are very particular and sensitive to faces changing 'cause we know, we know what, you know-

**Alessio** [22:17]
The faces

**Suhail Doshi** [22:17]
... we're, we're trained on faces-

**Alessio** [22:18]
Yeah.

**Suhail Doshi** [22:18]
... as humans. Um, and you know, that's something, um, I don't- I'm not seeing a lot of innovation, enough innovation around multitask editing. You know, there are two things like Instruct Pix2Pix and then the EmuEdit paper that are maybe very interesting, um, but, uh, we, we certainly are not pushing the fold on that in that regard.

Um, yeah, yeah just all kinds of things, uh, like around that. Rotation, um, you know, being able to keep coherence across images, style transfer is still very limited. Um, just even reasoning around images, you know, what's going on in an image-

**Alessio** [22:53]
Mm-hmm

**Suhail Doshi** [22:54]
... that kind of thing. Uh, things are still very, very, uh, underpowered, very nascent, so therefore the, the l- utility is very, very limited.

**Alessio** [23:01]
Mm-hmm. On the 1K prompt benchmark, you are 2.5x prefer-

**Suhail Doshi** [23:05]
Yep

**Alessio** [23:05]
... to Stable Diffusion XL. How do you get there? Is it better images in the training corpus? Is it... Yeah, can you, can you maybe talk through the improvements in the model?

**Suhail Doshi** [23:16]
I think we're still very early on in the recipe, um, but I think it's a lot of like little things and, you know, every now and then there are some big important things. Like certainly, uh, your data quality is really, really important, so we spend a lot of time, uh, thinking about that.

Um, but it-- I would say it's a lot of, lot of things that you kind of clean up along the way as you train your model, everything from captions to the data that you align with, um, after pre-train, to how you're picking your data sets, um, how you filter your data sets.

Uh, there's a lo- I feel like there's a lot of work in AI that's like doesn't really feel like AI, it just really feels like-

**Alessio** [23:54]
Mm

**Suhail Doshi** [23:54]
... just data set filtering and systems engineering and just like, you know, and the recipe is all there, but it's like a lot of extra work to do that. Um, so I think these, these models, I think, I think whatever version, I think we, we plan to do a Playground v-, uh, 2.1 maybe either by the end of the year or early next year, and we're just like watching what the community does with the model.

And then we're just gonna take a, a lot of the things that they're unhappy about and just like fix them. Um, you know, so for example, like maybe the eyes of people in an image don't feel right. They feel like they're a little misshapen or they're kind of blurry feeling.

That's something that we already know we wanna fix, so I think in that case it's gonna be about data quality. Um, or maybe we wanna improve the kind of the dynamic range of color. You know, we wanna make sure that that's like got a good range in any image, so what technique can we use there?

There's different things like offset noise, pyramid noise, uh, terminal zero SNR. Like there are all these various interesting things that you can do. So I think it's like a lot of just like tricks. Some are tricks, some are data, and some is just like cleaning.

Yeah.

**Swyx** [24:57]
If specifically for faces, it's very common to use a pipeline rather than just train, train the base model more.

**Suhail Doshi** [25:05]
Mm-hmm.

**Swyx** [25:06]
Uh, do you have a strong belief either way on like, oh, they should be separated out to different stages for like improving the eyes, improving the face or-

**Suhail Doshi** [25:13]
Mm-hmm

**Swyx** [25:13]
... even the hands or whatever, or do you think like it can all be done in one model?

**Suhail Doshi** [25:18]
I think we will make a unified model.

**Swyx** [25:19]
Okay.

**Suhail Doshi** [25:20]
Yeah, I think it will, I think we'll certainly in the end ultimately make a unified model. Um, you know, i- there, there, there's not enough research about this. There's pro- maybe there is something out there that we haven't read.

There are some bottlenecks, like for example, in the VAE. Um, like the VAEs are ultimately like compressing these things, and so you don't know, and then you might have like a big informational, information bottleneck. Or so maybe you would use a pixel-based model perhaps.

Um, you know, there's a lot of belief, I think we've talked from, to people, everyone from like Rombach to various people. Uh, Rombach trained Stable Diffusion. You know, th- there's, I think there's like a big question around the architecture of these things.

It's still kind of unknown, right? Like we've got transformers and we've got g- like a GPT architecture model, but then there's this like weird thing that's also seemingly working with diffusion. And so we, you know, are we gonna use vision transformers?

Are we gonna move to pixel-based models? Is there a different kind of architecture? We don't really-- I don't think there have been enough experiments in this area.

**Swyx** [26:19]
Still? Oh my God.

**Suhail Doshi** [26:21]
Yeah.

**Swyx** [26:22]
That's surprising.

**Suhail Doshi** [26:23]
Yeah. I think it's very computationally expensive to do a pipeline model where you're like fixing the eyes and you're fixing the mouth and you're fixing the hands.

**Swyx** [26:29]
That's what everyone does, as far as I understand.

**Suhail Doshi** [26:31]
Well, I, I'm not sure, I'm not exactly sure what you mean, but if you mean like you get an image and then you will like make another model specifically to fix a face, yeah, I don- I think that's very computationally-- that's fairly computationally expensive, and I think it's like not probably not the right thi- right way.

**Swyx** [26:45]
Yeah.

**Suhail Doshi** [26:46]
Yeah, and it, it doesn't generalize very well.

**Swyx** [26:48]
It doesn't.

**Suhail Doshi** [26:48]
Now you have to pick all these different things.

**Swyx** [26:49]
It's, yeah, you're just kind of glomming things on together.

**Suhail Doshi** [26:51]
Yeah.

**Swyx** [26:51]
But like when I look at AI artists, like that's what they do, so.

**Suhail Doshi** [26:54]
Ah, yeah, yeah, yeah. They'll do things like, you know, um, I think a lot of ARs will do, you know, control net tiling to do kind of generative upscaling-

**Swyx** [27:02]
Yeah. Yeah

**Suhail Doshi** [27:02]
... of all these different pieces of the image. Yeah, I mean, to me these are all just like, they're all hacks. Yeah, ultimately in the end. I mean, it just, to me it's like let's go back to where we were just three years, four years ago with where deep learning was at and where language was at.

You know, it's the same thing. It's like we were like, "Okay, well, I'll just train these very narrow models to try to do these things and kind of ensemble them or pipeline them to try to get to a best-in-class result" and, uh, and, and here we are with like where the, the models are gigantic and like very capable of solving huge amounts of tasks, uh, when given like lots of great data.

**Swyx** [27:36]
Yeah.

**Suhail Doshi** [27:36]
So yeah.

**Swyx** [27:37]
Makes sense.

**Alessio** [27:38]
Um, you also released a new benchmark called MJHQ 30K for automatic evaluation of a model's aesthetic quality.

**Suhail Doshi** [27:46]
Mm-hmm.

**Alessio** [27:46]
Um, I have one question. Um, the data set that you used for the benchmark is from Midjourney.

### Aesthetic Benchmarks

**Suhail Doshi** [27:52]
Yes.

**Alessio** [27:52]
You have-

**Suhail Doshi** [27:53]
Yeah

**Alessio** [27:53]
... ten categories. Um, how do you think about the Playground model, Midjourney?

**Suhail Doshi** [27:59]
You know, there are a lot of people-- a lot of people in research like to come up with, uh, they like to compare themselves to something they know they can, uh, beat.

**Alessio** [28:06]
Mm.

**Suhail Doshi** [28:06]
Right? But, um, a- and maybe this is the best reason why, uh, it's, it can be helpful to not be a researcher also sometimes. Like I'm not, I'm not like trained as a researcher. I don't have a PhD in anything AI related, for example.

Um, but I, I think if you care about products and you care about your users, then the, the most important thing that you wanna figure out is like we, you know, we have-- everyone has to acknowledge that Midjourney is very good.

You know? They're, they are the best at this thing. We would, I would happily, I'm happy to admit that. I have no, no problem admitting that. Uh, it's just, it's just easy. It's very visual, um, to tell. So, you know, I think it's incumbent on us to try to compare ourselves to the thing that's best, even if we lose, even if we're not the best, right?

And, um, you know, at some point, if we are able to surpass Midjourney, then we, you know, we only have ourselves to compare ourselves to. But on first blush, you know, I think it's worth comparing yourself to maybe the best thing and try to find like a really fair way of, um, of doing that.

So I think, I think more people should try to do that. I definitely don't think you should be kind of comparing yourself on like some Google model or some old SD, you know, Stable Diffusion model and be like, "Look, we beat, you know, Stable Diffusion one point five."

I think, I think users uh, you know, ultimately want care, you know, how close are you getting to the thing that like I, I also, you mostly, you- people mostly agree with. So we put out that benchmark not because, and, and for no other reason to say like this seems like a worthy thing for us to at least try, you know, for people to try to get to.

Uh, and then if we surpass it, great, we'll come up with another one.

**Alessio** [29:40]
Yeah. No, that's awesome. And, um, you killed Stable Diffusion XL and everything. Um, in the benchmark chart, it says Playground V2 1024 pixel dash aesthetic.

**Suhail Doshi** [29:52]
Yes.

**Alessio** [29:52]
Do you have, um, kind of like, yeah, style fine tunes or like what's the dash aesthetic for?

**Suhail Doshi** [29:57]
Yeah. We, we debated this. Maybe we named it wrong or something, but we were like how do we help people realize, uh, you know, the model that's, that's aligned versus the models that weren't? So because, because we gave out pre-trained models, we didn't want people to like use those.

Um, so we-- that's why they're called base. And then the aesthetic model, yeah, we wanted people to pick up the thing that we thought would be like the thing that makes things pretty. Um, who wouldn't want the thing that's aesthetic?

**Alessio** [30:23]
Mm-hmm.

**Suhail Doshi** [30:23]
But if, if there's a better name Yeah, we're, we're-- we definitely are open to feedback

**Swyx** [30:27]
No, no, that's cool. Um, I was using the product. You also have the style filter-

**Suhail Doshi** [30:31]
Uh-huh

**Swyx** [30:32]
... and you have all these different style.

**Suhail Doshi** [30:33]
Yeah.

**Swyx** [30:33]
And it seems like the styles are tied to the model, so there's some, like, SDXL styles-

**Suhail Doshi** [30:39]
Yeah

**Swyx** [30:39]
... there's some Playground v2 styles. Um, can you maybe give listeners a overview of how that works? Because in, in language, there's not this idea of, like, style, right?

**Suhail Doshi** [30:49]
Right.

**Swyx** [30:50]
Versus like in, in vision model there is, and you cannot get certain styles in different models. Um-

**Suhail Doshi** [30:55]
Mm-hmm

**Swyx** [30:56]
... how do styles emerge and how do you categorize them and find them?

### Emergent Styles

**Suhail Doshi** [30:59]
Yeah, I mean it's, it's so fun having a community where they-- people are just trying a model. Like we- we-- it's only been two days for Playground v2, and we actually don't know what the model's capable of and not capable of.

You know, we certainly see problems with it, but we have yet to see, uh, what emergent behavior is. And we, we've just sort of discovered that it takes about like a week before you start to see like new things.

But uh, I think like a lot of that style kind of emerges after that week, where you start to see, you know... You know, there's some styles that are very like well known to us, like maybe like pixel art is a well-known style.

But then there's some style-- you know, photorealism is like another one that's like well known to us. But there are some styles that cannot be easily named. Uh, you know, it's not as simple as like, okay, that's an anime style.

**Swyx** [31:46]
Yeah.

**Suhail Doshi** [31:46]
Uh, it's very visual. And, and in, and in the end you end up making up the name, uh, for what that style represents. And so the community kind of shapes itself around these different things. And so if you-- if anyone that's in- into stable diffusion and into any- building anything with graphics and stuff with these models, you know, you might have heard of like Protovision or DreamShaper- ...

some of these weird names. Uh, but they're just, you know, invented by these authors, but they have a sort of je ne sais quoi that, you know, appeals to users. Um, yeah, there's this-

**Swyx** [32:16]
Because it like roughly embeds to what you, what you want.

**Suhail Doshi** [32:19]
I, it- I guess so. I mean it's like, you know, there's this, uh, one of my favorite ones that's fine-tuned, it's not made by us, it's called like Starlight XL, um, it's just this beautiful model. It's got really great color contrast and, and visual elements, and the users love it.

I love it. And yeah, it's, it's so hard. I think that's like a very big open question with graphics that I'm not totally sure how we'll solve. Um, yeah, I think a lot of styles are sort of... I don't know, it's, it's like an evolving situation too, 'cause styles get boring-

**Swyx** [32:51]
Mm

**Suhail Doshi** [32:51]
... right? They get fatigue. Like it's like listening to the same style of pop song. I kind of-- I try to relate to graphics a little bit like with music because I think it gives you a little bit of a different shape to things.

Like in music it's not just-- it's not as if we just have pop music and-

**Swyx** [33:07]
Mm

**Suhail Doshi** [33:07]
... you know, rap music and country music. Like there-- all of these-- Like the EDM genre alone has like subgenres. And I think that's very true in, in graphics and painting and art and anything that we're doing. There's just these subgenres even if we can't quite always name them.

**Swyx** [33:22]
Yeah.

**Suhail Doshi** [33:23]
Uh, but I think they are emergent from the community which is why we're so always happy to work with the community.

**Swyx** [33:27]
Yeah. That is a struggle, you know, coming back to this like B2B versus B2C thing. Uh, B to- B2C you're gonna have a huge amount of diversity and then it's gonna reduce as you get towards more sort of B2B type use cases.

I, I, I'm making this up here. I'm-

**Suhail Doshi** [33:40]
Yeah, yeah

**Swyx** [33:40]
... tell me if you disagree. Um, so like you might be optimizing for a thing that you m- may eventually not need.

**Suhail Doshi** [33:46]
Yeah, possibly. Yeah, possibly. Um, yeah, I try not to share-- I think like a simple thing with startups is that I worry sometimes by like, like by trying to be, uh, overly ambitious and like really scrutinizing like what something is in its most nascent phase that you miss the most ambitious thing you could have done.

Like just having like very basic curiosity-

**Swyx** [34:07]
Okay

**Suhail Doshi** [34:07]
... um, with something very small, um, can like ki- kind of lead you to something amazing. Like Einstein definitely did that and then when-- and then he like, you know, he basically won all the prizes and- ... got everything he wanted and then basically did like kind-- didn't really-

**Swyx** [34:21]
Nothing Yeah. He was done.

**Suhail Doshi** [34:22]
He kind of dismissed quantum-

**Swyx** [34:24]
Yeah.

**Suhail Doshi** [34:24]
... uh, and then just kind of was still searching, you know, for the unifying theory and he like had this quest, and I think that happens a lot with like Nobel Prize people. I think there's like a, a term for it that I forget.

Um, I actually wanted to go after a toy almost intentionally.

**Swyx** [34:39]
Huh.

**Suhail Doshi** [34:40]
Um, so long as that I could see, I could imagine that it would lead to something, uh, very, very large later. And so yeah, it's a very-- like I said, it's very hobbyist but you need to start somewhere.

You need to start with something that there ha- has a m- big gravitational pull, um, even if these hobbyists are, aren't likely to be the people that you know ha- have a way to monetize it or whatever. Even if they're-- But they're doing it for fun so there's something, something there that I think is really important.

But I agree with you that, you know, in time, uh, we're gonna have to foc-- we, we will absolutely focus on, um, more u- utilitarian things, like things that are more related to editing feats that are much harder but...

And so I think like a very simple use case is just, you know, I'm not a graphics designer. Um, I don't know if, I don't know if you guys are.

**Swyx** [35:27]
Mm-mm.

**Suhail Doshi** [35:29]
Um, but it su- you know, it, it seems like very simple that like you-- if we could give you the ability to do really complex graphics w- without skill, wouldn't you want that? You know, like my wife the other day was said, you know, said, "Ah, I wish Playground was better because I wish that...

You know, don't you-- D- d-- When are you guys ha- gonna have a feature where like we could make my son, his name's Devin, smile when he was not smiling in the picture for the holiday card?" Right? You know, just being able to highlight his, his mouth and just say like, "Make him smile."

Like why can't we do that with like high fidelity and coherence?

**Swyx** [36:00]
Mm-hmm.

**Suhail Doshi** [36:00]
Uh, little things like that all the way to, um, you know, uh, putting you in completely different scenarios.

**Swyx** [36:06]
Is that true? Can we not do that in painting?

**Suhail Doshi** [36:09]
You can do in painting but it's-- the quality is just so bad. Yeah.

**Swyx** [36:14]
Oh.

**Suhail Doshi** [36:15]
It's just really terrible quality. You know, it's, it's like you'll do it five times and it'll still like kinda look like crooked or just the artifact.

**Swyx** [36:22]
Mm.

**Suhail Doshi** [36:22]
Part of it's like, you know, the lips on the face are so-- there's such, it gi- there's such little information there. It's so small that the models really struggle with it. Yeah.

**Swyx** [36:31]
Make the picture smaller and you won't see it.

**Suhail Doshi** [36:32]
Well, I think, I think one-

**Swyx** [36:33]
That's my trick. I don't know.

**Suhail Doshi** [36:35]
Well, uh, yeah, yeah, that's true. Or y- you know, you could take that region and make it re- get really big and then like say it's a mouth and then like shrink it. It-

**Swyx** [36:41]
Yeah

**Suhail Doshi** [36:41]
... it feels like you're wrestling with it, um, more than it's doing something that kind of-

**Swyx** [36:46]
Yeah

**Suhail Doshi** [36:46]
... uh, surprises you. Yeah.

**Swyx** [36:48]
It, it feels like you are very much the internal tastemaker, like you carry in your head this vision for what a, a good art model should look like.

**Suhail Doshi** [36:56]
Mm-hmm.

**Swyx** [36:57]
Um, is it-- do you find it hard to like communicate it to like your team and, and, you know, other, other people just because it's obviously it's, it's hard to put into words like we just said.

**Suhail Doshi** [37:06]
Yeah. It's, uh, it's very hard to explain, uh, like images have such, like such high bit rate compared to just words.

**Swyx** [37:15]
Yeah.

**Suhail Doshi** [37:16]
And word-- we don't have enough words to describe-

**Swyx** [37:18]
Yep

**Suhail Doshi** [37:18]
... um, these, these things. Difficult. I think everyone on the team, if, if they don't have good kind of like judgment taste or like an eye for some of these things, they're like subtly building it because they have no choice.

Right? So in that realm, I don't worry too much-

**Swyx** [37:33]
Sure

**Suhail Doshi** [37:33]
... actually. Like everyone is kind of like l- learning, uh, to, to get the eye, is what I would call it. But I also have, you know, my own narrow taste, like I'm at my, you know, I'm not-- I don't represent the whole population either.

**Swyx** [37:45]
True, true.

**Suhail Doshi** [37:46]
So...

**Swyx** [37:46]
Um, y- when, when you benchmark models, you know, like this benchmark we're talking about, we use FID, uh-

**Suhail Doshi** [37:52]
Yeah

**Swyx** [37:53]
... Fisher Input Distance. Um, okay, that's one, one measure, but like doesn't capture anything you just said about smiles.

**Suhail Doshi** [38:00]
Yeah. FID, FID is, FID is generally a bad metric. Um-

**Swyx** [38:03]
So-

**Suhail Doshi** [38:03]
You know, it's good up, up to a point, and then it kind of like is irrelevant.

**Swyx** [38:07]
Yeah.

**Suhail Doshi** [38:07]
Yeah.

**Swyx** [38:07]
And then so w- are there any other metrics that you like, um, apart from vibes? I'm, I'm always looking for alternatives to vibes.

**Suhail Doshi** [38:14]
Apart from vibes.

**Swyx** [38:14]
Because vibes don't scale, you know?

**Suhail Doshi** [38:16]
You know, it might be fun to kind of talk about this, um, because it's actually kind of fresh. So up till now, we haven't needed to do a ton of like benchmarking because it's-- we hadn't trained our own model, and now we have.

So now what? What does that mean? How do we evaluate it? You know, we're kind of like living with the last forty-eight, seventy-two hours of going, "Did the way that we benchmark actually succeed? Did it deliver?"

**Swyx** [38:38]
Yeah.

**Suhail Doshi** [38:38]
Right? You know, like I think Gemini just came out. They just put out a bunch of benchmarks, but all these benchmarks are, are just an approximation of how you think it's going to end up with real world performance, and I think that's like very fascinating to me.

Um, so if you fake that benchmark, you'll, you'll still end up in a really bad scenario at the, at the end of the day. And so, you know, what, one of the benchmarks we did was we did a-- we kind of curated like a thousand prompts.

That's what, that's what we published in our blog post, you know. Of all these tasks that we-- a lot of the-- some of them are curated by our team where we know the models all suck at it. Like my favorite prompt that no model is really capable of is a, a horse riding an astronaut.

**Swyx** [39:16]
Yep.

**Suhail Doshi** [39:16]
The inverse one, and it's really, really hard-

**Swyx** [39:19]
Yep

**Suhail Doshi** [39:19]
... to do. Um-

**Swyx** [39:20]
Not in data.

**Suhail Doshi** [39:21]
You know, another one is like a giraffe underneath a microwave. How does that work? Right. There's so many of these little funny ones. We do-- we have prompts that are just like misspellings of things.

**Swyx** [39:32]
Yeah.

**Suhail Doshi** [39:33]
Right? Just to see if the models will figure it out. Uh, so sp-

**Swyx** [39:35]
That's easy. That's, that should, uh, embed to the same space.

**Suhail Doshi** [39:39]
Yeah.

**Swyx** [39:39]
Yeah.

**Suhail Doshi** [39:39]
And, and, and, and just like all these very interesting, weird, weirdo things. And so we have so many of these, and then we kind of like evaluate whether the models are any good at it, and the reality is that they're all bad at it, and so then you're just picking the most aesthetic image.

But, uh, but I think, you know, we're just-- we're still at the beginning of building like our, uh, like the best benchmark we can that aligns most with just user happiness-

**Swyx** [40:01]
Mm-hmm

**Suhail Doshi** [40:01]
... I think. 'Cause we're not, we're not like putting these in papers and trying to like win, you know, I don't know, awards at ICCV or something, if they have awards. Sorry if they don't. Um, and um-

**Swyx** [40:11]
You could. Well, that's absolutely a valid strategy.

**Suhail Doshi** [40:13]
Yeah, y- you could. I don't think it would correlate necessarily with the impact we want to have on humanity. I think we're still evolving whatever our benchmarks are. So the first benchmark was just like very difficult tasks that we know the models are bad at.

Can we come up with a thousand of these? Um, whether they're hand-rated and some of them are generated, uh, and then can we ask the users like, "How do we do?" Um, and then we wanted to use a benchmark like party prompts so that people in academi- we mostly did that so people in academia could measure their models against ours versus others.

Um, and, uh, but yeah, I mean, FID, FID is pretty bad, and I think, yeah, you-- in terms of vibes, it's like you put out the model, and then you try to see like what users make. And I think my sense is that we're gonna take all the things that we notice that the users kind of were failing at, um, and try to find like new ways to measure that, whether that's like a smile or, you know, color contrast or lighting.

Um, one benefit of Playground is that we have users making millions of images, um, every single day, and so we can just ask them. Um, and that-

**Swyx** [41:18]
Like go for like a post-generation feedback.

**Suhail Doshi** [41:20]
Yeah. We can just ask them. We can just say like, "How, how good was the lighting here? How was, um, how was the subject? How was the background?"

**Swyx** [41:27]
Yeah.

**Suhail Doshi** [41:27]
Uh-

**Swyx** [41:28]
Oh, like a for- like a proper form of like-

**Suhail Doshi** [41:31]
It's just like you make it-

**Swyx** [41:32]
... some like six-

**Suhail Doshi** [41:32]
You come to our site, you make an image, and then we say-

**Swyx** [41:34]
Yeah

**Suhail Doshi** [41:34]
... and then maybe randomly you just say, "Hey, you know, like how was, how was the color and contrast of this image?" And you say, "It was, it was not very good," and then you just tell us. So I think, I think we can get like tens, uh, tens of thousands of these, uh, evaluations every single day to, to truly measure real world performance-

**Swyx** [41:52]
Yeah

**Suhail Doshi** [41:52]
... as opposed to just like benchmark performance. Hopefully next year, I think we will try to publish kind of like a, like a, a benchmark that anyone could use, that we evaluate ourselves on-

**Swyx** [42:03]
Yep

**Suhail Doshi** [42:03]
... and that other people can-

**Swyx** [42:04]
That's an ideal goal

**Suhail Doshi** [42:05]
... that we think does a good job of approximating real world performance because we've tried it and done it and noticed that it did. Yeah. I think, I think we will do that.

**Swyx** [42:13]
Yeah. Um, we're-- I think we're gonna ask a few more like sort of product-y questions. Um, I, I, and I, I personally have a few like categories that I consider special-

**Suhail Doshi** [42:22]
Mm-hmm

**Swyx** [42:22]
... among, you know, you know, you have like animals, art, fashion, food. Um- There are some s- categories which I consider like a different tier of image. Uh, so f- the top among them is text in images.

**Suhail Doshi** [42:33]
Mm. Mm-hmm.

**Swyx** [42:34]
Um, how do you think about that? Um, so one of the big wow moments for me, or something I've been looking out for the entire year is just the progress of text in images. Like do you, can you write in an image?

**Suhail Doshi** [42:45]
Yeah.

**Swyx** [42:45]
Or, um, and Ideogram-

**Suhail Doshi** [42:47]
Mm-hmm

**Swyx** [42:47]
... I think came out recently-

**Suhail Doshi** [42:48]
Mm-hmm

**Swyx** [42:48]
... which had decent, uh, but not perfect text in images. Um, DALL-E 3 had improves, uh, some, and all they said in their, their, uh, paper was that they just included more text in the dataset and it just worked.

I was like, "That's just, that's just lazy." All right. But anyway, uh, do you care about that? Uh, because I don't see any, any of that in like your samples.

**Suhail Doshi** [43:08]
Yeah, yeah. We're, our, yeah, the, the V2 model is, um, was, was mostly focused on, um, image quality versus like the feature of, uh, y- text synthesis.

**Swyx** [43:18]
Yeah. 'Cause I, well, as a business user, I care a lot about that.

### Safety and Limits

**Suhail Doshi** [43:21]
Yeah.

**Swyx** [43:21]
Right.

**Suhail Doshi** [43:22]
Yeah. I'm very excited about text synthesis, and yeah, I think Ideogram has done a good job of may- maybe the best job. Uh, DALL-E kind of ha- it's like a, it has like a hit rate. You know, you don't want just text effects.

I think where this has to go is h- it has to be like you could like write little tiny pieces of text like on like a milk carton.

**Swyx** [43:41]
Yeah.

**Suhail Doshi** [43:41]
That's maybe not even the focal point of a scene.

**Swyx** [43:44]
Yeah.

**Suhail Doshi** [43:44]
I think that's like a very hard task that, um, you know, if you could do something like that, then there's a lot of other possibilities.

**Swyx** [43:50]
Well, you don't have to zero shot it. You can just be like, "Here, f- focus on this."

**Suhail Doshi** [43:54]
Sure, yeah, yeah. Definitely. Yeah, yeah. So I think text synthesis would be very exciting.

**Swyx** [43:58]
Yeah. Uh, and then also, I'll also flag that, um, Max Wolf, Minimaxir, which you must have come across his work, um, he's done a lot of stuff about w- using like logo masks-

**Suhail Doshi** [44:08]
Mm-hmm

**Swyx** [44:09]
... that then map onto like a, like food or-

**Suhail Doshi** [44:13]
Mm-hmm

**Swyx** [44:13]
... vegetables and it, and it looks, looks like text, uh, which, which can be pretty funny.

**Suhail Doshi** [44:17]
Yeah. Yeah, I mean, you, you... It's very interesting to-- That, that's the wonderful thing about like the open source community is that you get things like ControlNet-

**Swyx** [44:25]
Yeah

**Suhail Doshi** [44:25]
... and then you see all these people do these just amazing things with ControlNet, and then you wonder, uh, I think from our point of view, we, we sort of go that, that's really wonderful, but how, how do we end up with like a unified model that can do that?

What are the bottlenecks? What are the issues? Um, because the community ultimately has very limited resources.

**Swyx** [44:42]
Yeah.

**Suhail Doshi** [44:42]
And so they, they need these kinds of like workaround, um, workaround research ideas to get there. Um, but yeah.

**Swyx** [44:50]
Yeah. Are, are techniques like ControlNet portable to your architecture?

**Suhail Doshi** [44:54]
Definitely.

**Swyx** [44:55]
Okay.

**Suhail Doshi** [44:55]
Yeah. It, we kept the Playground v2 arc exactly the same as SDXL, not because, not out of laziness, but just because we wanted-- we knew that the community already had tools.

**Swyx** [45:04]
Yeah.

**Suhail Doshi** [45:05]
It's, you know, all you have to do is maybe change a string in your code and then, you know, retrain a ControlNet for it, so that it was very intentional to do that. We didn't wanna fragment the community with different architectures.

**Swyx** [45:15]
Yeah. Yeah.

**Suhail Doshi** [45:15]
Yeah.

**Swyx** [45:15]
Uh, I, I have more questions about that. I, I don't know. I don't, I don't wanna DDoS you with, uh- ... with topics. But, but okay, I was basically gonna, gonna go over three more categories.

**Suhail Doshi** [45:24]
All right.

**Swyx** [45:24]
One is, uh, UIs, like, um, app UIs, like mock UIs. Uh, th- third is, uh, not safe for work.

**Suhail Doshi** [45:31]
Mm-hmm.

**Swyx** [45:31]
Um, obviously. Uh, and then copyrighted stuff.

**Suhail Doshi** [45:34]
Mm-hmm.

**Swyx** [45:34]
Um, I don't know if you care to comment on any, any of those.

**Suhail Doshi** [45:37]
The NSFW kind of like safety stuff is really important. Um, part, part of-- I, I kind of think that one of the biggest risks kind of going into maybe the US election year will probably be inter, very interrelated with like graphics, audio, um, video.

I think it's gonna be very hard to explain, you know, to a family relative who's not kind of in our world, and our w- our world is like sometimes very, you know, we think it's very big, but it's very tiny-

**Swyx** [46:05]
Yeah

**Suhail Doshi** [46:05]
... compared to the rest of the world-

**Swyx** [46:05]
Absolutely

**Suhail Doshi** [46:06]
... sometimes like there's still lots of humanity have no idea what ChatGPT is. And I think it's gonna be very hard to explain, you know, to your uncle, aunt, whoever, you know, "Hey, I saw, you know, I saw President Biden say this thing on a video," you know, "I, I can't believe, you know, he said that."

I think that's gonna be a very troubling thing going into, um, going into the world next year or the year after.

**Swyx** [46:30]
Oh, uh, uh, I didn't, that, that's more of like a risk thing-

**Suhail Doshi** [46:32]
Yeah

**Swyx** [46:32]
... or like deepfakes, uh, well, faking, political faking. But, uh, there's just, there's a lot of, um, studies on how, um, yeah, for most businesses you don't wanna train on not safe for work images-

**Suhail Doshi** [46:44]
Mm-hmm

**Swyx** [46:44]
... except that it makes you v- really good at bodies.

**Suhail Doshi** [46:48]
Yeah, I mean- ... uh, yeah, I mean, we, we personally, we filter out, um, NSFW type of, uh, images in our dataset so that it's, you know, so our safety filter stuff doesn't have to work as hard.

**Swyx** [47:00]
But you've, you've heard this argument that it get, it makes you worse at, uh, because obviously not safe for work images are very good at, um, human anatomy-

**Suhail Doshi** [47:08]
Mm

**Swyx** [47:08]
... which you do wanna be good at.

**Suhail Doshi** [47:10]
Yeah, it's, it's not about like, it's not like necessarily a bad thing to train on that data. It's more about like how you go and use it. That's why I was kind of talking about safety, um-

**Swyx** [47:18]
Yeah, yeah, I see

**Suhail Doshi** [47:19]
... you know, in part because there are very terrible things that can happen in the world. If you have a sufficiently, you know, extremely powerful graphics model, you know, suddenly like you can kind of imagine, you know, now if you can like generate nudes and then there's like you could do very character consistent things with faces, like what does that lead to?

**Swyx** [47:34]
Yeah.

**Suhail Doshi** [47:34]
Yeah. I think it's like more what occurs after that, right? Even if you train on, let's say, you know, new data, if it does something to kind of help, there's nothing wrong with the human anatomy, um, it's very valid for a model to learn that, uh, but then it's kind of like how does that get used?

And, uh, you know, I, I, I won't bring up all of the very, very unsavory terrible things that we see, uh, on, on a daily basis-

**Swyx** [47:57]
Oh, God

**Suhail Doshi** [47:57]
... on the site. I think it's more about what, what occurs. And so we, you know, we just recently did like a big sprint on safety internally around... And, and it's very, it's very difficult with graphics and art, right?

Because there is tasteful art that has nudity.

**Swyx** [48:12]
Yeah.

**Suhail Doshi** [48:13]
Right? They're all over in museums, like, you know, it, it's very, very valid situations for that, and then there's, you know, there's the things that are the gray line of that. You know, what I might not find tasteful, someone might fi- be like, "That is completely tasteful," right?

And then, and then there are things that are way over the line.

**Swyx** [48:29]
Yeah.

**Suhail Doshi** [48:30]
Um, and then there are things that are, you know, maybe, maybe you or, you know, maybe I would, you know, be okay with, but society isn't.

**Swyx** [48:37]
Yeah.

**Suhail Doshi** [48:38]
I think it's really hard with art. Think it's really, really hard. Sometimes if even if you have like even if you have, um, things that are not new, if, if a child goes to, to your site, scrolls down some images, you know, classrooms of kids, you know, using our product, it's a really difficult problem.

And, um, and it, and it stretches mostly culture, society, politics, everything. Yeah.

**Alessio** [49:00]
Okay. Um, another favorite topic of our listeners is, um, UX in AI, and I think you're probably one of the best all-inclusive editors for these things.

**Suhail Doshi** [49:12]
Mm.

**Alessio** [49:12]
So you don't just have the, you know, prompt images come out, you pray and if no, you do it again. Uh, first you let people, um, pick a seed so they can kind of have semi-repeatable generation. Um-

**Suhail Doshi** [49:26]
Yeah.

**Alessio** [49:26]
You also have-

**Suhail Doshi** [49:27]
Absolutely.

**Alessio** [49:27]
Yeah, you can pick how many images and then you leave all of them in the canvas, and then you have kind of like this box, the generation box, and you can even cross between them and outpaint.

**Suhail Doshi** [49:38]
Yeah.

**Alessio** [49:38]
There's all these things.

How did you get here? You know?

**Suhail Doshi** [49:42]
Yeah.

**Alessio** [49:43]
Most people, most people are kind of like, "Give me text, I give you image," you know?

**Suhail Doshi** [49:46]
Yeah.

**Alessio** [49:46]
And you're like, "These are all the tools for you."

### AI-First UX

**Suhail Doshi** [49:48]
Even though we are trying to make, um, a, a graphics foundation model, I think we think that we're also trying to f- like reimagine like what a graphics editor might look like given the change in technology. So you know, we-- I don't think we're trying to build Photoshop, but it's the only thing that we could say that people are fam- you know, largely familiar with.

"Oh, okay. There's Photoshop." Uh, I think... You know, I don't think you would think of Photoshop without like the com- you know, you don't, you wouldn't think what would Photoshop compare itself to pre, pre-computer? I don't know, right?

**Alessio** [50:22]
Mm-hmm. Mm-hmm.

**Suhail Doshi** [50:22]
It's like, oh, we're kind of like a, a canvas.

**Alessio** [50:25]
Mm-hmm.

**Suhail Doshi** [50:25]
But you know, there's these menu options, and you can use your mouse. What's a mouse? Um, so I, I think that we're trying to make like-- we're trying to reimagine what a graphics editor might look like, not, not just for the fun of it, but because we kind of have no choice.

Like, there's this idea in, in image generation where you can gen-generate images. That's like a super weird thing. What is that in Photoshop, right? You have to wait right now for the time being, um, but the wait is worth it often for a lot of people because they can't make that with their own skills.

So I, I think it goes back to, you know, how we started the company, which was kind of looking at GPT-3's playground. The, that the reason why we're named Playground is, is a homage to that actually. Um, and you know, it's like shouldn't these products be more visual?

Shouldn't, you know, shouldn't they... These prompt boxes are like, like a terminal window, right?

**Alessio** [51:14]
Mm-hmm.

**Suhail Doshi** [51:15]
We're kind of at this weird point where it's just like CLI. It's like MS-DOS. I remember my mom using MS-DOS, and I memorized the keywords like dir, ls, all those things. Right? It feels a little like we're there, right?

Prompt engineering is this like-

**Alessio** [51:28]
The, the shirt I'm wearing, you know, it's, it's-

**Suhail Doshi** [51:29]
Yeah

**Alessio** [51:29]
... it's a bug, not a feature.

**Suhail Doshi** [51:30]
Yeah. Exactly. Parentheses to say beautiful or whatever, which weights the word token more in the model or whatever. Um, yeah, it's that- that's like super strange. I think that's not... I think everybody, I think a large portion of humanity would agree that that's not user-friendly, right?

So how do we think about the products to be more user-friendly? Well, sure. You know, sure would be nice if I could like, you know, uh, if I wanted to get rid of like the headphones on my, my head, you know, it'd be nice to mask it and then say, you know, "Can you remove the headphones?"

Um, you know, if I want to grow the, expand the image, it should... You know, how can we make that feel easier without typing lots of words and being really confused? And by no, by no m- stretch of the imagination, I don't even think we've nailed the UI/UX yet.

Um, part of that is because we don't-- we're still experimenting, and part of that is because the model and the technology's gonna get better. And whatever felt like the right UX six months ago is gonna feel very broken now.

Um, and, uh, so that, that's a little bit of how we got there is kind of saying, "Does everything have to be like a prompt in a box, or can we do, can we do things that make it very intuitive for users?"

**Alessio** [52:42]
How do you s- decide what to give access to? So you have things like, um, expand prompt-

**Suhail Doshi** [52:47]
Mm-hmm

**Alessio** [52:48]
... uh, which DALL-E 3 just does.

**Suhail Doshi** [52:50]
Mm-hmm.

**Alessio** [52:50]
It doesn't let you decide-

**Suhail Doshi** [52:51]
Yeah

**Alessio** [52:51]
... whether you should or not. Um, yeah.

**Swyx** [52:54]
It as in like, uh, rewrites your prompts for you.

**Alessio** [52:56]
Yeah.

**Suhail Doshi** [52:56]
Mm-hmm.

**Swyx** [52:56]
Yeah.

**Suhail Doshi** [52:57]
Yeah, for that feature, I, I think we'll probably... I, I think once we get it to be, uh, cheaper, we'll probably just give it out. We'll probably just give it away. But we also decided something that we-- that might be a little bit different.

We noticed that most of image generation is just like kind of casual. You know, it's in WhatsApp, it's, you know, it's in a Discord bot somewhere with Midjourney. It's in ChatGPT. One of the differentiators I think we provide is at the expense of just lots of users necessarily, mainstream consumers, is that we provide as much like power and tweakability and configurability as possible.

So the only reason why it's a tog- it's a toggle because we know that users might wanna use it and might not wanna use it, right? There are some us- there are some really powerful power user hobbyists that know what they're doing, and then there's a lot of people that, um, uh, you know, just want something that looks cool, but they don't know how to prompt.

And so I think a lot of Playground is more about, um, going after that core user base that like knows-- has a little bit more savviness, uh, in how to use these tools. Yeah. So they might not use like these users probably...

You know, the average DALL-E user is probably not gonna use ControlNet. They probably don't even know what that is.

**Alessio** [54:07]
Hmm.

**Suhail Doshi** [54:08]
Um, and so I think that like as the models get more powerful, as the, the-- there's more tooling, um, yeah, I think you could imagine it. Hopefully, you'll imagine a new sort of AI first graphics editor that's just as like powerful and configurable as Photoshop.

Uh, and you might have to master a new kind of tool.

**Swyx** [54:27]
Yeah.

**Suhail Doshi** [54:28]
Yeah.

**Swyx** [54:28]
Well, um, th- uh, there are so many things I could, I could go bounce off of that. Um, one, one you, you mentioned about waiting.

**Suhail Doshi** [54:36]
Mm-hmm.

**Swyx** [54:37]
Um, w- we have to kind of somewhat address the elephant in the room. Uh, uh, consistency models have been blowing up, uh, uh, the past month. Um, is that... Like how do you think about integrating that? Um, obviously there's, there's a lot of other companies also trying to, uh, beat you to that space as well.

**Suhail Doshi** [54:53]
I think we were the first company to integrate it. Well, we integrated it in a different way. There are like 10 companies right now that have kind of tried to do like interactive editing where-

**Swyx** [55:01]
Yeah

**Suhail Doshi** [55:02]
... you can like draw on the left side and then you get an image on the right side. We decided to kind of like wait and see whether there's like true utility on that. Um, we have a different feature that's like unique, uh, in our product that, um, that, that's called preview rendering.

And so you go to the product and you, and you say... You know, we, we're like, "What is the most common use case?" The most common use case is you write a prompt and then you get an image.

But what's the most annoying thing about that? The most annoying thing is like it's like feels like a slot machine, right?

**Swyx** [55:28]
Mm-hmm.

**Suhail Doshi** [55:28]
You're like, "Okay, I'm gonna put it in and I'm gonna... Maybe I'll get something cool." So we did something that seemed a lot simpler but a lot more relevant to how users already use these products, which is preview rendering.

You toggle it on and it will show you a render of the image, and then it's just like a pr- graphics tools already have this. Like if you use Cinema 4D or After Effects or something, it's called viewport rendering.

**Swyx** [55:49]
Mm-hmm.

**Suhail Doshi** [55:50]
And so we try to take, take something that exists in the real world that has familiarity and say, "Okay, you're gonna get a rough sense of an early preview of this thing, and then when you're ready to generate, it's...

We're gonna try to be as coherent about that image that you saw." That way you're not spending so much time just like, you know, uh, pulling down the slot machine lever. So I, I... We were, we were actually the first company, I think we were the first company to actually ship-

**Swyx** [56:13]
My bad

**Suhail Doshi** [56:13]
... that quick LCM-

**Swyx** [56:14]
Yeah

**Suhail Doshi** [56:15]
... uh, thing. Yeah.

**Swyx** [56:16]
Okay.

**Suhail Doshi** [56:18]
Yeah. We were very excited about it, so we shipped it very quick. Yeah.

**Swyx** [56:21]
Yeah, yeah. I, I think m- like the other, um, the... Well, the demos I've been seeing, it, it's, it's also, I guess, it's not like a preview necessarily. They're almost using it to animate theirs, their, their generations. Like to, you can- because you can kind of move shapes over-

**Suhail Doshi** [56:36]
Yeah, yeah. They're, they're like doing it. They're like animating it, but they're sort of showing like if I move a moon, you know, can I... Yeah.

**Swyx** [56:43]
Yeah. I don't know. It, it, it to me unlock, un-unlocks video in a way.

**Suhail Doshi** [56:47]
Yeah.

**Swyx** [56:47]
Um, that, uh

**Suhail Doshi** [56:48]
But the video-

**Swyx** [56:49]
I've seen

**Suhail Doshi** [56:49]
... the video models are already so much better than that.

**Swyx** [56:51]
Yeah.

**Suhail Doshi** [56:51]
Yeah. So.

**Swyx** [56:52]
So . Uh, there's a- there's another one which I, I think is, uh, um, um, like h- how about like the just g- general ecosystem of LoRAs?

**Suhail Doshi** [57:01]
Uh-huh.

**Swyx** [57:02]
Right? That, um, Civit is obviously the most popular, um, repository of LoRAs. Um, how do you t- think about sort of interacting with that ecosystem?

**Suhail Doshi** [57:12]
Yeah, I mean, uh, the guy that, that did LoRA, not the guy that invented LoRAs, but the person that brought LoRAs to Stable Diffusion, uh, actually works with us, um, on, on some projects. Uh, his name is Simu.

### LoRA Ecosystem

**Suhail Doshi** [57:24]
Um, shout out to Simu. Um, and I, I think LoRAs are, are wonderful. Um, obviously fine-tuning all these DreamBooth models and such are j- is just so heavy and giving... And I- it's obvious in our conversation around styles and vibes and, you know, it's very hard to evaluate the artistry of these things.

LoRAs give people, uh, this wonderful like opportunity, uh, to create like subgenres of art, and I think they're amazing. And so any graphics tool, any kind of thing that's expressing art has to provide some level of customization to its, its user base that goes beyond, you know, just typing like Greg Rutkowski in a prompt.

Right? We have to give more than that. Um, y- it's not like users wanna type these, you know, art- real artist names. It's that they don't know how else to get an image that looks interesting. They, they truly want like originality and uniqueness, and I think LoRAs provide that, and they provide it in a very nice scalable way.

Um, I hope that we find something even better than LoRAs in the, in the long term, in the long term. Um, 'cause there, there are still weaknesses to LoRAs, um, but I think they do a good job for now.

**Swyx** [58:33]
Yeah. And so you don't want to be the... Like you don't, you would never compete with Civit. You would just kind of let people in.

**Suhail Doshi** [58:38]
Civit's a site where like all these things get kind of hosted by the community, right?

**Swyx** [58:41]
Yeah.

**Suhail Doshi** [58:41]
Um, and so yeah, we'll often pull down the thing, like some of the best things there. Um, I think, I think when we have a significantly better model, uh, we will certainly build something-

**Swyx** [58:53]
I see

**Suhail Doshi** [58:54]
... that gets closer to that. I still, again, I go back to saying just I still think this is like very nascent. Things are very underpowered, right? We, you know, LoRAs are not easy for people to train. You know, they're easy for an engineer.

**Swyx** [59:07]
Okay.

**Suhail Doshi** [59:08]
But they're not e- they're not easy, you know... It sure would be nicer if I could just pick, you know, five or six reference images-

**Swyx** [59:14]
Yeah

**Suhail Doshi** [59:14]
... right? And, and then say, "Hey, you know, this is, this is..." And it, and there might even be five or six different reference images that are not... They're just very different actually. Like they're, they're, they communicate a style, but they're actually like, it's like a mood board, right?

And it takes, you have to be kind of an engineer almost to train these LoRAs or go to some site and be technically savvy at least. Um, it seems like it'd be much better if I could say, "I love this style.

I love, I love this, this style. Here are five images" And you tell the model like, "This is what I want" and the model gives you, gives you something that's very aligned with what your style is, what you're talking about.

And it's a style you couldn't even communicate, right? There's no word. You know, this is, you know, if you have a Tron image, it's not just Tron. It's like Tron plus like s- four or five different weird things.

**Swyx** [59:59]
And cyberpunk. Yeah.

**Suhail Doshi** [1:00:00]
Yeah. Um, even cyberpunk can have its like subgenre, right? But I just think training LoRAs and doing that is very heavy, so I hope we can do better than that.

**Swyx** [1:00:09]
Cool.

**Suhail Doshi** [1:00:09]
Yeah.

**Swyx** [1:00:10]
Yeah. Um, we had Sharif from Lexica on the podcast before.

**Suhail Doshi** [1:00:14]
Oh, nice.

**Swyx** [1:00:15]
And both of you have like a landing page with just a bunch of images where you can like explore things.

**Suhail Doshi** [1:00:21]
Yeah.

**Swyx** [1:00:21]
Um-

**Suhail Doshi** [1:00:22]
Yeah, we have a feed.

**Swyx** [1:00:23]
Yeah, yeah. Is that something you see more and more of in terms of like coming up with these styles? Is that why you, you have that as the starting point versus a lot of other products you just go in, you have the generation prompt, you don't see a lot of examples.

**Suhail Doshi** [1:00:36]
Right. Our feed is a little different than, than their feed. Our feed is more about community, so we have kind of like a Reddit thing going on where it's a kind of a competition like every day, loose competition, mo- mostly fun competition of like making things, and there's just this wonderful community of people where they're liking each other's images and just showing their like, their genuine interest in each other, and I think we definitely learn about styles that way.

One of the funniest polls, uh, i- if you go to the Midjourney c- c- polls, they'll sometimes put these polls out and they'll say, "You know, what do you wish you could like learn more from?" And like one of the, one of the things that people vote the most for is like learning how to prompt.

Right? And so I think like, you know, if you, if you put away your research hat for a minute and you just put on like your product hat for, for a second, you're kind of like, "Well, why do people wanna learn how to prompt?"

Right? It's because they wanna get higher quality images. Well, what's higher quality? Composition, lighting, aesthetics, so on and so forth. And I think that the community on our feed, I think, I think we have-- I think we might have the biggest community and, uh, and it gives all of the users a way to learn how to prompt.

Because they're just seeing this huge rising tide of all these images that are super cool and interesting, and they can kinda like take each other's prompts and like kind of learn how to do that. Um, I, I think that'll be short-lived because I think the complexity of these things is gonna get higher.

Um, but, um, but that, that's more about why we have that feed, is to help each other. To help teach us-users, and then also just, you know, celebrate people's art.

### GPU Engineering

**Swyx** [1:02:09]
You run your own infra.

**Suhail Doshi** [1:02:10]
We do.

**Swyx** [1:02:11]
Yeah. That's unusual.

**Suhail Doshi** [1:02:14]
Uh, it's necessary.

**Swyx** [1:02:16]
It's necessary.

**Suhail Doshi** [1:02:16]
Yeah.

**Swyx** [1:02:17]
Uh, what have you learned running DevOps for GPUs? Uh, you-- I, I... You had a tweet about like how many A100s you have, but I feel like it's out of date probably.

**Suhail Doshi** [1:02:26]
Uh, yeah, we, uh... I think, I mean, it just comes down to cost. These things are very expensive, so, uh, we just wanna make it as affordable for everybody as possible. Um, I don't find... I find the DevOps for inference to be rel-relatively easy.

**Swyx** [1:02:40]
Okay.

**Suhail Doshi** [1:02:40]
It doesn't feel that different than, you know... I think we had thousands and thousands of servers at Mixpanel, uh, just for dealing with the, the API. It had such huge quantities of volume that I didn't find it, I don't find it particularly very different.

Um, I do find, uh, GP, model optimization performance is very new to me, so I think that, I find that very difficult at the moment. So that's very interesting. But, uh, scaling inference is not, not, not terrible. Tr- scaling a training cluster is very, much, much harder, um, than I perhaps anticipated.

**Swyx** [1:03:12]
Why is that?

**Suhail Doshi** [1:03:13]
Well, you have, you know, you have to, it's just like a very large distributed system with, um, you know, if you have like a, a node that goes down, then your tr- you know, training run crashes, and then you have to somehow be resilient to that.

And I would say training infra software is very early. It feels very broken. Feel-

**Swyx** [1:03:30]
Like a, like a-

**Suhail Doshi** [1:03:30]
I can tell in 10 years it would be a lot better

**Swyx** [1:03:32]
... like a Mosaic or whatever.

**Suhail Doshi** [1:03:34]
We don't, yeah, we don't even... No, we don't. We, we think we use very basic tools like, you know, Slurm for scheduling and just normal PyTorch, PyTorch Lightning, that kind of thing. I think our tooling is nascent. I think I talked to a friend that's over at xAI.

They just, they like built their own scheduler, you know, and doing things with Kubernetes. Like, when people are building out tools because the existing open source stuff doesn't work and everyone's doing their own bespoke thing, you know there's a ma- there's a valuable company to be formed.

**Swyx** [1:03:58]
Yeah. Uh, I think it's Mosaic. I don't know.

**Suhail Doshi** [1:04:01]
Well, Mos- with Mosaic, yeah. It's t- it's tough with Mosaic 'cause, um... Anyway, I won't, I won't go into the details why, but yeah, we, we found it, uh, difficult to do. You, you... It might be worth like wondering like why, why, why not everyone is going to Mosaic.

**Swyx** [1:04:13]
Yeah.

**Suhail Doshi** [1:04:13]
And perhaps it's still, it's, I, I just think it's nascent.

**Swyx** [1:04:16]
Cool.

**Suhail Doshi** [1:04:16]
And perhaps Mosaic will come through.

**Swyx** [1:04:18]
Cool. Anything for you?

**Alessio** [1:04:19]
Um, no, no. This was great. And just to wrap, we, we talked about some of the pivotal moments in your mind with like DALL-E and, and whatnot. If you were not doing this, what's the most interesting unsolved question in AI that you would try and build then?

**Suhail Doshi** [1:04:36]
Oh, man. Coming up with startup ideas is very hard on the spot. Uh... And-

**Swyx** [1:04:41]
You sh- you have to have them. I mean, you're a founder. You're a repeat founder, like...

### Lightning Round

**Suhail Doshi** [1:04:46]
I, I'm very picky about my startup ideas. Um, so I don't, I, you know, don't have any great ones. Uh, the only thing that I... I, I don't have an idea per se, as much as a curiosity. Uh, and I'll po- I suppose I'll pose it to you guys.

Right now, we sort of think that a lot of the modalities just kinda feel like they're, you know, vision, language, audio, like that's roughly it. And somehow all this will like turn into something. It'll be multimodal, and then we'll end up with AGI, um, perhaps.

And I just think that there are probably far more modalities than maybe we, than meets the eye. And it just seems hard for us to see it right now because it's sort, sort of like we have tunnel vision on the moment.

**Swyx** [1:05:32]
We're, we're just like code, image, audio, video.

**Suhail Doshi** [1:05:35]
Yeah. I think I-

**Swyx** [1:05:36]
Very, very broad categories

**Suhail Doshi** [1:05:37]
... I think we are lacking imagination as a species in this regard.

**Swyx** [1:05:41]
Yeah.

**Alessio** [1:05:41]
I see it. I see it.

**Suhail Doshi** [1:05:42]
And, and I think like, you know, just like, uh, you know, it's not, I don't know what company would, would form as a result of this, but you know, like there's some, some very difficult problems like just tr- like a true actual, like not a meta world model, but an actual world model that truly maps everything that's going in terms of like physics and fluids and all these various kinds of interactions.

And y- what does that kind of model, like a true physics foundation model of sorts that represents Earth. And that in of itself seems very difficult, you know? But we just think of, but we're, but we're kinda stuck on like thinking that we can approximate everything with like, you know, a word or a token, if you will.

And I went, you know, I had a dinner last night where we were kind of debating this philosophically, and I think someone, you know, said something that I also believe in, which is like, at the end of the day, it doesn't really matter that it's like a token or a byte.

At the end of the day, it's just like some, you know, unit of information that it emits. But you know, you do, I do wonder if there are more, far more modalities than, um, than meets the eye. And if, if you could create that, then what would that, what would, what would that company become?

What problems could you solve? So I, I don't, I don't know yet, so I don't have a great company for it.

**Swyx** [1:06:53]
I don't know.

**Suhail Doshi** [1:06:53]
But-

**Swyx** [1:06:54]
Maybe you just inspire somebody to, to try, so.

**Suhail Doshi** [1:06:56]
Yeah. Hopefully.

**Swyx** [1:06:57]
Yeah. Uh, my personal response to that is I'm, I'm less interested in physics and more interested in people. Uh-

**Suhail Doshi** [1:07:02]
Mm-hmm

**Swyx** [1:07:02]
... like, like how do I, how do I mind upload? Because that is-

**Suhail Doshi** [1:07:05]
Right. Exactly

**Swyx** [1:07:06]
... teleportation, that is immortality, that is everything.

**Suhail Doshi** [1:07:09]
Yeah. Yeah. Can we, can we model our own... Rather than trying to create consciousness, could we model our own- Um-

**Swyx** [1:07:16]
Yeah

**Suhail Doshi** [1:07:16]
... even if it was lossy-

**Swyx** [1:07:17]
Yeah

**Suhail Doshi** [1:07:17]
... to some extent. Yeah.

**Swyx** [1:07:19]
Yeah. Um, well, we won't solve that here.

**Suhail Doshi** [1:07:21]
Yeah.

**Swyx** [1:07:22]
Um, if I were to take a Bill Gates book trip- ... uh, and had a week, uh, what should I take with me to learn AI?

**Suhail Doshi** [1:07:30]
Oh, man. Oh, gosh. You shouldn't take a book. You should just go to- ... uh, YouTube and visit Karpathy's, uh- ... class-

**Swyx** [1:07:38]
Zero to hero

**Suhail Doshi** [1:07:39]
... and just do it, do it. Grind through it.

**Swyx** [1:07:42]
Is, was that actually the most useful thing for you?

**Suhail Doshi** [1:07:44]
I wish it came out when I started-

**Swyx** [1:07:45]
Wow

**Suhail Doshi** [1:07:45]
... back last year. I, I'm, I'm as, as bummed that I didn't get to take it at the beginning. Um, but I did, I did do a few of his classes regardless. I, I don't think books-- Every time I buy a programming book, I never read it.

I always find that just writing code helps cement my internal understanding.

**Swyx** [1:08:02]
Yeah. So, so, so more generally, advice for founders who are not PhDs and are effectively self-taught, like, like you are. Like, what should they do? What should they avoid?

**Suhail Doshi** [1:08:11]
Same thing as if, as that I would advise if you were programming. Pick a project that seems very exciting to you, but don't, you know, it doesn't have to be too serious, and build it and learn every detail of it while you do it.

**Swyx** [1:08:22]
And it must be, uh, like, should you train, or can you, can you go f-far enough not training, just-

**Suhail Doshi** [1:08:29]
It-

**Swyx** [1:08:29]
... fine-tuning?

**Suhail Doshi** [1:08:29]
It depends. I would s- I would just follow your curiosity. If, like, you want, if what you wanna do is something that requires fundamental understanding of training models, then you should learn it. You don't have to be a P- you don't have to get to become a five, you know, five-year whatever PhD, but if that's necessary, I would do it.

If it's not necessary, then go as far as you need to go. But I would learn, you know, pick something that motivates. I think most people tap out on motivation, but they're deeply curious.

**Swyx** [1:08:53]
Yeah. Cool.

**Suhail Doshi** [1:08:54]
Yeah.

**Swyx** [1:08:54]
Cool. Excellent. Thank you so much for coming out, man.

**Suhail Doshi** [1:08:56]
Thank you. Thank you for having me.

**Swyx** [1:08:57]
This was fun.

**Suhail Doshi** [1:08:57]
Appreciate it.

---

This library is powered by PodHood (https://podhood.com), the podcast website platform.
