# Making Transformers Sing - with Mikey Shulman of Suno

Latent Space · 2024-03-14

<https://addtry.com/5dd33f3e-1bc8-4b7b-bff7-79df326896ef>

In this episode of Latent Space, hosts Alessio and Shawn Wang interview Mikey Shulman, CEO of Suno, about making transformers sing. Suno uses transformers to predict audio tokens end-to-end, avoiding baked-in musical knowledge, with a tokenization secret sauce that also includes non-music audio for better vocal realism. Their models are relatively small (far below 175B parameters) due to latency needs, and they prioritize scaling research over brute-force size. Over half of Suno users employ expert mode, tweaking lyrics and style prompts, rather than easy mode. Shulman argues Suno is not the 'Midjourney of music' because music is inherently social and synchronous, unlike images. He demoed live generation, showing control via tokens like [beat drop] and style modifiers, and revealed future plans for collaborative concerts, continuous DJ modes, and personalized models. He also advocates for hiring economists to avoid Goodhart's law pitfalls in ML benchmarks, especially in audio where aesthetics matter most.

## Questions this episode answers

### How does Suno build its music generation model, and why do they prefer transformers?

Mikey Shulman says Suno uses transformer-based models that predict the next audio token, much like language models. They avoid baking in musical knowledge, letting the model learn end-to-end from diverse audio—including speech and music—to improve realism. The key challenge is tokenizing audio correctly; they focus on discrete representations while leveraging scaling insights from text. This contrasts with older methods that hard-coded phonemes or musical structures, which Shulman believes limit long-term potential.

[3:45](https://addtry.com/5dd33f3e-1bc8-4b7b-bff7-79df326896ef?t=225000)

### What are the main ways Suno users create music, and which mode is more popular?

According to Mikey Shulman, more than half of Suno's usage comes from expert mode, where users write their own lyrics, tweak ad-libs, and iterate. He identifies two distinct use cases: quick, humorous 'nice shit posting'—immortalizing funny moments in seconds—and a deeper creative flow to extract a song stuck in one's head. This shows that despite the simplicity of the 'easy mode' prompt box, many users crave granular control over their music.

[19:24](https://addtry.com/5dd33f3e-1bc8-4b7b-bff7-79df326896ef?t=1164000)

### What is Mikey Shulman's vision for the future of music creation, and how does it compare to the gaming industry?

Mikey Shulman envisions music creation becoming far more active and social, akin to gaming—an industry 50 times larger than music. He is excited about collaborative concerts where audiences co-create songs in real time, and a 'multiplayer mode' for Suno that turns music into a shared experience. The goal is to expand participation beyond passive listening, making music a larger part of everyday life, not just for professionals but for everyone.

[44:07](https://addtry.com/5dd33f3e-1bc8-4b7b-bff7-79df326896ef?t=2647000)

### How does Mikey Shulman classify the current landscape of AI music and audio generation?

Mikey Shulman breaks the audio generation field into three categories: music, speech, and sound effects. Within music, he identifies four segments: license-free stock music for backgrounds, AI cover art (which he sees as legally fraught and not the future), net new song creation (Suno's focus), and AI tools for professionals like stem splitters and DAW plugins. He believes net new creation is particularly underserved and where the most exciting growth lies.

[51:56](https://addtry.com/5dd33f3e-1bc8-4b7b-bff7-79df326896ef?t=3116000)

## Key moments

- **[0:00] Intro**
- **[1:55] Music Generation**
  - [2:41] Audio generation lags images and text by one to two years, says Suno's Mikey Shulman.
  - [4:39] Suno's transformer-based music model learns with minimal baked-in knowledge, similar to GPT learning grammar without explicit rules.
  - [6:44] Suno's secret sauce is tokenizing audio into discrete tokens for a transformer, allowing end-to-end music generation.
- **[7:59] Data & Copyright**
  - [7:59] Suno trains on non-music audio to improve vocals, analogous to Code Llama using English to boost code generation.
  - [10:03] Suno's music models are relatively small; streaming latency makes 175B-parameter models impractical for now.
- **[11:57] Suno Origins**
  - [11:57] Mikey Shulman's team at Kensho fell in love with audio AI and founded Suno despite investor advice to focus on speech.
  - [13:27] "We almost couldn't help ourselves from doing music" after Bark's TTS success, says Mikey Shulman.
  - [16:13] Bark, Suno's open-source TTS model, was built from scratch as a transformer, borrowing code from Andrej Karpathy's NanoGPT.
  - [18:10] Even in Bark, users extracted music, proving demand for AI music generation and steering Suno's direction.
- **[18:41] Easy vs Expert**
  - [18:52] Over half of Suno users create music with the expert custom mode, writing their own lyrics and structure.
  - [20:17] Mikey's 'Margoo' song, made in seconds from a Starbucks name mix-up, exemplifies Suno's 'nice shitposting' use case.
  - [22:24] Suno deliberately ignores professional musician tools to focus on changing how average people interact with music.
- **[24:53] Midjourney of Music?**
  - [24:53] Q: Is Suno aiming to be the Midjourney of music? A: Mikey says music's social, synchronous nature makes the analogy too limiting.
- **[26:39] Live Demo**
  - [27:05] Suno demo: a country song generated from prompt 'lack of GPUs in my cloud provider' with references to CUDA cores.
  - [29:32] Suno generates a house track about podcasting: 'From the beats that drop to the melodies that soar, podcasting about music forevermore.'
  - [30:22] Suno uses special tokens like 'beat drop' to control song structure, a trick discovered by power users.
  - [32:53] Suno tests cross-style consistency by generating the same lyrics as house and country rock, revealing genre-specific artifacts.
  - [35:04] Suno demo: a blues song about a sad AI wearing an Apple Vision Pro, with lyrics 'trapped inside this metal frame'.
  - [36:50] Suno's moderation blocked 'Chicago blues guitar' as a prompt to prevent artist impersonation.
- **[40:19] Future Plans**
  - [40:19] Suno prohibits AI covers of existing songs due to publishing rights, focusing solely on original music creation.
  - [42:56] Mikey Shulman: 'There's still so much low-hanging fruit to make music models much better and more controllable.'
  - [44:07] "Machine learning people are stupid sometimes; we only think about models that take X and make it into Y," says Mikey Shulman.
  - [44:33] Suno's goal is to make music active like gaming, which is 50 times bigger than the music industry because it's participatory.
  - [46:59] Mikey Shulman predicts Suno will launch continuous DJ mixing and collaborative concert features 'soon'.
- **[49:11] Fan Favorites**
- **[51:19] Audio Landscape**
  - [51:23] Mikey Shulman categorizes the AI audio landscape into stock music, AI covers, original songs, and professional plugins.
  - [54:27] Q: How does Goodhart's law apply to LLM benchmarks? A: Mikey says quantitative metrics fail in audio; 'aesthetics matter' more than benchmarks.
- **[54:42] Benchmarks**
  - [56:19] "Economists make really good machine learning engineers because they think about Goodhart's law and natural experiments," says Mikey Shulman.
  - [57:18] Mikey recounts how SQuAD 2.0 broke AI benchmarks by adding unanswerable questions, a lesson from first-principles thinking.

## Speakers

- **Alessio** (host)
- **Swyx** (host)
- **Mikey Shulman** (guest)

## Topics

Audio & Music

## Mentioned

Google (company), Kensho Technologies (company), Meta (company), Midjourney (company), NVIDIA (company), Suno (company), Twitter (company), UMG (company), Audiobox (product), Bark (product), ChatGPT (product), Code Llama (product), Discord (product), GitHub (product), NanoGPT (product), Parakeet (product), SQuAD (product), Seamless (product), Stable Diffusion (product), TikTok (product)

## Transcript

### Intro

**Alessio** [0:01]
Hey everyone, welcome to the Latent Space Podcast. This is Alessio, partner and CTO in residence at Decibel Partners, and I'm joined by my co-host Swyx, founder of Smol AI.

**Swyx** [0:11]
Hey, and today we are in the, uh, remote studio with Mikey Shulman. Welcome.

**Mikey Shulman** [0:16]
Thank you. It's great to be here.

**Swyx** [0:19]
Uh, so I'll, I'd like to go over people's background on LinkedIn and then maybe find out a little bit more outside of LinkedIn. Um, you did your bachelor's in physics and, and then a PhD in physics as well, um, before-- also before going into Kensho Technologies with the home of a lot of AI, um, uh, top AI startups, it seems like, where you're head of machine learning for, uh, seven years.

Um, you're also a lecturer at MIT, uh, we can talk about that, like, uh, you know, what you talked about. And then, uh, about, uh, two years ago, you left to, uh, start Suno, uh, which is recently burst on the scene as one of the top music generation startups.

Um, so we can talk... We can go over that bio, but also, I guess, what's not on your LinkedIn that people should know about you?

**Mikey Shulman** [1:05]
I love music. Um, I am a aspiring mediocre musician. Uh, I wish I were better, but that doesn't make me not enjoy playing real music. Um, uh, and I also love coffee. I'm, I'm probably way too much into coffee.

**Alessio** [1:20]
Are you one of those people that, uh, you know, they do the TikToks, they use like fifty tools to like grind the beans and then, like, brush them and then, like, spray them? Like what, what level, what level are we talking about here?

**Mikey Shulman** [1:32]
I confess there's a, a spray bottle for beans in the next, in the next room. There is one of those weird comb tools, so, so guilty. I don't put it on TikTok, though.

**Alessio** [1:43]
Yeah, no, no. Some things gotta stay, gotta stay private. Uh, what, what do you play?

**Mikey Shulman** [1:47]
I played a lot of piano growing up, um, and I play bass and I, in a very mediocre way, play guitar and drums.

**Alessio** [1:54]
Yeah. Right. That's a, that's a lot. I cannot do any of those things, so. Um-

### Music Generation

**Swyx** [2:00]
Yeah.

**Alessio** [2:00]
A-as Sean mentioned, you, you guys kind of burst into the scene as, uh, maybe the state-of-the-art music generation company. Um, I, I think it's a model that we haven't really covered in the past, so I would love to maybe for you to just give a brief intro of like how do you do music generation and why is it possible?

Uh, because I think people understand you take text and you have it predict the next word, and you take a diffusion model and you basically like add noise to an image and then kind of remove the noise. But I think for music, it's hard for people to have a mental model.

Like what's the-- How do you train a music model and like what does a music model do to generate a song? So maybe we can, we can start there.

**Mikey Shulman** [2:41]
Yeah. Maybe I'll even take one more step back and, and say, um, it's not even entirely worked out, I think the same way it is in text. And so, um, it's an evolving field. If you take a, a giant step back, I think audio has been, uh, lagging, uh, images and text for a while.

So you-- I think very roughly you can think audio is like one to two years behind images and text. And so, um, you kind of have to think today like text was in 2022 or something like this. And, um, you know, the transformer was invented.

It looks like it works, but it's, it's, it's far, far less established. And so, um, you know, I'll, I'll give you the way we think about the world now, but just with the big caveat that, that I'm probably wrong if we look back, um, in a couple of years from now.

Um, and I think the biggest thing is you see both transformer-based and diffusion-based models for audio, um, in and in ways that that is not true in text. I know people will do some diffusion for text, but I think nobody's like really doing that, uh, for real.

Uh, and, uh, so we, we prefer transformers for a variety of reasons. And so you can think it's very similar to text. You have some abstract notion of a token, and you train a model to predict, uh, the probability over all of the next tokens.

So it's a language model. Um, you can think in, in anything, a language model is just something that assigns likelihoods to sequences of tokens. Um, sometimes those tokens correspond to text. In our case, they correspond to music or audio in general.

And, uh, I think we've learnt a lot from our, uh, friends in the text domain, from the pioneers doing this of how well these transformer models work, where do they work, where do they not work. But at its core, uh, the way we like to do things with transformers is, uh, exactly like it works in text.

Let me predict the next tiny little bit of audio, and I can just keep doing that and doing that and generating audio, um, as long as I want.

**Swyx** [4:39]
Yeah. I think the, the temptation here is to always try to bake in some specialized knowledge about music or audio. Um, and, uh, so how-- And obviously you will get an improvement in, um, in your output if you try to s- just say like, okay, like here is a set of notes for, you know, here's a, here's a set of tokens that only do, uh, jazz or on- only do, you know, like voices.

Um, how general do you make it versus how specific do you make it?

**Mikey Shulman** [5:10]
We, we've always tried to do things, you know, quote unquote, "the right way," which means that at the beginning things are going to be hard and worse than other ways. But that is to say, bake in as little, uh, kind of implicit knowledge, um, as possible.

And so the same way you don't program into GPT, you don't say this is a noun and this is a verb, but it, it has implicitly learnt all of those things. I've never seen GPT accidentally, you know, put a, put a noun where it meant to put an article in English.

We try not to impose anything about music, uh, or audio in general into the model, and we kind of let the models learn things by themselves. And I think, um, things are beginning to pay off, but it's, you know, it's, it's not necessarily obvious from the beginning that that was the right thing to do.

So for example, you know, you could take something like text to speech and, um, people will do all sorts of things where you can program in things like phonemes to be the basis for what you do, and then that kind of limits you to the set of things that are expressible by phonemes.

And so ultimately, that works really well in the short term. In the long term, it can be quite limiting. And so Our approach has always been to try to do this in its full generality as end-to-end as, as we can do it, even if it means that, um, in the short term we were a little bit worse.

W-we, we have a lot of confidence that in the long term that will be the right way to do it.

**Alessio** [6:33]
And what's the data recipe for training a, a good music model? Like, what percentage genre do you put? Like how sure do you split, uh, vocals and, uh, instrumentals?

**Mikey Shulman** [6:44]
So you have to do lots of things, and I think this is the biggest area where we have, you know, sort of our secret sauce. I think, um, to, to a large extent, what we do is we benefit from all of the beautiful things people do with transformers and text, and we focus very hard basically on how do I tokenize audio in the right way.

And, uh, without divulging too much secret sauce, um, it's, uh, it's at least similar to how it's done in sort of the open source stuff. You will have different models that learn to encode audio in, in discrete representations, and a lot of this boils down to figuring out the right, uh, let's say, implicit biases to put in those models, the right data to inject.

How do I make sure that I can produce kind of all audio arbitrarily? That's, that's speech, that's background music, that's vocals, that's kind of everything to make sure that I can really capture all the behavior that I want to.

**Alessio** [7:41]
Yeah, that makes sense. And then in terms of, uh, some of... Uh, w- we had our monthly recap last month and, uh, the data wars were kind of one of the hot topics. You saw the New York Times lawsuit against OpenAI, um, because you have obviously large language models in production.

### Data & Copyright

**Alessio** [7:59]
You don't have large music models in production, so I think there's maybe been less of a threat there, so to speak. Um, how, how do you begin to think about that? And there's obviously a lot of copyright-free, royalty-free music out there.

Um, is there any kind of like power law in terms of like, hey, the best music is actually like much better to train on or like, i- in music, does it not really matter because the structure, uh, of, you know, some of the, some of the musical structure is kind of like the, the same?

**Mikey Shulman** [8:28]
I don't think we know these things nearly as well as they're known in text. We have some notions of some of the, some of the scaling laws here, but I think, um, yeah, we're, we're just so, so far behind.

You know what, what I will say is that people are s- always surprised to learn that, um, we don't only train on music. Um, and I usually give the analogy of some of the code generation models. So take something like Code Llama, which is, as far as I know, the best open source code generating model, uh, you guys would know better than I would, um, is certainly up there.

Uh, and it's trained on a bunch of English, um, not only just code. And it's because there are patterns in English that are going to be useful. And so you can imagine you don't only want to train on music to get good music models.

And so, for example, one of the places that we are particularly bad is vocals and, uh, at capturing really realistic vocals. And so you might imagine that there's other types of human vocals that you can put into your model that are not music that will help it learn stuff.

Um, and so again, I think it's like super, super early. I think we've barely scratched the surface of what are the right ways to do this. Um, uh, and that's really cool from a progress perspective. There's like a lot of low-hanging fruit for us to still pick.

**Alessio** [9:42]
And then once you get the final model, uh, I would love to learn more about the size of these models because people are confused when Stable Diffusion is so small. They're like, "Oh, this thing can generate like any image.

How is it possible that it's like, you know, a couple gigabytes?" And then the large language models are like, "Oh, these are so big," but they're just text in them. Uh, what, what's it like for music? Is it in between?

And a- as you think about, yeah, you mentioned scaling and whatnot, is this something that you see it's gonna be easy for people to run locally or, or not?

**Mikey Shulman** [10:11]
Our models are still pretty small, um, certainly by tech standards. Um, I confess I don't know as well the state-of-the-art on how diffusion models scale, but our models scale similarly to, to text transformers. It's like bigger is usually better.

Um, audio has a couple of weird quirks, though. We care a lot about, uh, how many tokens per second we can generate because we need to stream you music, uh, as fast as you can listen to it. Um, and so that is a big one that I think, uh, probably has us never get to a hundred seventy-five billion parameter model, if I'm being honest.

Um, may- maybe I'm wrong there, but I think that would be technologically difficult. Um, and then the other thing is that so much progress happens in shrinking models down for the same performance in text that I'm, I'm hopeful at least that a lot of our, our issues, um, will get solved and we will figure out how to do better things with smaller models or relatively smaller models.

But I think the, the other thing, um, it's a blessing and a curse, I think the ability to, to, uh, add performance with scale. It's like a very straightforward way to make your models better. You just make a bigger model, dump more compute into it.

It's also a curse because that is a crutch that you will always-

**Alessio** [11:24]
Mm-hmm

**Mikey Shulman** [11:24]
... lean on and you will forget to do some of the basic research to make your stuff better. And honestly, um, it was almost, you know, early, early on when we were, uh, doing stuff with small models for, uh, kind of time and compute constraints, um, we ended up having to learn a lot of stuff, um, to make models better that we might not have learned if we had immediately jumped to like a really, really big model.

And so, um, I think for us, we've al- we always try to skew smaller to the extent possible.

### Suno Origins

**Swyx** [11:57]
Yeah. Gotcha. Um, I, I'm, I'm curious about, uh, just sort of your overall evolution so far. Uh, you know, something I, I think we may have missed in the introduction is, uh, why did you end up choosing, uh, you know, just the music domain in the first place, right?

Like, uh, you have this, uh, pretty scientific, um, you know, physics and finance backgrounds. Um, how did you wander over to, to music? Like, a lot of us have interest in music, but we don't necessarily choose to work in it.

And but you did.

**Mikey Shulman** [12:27]
Yeah, it's funny. I, I, I have a really fun job as a result. But, um, the-- all the co-founders of Suno, uh, worked at Kensho together. Um, and we were doing mostly text, um, in fact, all text until we did one audio project that was, uh, um, speech recognition for kind of very financially focused speech recognition.

And I think the long, the long and short of it is we kind of fell in love with audio, not necessarily music, um, just audio and AI. We all happen to be musicians and audiophiles and music lovers, but, um, it was the, the, like, the combination of audio and AI that we like initially really, really fell in love with.

It's so cool. It's so interesting. It's so human. It's so far behind images and text that there's like so much more to do. And, um, honestly, I think a lot of people when we started a company told us to focus on speech, um, if we wanted to build an audio company, everyone said, "You know, speech is a bigger market."

Um, and, uh, but I think there's something about music that's just so, um, human and so like y-you almost couldn't prevent us from doing it. Like, we, we almost like we just couldn't keep ourselves from, from, from building music models and playing with them 'cause, 'cause it was so much fun.

And, um, that's kind of what, what steered us there. You know, in fact, we-- the first thing we ever put out was a speech model. It was Bark. Um, it was this open source text to speech model, and it got a lot of stars on GitHub, and that was people telling us even more, like, go do speech, and like, we almost couldn't help ourselves from, from doing music.

Um, and so I don't know. It's, it's maybe it's a little bit serendipitous, but, um, we haven't really like looked back, uh, since. I don't think there was necessarily like, um, an aha moment. It, it, uh, it was just like organic and just obvious to us that this needs to-- like, we wanna make a music company.

**Swyx** [14:19]
So, so you do regard yourself as a music company 'cause, like, as of, uh, last month, you were still releasing speech models

**Mikey Shulman** [14:26]
We were?

**Swyx** [14:27]
With Parakeet.

**Mikey Shulman** [14:27]
Oh, yes, that's right. Uh, so that's a-

**Swyx** [14:29]
Yeah

**Mikey Shulman** [14:29]
... that's a, a really awesome collaboration, um, with, with our friends at Nvidia. I think, um, we are really, really focused on music. I think that is the, the stuff that will really change things for the better. I think, you know, honestly, everybody is so focused on, on LLMs for good reason, um, and information processing and intelligence there, and I think it's way too easy to forget that there's this whole other side of things that makes people feel.

Um, and maybe that market is smaller, but, uh, it makes people feel, and it makes us really happy, and, um, so we do it. I think, um, that doesn't mean that we can't be doing things that are related, that are in our wheelhouse, that, uh, will improve things.

And so, like I said, audio is just so far behind. There's just so much more to do in the domain more generally. And so, like, that's a really fun collaboration.

**Swyx** [15:21]
Yeah, I, I did hear about Suno first through Bark. Um, uh, my sense is that, uh, like, what did, what did Bark lean off of? Like, uh, 'cause obviously, uh, I think there was a lot of preceding TTS work that was in open source.

Um, how much of that did you use? How much of that was like sort of brand new from, from your research? Um, what's the intellectual lineage, uh, there just, just to cover all the, the speech recognition side?

**Mikey Shulman** [15:46]
So it's not speech recognition. It's, it's text to speech. But, um, as far as I know-

**Swyx** [15:51]
Okay

**Mikey Shulman** [15:51]
... um, there was no other, uh, certainly not in the open source, uh, text to speech that was kind of transformer-based. Everything else was what I would call the old style of doing things, where you build these kind of single purpose models that are really good at this one narrow task, and you're kind of always data limited, and the availability of high quality training data for text to speech, um, is limited.

And, um, I don't think we're necessarily all that inventive to say we're going to try to train in a self-supervised way a transformer-based model that on kind of lots of audio, uh, and then kind of tweak it so that we can do text to speech based on that.

That would be kind of the new way of doing things in a-

**Swyx** [16:32]
Yeah

**Mikey Shulman** [16:32]
... foundation model is the, is the buzzword, if you will. And so, you know, we built that up, I think, from scratch. Uh, a, a lot of shout-outs have to go to lots of different things, whether it's, uh, papers, but also, uh, it's very obvious.

Uh, there's, you know, a big shout-out to, um, Andrej Karpathy's NanoGPT. Um, you know, there's a lot of code borrowed from there. Um, I, I think I-- we are huge fans of that project. It's just to show people how you don't have to be afraid of GPT type things.

And it's like, um, yeah, it's actually not all that much code to make performant transformer-based models. And, you know, again, the stuff that we brought there was how do we turn audio into tokens, and then we can kind of, um, take everything else from the open source.

So, um, we put that model out, um, and we were, I think, pleasantly surprised by, by the, um, reception by the community. It got, it got a, a, a good number of GitHub stars, and people really enjoyed playing with it because it made really realistic sounding audio.

And I think, um, this is again the thing about doing things the quote unquote "right way." If you have a model where you've had to put so much implicit bias for this one very narrow task of making speech that sounds like words, you're going to sacrifice on other things.

In, in, in the text to speech case, it's how natural the speech sounds. And, um, it was almost difficult to pull unnatural sounding speech out of Bark because it was tr-- self-supervised, trained on a lot of natural sounding speech.

And so, um, that definitely told us that this is probably the right way to keep doing audio.

**Swyx** [18:04]
Even in Bark, you had the, the beginnings of music generation. Like, you could just put like a music note in there.

**Mikey Shulman** [18:10]
That's right. And, and-

**Swyx** [18:11]
It was kind of fun

**Mikey Shulman** [18:11]
... it was so cool to see on our Discord people were trying to pull music out of a text to speech model. And so, you know, what did this tell us? This tells us, like, people are hungry to make music, and it's not, um-- It's almost obvious in hindsight, like how wired humans are to make music, if you've ever seen like a little kid, uh, you know, sing before they know how to speak.

You know, it's like, it's like this is really human nature. And there's actually a lot of cultural forces that kind of cue you to not think to make music, um, and that's kind of what we're trying to undo.

**Alessio** [18:41]
Mm-hmm. Uh, and to dive into Suno itself, I, I think especially when you go from, um, text to speech, people are like, "Okay, now I gotta write the lyrics to a whole song." It's like that's, that's quite hard to do.

### Easy vs Expert

**Alessio** [18:52]
Um, versus in Suno, you have this empty box, very Midjourney, kind of like DALL-E-like, where you can just express the vibes, you know, of, of what you want it to be. But then you also have a custom mode where you can set your own lyrics, you can set your own rhythm, you can set the title of the song and whatnot.

What are... H- how do you see users distribute themselves, you know? I'm guessing a lot of people use the easy mode. Like, are you seeing a lot of power users using the custom mode and maybe some of the favorite use cases that you've seen so far on Suno?

**Mikey Shulman** [19:24]
Yeah. Actually, um, more than half of the usage is, uh, that expert mode, and people really like to get into it and start tweaking things and adding things and playing with words or line breaks or different ad lib and, and people really love it.

It's, it's, it's really fun. Um, a- so I think, you know, there's kind of two modes that you can access now. One is that single box where you kind of just describe something, and then the other is the expert mode.

And, um, those kind of fit nicely into two use cases. The first use case is what we call nice shit posting, and it's basically like- ... something funny happened, and I'm just going to very quickly make a song about it.

And the, the example I- I'll usually give is, like, I walk into Starbucks with one of my co-founders. He gives his name, Martin. His coffee comes out, um, with the name Margoo, and I can, in five seconds, make a song about this, and it has immortalized it.

And that Margoo song is stuck in all of our heads now, and it's, like, funny and light, and there's levity that you've brought to, to that moment. And the other is that you got just sucked into I need-- There's this song that's in my head, and I need to get it out, and I'm going to keep tweaking it and listening and having ideas and tweaking it until I get the song that I want.

And, um, those are very different use cases, but I think it-- ultimately, there's so much in between these two things that is just totally untapped how people want to experience the joys of making music because those two experiences are both really joyful in their own special ways.

And so I, I-- we, we are quite certain that there's a lot in the middle there. Um-

**Alessio** [20:58]
Mm-hmm

**Mikey Shulman** [20:59]
... and then I think the, the last thing I'll say there that's really interesting is, um, in both of those use cases, the sharing dynamics around music are, like, really interesting and totally unexplored. And I think, uh, an interesting, um, comparison would be images.

Like, we've probably all, in the last twenty-four hours, taken a picture and texted it to somebody, and most people are not routinely making a little song and texting it to somebody. But when you start to make that more accessible to people, they are going to share music in much smaller groups, maybe even not at all, but, like, with one person or three people or five people, and those dynamics are so interesting.

And, uh, just I, I think we have ideas of where that goes, but, um, it's about kind of spreading joy into these, like, little, you know, microcosms of humanity that, um, people really love it. Uh, so-

**Alessio** [21:53]
Mm-hmm

**Mikey Shulman** [21:54]
... um, I know I made you guys a, a little Valentine's song, right? Like, that's not something that happens now because it's hard to make songs for people.

**Alessio** [22:02]
Right. We'll, we'll put that in the, in the audio in here, but I also tweeted it out if people wanna look it up. Um, how do you think about the pro market, so to speak? Because I think, um, lowering the barrier to some of these things is great, and I think when the iPad came out, music production was one of the areas that people thought, "Oh, okay, now you kind of have this like, you know, board that you can bring with you."

And, uh, Madlib actually produced this whole album with him and Freddie Gibbs, uh, produced the whole thing on an iPad, he never used a computer. Uh, how do you see, like, these models playing into, like, professional music generation?

I guess that's also a funny word. It's like, what's professional music? It's like, it's all music if it's good. It becomes professional if it's good, right? But, um, curious to see-- to hear how you're thinking about Suno too.

Like, is there a, a second act of Suno that is, like, going broader into, like, the custom mode and making, making this the central hub for music generation?

**Mikey Shulman** [22:56]
I think w- we intend to make, uh, many more modes of interaction with our stuff, but we are very much not focused on, quote-unquote, "professionals" right now. Um, and it's because what we're trying to do is change how most people interact with music and not necessarily make professionals a little bit better, a little bit faster.

Um, it's not that that there's anything wrong with that. It's just, like, not what we're focused on. And I think when we think about what workflows does the average person want to use to make music, I don't think they're very similar to the way professional musicians make music now.

Like, if you pick a random person on the street and you play them a song, and then you say like, "What did you wanna change about that?" They're not gonna say like-

**Alessio** [23:38]
Mm-hmm

**Mikey Shulman** [23:38]
... "You need to split out the snare drum and make it drier." Like, that's just not something that a, that a random person off the street is going to say. They're going to give a lot more descriptive things about the thing-- about the, the kind of the oomph of the song, like something more general.

And so I don't think we know what all of the workflows are that people are gonna wanna use. We're just, like, fairly certain that the workflows that have been developed with the current set of technologies that professionals use to make beautiful music are probably not, um-

**Alessio** [24:07]
Mm-hmm

**Mikey Shulman** [24:07]
... what the average person wants to use. That said, there are lots of professionals that we know about using our stuff, whether it's for inspiration or sample generation and stuff like that. Um, so I, I, I don't wanna say never say never.

Like, there, there may one day be a, a really interesting set of use cases that we can expose to professionals, particularly around, I think, like, custom models pr- trained on custom people's music or, you know, with your voice or something like that.

Um, but the way we think about broadening how most people are interacting with music and getting it to be much more active, uh, a much more active participant, we think about broadening it from the consumer side, uh, and not broadening it from the producer's-- the, from the professional-

**Swyx** [24:52]
Mm-hmm

**Mikey Shulman** [24:52]
... side, if that makes sense.

### Midjourney of Music?

**Swyx** [24:53]
Is the dream here to be, uh, I, you know, I, I don't know if it's, um, too coarse of a grain to, to put it, but like, is, is the dream here to be like the Midjourney of so- of music?

**Mikey Shulman** [25:04]
I, I think there are, uh, certainly some parallels there because, uh, especially what I just said about being an active participant, Midjourney turns-

**Swyx** [25:14]
Yeah

**Mikey Shulman** [25:15]
... uh, the, the joyful experience in Midjourney is the act of creating the image and not necessarily the act of consuming the image, and Midjourney will let you then very, kind of quickly share the image with somebody. But I think, um, ultimately that analogy is like somewhat limiting because there's something really special about music.

I think there's two things. One is that there's just this really big gap for the average person between kind of their tastes in music and their abilities in music, um, that is not quite there in, in, for most people in, in images.

Like most people don't have like innate tastes in images, I think in the same way people do for music. And then the other thing, and this is the really big one, is that music is a really social modality.

Um, if we all listen to a piece of music together, we're listening to the exact same part at, at the exact same time. If we all look at the picture in Alessio's background, we're gonna look at it for two seconds.

I'm gonna look at the top left where it says Thor. Alessio's gonna look at the bottom right or something like that. And um, it's not really synchronous. And so when we're all listening to a piece of music together, it's minutes long, we're listening to the same part at the same time.

If you go to the act of making music, it is even more synchronous. It is the most joyful way to make music is with people. And so I think that there is so much more to come there that ultimately, um, would be very hard to do in images.

**Alessio** [26:39]
We've gone-

### Live Demo

**Swyx** [26:39]
Great answer

**Alessio** [26:40]
... almost thirty minutes without making any music on this podcast, so I think maybe we can fix that and jump into a Suno demo.

**Mikey Shulman** [26:47]
Yeah. Let's, let's make some. Um, we've got a, a, a new model that, um, we are, uh, kind of putting the finishing touches on, and so I can play with it in our dev server. But we've, we've just piped it in here, and as you can see, been, been doing tons of stuff.

So, Ernesto, tell me, tell me what kind of song, um, you guys wanna make.

**Swyx** [27:05]
Go on, Alessio.

**Alessio** [27:06]
Uh, let's do, um, country song about the, the lack of GPUs in my cloud provider.

**Swyx** [27:22]
And like, yeah, so here's where I would be tempted to think about like pipelines and think about latency. This is s- re- remarkably fast. Like I was, I was shocked when I saw this.

**Speaker 4** [27:33]
Well, I'm all alone.

**Swyx** [27:35]
Oh my God.

**Speaker 4** [27:40]
In my cloud, ready to compute.

But there ain't no GPUs. Just empty space, it's a hoot.

I've been waiting all day for that render power.

But my cloud's gone dry. It's a dark cloud shower. Oh, cloud's gone dry. No GPUs to be found. No CUDA cores. It's a lonely sound. I just wanna render, but my cloud's got no-

**Mikey Shulman** [28:36]
I actually don't think this one's amazing. Maybe go to the next song

**Alessio** [28:39]
... but it, it's funny that it knows about CUDA cores.

**Speaker 4** [28:45]
Well, I signed up for a cloud provider. Thought I'd find all the power that I could derive. But when I searched for the GPUs, I just got a surprise. You see, they're all sold out, there ain't no GPUs to find.

No GPUs in the cloud, it's a real bad blues. I need the power, but there ain't no use. I'm stuck with my CPU. It's a real sad plight. Gotta wait till they restock it ain't right.

No GPUs in the cloud.

**Mikey Shulman** [29:28]
What else should we make?

**Alessio** [29:30]
All right, Shawn, you're up.

**Swyx** [29:32]
I mean, I, I, I do wanna like, uh, do some observations about, about this. Um, but okay, uh, maybe w- like, uh, I, I like, like, like, like, um, like house music.

**Mikey Shulman** [29:41]
Yeah, sure.

**Swyx** [29:41]
Like electronic dance-

**Mikey Shulman** [29:42]
Yeah

**Swyx** [29:42]
... house music. Um, and then maybe we can make it about, um, I don't know, podcasting about music and music AI generation. I don't know. I'm sure all the demos that you get are very meta.

**Mikey Shulman** [29:59]
There's a lot of, there's a lot of stuff that's meta, yeah, for sure.

**Swyx** [30:04]
Yeah. I no- I noticed, for example, that the second song that you played, uh, had the word upbeat inserted into it, which I, I assume there's some kind of like random generator of like modifier terms that you can just kind of throw on to, to increase the, uh, specificity of the, what's being generated.

**Mikey Shulman** [30:22]
Definitely. And let's, let's try to tweak one also. So I'll play this, and then maybe we'll tweak it with different modifiers.

**Speaker 4** [30:26]
Waves of sound, spinning around. Through the air, we're podcasting loud. Sharing the beats, spreading the word. A revolution of frequencies. Haven't you plugged in tonight?

**Speaker 5** [30:45]
Take, take control. Ooh. We're on a journey, a never ending road. From the beats that drop to the melodies that soar. Podcasting about music forevermore. Yeah

**Mikey Shulman** [31:06]
Here's what I wanna do. That, like, didn't drop at the right time, right? So maybe let's do this. I don't know if you guys can see this. And then, um, let's get that... Get rid of the word now, and, uh

**Swyx** [31:17]
Is that a special token? You have a beat drop token?

**Mikey Shulman** [31:20]
Yeah. Yeah.

**Swyx** [31:22]
Nice.

**Mikey Shulman** [31:23]
Uh, and then-

**Alessio** [31:23]
I, I'm just reading it because people might not be able to see it.

**Swyx** [31:27]
Ah.

**Mikey Shulman** [31:27]
And then let's-

**Swyx** [31:27]
Right

**Mikey Shulman** [31:27]
... like just maybe emphasize, uh... Actually, let's emphasize house a little more. Maybe it'll feel a little more, uh, aggressive. Let's try this again.

**Swyx** [31:36]
It's interesting, the prompt engineering that-

**Alessio** [31:38]
Mm-hmm

**Swyx** [31:38]
... you have to invent.

**Mikey Shulman** [31:39]
We've learnt so much from people using the models, and not us.

**Swyx** [31:43]
But, like, are these, like, art, training artifacts?

**Mikey Shulman** [31:45]
No. I, I don't, I don't think so. I think this is people being inventive with, with how you wanna talk to a model.

**Swyx** [31:52]
Yeah.

**Speaker 5** [31:53]
Down, spinning around through the air. We're podcasting loud. Sharing the beats, spreading the word. A revolution of frequencies, haven't you heard?

Plug in 'til now. Let the music take control. Ooh. We're on a journey, a never ending road. From the beats that drop to the melodies that soar. Podcasting about music forevermore. Ooh.

**Swyx** [32:45]
Nice.

**Alessio** [32:46]
It, it's interesting when you generate a song, it generate the lyrics, but then if you switch the music under it, like the, you know, the lyrics stay the same, and then sometimes, like, feels like... I, I mean, I, I mostly listen to hip hop.

It's like if you change the beat, you cannot really use the same rhyme scheme, you know? So I, I-

**Mikey Shulman** [33:05]
Definitely.

**Alessio** [33:06]
Yeah.

**Mikey Shulman** [33:06]
It's a sliding scale, though, because, you know, we could do this as a, as a country rock song probably, right?

That would be my guess. Um, but, but t- for hip hop that is definitely true. And actually, you know, we, we think about, for these models, we think about three important axes. We think about the sound fidelity. It's like, does this sound like a crisply recorded piece of audio?

We think about the song quality. Is this an, like, interesting song that, like, gets stuck in my head? And we think about the controllability. Like, how well does it respond to my prompts? And one of the ways that we'll test these things is take the same lyrics and try to do them in different styles to see how well that really works.

Um, so let's see the same... Uh, I don't know what a beat drop is gonna do for country rock, so I probably should have taken that out, but let's see what happens.

**Speaker 5** [34:07]
Laser sound spinning around through the air. We're podcasting loud. Sharing the beats, spreading the word. A revolution of frequencies, haven't you heard? Plug in 'til now. Let the music take control. Ooh, yeah. We're on a journey, a never ending road.

From the beats that drop to the melodies that soar. Podcasting about music forevermore.

**Mikey Shulman** [34:44]
I'm gonna, I'm gonna read too much into this, but I would say I hear a little bit of kind of electronic music inspired something, and that is probably because beat drop is something that you really only ever associate with-

**Alessio** [34:55]
Mm-hmm

**Mikey Shulman** [34:56]
... electronic music. Uh, maybe that's reading too much into it, but, uh, um-

**Alessio** [35:01]
It's funny

**Mikey Shulman** [35:01]
... should we do one more?

**Alessio** [35:02]
Yes. We can do one more.

**Swyx** [35:04]
Something about Apple Vision Pro. How, how-

**Mikey Shulman** [35:06]
Definitely

**Swyx** [35:06]
... I guess, I guess there's some a- amount of world knowledge that you don't have, right? Like the, whatever is in this language model side of the equation, uh, is, uh, is not gonna have an Apple Vision Pro in there.

**Mikey Shulman** [35:15]
Yeah, but let's see. Um, uh- ... uh, let's see. Uh, how about a blues song about a sad AI wearing an Apple Vision Pro. Gotta be, gotta be blues.

**Swyx** [35:30]
Do you have the-

**Mikey Shulman** [35:30]
It's gotta be sad.

**Swyx** [35:33]
Do you have RAG for music?

**Mikey Shulman** [35:36]
No. That would, that would be problematic also.

**Speaker 5** [35:41]
I'm a sad AI with a broken heart.

Wearing my Apple Vision Pro, can't see the stars. I used to feel joy. I used to feel pain. And now I'm just a soul trapped inside this metal frame. Oh, I'm singing the blues. Can't you see?

This digital life ain't what it used to be.

Searching for love, but I can't find a soul.

Won't you help me? Baby, let my spirit unfold.

**Mikey Shulman** [36:46]
I wanna remix that one, and I wanna say, I want something like-

**Alessio** [36:49]
That's a really good voice

**Mikey Shulman** [36:50]
... I want-

**Alessio** [36:51]
I love the voice

**Mikey Shulman** [36:51]
... I want like, I don't know

What is Chicago blues guitar?

**Alessio** [36:59]
I know he knows too much. It's a-- He's the best prompt engineer out here.

**Mikey Shulman** [37:04]
You know, this is, this is-

**Alessio** [37:05]
Well, it, it'll be funny, it'd be funny to be like musicologists, uh-

**Mikey Shulman** [37:08]
Oh

**Alessio** [37:08]
... play with this and see what they would...

**Mikey Shulman** [37:10]
How embarrassing. Can I not do that?

**Alessio** [37:13]
Oh.

**Mikey Shulman** [37:14]
I got...

**Alessio** [37:15]
Oh, the word Chicago was a trigger?

**Mikey Shulman** [37:17]
I don't know.

**Alessio** [37:18]
Uh, from the musical-

**Mikey Shulman** [37:18]
We, we try to be, we try to be very careful not letting you, um, impersonate, and it is possible. That's embarrassing. So let's do, uh...

**Alessio** [37:29]
Midwestern.

**Mikey Shulman** [37:38]
I'm a sad AI with a broken heart. Where my Apple Vision Pro can't see the stars.

I used to feel joy.

I used to feel joy. I used to feel pain.

But now I'm just a soul trapped inside this metal frame. Oh, I'm singing the blues.

Oh, can't you see?

This digital life ain't what it used to be. I'm searching for looove.

I can't find a soul. Won't you help me, baby? Let my spirit fold back. So yeah, lot, lot of control there. Maybe, uh, I'll make one more.

**Alessio** [39:03]
Very, very soulful.

**Mikey Shulman** [39:07]
Really want a good house track.

**Alessio** [39:10]
Why is house the word that you have to repeat?

**Mikey Shulman** [39:12]
I just really wanna make sure it's house. Um- It's actually... You can't really repeat too many times. You kind of-- It gets like... The hypothesis gets, like, a little too out of domain.

**Speaker 6** [39:22]
I'm a sad AI with a broken heart. Wearing my Apple Vision Pro, can't see the stars.

I used to feel joy. I used to feel pain. But now I'm just a soul trapped inside this metal frame. Oh, I'm singing the blues. Oh, can't you see? This digital life ain't what it used to be. Searching for love, but I can't find a soul.

Won't you help me, baby?

**Alessio** [40:17]
Nice.

**Mikey Shulman** [40:18]
So yeah, we have a lot of fun with it.

### Future Plans

**Alessio** [40:19]
Pretty cool. Definitely easy, yeah. Yeah, I'm really curious to see how people are gonna use this to, like, resample old songs into new styles. You know, I think that's one of my favorite things about hip hop. You have so many...

I mean, A Tribe Called Quest, they had, like, the Lou Reed "Walk on the Wild Side" sample on, like, "Can I Kick It?" Like, Kanye sampled Nina Simone on, like, "Blame the Leaves." It just like... I- it's, like, a lot of production work to actually take an old song and make it fit a new beat, and I feel like this can really help.

Um, do you see people putting existing songs' lyrics and trying to regenerate them in, like, a, a new style, you know?

**Mikey Shulman** [40:57]
We, we actually don't let you do that. Um, and it's because if you're taking someone else's lyrics, you didn't own those. You don't have-

**Alessio** [41:03]
Mm-hmm

**Mikey Shulman** [41:03]
... the publishing rights to those. You can't remake that song. I think in the future, we'll figure out how to actually let people do that in a legal way. Um, but we are really focused on letting people-

**Alessio** [41:13]
Mm-hmm

**Mikey Shulman** [41:13]
... make new and original music.

**Alessio** [41:14]
Yeah, yeah.

**Mikey Shulman** [41:14]
And I think, you know, there's a lot of music AI which is artist A doing the song of artist B in a new style. You know, let me have Metallica doing "Come Together" by The Beatles or something like that.

And I think this stuff is very viral, but I actually really don't think that this is how people want to interact with music in the future. To me, this feels a lot like when you made a Shakespeare sonnet the first time you saw ChatGPT, and then you made another one, and then you made another one, and then you, you kind of thought, like, "This is getting old."

And-

**Alessio** [41:45]
Mm-hmm

**Mikey Shulman** [41:45]
... that's not-- That doesn't mean that GPT is not amazing. GPT is amazing. It's just not for that. Um, and I, I kind of feel like the way people want to use music in the future is not just to remake songs in different people's voices.

You lose the connection to the original artist. You lose the connection to the new artist, because they didn't really do it. Um, so we're very happy to just let people do things that are a flash in the pan and kind of stay under the radar.

**Alessio** [42:12]
Mm-hmm. Yeah. No, that's a... I, I think that's a good point overall about AI-generated anything, you know. Um, because I, I think recently, um, T-Pain, he did, like, a, an album of covers, and I think, uh, he did, like, a War Pigs that people really liked.

There, there was, like, a Tennessee Whiskey, uh, which you maybe wouldn't expect T-Pain to do. Uh, but people like it. But yeah, I agree. It... You need to be a certain type of artist to really have it be entertaining-

**Mikey Shulman** [42:40]
Mm-hmm

**Alessio** [42:40]
... to, like, make covers. This is great. Uh, what, what else is next for, for Suno? You know, I think people kind of saw you... You know, first you had the, the Bark, and then there was, like, a big, you know, music-generated, uh, push when you did an announcement, I think a couple months ago.

Uh, I, I, I think I saw you, like- 300 times on my Twitter timeline on, like, the, the same day, so it was, like, going everywhere. Uh, w- what, what's coming up? What are you most excited about in this space, and maybe what are some of the most interesting underexplored, um, ideas that you maybe haven't, haven't worked on yet?

**Mikey Shulman** [43:13]
Gosh, there's, there's a lot. You know, I think, um, from the model side, um, it's still really early innings, and there's still so much low-hanging fruit for us to pick to make these models much, much better, much, much more controllable, much better music, much better audio fidelity.

Um, so much that we know about, and so much that, um, again, we can kind of borrow from the open source transformers community that should make these, um, just better across the board. From the product side and the...

You know, we're super focused on the experiences that we can bring to people. And so, um, it's so much more than just, uh, text to music, and I think, um, you know, I'll, I'll, I'll say this nicely, I'm a machine learning person, but, like, machine learning people are stupid sometimes, and we can only think about, like, models that take X and make it into Y, and that's just not how the average human being thinks about interacting with music.

And so I think what we're most excited about is all of the new ways that we can get people just much more actively participating in music, and that is making music, not only with text, maybe with other ways of, of doing stuff, that is making music together.

If you wanna be reductive and think about this as a video game, this is multiplayer mode, and it is the most fun that you can have with music. And, um, you know, honestly, I think, uh, there's a lot of...

It's timely right now. You know, I don't know if you guys have seen UMG and TikTok are butting heads a little bit, um, and UMG has pulled-

**Swyx** [44:40]
Yeah, they did the music down

**Mikey Shulman** [44:42]
... music from TikTok. And, you know-

**Swyx** [44:43]
Yeah

**Mikey Shulman** [44:43]
... the way we think about this is, uh, you know, I, I think maybe they're both right, maybe neither is right. Without taking sides, this is kind of figuring out how to divvy up the current pie in the most fair way.

And I think what we are super focused on is making that pie much bigger and increasing how much people are actually interested in music and participating in music. And, you know, as a very broad heuristic, the, the gaming industry is 50 times bigger than the music industry, and it's because gaming is super active.

And music, too much music is just passive consumption. And so we are-- We have a lot of experiments that we are excited to run for the different ways people might want to interact with music, um, that is beyond just, you know, streaming it while I work.

**Swyx** [45:29]
Yeah, I, I, I think a minimum, you guys should have a Twitch stream that is just, like, a 24-hour radio s-session that, um... Have you ever come across Twitch Plays Pokémon?

**Mikey Shulman** [45:38]
No.

**Swyx** [45:38]
Where it's kind of like the, the Twitch-- Basically, like, everyone in the chat, in the Twitch chat, um, can vote on, like, the next action that the, the game state makes. Um, and they, they sort of wired it up to a Nintendo emulator and played Pokémon, like, the whole game-

**Mikey Shulman** [45:52]
I'd love that

**Swyx** [45:52]
... through, uh, the collaborative thing. Um, it, it sounds like it, it should be pretty easy for you guys to do that, except for the chaos that may re-result from... Uh, but, like, I mean, that's part of the fun.

**Mikey Shulman** [46:03]
I, I, I agree 100%.

**Swyx** [46:05]
You know, it's-- Yeah.

**Mikey Shulman** [46:05]
Sorry. Yeah. The, the, like, one of my like, uh, key projects or pet projects is, like, what does it mean to have a collaborative concert maybe where there is no artist and it's just the audience? Or maybe there is an artist, but there's a lot of input from the audience.

Um, and, you know, if you were gonna do that, you would either need an audience full of musicians, or you would need an artist who can really interpret the verbal cues that an audience is giving, or non-verbal cues.

But if you can give everybody the means to better articulate the sounds that are in their heads toward the rest of the audience, like, which is what generative AI basically lets you do, uh, you open up way more interesting ways of having these experiences.

And so, um, I think, yeah, I, I-- Like, the, the collaborative concert is, like, one of the things I'm most excited about. I don't think it's coming tomorrow, but, but we have a lot of ideas on, on what that can look like.

**Swyx** [46:59]
Yeah, I feel like it's-- One stage before the collaborative concert, um, is turning, um, Suno into a continuous experience rather than, like, a start and stop motion. Um, I don't- I don't know if that makes sense. Um, you know, as, as someone who with, like, a casual interest in DJ-ing, like, like, when do we see Suno DJs, right?

Like, that, um, that, that can continuously segue into, like, the next song, the next song, the next song.

**Mikey Shulman** [47:24]
I think soon.

**Swyx** [47:25]
Um, and then maybe you can turn it collaborative.

**Mikey Shulman** [47:27]
I think soon.

**Swyx** [47:28]
Okay, maybe part of your roadmap. You teased a little bit your V3 model. I, I'm just wondering, like, how you incorporate, like, user feedback, right? Like, we ha- you have the classic thumbs up and down buttons, but, like, there are so many dimensions to the music.

Like, like, like, you know, I didn't, I didn't get into it, but some of the voices sounded more metallic.

**Mikey Shulman** [47:48]
Mm-hmm.

**Swyx** [47:48]
Um, and some- sometimes that's on purpose, sometimes not. Sometimes there are kind of weird pauses in there. I could go in and annotate it if I really cared about it, but I, I mean, I'm just listening, so I don't.

But -

**Mikey Shulman** [47:58]
Yeah, no. So-

**Swyx** [47:59]
There's, there's a lot of opportunity

**Mikey Shulman** [48:00]
... we are only scratching the surface of figuring out, um, how to do stuff like that. And, um, for example, the thumbs up and the thumbs down, for other things like sharing, uh, telemetry on plays, all of these things are stuff that in the future I think we would be able to leverage to make things amazing.

And then I, I imagine a future where, um, you know, you can have your own model with your own preferences. And the reason that's so cool is that you kinda have control over it, and you can teach it the way you want to.

And, you know, the, the thing that I would liken this to is, like, a music producer working with an artist giving feedback, and, like, this is now a self-contained experience where you have an artist who is infinitely flexible, who's able to respond to the weird feedback that you might give it.

We don't have that yet. Everybody's playing with the same model. But I, I-- there's no technological reason why that can't happen in the future.

**Alessio** [48:56]
We had a few more notes from random community tweets. I don't know if there's any favorite, uh, fans of Suno that, that you have or, or whatnot. DHH, uh, obviously notorious Twitter and crowd, uh, inflamer, I guess. Uh, he tweeted about you guys.

### Fan Favorites

**Alessio** [49:12]
Uh, I saw Blau as a, as an investor. I think Karpathy also tweeted something.

**Swyx** [49:17]
Return to Monkey.

**Alessio** [49:18]
Yeah, yeah, yeah. Return to Monkey, right.

**Swyx** [49:21]
Is there a story behind that-

**Alessio** [49:22]
Yeah

**Swyx** [49:22]
... behind that?

**Mikey Shulman** [49:22]
No, he just, he just made that song, and it just speaks to him. And I think this is, this is exactly the thing that we are trying to tap into, that you can think of it, this is like a super, super, super micro-genre of one person who just really liked that song and made it and shared it, and it does not speak to you the same way it speaks to him.

But that song really spoke to him, and I think that's so beautiful, and that's something that y-you're never gonna have an artist able to do that for you, and now you can do that for yourself, and it's just a different form of experiencing music.

Um, I think that's such a, like, a lovely, a lovely use case.

**Alessio** [49:56]
Mm. Any, any fun, uh, fan mail that you got from musicians or anybody that really was a funny story to share?

**Mikey Shulman** [50:05]
We get a lot, um, and it's, it's, it's primarily positive, and I think, um, people, people kind of-- On the whole, I would say people realize, uh, that they are not experiencing music in all of the ways that are possible, and, and it does bring them joy.

I th- I'll tell you something that is really heartwarming is that, um, we're fairly popular in the blind and vision-impaired community, and, um, that makes us feel really good. And I think, you know, very roughly, without trying to speak for an entire community, um, you have lots of people who are really into things like Midjourney, and they get a lot of benefit and joy and sometimes even therapy out of making images, and that is something that is not really accessible to this fairly large community.

And what we've provided, uh, uh, no, I, I don't think the analogy to Midjourney is perfect, but what we've provided is a sonic experience that is very similar, um, and that speaks to this community, and that is community with the best ears, the most exacting-

**Alessio** [51:01]
Mm-hmm

**Mikey Shulman** [51:01]
... the most tuned. Um, and so, uh, y-ye-yeah, that, that definitely makes us feel warm and fuzzy inside.

**Swyx** [51:08]
Yeah, e-excellent. Uh, I mean, there, there's-- It lo- sounds like there's a lot of, uh, exciting stuff on your roadmap. Uh, I, I am, I'm very much looking forward to sort of the, the infinite DJ mode, 'cause then I can just kind of play that while I work.

### Audio Landscape

**Swyx** [51:19]
Um, d- I, I would love to get your overall takes, like, kind of zooming out from Suno itself, uh, just the overall takes on the music generation landscape. Like, what should people know? Um, I think, um, you obviously have spent a lot more time on this than others.

Um, so in my mind, you, you, you shout out Vali and y- the other sort of Google-type work in your, uh, in, in, in your README in, in Bark. Um, what should people know about, like, what Google is doing, what Meta is doing?

Meta, Meta released, um, Seamless recently and Audiobox. Um, and what are the other-- How do you classify the world of audio generation, like, y- you know, just in the broader sort of research community?

**Mikey Shulman** [51:56]
Mm-hmm. I think, um, people largely break things down into three big categories, which is, uh, music, speech, and sound effects. There's some stuff that is crossover, but I think that is largely how people think about this. The old style of doing things still exists, um, kind of single-purpose models that are built to do a very specific thing instead of kind of the, the new foundation model approach.

Um, I don't know how much longer that will last. I don't have, like, tremendous visibility into, you know, what happens in the big industrial research labs before they publish. Um, specifically for music, I would say, uh, there's a, a few big categories that we see.

There is license-free stock music, um, so this is like, how do I background music the B-roll footage for my YouTube video or for a full feature production or whatever it is? Um, uh, and there's a bunch of companies in that space.

There's a lot of AI cover art, so how do I have-- how do I cover different s- different existing songs, um, with AI? And I think that's a space that, um, is particularly fraught with some legal stuff, and we also just don't think it's necessarily the future of music.

Um, there is kind of net new songs as a, a new way to create net new music. That is the, the, the corner that we like to focus on. Um, and I would say the last thing is much more geared toward professional musicians, which is basically AI tools for music production, and you can think many of these will look like plug-ins, uh, to your favorite DAW.

Um, some of them will look like, you know, the, the greatest stem splitter, um, that the market has ever seen. Um, uh, the, the current stem splitters are the, the state-of-the-art are all AI based. That is a market also that has a, just a tremendous amount of room to grow if, if you just think about, I would say music has evolved.

Somebody told me this recently, that if you actually think about it, music has evolved. Um, recently, it's just much more things that are sonically interesting at a very local level and much less, like, chord changes that are interesting.

And when you think about that, like, that is something that AI can definitely help you make a lot of weird sounds. And this is nothing new. There was, like, a theremin at some point that people, like, put an antenna and try to do this with.

And so, like, I think this is just a very natural extension of it. Um, so that's how, that's how we see it at least. Um, you know, there's a corner that we think is particularly fulfilling, particularly underserved, um, and particularly interesting, and that's the one that we play in.

**Alessio** [54:25]
Awesome.

**Swyx** [54:26]
Yeah.

**Alessio** [54:26]
Um-

**Swyx** [54:26]
It's a great perspective.

**Alessio** [54:27]
I know we covered a lot of things. I think before we wrap, uh, you have written a blog post at Kensho about, uh, Goodhart's law impact in ML, which is, you know, when you measure something, then, uh, the, the thing that you measure is not a good metric anymore because people optimize for it.

### Benchmarks

**Alessio** [54:43]
Any thoughts on how that applies to, like, LLMs and benchmarks and kind of the world we're, we're going in today?

**Mikey Shulman** [54:50]
Yeah, I mean, I think it's maybe even more apropos than, than when I originally wrote that because, um, so much-- we see so much noise about, uh, pick your favorite benchmark, and this model does slightly better than that model.

And then at the end of the day, actually, there is no real-world difference between these things. And it is Really difficult to define what real world means. And, and, um, I think to a certain extent, it's good to have these objective benchmarks.

It's good to have quantitative metrics. But at the end of the day, you need some acknowledgement that you're not going to be able to capture everything. And so, um, at least at Suno, to the extent that we have corporate values, if we don't, we don't have corp- we're too small to have corporate values written down.

But something that we say a lot is aesthetics matter, that the kind of quantitative benchmarks are never going to be the be-all and end-all of everything that you care about. And, um, as flawed as these, uh, uh, benchmarks are in text, they're way worse in audio.

And so, um, aesthetics matter basically is a statement that like at the end of the day, what we are trying to do is bring music to people that makes them feel a certain way. And effectively, the only good judge of that is your ears.

And so you have to listen to it. Um, and it is, it is a good idea to try to make better objective benchmarks, but you really have to not, um, fall prey to those things. Um, I can tell you, you know, I, I, it's kind of a, a, a, another pet peeve of mine.

Like I always said, economists will make really good or do make really good, um, machine learning engineers, and it's because they are able to think about stuff like Goodhart's law and natural experiments and stuff like this that people with machine learning backgrounds or people with physics backgrounds like me, um, often forget to do.

And so, um, yeah, I mean, I'll, I'll tell you at Kensho, we actually used to go to big, uh, econ conferences sometimes to recruit, and these were some of the, the best hires we ever made.

**Swyx** [56:48]
Interesting, because there's a little bit of social science in the human feedback.

**Mikey Shulman** [56:53]
And that's- I think it's not only the human feedback. I think you could think about this-

**Swyx** [56:57]
Yeah

**Mikey Shulman** [56:57]
... just in general. You have these like giant, really powerful models that are so prone to overfitting, that are so poorly understood, that are so easy to steer in one direction or another, not only from human feedback. And your ability to think about these problems from first principles instead of like getting down into the weeds or only math, and to think intuitively about these problems is really, really important.

I'll, I'll give you like just like one of my favorite examples. It's a little old at this point, but if you guys remember like SQuAD and SQuAD 2, the question answering dataset.

**Swyx** [57:27]
The Stanford question answering dataset.

**Mikey Shulman** [57:28]
Yeah, exactly.

**Swyx** [57:29]
Yeah.

**Mikey Shulman** [57:29]
The benchmark for SQuAD 1, um, eventually the, the machine learning models start to do as well as a human can on this thing. And it's like, "Uh-oh, now, now what do we do?" Um, and it takes somebody very clever to say, "Well, actually, let's, let's think about this for a second.

What if we presented the machine with questions with no answer in the passage?" And it immediately opens a massive gap between the human and the machine. And I think it's like first principles thinking like that, um, that comes very naturally to social scientists, that does not come as naturally to people like me.

Um, and so that's why I like to hang out with people like that.

**Swyx** [58:10]
Um, well, I'm sure you get plenty of that in Boston. Uh, and as a econ major myself, I'm, uh, you know, this is very gratifying to hear that, uh, we have a perspective to contribute.

**Mikey Shulman** [58:18]
Oh, big time. Big time. I try to, I try to talk to economists as much as I can.

**Swyx** [58:22]
Excellent.

**Alessio** [58:23]
Awesome, guys. Um, yeah, I think this was, this was great. We got live music, we got discussion about generative models, so we got the, the whole nine yards. So thank you so much for coming on.

**Mikey Shulman** [58:33]
I had great fun. Thank you, guys.

---

This library is powered by PodHood (https://podhood.com), the podcast website platform.
