Intro0:00
Hey, I'm here in New York with Kevin Ben-Smith of Snipd. Welcome.
Hi. Hi. Amazing to be here.
Yeah. This is our first ever, I think, outdoors, uh, podcast recording.
Well, it's quite a location for the first time- ... I have to say.
I was actually unsure because, you know, it's cold. It's like-- I checked the temperature. It's, like, kind of one, one degree Celsius, but it's not that bad with the sun.
No, it's quite nice.
Yeah, yeah.
Yeah, especially with our beautiful tea.
With the tea, yeah. Perfect. We're gonna talk about Snipd. Uh, I'm a Snipd user. I had to basically... You know, apart from Twitter, it's, like, the, the number one used app on my phone.
Nice.
Um, when I, when I wake up in the morning, I, I open Snipd and I, you know, see what's, what's new. And I, I, I think in terms of time spent or usage on my phone, like, it's, it's, uh, it's...
I think it's number one and number two.
Nice. Nice.
So, so I, I really had to talk about it also because I think, like, people interested in AI want to think about, like, how can they... And we're an AI podcast, we have to talk about the AI podcast app.
Um, but before we get there, we just finished the AI Engineer Summit, and you came for the, the two days. How was it?
AI Summit1:01
Uh, it was quite incredible. I mean, for me, the most valuable was just being in the same room with like-minded people who are building the future and who are seeing the future.
Yeah.
You know, especially when it comes to AI agents, it's-- So often I have conversations with friends who are not in the AI world, and it's, like, so quickly it happens that you-- it sounds like you're, you're talking in science fiction.
And, uh, it's just crazy talk. It was... You know, it's so refreshing to, to talk with so many other people who already see these things and, um, yeah, be inspired then by them and not always feel like, like, "Okay, I think I'm just crazy, and, like, this will never happen."
It really is happening, and, uh, for me, it was very valuable.
So day two more relevant, more relevant for you than day one?
Yeah, day two. So day two was the engineering track.
Yep.
Uh, that was definitely the most valuable for me, like, also as a practitioner myself.
Yeah.
Especially there were one or two talks that had to do with voice AI-
Oh.
... and AI agents with voice.
Okay.
So that was, uh, quite fascinating. Also spoke with the speakers afterwards.
Yeah.
And yeah, they were also very open and, and, you know, this, this sharing attitudes that's, uh, I think in general, quite prevalent in the AI community. I also learned a lot, like, really practical things that I can now take away with me.
Yeah. I mean, on, on my side, I, I think I watched only, like, half of the talks- ... 'cause I was running around, and I think people saw me, like, towards the end, I was kind of collapsing. I, I was on the floor, like, uh, towards the end because I, I, I needed to get, to get a rest.
But yeah, I, I'm excited to watch the voice AI talks myself.
Yeah, yeah, do that. And I mean, from my side, thanks a lot for organizing this conference-
Yeah
... for bringing everyone together.
Do you have anything like this in Switzerland?
The sh- The short answer is no. Um, I mean, I have to say the AI community in especially Zurich-
Yeah
... where we're, where we're based-
Yeah
... it is quite good, and it's, uh, growing, uh, especially driven by ETH, the, the technical university there, and all of the big companies, they have AI teams there. Google, like, Google has the biggest tech hub outside of the US in Zurich.
Yeah.
Facebook is doing a lot in Reality Labs. Uh, Apple has a secret AI team. OpenAI and Anthropic just announced that they're coming to Zurich.
Yeah.
Um, so there's a lot happening. Yeah, so.
Yeah. Uh, I think the most recent notable move, I think the entire vision team from Google, uh, Lucas Beyer, um, and, and all the other authors of SigLIP left Google to join OpenAI, which I thought was like... It's, like, a big move for a whole team to move all at once.
Yeah.
Also, at the same time, so I've been to Zurich, and it just feels expensive. Like, it's a great city.
Yeah.
A great university. But I don't see it as, like, a business hub. Is it a business hub? Uh, I guess it is, right? Like, it's kind of-
Well, historically, it's, uh, it's a finance hub.
Finance hub? Yeah.
Yeah. I mean, there are some, some large banks there, right? Especially UBS, uh, the, the largest wealth manager in the world.
Yeah.
But it's really becoming more of a tech hub now with all of the big, uh, tech companies there.
Right. I guess, yeah.
Yeah.
And but re- and research-wise, it's all ETH.
Yeah.
Maybe some other things.
I mean-
Yeah, yeah.
Yeah, it's all driven by ETH, and then, uh, its sister university, EPFL, which is in Lausanne.
Okay.
Um, which they're also doing a lot, but, uh, it's, it's really ETH. Uh, and otherwise, no, I mean, it's a beautiful, really beautiful city. I can recommend to anyone- ... to come, uh, visit Zurich. Uh, uh, let me know.
Happy to-
Yeah
... show you around. And of course, you know, you, you have the nature so close, you have the mountains so close, you have so b- so beautiful lakes.
Yeah.
Um, I think that's what makes it such a livable city.
Yeah.
Um, and the cost is not, it's not cheap, but I mean, we're in New York City right now, and, uh, I don't know, I paid $8- ... for a coffee this morning, so, uh, the coffee is cheaper in-
It's variable
... Zurich than the New York City, so.
Okay. Okay. Let's talk about Snipd. What is Snipd? And, you know, then we'll talk about your origin story. But just let's, let's get it crisp. What is Snipd?
Snipd Defined5:03
Yeah. I always see two definitions of Snipd, so I'll give you one really simple, straightforward one and then a second more nuanced, um, which I think will be valuable for the rest of, uh, our conversation. So the most simple one is just to say, look, we're an AI-powered podcast app.
So if you listen to podcasts, we're now providing this AI-enhanced experience. But if you look at the more nuanced, uh, perspective, it's actually we, we've have a very big focus on people who, like your audience, who listen to podcasts to learn something new.
Like your audience, you want-- they want to learn about AI, what's happening, what's, what's, what's the latest research, what's going on. And we want to provide a, a spoken audio platform where you can do that most effectively, and AI is basically the way that we can achieve that.
Yeah. Means to an end.
Yeah, exactly.
When you started, was it always meant to be AI, or was it, was it more about the social sharing?
So the first version that we ever released was, like, three and a half years ago.
Okay.
Yeah. So this was before ChatGPT, um-
Before Whisper
Yeah, before Whisper.
Yeah.
So I think a lot of the features that we now have in the app, they weren't really possible yet back then. But we-
Oh
... already from the beginning, we always had the focus on knowledge. That's the reason why, you know, we and our team, why we listen to podcasts. But we did have a bit of a different approach. Like, the idea in the very beginning was, so the name is Snipd, and you can create these, what we call snips.
Uh, which is basically a small snippet, like a clip from a, from a podcast. Um, and we did envision sort of like a, like a social TikTok platform where some people would listen to full episodes, and they would snip certain, like, the best parts of it, and they would post that in a feed.
And other users would consume this feed of snips, um, and use that as a discovery tool or just as, as a means to an end. And yeah, so you would have both people who create snips and people who listen to snips.
So our big hypothesis in the beginning was, you know, it would be easy to get people to listen to these snips, but super difficult to actually get them to create them. So we focused a lot of, uh-
Interesting
... a lot of our effort on making it as seamless and easy as possible to create a snip.
Yeah. It's, it's similar to TikTok. Uh, you, you need CapCut for, for there to be videos on TikTok.
Exactly. Exactly. And so for, for Snipd, basically, whenever you hear an, an amazing insight, a great moment, you can just triple-tap your headphones, and our AI actually then saves the moment that you just, uh, listened to and summarizes it to, to create a note, and this is then basically a snip.
So yeah, we built, we built all of this, uh, launched it, and what we found out was basically the exact opposite. So we saw that people used the snips to discover podcasts, but they really, you know, really love listening to long form- ...
uh, podcasts. But they were creating snips like crazy.
Mm.
And this was, this was definitely one of these aha moments when we realized, like, hey, we should be really doubling down on the knowledge of learning, of, yeah, helping you learn most effectively and helping you capture the knowledge that you listen to and actually do something with it.
'Cause this is, in general, you know, we, we live in this world where there's so much content, and we consume and consume and consume, and it's so easy to just, at the end of the podcast, you just start listening to the next podcast.
And five minutes later, you've forgotten-
You've forgotten everything
... 90%, 99%-
Yeah
... of what you've actually just learned.
Yeah. You don't know this, but, and, and most people don't know this, but this is my fourth podcast. My third podcast was a personal mixtape podcast where I snipped, manually, sections of podcasts that I liked and added my own commentary on top of them and published them as small episodes.
Nice.
So those would be maybe five to 10-minute snips-
Yeah
... of something that I thought was a good story or, like, a good insight, and then I added my own commentary and published it as a separate podcast.
It's cool. Is that still live?
It's still live, but it's not active. But you can go back and find it. If you're, if, if you're curious enough, you'll see it.
Nice. Nice. Yeah, you have to show me later.
But it was so manual, uh, because basically what my process would be, I hear something interesting, I note down the timestamp, and I note down the URL of the, the podcast. I, I used to use Overcast, so it would just link to the Overcast page.
And then put it in my note-taking app, go home. Uh, whenever I feel like publishing, I will take one of those things and then, uh, download the MP3, clip out the MP3, and record my intro, outro, and then publish it as a, as a podcast.
But now Snipd, I mean, I can just kinda double-click or triple-tap.
I mean, th- th- those are very similar stories to what we hear from our users.
Mm.
You know, it's, it's normal that you're doing, y- you're doing something else while you're listening to your podcast. So-
Yeah
... a lot of our users, they're driving, they're working out, walking their dog. So in those moments when you hear something amazing, it's difficult to just write them down or, you know, you have to take out your phone.
Some people take a screenshot, write down the timestamp, and then later on you have to go back and try to find it again. Of course, you can't find it anymore- ... 'cause there's no search. There's no Command+F. And, um, these, these were all of the issues that, that, that we encountered also ourselves as users.
And-
Yeah
... given that our background was in AI, we realized like, wait, hey, this is... This should not be the case. Like, podcast apps today, they're still-- They're basically repurposed music players. But we actually look at podcasts as one of the largest sources of knowledge in the world.
Mm.
And once you have that different angle of looking at it, together with everything that AI is now enabling, you realize, like, hey, this is not the way that we, that podcast apps should be.
Yeah. Yeah. Agreed. Uh, you mentioned something there. You said your background's in AI. First of all, who's the team, and what do you mean your background's in AI?
Kevin's Roots10:48
Those are two very different questions. Um, maybe starting with, with my backstory.
Yeah.
My backstory ac- actually goes back, like, let's say 12 years ago or something like that. I moved to Zurich to study at ETH. And actually I studied something completely different. I studied mathematics and economics.
Mm-hmm. Same.
Basically with this, this, the specialization for quant finance.
Same. Okay. Wow. All right.
So yeah. And then, as you know, all of these mathematical models for, um, asset pricing, derivative pricing, quantitative trading. And for me, the thing that, that fascinated me the most was the mathematical modeling behind it, uh, mathematics, uh, statistics.
But I was never really that passionate about the finance side of things.
Oh, really? Oh, okay.
Yeah, I mean it's-
Okay, we're, we're different there.
I mean, one just, let's say, symptom that I notice now, like, like looking back, during that time, I think I never read an academic paper about the subject in my free time. And then it was towards the end of my studies, I was already working for a big bank.
One of my best friends, he comes to me and says, "Hey, I just took this course. You have to, you have to do this. You have to take this lecture."
Okay.
And I'm like, "What? What- "... what is it about?" "It's called machine learning." And I'm like, "What? What? What kind of stupid name is that?" Uh, so he sent me the slides and, like, over a weekend I went through all of the slides and I just, I, I just knew, like, freaking hell, like, this is it.
I'm, I'm in love.
Wow.
Yeah.
Okay.
And that was then over the course of the next, I think like 12 months, I just really got into it. Started reading all about it, like reading blog posts, starting building my own models.
Was this course by a famous person? Famous university? Was it like the-
So-
... Andrew, Andrew Ng Coursera thing or?
Uh, no. So this was a ETH course.
Oh, it was ETH.
Uh, so a professor at ETH. Um-
Did he teach in English by the way or...?
Yeah. Yeah, yeah.
Oh, okay. So these slides are somewhere available.
Yeah, yeah.
Okay.
Definitely. I mean, now they're quite outdated.
Yeah, yeah. Sure, sure, sure. Well, I think, you know, reflecting on the finance thing for a bit. So I, I was, used to be a trader.
Yeah.
Uh, sell side and buy side. I was options trader first and then I, I was, uh, more of like a quantitative, uh, hedge fund, uh, analyst. We never really used machine learning.
Yeah.
It was more like y- a little bit of statistical modeling but really like you, you fit, you know, your regressions.
Yeah. No, I mean that's, that's what it is. And, uh, or you, you solve partial differential equations and have then numerical methods to, to, to solve these.
That's, that's for your degree. That's, that's not really what you do at work, right? Unless... Well, I don't know what you do at work.
In my job? No. No, we weren't solving the partial differential equations.
Yeah, you learn all this in school and then you not use it.
Yeah. I mean, we, we-- Well, let's put it like that. Um, in some things, yeah, I mean, I did code algorithms that would do it, but it was basically like it was the, the most basic algorithms and then you just like slightly improve them a little bit.
Like you just tweak them here and there.
Yeah.
It wasn't like starting from scratch like, "Oh, here's this new partial differential equation." Like how do we attack? No. Um...
Yeah, yeah. I mean, that's, that's real life, right? Most-
Yeah
... most of it's kinda boring or you're, you're using established things because they're established because, uh, they tackle the most important topics. Um, yeah, portfolio management was more interesting for me. Um, and, uh, we, we were sort of the first to combine like social data with, with quantitative trading and I think, uh-
Nice. Nice
... I think pe- I think now it's very common but, um, yeah. Then you s- you went, you went deep on machine learning and then what? You quit your job?
Yeah. Yeah.
Wow.
I quit my job because, uh, um, I mean I started using it at, at the bank as well.
Okay.
Like try-- Like, you know, I'd like desperately try to find any kind of excuse to like use it here or there or there. But it just was clear to me like, no, if I wanna do this, um, like I just have to like make a real cut.
So I quit my job and joined an early stage, uh, tech startup in Zurich-
Ooh
... where I then built up the AI team over five years.
Wow.
Yeah, we built various machine learning, uh, things for, for banks from like models for, for sales teams to identify which clients, like which product to sell to them and, and with what reasons, all the way to... We did a lot, lot with bank transactions.
One of the actually most fun projects for me was we had an, an NLP model that would take the booking text of a transaction, like a credit card transaction and prettify it.
Yeah.
Because it had all of these, you know, like numbers in there and abbreviations and whatnot, and sometimes you look at it like, "What, what is this?"
Yeah.
And it was just, you know, it would just change it to, I don't know, CVS.
Yeah, yeah. But I mean, would you have hallucinations?
No. No, no. The way that everything was set up, it wasn't, like it wasn't yet-
That smart
... this fully-
Yeah
... end-to-end generative, uh, neural network as what you would use today.
Okay. Okay. Awesome.
Yeah.
And, and then when did you go like full-time on Snipd?
The Hackathon15:32
Yeah. So basically that was, that was afterwards. I mean, how that started was the friend of mine who got me into machine learning, uh, him and I-- uh, like he also got me interested into startups. He's had a big impact on my life.
And the two of us would just, uh, jam on, on like ideas for startups every now and then, and his background was also in AI data science. And we had a couple of ideas, but given that we were working full times, we were thinking about...
Uh, so we participated in Hack Zurich.
Okay.
That's, uh, Europe's biggest hackathon, um, or at least was at the time. And we said, "Hey, this is just a weekend. Let, let's just try out an idea, like hack something together and see how it works." And the idea was that we'd be able to search through podcast episodes-
Yeah
... like within a podcast.
Yeah.
So we did that. Long story short, uh, we managed to do it. Like to build something that we realized, hey, this actually works. You can, you can find things again in podcasts via like a natural language search. And we pitched it on stage, and we actually won the hackathon-
Ooh
... which was cool. I mean, we, we also, I think we had a good, um, like a good, good pitch or good example. So we, we used the famous Joe Rogan episode with Elon Musk where Elon Musk smokes a joint.
Okay.
Um, it's like a two and a half hour episode.
Uh-huh.
So we were on stage and then we just searched for like smoking weed.
Uh-huh.
And it would find that exact moment, it would play it and it just like come on with Elon Musk just like smoking.
Oh, so it was video as well?
No, it was actually completely based on audio, but we did have the video for the presentation-
Yeah, yeah
... which had a, had of course, an amazing effect.
Yeah.
Like this gave us a lot of activation energy, but it wasn't actually about winning the hackathon.
Yeah, yeah.
But the interesting thing that happened was after we pitched on stage, several of the other participants, like a lot of them came up to us and started saying like, "Hey, can I use this?" "Like I, I have this issue."
And like some also came up and told us about other problems that they have, like very adjacent to this with the podcast was like, like, "Could, could I use this for that as well?" And that was basically the, the moment where I realized, hey, it's actually not just us who are think- who are having these issues with, with podcasts and getting to the making the most out of this knowledge.
Yeah.
Um, there are other people.
Yeah.
That was now I guess like four years ago or something like that. And then yeah, we decided to quit our jobs and start, start this whole Snipd thing.
Yeah. How big is the team now?
Uh, we're just four people.
Yeah.
We're just four people. Yeah, like four. We're all, uh, technical.
Yeah.
Basically two on the, the back-end side. So one of my co-founders is this person who got me into-
Yeah
machine learning-
Yeah, yeah
... startups, and we won the hackathon together. So we have two people for the back-end side with the AI and, and all of the other back-end things, and two for the front-end side building the app.
Which is mostly Android and, uh, iOS.
Yeah, it's iOS and Android. We also have a watch-
Yeah
... uh, app for, for Apple, but yeah, it's mostly iOS too.
Yeah. The watch thing, uh, it was very funny 'cause in the, in the Lean Space Discord, you know, most of us have been slowly adopting Snips. You came to me, like, a year ago and, uh, you introduced Snipd to me, and I was like, "I don't know."
I'm, you know, I'm very sticky to Overcast. Taking it slowly with switch. Um, why watch?
So it goes back to a lot of our users, they do something else while, while listening to a podcast, right?
Uh-huh.
And one of the... Us giving them the ability to then capture this knowledge even though they're doing something else at the same time is one of the killer features.
Yeah.
Um, maybe I can actually, maybe at some point I should maybe give a bit more of an overview of what the f- all of the features that we have.
Sure.
So this is one of the, the, the killer features, and for one big use case-
Uh-huh
... that people, um, use this for is for running.
Yeah.
So if you're a big runner, a big jogger or cycling, like, really, really cycling, um, competitively, and a lot of the people, they don't wanna take their phone with them when they go running.
So you load everything onto the watch.
So you can download episodes.
Oh.
I mean, if you, if you have an Apple Watch that has internet access, like with a SIM card, you can also directly stream.
Oh.
Um, that's also possible.
Uh-huh.
Yeah. Of course, it's a... It's basically very limited to just listening and snipping.
Yeah.
And then you can see all of your snips later on your phone.
Let me tell you this error I just got. "Error playing episode. Substack, the host of this podcast, does not allow this podcast to be played on an Apple Watch."
Yeah, that's a very beautiful thing. So we found out that all of the podcasts hosted on Substack, uh, you cannot play them on an Apple Watch.
Why is this restriction? What-
Like, don't ask me. We tried to reach out to Substack. We tried to reach out to some of the bigger podcasters who are hosting their, their podcast on Substack to also let them know.
Uh-huh.
Um, Substack doesn't seem to care. This is not specific to our app. You can also check out the Apple Podcasts app.
Yeah.
It's the same problem. It's just that we actually have identified it, and we, we tell the user what's going on.
I will say I've been, you know, we host it, we host our podcast on Sub- on Substack, but they're not very serious about their podcasting tools. I've told them before. I've been very upfront with them, so I don't feel like I'm, you know, shitting on them in any way.
And, uh, it's kinda, it's kinda sad because otherwise it's a perfect creative platform. But the way that they treat podcasting as a s- as an afterthought, uh, I think it's really disappointing.
Maybe, uh, given that you mentioned all these features, maybe I can give a bit of a better overview of the features-
Let's do that. Let's do that
AI Features20:50
... that, that, that, that we have.
Let's do that. Yeah.
Because for us, it's clear in our minds. Maybe for, for some of the-
I mean, okay
... uh, listeners.
I'll tell you, I'll tell you my version.
Yeah.
You can correct me, right? So first of all, the, I think the main job is for it to be a podcast listening app. Um, it should be basically a complete superset of what you normally get on Overcast or Apple Podcasts, anything like that.
You pull your show list from, from ListenNotes? Like, how do you, how do you find shows? Like, I type in anything and you f- you find them, right?
Uh, yeah. We have, we have a search engine that is powered by ListenNotes.
Yeah.
But, I mean, in the meantime, we have a huge database of-
Huge, huge amount, yeah
... like 99% of all podcasts out there ourselves. So...
Yeah. What I noticed, the, the default experience is you do not auto-download shows, and that's, that's one very big difference for you guys versus other apps. Uh, where like, you know, if I'm su- subscribed to a thing, it auto-downloads, and I already have the MP3 downloaded over- overnight.
For me, I have to put- actively put it onto my queue, then it auto-downloads. And actually, I initially didn't like that. I think I maybe told you that I was like, "Oh, this is like a feature that I don't like," 'cause it means that I have to choose to listen to it in order to download and not to...
It's just like opt-in. There's a difference between opt-in and opt-out. So I opt into every episode that I listen to. And then, like, you know, you open it and depends on whether or not you have the AI st- AI stuff enabled, but the default experience is no, no AI stuff enabled.
You can listen to it. You can see the snips, the number of snips, and where people snip, uh, during the episode, which roughly correlates to interest level, and obviously you can snip there. I think that's the default ex- experience.
Uh, I think snipping is really cool. Like, I use it to share a lot on our Discord. I think we have tons and tons of just people sharing snips of stuff. And tweeting stuff is, is also like a nice, pleasant experience.
But, like, the real features come when you actually turn on the AI stuff. And so the reason I got Snips because I got fed up with Overcast not implementing any AI features at all. In- instead, they spent two years rewriting their app to be a little bit faster.
And I'm like- Like, it's 2025, I should have a podcast that has transcripts that I can search. Very, very basic thing. Overcast will basically never have it.
Yeah, I think that was a, was a good, like, basic overview. Maybe I can, um-
Yeah, please
... add a bit to it with, uh, with the, the AI features that we have. So one thing that we do every time a new podcast comes out, we, uh, transcribe the episode. We do speaker diarization. We identify the speaker names.
Each guest, we extract a mini bio of the guest. Uh, try to find a picture of the guest online, add it. We break the, the podcast down into chapters, uh, as in AI-
Mm.
generated chapters with quick-
That one's very handy
... with a quick description per, uh, title and quick description per each, uh, chapter. We identify all books that get mentioned-
Yeah
... on a, on a podcast. Uh, which is a-
You can tell I don't use that one.
It depends on the podcast. There, there are some podcasts where the guests o- often recommend, like, an amazing book. And so later on you can, you can find that again. We-
So literally you search for the word book or-
No, so it's-
... like, "I just read," blah, blah, blah.
Uh, no, I mean, it's, it's all LLM based.
Mm.
So basically we have, we have an LLM that goes through the entire transcript and identifies if a user, uh, mentions a book. Then we use Perplexity API together with various other LLM orchestration to go out there on the internet, find everything that there is to know about the book, find the cover, find who auth- what the au- who the author is, uh, get a quick description of it.
For the author, we then check on which other episodes the author appeared on.
Yeah, that is killer.
Because that, like, for me, if, if, if there's an interesting book, the first thing I do is I actually listen to a podcast episode with the, with the writer. Because he usually gives a really great overview already on, on a, on a podcast.
Sometimes the podcast is with the person as a guest. Sometimes his podcast is about the person without him there. Do you pick up both?
So yes, we pick up both-
Okay
... in, like, our latest models. But actually what we show you in the app, the goal is to currently only show you the guest to separate that. Uh, in the future, we wanna show the other things more. But that says-
For what, for what it's worth, I don't mind.
Yeah.
I don't think... Like, if, if I like, if I like somebody, I'll just learn about them regardless of whether they're, they're not.
Yeah. I mean, yes and no. We, we, we have seen there are some personalities where this can break down. So for example, the first version that we released with this feature, it picked up much more often a person, even if it was not a guest.
Yeah.
For example, the, the best examples for me is Sam Altman and Elon Musk. Like, they're just mentioned on every second podcast, and it has-- Like, they're not on there. And if you're interested in actually, like, learning from them-
Yeah, I see.
Um, yeah, we updated, uh, our, our algorithms improved that a lot, and now it's gotten much better to only pick it up-
Yeah
... if they're a guest. Um, yeah, so this, this is, um, maybe to come back to the features, two more important features. Like, we have the ability to chat with an episode.
Yes.
Of course, you can do the old style of searching through a transcript with a keyword search, but I think for me, this is, this is how you used to do search and extracting knowledge in the, in the past.
Old school.
Um, the AI way is, is basically an LLM. Uh, so you can ask the LLM, "Hey, when do they talk about topic X?" If you're interested in only a certain part of the episode.
Yeah.
You can ask them for, for to give a, um, quick overview of the episode, key takeaways. Um, afterwards also to create a note for you. So this is really, like, very open, open-ended. And yeah. And then finally, the snipping feature that we mentioned.
Just to reiterate, yeah, I mean, here the, the feature is that whenever you hear an amazing idea, you can triple tap your headphones or click a button in the app, and the AI summarizes the insight you just heard, uh, and saves that together with the original transcript and audio in your knowledge library.
I also noticed that you, you skip dynamic content.
Um, so dynamic content, we do not skip it automatically.
Oh, sorry. You-
Um-
You detect
... but we detect it. Yeah. I mean, that's one of the thing that most people don't, don't actually know that. Like, the way that ads get inserted into podcasts or into most podcasts is actually that every time you listen to a podcast, you actually get access to a different audio file, and on the server, uh, a different ad is inserted into the MP3 file automatically.
Yeah, based on IP.
Exactly. And, um, that, what that means is if we transcribe an episode and have a transcript with timestamps, like word, word specific timestamps, if you suddenly get a different audio file- ... like, the whole timestamps are messed up.
And that's, like, a huge issue. And for that, we actually had to build another algorithm that would dynamically, on the fly, resync the audio that you're listening to the transcript that we have.
Yeah.
Which is a fascinating problem in and of itself. Um-
You, you sync by matching up the sound waves or like... Or do you ma- sync by matching up words? Like, you basically you do partial transcription.
We are not matching up words. It's, it's happening on the basically like a bytes level- ... matching. Yeah.
Okay.
Uh, as in-
So it relies on this, it relies on the, uh, there being exact match at some point.
Uh, so it's actually not, uh, we're actually not doing exact matches-
Okay
... but we're doing fuzzy matches-
Wow
... to, to identify the, the moment. It's basically, um, we basically built Shazam for podcasts. Uh, just as a little side project to, to solve this issue.
Yeah, yeah. A- actually, fun, fun fact, uh, apparently the Shazam algorithm is open. It's this, they published a paper.
Yeah.
Like, it's talked about it.
Yeah.
Yeah. I've, I haven't really dived into the paper. I thought it was kinda, kinda interesting that basically no one else has built Shazam. Like
Yeah. I mean, well, the one thing is the algorithm. Like, if you now talk about Shazam, right? The other thing is also having the, uh, the database behind it-
Yes
... and having the user mindset that-
Yeah
... if they have this problem, they come to you, right?
Yeah, yeah. Yeah, I'm very interested in the tech stack. There's a big data pipeline. Could you share like, you know, h- what is the tech stack? What are the, you know, the most interesting or challenging pieces of it?
So the general tech stack is our entire backend is, or 90% of our backend is written in Python.
Under the Hood28:39
Okay.
Hosting everything on, uh, Google Cloud, uh, platform. And our front end is, uh, written with, well, we're using the Flutter, um, framework.
Ah.
So it's written in Dart and then, but compiled natively. So we have one code base for, that handles both Android and iOS.
You think that was a good decision? That's something that a lot of people are exploring.
Um, so up until now, yes.
Okay.
Look, it, it has its pros and cons. Some of the, you know, for example, earlier I mentioned we have a Apple Watch app.
Yeah.
I mean, that, there, there's no Flutter for that, right? So that you build native and then, of course, you have to sort of like sync these things together. I mean, I'm not the front-end engineer, so I'm not just-
Yeah
... relaying this information, but our front, front-end engineers are very happy with it. It's enabled us to be quite fast and be on both platforms from, from the very beginning. And when I, when I talk with people and they hear that, that we are using Flutter, usually they s- you know, they, they think like, "Ah, it's not performant, it's super junk- janky and, and everything."
And then they use our app, and they're always super surprised.
Yeah.
Or if they've already used our app-
I couldn't tell
... and you tell them-
Yeah
... they're like, "What?"
Yeah.
Um, so there is actually a lot that you can do there.
And then the danger, the, the, the concern... There's a few concerns, right? One, it's Google, so when would they, when are they gonna abandon it? Two, you know, they're, they optimize for Android first, so iOS is like a second, second thought, or like you can feel-
Yeah
... that it is not a native iOS app.
Yeah.
Uh, but you guys put a lot of care into it. And then maybe three, from my point of view, JavaScript's, uh, as a JavaScript guy, React Native was supposed to be that dream, and I think that it hasn't really fulfilled that dream Uh, maybe Expo is trying to do that, but, um, again, it is not-- does not feel as productive as Flutter.
And I've-- I spent a week on Flutter and Dart, and I'm an investor in FlutterFlow-
Oh, cool
... which is the low-code, uh, Flutter, Flutter startup that's doing very, very well. I think a lot of people are still Flutter skeptics.
Yeah.
Wait, so are, are you moving away from Flutter?
Uh, no.
Okay.
We don't have plans to do that.
Yeah, you were just saying about the w- the watch app. Okay. Let, let's go back to the stack.
Yeah.
Um-
You know, that was just to give you a bit of an overview. I think the more interesting things are, of course, on the AI side.
Yeah.
So we-- like, as I mentioned earlier, when we started out, it was before ChatGPT, before the ChatGPT moment, before there was the GPT-3.5 Turbo, uh, API. So in the beginning, we actually were running everything ourselves.
Mm-hmm.
Open source models, try to fine-tune them. They worked, the results, but let's, let's be honest, they weren't...
What was the soda before Whisper?
The transcription?
Yeah.
Uh, we were using Wave2vec.
I s-
Um-
That was a Google one, right?
No, it was a Facebook, Facebook one. That was actually one of the papers, like, when that came out, for me, that was one of the reasons why I said we, we should try something to start a startup in the audio space.
For me, it was a bit like, before that, I'd been following the NLP space, uh, quite closely, and as, as I mentioned earlier, we, we did some stuff at, at the startup as well that I was working at before.
And Wave2vec was the first paper that I had at least seen where the whole transformer architecture moved over to-
Audio
... audio.
Yeah.
And bit more general way of saying it is, like, it was the first time that I saw the transformer architecture being applied to continuous data-
Hmm
... instead of discrete tokens.
Okay.
And it worked amazingly. Uh, and, like, the transformer architecture plus self-supervised learning, like, these two things moved over. And then for me, it was like, hey, this is now gonna take off similarly as the text space has taken off, and with these two things in place, even if some features that we wanna build are not possible yet, they will be possible in the near term, uh, with this tr-
Yeah
... uh, trajectory. So that's a little side, side note. No, so in the meantime, yeah, we're using Whisper. We're still hosting some of the models ourselves, so for example, the whole transcription speaker diarization pipeline, uh-
You need it to be as cheap as possible.
Yeah, exactly. I mean, we're doing this at scale. We're-
Yeah
... we have a lot of audio that we're-
What, what numbers can you disclose? Like, what, what are-- just to give people an idea.
Uh-
'Cause it's a lot.
So we have more than a million podcasts that, that we've already processed.
When you say a million, so processing is basically you have some kind of list of podcasts that you will auto process, and others where a paying s- pay member can choose to-
Yeah
... press a button and, and, and transcribe it, right? Is that the rough idea?
Yeah, yeah.
Yeah, yeah.
Yeah, exactly. Yeah, and we-- if-- when you press that button or we auto transcribe it, yeah, so first we do the, we do the transcription, we do the, the speaker diarization, so basically you identify speech blocks that belong to the same speaker.
This is then all orchestrated within, within LLM to identify which speech, speech block belongs to which speaker together with, you know, we ident... as I mentioned earlier, we identify the guest name and the bio. So all of that comes together within LLM to actually then assign, assign speaker names to, to each block.
Yeah.
And then most of the rest of the, the pipeline we've now used-- we've now migrated to LLM APIs. Uh, so we use mainly OpenAI, uh, Google models, so the Gemini models and then the OpenAI models, and we use some Perplexity basically for those-
For search
... things where we need-
For book search
... where we need web search.
Yeah, yeah.
That's something that I'm still hoping, especially OpenAI, will also provide as an API. Um-
Oh, why?
Well, basically for us as a consumer, the more providers there are-
The more downtime
... you know, the more competition, and it will, um, lead to better, better, uh, results and, um, lower costs over time.
I don't, I don't see Perplexity as expensive.
If you use the web search, uh, the price is, like, five dollars per a thousand queries.
Okay.
Which is affordable, but, uh, if you compare that to just a normal LLM call-
Okay
... um, it's, it's, uh, much more expensive.
Have you tried Eksa?
We've, uh, looked into it, but we haven't really tried it.
Yeah.
Um, I mean, we, we started with Perplexity, and, uh, it works, it works well. And if I remember correctly, Eksa is also a bit more expensive.
I don't, I don't know. Uh, they seem focused on the search thing-
Yeah. Yeah
... as a search API, whereas Perplexity may be more consumery business that is higher, higher margin. Like, uh, I'll put it like Perplexity is trying to be a product, Eksa is trying to be infrastructure.
Yeah. Yeah.
So that, that would be my dif- distinction there. And then the other thing I will mention is Google has a search grounding feature-
Yeah. Yeah
... which you, uh, you might-
Yeah, yeah. We've, uh, we've also tried that out. Um-
Not as good?
So we, we didn't sh- we didn't go into too much detail in, like, really comparing it, like, quality-wise.
Yeah.
Because we actually already had the Perplexity one, and it, and it's, and it's working.
Yeah.
Um, I think also there the price is actually higher than Perplexity.
Ooh.
Yeah.
Really?
Yeah.
Google should cut their prices.
Maybe it was the same price. I don't wanna say something incorrect.
Sure.
But it wasn't cheaper.
It wasn't, like, compelling.
And, and then-
Yeah
... then there, there was no, uh, reason to switch. So I mean, maybe, like, in general, like, for us, given that we do work with a lot of content, price is actually something that we do look at. Like, for us, it's not just about taking the best model for every task, but it's really getting the best...
like identifying what kind of intelligence level you need and then getting the best price for that to be able to really scale this and, and provide us, uh, um, yeah, let our users use these features with as many podcasts as possible.
Yeah.
Uh-
I wanted to double, double-click on diarization.
Yeah.
Uh, it's something that I don't think people do very well. So, you know, I'm a, I'm a, I'm a B user. I don't have it right now. But... and, and they were supposed to speak, but they dropped out.
Last minute. Um, but, uh, we've had them on the podcast before, and I-- it's not great yet. Do you use just Pyannote, the, the default stuff, or do you find any tricks for diarization?
So we do use the, the open source packages But we have tweaked it a bit here and there. For example, if you mentioned the BAI guys.
Yeah.
I actually listened to the podcast episode-
Yeah
... which was super nice.
Thank you.
And when you started talking about speaker diarization, and I just had to think about their use case, like with all of the different environments.
Yeah.
Um, it could b- basically be anything.
It's completely out of domain. Like-
Um-
There's no, there's no data for this.
Yeah. I mean, I was feeling with them. Because, like, our advantage is that we're working with very high quality audio.
Yeah.
It's very controlled, uh, usually recorded in a studio. This is quite an exception, I guess.
It is kind of a studio. It's, like, pretty quiet. There's consistent background noise, which you can edit out.
Yeah, yeah.
Uh, just New York.
Yeah.
It's nice.
Yeah.
It's a character.
Um, no, so that, that of course, uh, helps us. Uh, another thing that helps us is that we know certain structural aspects of the podcast. For example, how often does someone speak? Like, if someone... Like, let's say there's a one-hour episode and someone speaks for 30 seconds.
That person is most probably not the guest and not the host. It's probably some ad, uh, like some speaker from an ad, you know.
Okay.
So we have, like, certain of these, uh-
Heuristics
... heuristics, yeah, exactly, that we can use and we leverage to, like, im- im- improve things. And in the past, we've, we've also changed the clustering algorithm. So basically how, how a lot of this, the, the speaker diarization works is you basically create an embedding for the speech that's happening, and then you try to somehow cluster these, these, um, embeddings, uh, and then find, ah, this is all one speaker, this is all another speaker.
And there we've also tweaked a couple of things where we again used heuristics that we could apply from knowing how podcasts function. Um, and that's also actually why I was feeling so much for the BAI guys because, like, all of these heuristics, like, they, like, for them it's probably almost impossible to use any heuristics because it can just be any- ...
any situation, any, uh, anything. Um, so that's, that's, uh, one thing that we do. Yeah, another thing is that we actually combine it with LLMs. So the transcript LLMs and, and the speaker diarization, like bringing all of these together to recalibrate some of the switching points, like when does the speaker stop, when does the next one start.
Um-
But the, the LLMs can add errors as well. You know, I don't, I wouldn't feel safe using them to be so precise.
I mean, at the end of the day, like also just to not give a wrong impression, like the speaker diarization is also not perfect, uh, that we're doing, right? Um-
I basically don't really notice it. Like I use it for search.
Yeah.
Like it... Yeah.
Yeah. It's not perfect yet, but it's, it's, uh, it's gotten quite good. Like, especially if you compare... If you, if you look at some of the... Like if you take a latest episode and you compare it to an episode that came out a year ago, we've improved it quite a bit, yeah.
Well, it's beautifully presented. Oh, I love that I can click on the ti- the transcript, uh, and it goes to the timestamp. So simple. But it, you know, it should exist.
Yeah. I agree. I agree.
So this, uh, I'm loading a two-hour episode of the Tech Meme Ride Home where there's, there's a lot of different guests calling in, and you've identified the guest name and, uh, yeah, indeed. Yeah, so these are all LLM based.
Yeah, it's really nice.
Yeah, yeah.
Yeah.
Like the speaker names.
I would say, uh, um, I would say that, you know, obviously I'm a power user of all these tools. Uh, you have done a better job than Descript.
Okay, wow.
Descript has so much funding. You have... They had, they had OpenAI invested in them, and they still suck. So I don't know, like, you know, keep going. Like you, you're doing, you're doing great.
Yeah. Thanks. Thanks. Um, I mean, I would, I would say that especially for anyone listening who's interested in building a consumer app with AI-
Right
... I think the, like, especially if your background is in AI and you love working with AI and doing all of that, I think the most important thing is just to keep reminding yourself of what's actually the job to be done here.
Like, what does actually the consumer want? Like for example, you now were just delighted by the ability to click on this word-
Always
... and it jumps there.
Yeah.
Like this is not, this is not rocket science. This is... Like you don't have to be like, I don't know, Andrej Karpathy to come up with that and build that, right? And I think that's, that's, uh, something that's super important, uh, to keep in mind.
Yeah, yeah. Uh, amazing. I mean, there's so many features, right? It's, it's so packed. There's quotes that you pick out.
Mm.
There's summarization. Oh, by the way, uh, I'm gonna use this as my official feature request. Um, I wanna customize what, how it's summarized.
Yeah.
I wanna, I wanna have a custom prompt.
Yeah.
Uh, because your summarization good, but you know, I, I have different preferences, right?
Yeah.
Like, you know.
So one thing that you can already do today, I completely get your feature request, and I think it just-
I'm sure people have asked it.
I mean, maybe just in general as a, as a, how I see the future. You know, like in the future, I think all, everything will be personalized.
Yeah, yeah.
Like n- not... That this is not specific to us.
Yeah.
Um, and today we're still in a, in a phase where the cost of LLMs, at least if you're working with, like, such long context windows as, uh, as us, I mean, there, there's a lot of tokens in if you take an entire podcast.
So you still have to take that cost into consideration. So for every single user, we regenerate it entirely, it, it gets expensive. But in the future, this, you know, cost, like, will continue to go down-
Mm-hmm
... and then it will just be per- uh, personalized. So that being said, you can already today, if you go to the player screen-
Okay
... um, and open up the, the chat.
Yeah.
You can go to the, to, um, to the chat-
Yes
... and just ask for a summary in your style.
Yeah. Okay. I mean-
I-
... I, I listen to consume, you know?
Yeah, yeah.
I, I've, I've never really used this feature. I don't know. I think that's a, that's me being a slow adopter.
No, no, I mean, that's-
It has... When does the conversation start? Okay.
I mean, you can just type anything. I think what you're-
Chat Craft42:26
This all pieces
... what you're describing, I mean, that maybe that is also an interesting topic to talk about-
Yes
... where, like basically, I told you, like, look, we have this chat, you can just ask for it.
Yeah.
And this is- This is how ChatGPT works today. But if you're building a consumer app, you have to move beyond the chat box. Uh, people do not want to always type out what they want. So your feature request was, even though theoretically it's already possible, what you are actually asking for is, "Hey, I just want to open up the app, and it should just be there in a nicely formatted way."
Yeah, yeah.
"Beautiful way such that I can read it or consume it without any issues." And-
Interesting
... um, I think that's in general where a lot of the, the, the opportunities lie currently in the market if you wanna build a, um, a consumer app. Taking the capability and the intelligence, but finding out what the actual user interface is the best way how a user can engage with this, uh, intelligence in a natural way.
This is something I've been thinking about as kind of like AI that's not in your face.
Mm.
Because right now, you know, we, we like to say like, "Oh, use... Notion has Notion AI," and we have the little thing there, and it's... Or like some other, any other platform has like the sparkle magic wand emoji.
Like, "That's our AI feature. Use this." And it's like really in your face. A lot of people don't like it. You know, it should just kind of be become invisible. Kind of like an invisible AI.
100%. I mean, the, the way I see it is AI is, is the electricity of, of the future. And-
Ooh
... like no one... Like, like we don't talk about, I don't know, this, this microphone uses electricity, this phone. You don't think about it that way. It's just in there, right? It's not an electricity-enabled product. No, it's just a product.
Yeah.
It will be the same with AI. I mean, now it's still a something that you use to market your product. I mean, we do, we do the same, right? Um, because it's still something that, uh, people realize, ah, they're doing something new.
But at some point, no, it'll just be a podcast app.
Yeah.
And it would be normal that it has all these-
It's a normal
... this AI in there.
I noticed you do something interesting in your chat where you source the timestamps.
Yeah.
Is that part of this prompt? Is there a separate pipeline that adds source- sources?
This is, uh, actually part of the prompt. Um, so this is all prompt engineering. Um, uh, you should be able to click on it.
Yeah, yeah, I clicked on it.
Um, this is all prompt engineering with how to provide the, the context. You know, we- because we provide all of the transcript, how to provide the context and then, yeah, get the model to respond in a correct way with a certain format, and then rendering that on the front end.
This is one of the examples where I would say it's so easy to create like a quick demo of this. I mean, you can just go to ChatGPT, paste this thing in and say like, "Yeah, do this."
Okay.
Like 15 minutes and you're done.
Yeah.
But getting this to like then production level that it actually works 99% of the time-
Okay
... this is then where, where the difference lies.
Yeah.
So, um, for this specific feature, like we actually also have like countless regexes.
Ah.
That, that are... They're just there to correct certain things that the LLM is doing because it doesn't always adhere to the format correctly, and then it looks super ugly on the front end. So yeah, we have certain regexes that correct that.
And maybe you'd ask like, "Why don't you use an LLM for that?" Because that's sort of the, again, the AI-native way. Like who uses regexes anymore? But with the chat, for user experience, it's very important that you have the streaming because otherwise you need to wait so long until your message has arrived.
So we're streaming live the... Like, just like ChatGPT, right? You get the answer, and it's streaming the text. So if you're streaming the text and something is like incorrect, it's currently not easy to just like pipe, like stream this into-
Stream it into another stream.
Yeah.
Yeah, yeah, yeah.
Yeah, stream this into another stream and get the stream back-
Yeah, yeah, yeah
... which corrects it.
Yeah, yeah.
That would be amazing. I don't know. Maybe you can answer that. Do you know of any, um-
There's no API that does this.
Yeah.
But-
Like you cannot stream in. You cannot stream in
... if you own the models, you can-
Yeah
... uh, you know, whatever token sequences has been emitted, start loading that into the next one-
Yeah
... if you fully own the models.
Yeah.
Uh, I don't... It's probably not worth it. That's what, what you think is better.
Yeah.
And I think most engineers who are n- new to AI research and benchmarking actually don't know how much re- regexing there is- ... that goes on in normal benchmarks. I- it's just like this ugly list of like 100 different, you know, matches for some criteria that you're looking for.
Yeah.
Um, no, it's very cool. I think it's, it's an example of like real world engineering.
Yeah.
Do you have a tooling that you're proud of, that you developed for yourself? Is it just a test script or is it... You know?
I think it's a bit more... I guess the term that has come up is, uh, vibe coding.
Okay.
Well, well, vibe coding is some- no, sorry, that's actually something else in this-
Vibe coding, yeah
... in this case. But, uh, no, no, yes, um, vibe evals-
I see
... was the term that w- in one of the talks actually on, on, um... I think it, it might have been the first, the first or the, the first day in, at the conference, someone brought that up.
Is that right? Yeah, yeah.
Uh, because yeah, a lot of the talks were about evals, right? Which is so important. And yeah, I think for us it's a bit more vibe evals. You know, that's also part of, you know, being a startup, we can take risks.
Like we can take the cost of maybe sometimes it failing a little bit or being a little bit off, and our users know that, and they appreciate that in return. Like we're moving fast and iterating and building-
Okay
... building amazing things. But you know, a Spotify or something like that, half of our features will probably be in a six months review through legal or I don't know what- ... uh, before they could send them out.
Um, let's just say Spotify is not very good at podcasting. Um, I have a documented, uh, dislike for, for their podcast features. Just overall really, really well integrated. Any other like sort of LLM-focused engineering challenges or problems that, that you, that you wanna highlight?
I think it's not unique to us, but it goes again in the direction of handling the uncertainty of LLMs. So for example, with last year, at the end of the year, we did sort of a Snipd wrapped, and one of the things we thought it would be fun to just, to do something with a, with an LLM and something with the Snipd that, that a user has.
And uh, three, let's say, unique LLM features were that we Assigned a personality to you based-
Mm
... on the, the snips that, that you have. It w-- I mean, it was just all a, like, as a bit of a fun-
I think it-
... playful way.
I'm gonna look up mine. I, I forgot mine already.
Um, yeah, I don't know whether it's actually still in the, in the app.
No, no, we, we all took s-screenshots of it.
Ah, okay.
And we posted it in the, in the Discord.
And the, the second one was, uh, we had a learning scorecard where we identified the topics that you snipped on the most, and you got, like, a little score for that. And the third one was a, a quote that stood out.
And the quote is actually a very good example where we would run that for user, and most of the time it was an interesting quote, but every now and then it was, like, a super boring quote that you think like, like how...
Like, why did you select that? Like, come on. For there, the solution was actually just to say, "Hey, give me five candidates." So it extracted five quotes as a candidate, and then we piped it into a different model as a judge, LLM as a judge.
And there we used a, um, a much better model.
Okay.
Because with the, the initial model, again, as, as I mentioned also earlier, we do have to look at the, like, the, the cost.
Everything.
Because, like, we have so much text that goes into it, so we-- There we use a bit more cheaper model, but then the judge can be, like, a really good model to then just choose one out of five.
Okay.
So this is a, the-
Yeah
... practical example.
I can't find it. Bad search in Discord. Um, so, so you do recommend having a much smarter model as a judge?
Uh, yeah. Yeah.
And that works for you?
Yeah. Yeah.
Interesting. I think this year I'm very interested in LLM as a judge being more developed as a concept. I think for things like, you know, Snips Rap, like, it's, it's fine. Like, you know, it's, it's, it's, it's entertaining.
There's no right answer.
I mean, we also have it, um-- we also use the same concept for our books feature where we identify the, the mentioned books.
Yeah.
Because there it's the same thing. Like, ninety percent of the time, it, it works perfectly out of the box, one shot, and every now and then it just, uh, starts identifying books that were not really mentioned or that are not books or-
Yeah
... made... Yeah, starting to make up books. And, uh, there basically we have the same thing of, like, another LLM challenging it. Um, yeah, and actually with the speakers we do the same now that I, now that I think about it.
Yeah.
Um, so I'm... I, I think it's a, it's a great technique.
Interesting. You run a lot of calls.
Yeah.
Okay, you know, you mentioned costs. You moved from self-hosting a lot of models to the, to the, you know, big lab models, OpenAI, uh, and Google. Uh, no topic?
Cost Crunch51:08
Um, no, we love Claude. Like, in my opinion, Claude is the, the best one when it comes to the way it formulates things.
The personality.
Yeah, the personality.
Okay.
Uh, I actually really love it. But yeah, the cost is, is still high.
Okay. So you cannot... You tried Haiku, but you're, you're like, you have to have Sonnet.
Uh, like, basically, we... Like, with Haiku we haven't experimented too much. We obviously work a lot with three point five Sonnet. Uh, also, you know, like-
For coding.
Yeah, for coding, like in Cursor, just in general. Also brainstorming, we use it a lot. Um, I think it's a great brainstorm partner. But yeah, with, uh, with, with a lot of things that we've then done-
Yeah
... we, we opted for different models.
What, what I'm trying to drive at is how much cheaper can you get if you go from closed models to open models? And maybe it's, like, zero percent cheaper, maybe it's five percent cheaper, or maybe it's, like, fifty percent cheaper.
Do you have a sense?
It's very difficult to, to judge that. I don't really have a sense, but I can, I can give you a couple of thoughts that have gone through our minds-
Yeah
... over the time. Because obviously we, we do realize, like, given that we, we have a couple of tasks where there are just so many tokens going in. Um, at some point it will make sense to, to offload some of that, uh, to an open source model.
But going back to, like, we're, we're a startup, right? Like, we're not an AI lab or, or whatever. Like, for us, actually, the most important thing is to iterate fast because we need to learn from our users, improve that and, yeah, just this velocity of this, these iterations.
And for that, the closed models hosted by OpenAI, Google is, uh, and Topic, they're just unbeatable because y-you just... it's just an API call.
Yeah.
Um, so you don't need to worry about so much complexity behind that. So this is, I would say, the biggest reason why we're not doing more in this space. But there are other thoughts, uh, also for the future.
Like, I see two different... Like, we basically have two different usage patterns of LLMs where one is this, this pre-processing of a podcast episode, like, this initial processing, like the transcription, speaker diarization, chapterization. We do that once. And this, this usage pattern, it's, it's quite predictable because we know how many podcasts get released when.
Um, so we can sort of have a certain capacity, and we can... we, we are running that twenty-four seven. It's one big queue running twenty-four seven.
What's the queue job runner? Uh, is it jang- uh, Django? Is this, like, the Python one?
No. That, that's just our own, like-
Your own. Okay
... we, in our database-
Okay
... uh, and the back end talking to the database, picking up jobs-
Okay
... piping it back.
I'm just curious in orchestration and queues.
I mean, we, we of course have, like, uh, a lot of other orchestration where we use, uh, the Google Pub/Sub-
Okay
... uh, thing. But okay, so we have this, this, this usage pattern of, like, very predictable, uh, usage, and we can max out the, the usage. And then there's this other pattern where it's, for example, with Snipd, where it's, like, a user...
it's a user action that triggers an LLM call, and it has to be real time, and there can be moments where it spikes the usage, and there can be moments when there's very little usage. For that, there is...
it's, that's basically where these LLM API calls are just perfect because you don't need to worry about scaling this up, scaling this down, um, handling-
I see
... handling these issues.
Serverless versus serverful.
Yeah, exactly.
I see. Okay.
Like, I see them a bit... Like, I see OpenAI and all of these other providers, I see them a bit as the Like as the Amazon, sorry, AWS of, of AI. So it's a bit similar how like- ...
back before AWS you would have to have your, your servers and, and, and like buy new servers or get rid of servers, and then with AWS-
Mm-hmm
... it just became so much easier to just ramp stuff up and down.
Yeah.
And this is like the taking it even, even, uh, to the next level for AI.
Yeah. I, I am a big believer in this. Basically, it's, you know, intelligence on demand.
Yeah.
We're probably not using it enough in our daily lives to do things. Actually, we should be able to spin up 100 things at once and like go do things and then, you know, stop, and I feel like we're still trying to figure out how to use LLMs in our lives effectively.
Yeah.
Yeah.
Yeah.
100%. I think that goes back to the whole-- Like that, that's for me where the big opportunity is for if you wanna do a startup. Um, it's not about... Like you can let the big labs handle the challenge of more intelligence.
Yeah.
But, um, it's the-
Existing intelligence, how do you integrate?
Yeah.
Yeah.
How do you actually incorporate it into your life?
It's AI engineering.
Future Gaze55:59
Okay. Cool, cool, cool. Um, the one, one other thing I wanted to touch on was multimodality in frontier models. Dwarkesh had a interesting application of Gemini recently where he just fed raw audio in and got diarized transcription out.
Yeah, yeah.
Or, or timestamps out.
Yeah.
And I think that'll come. So basically what, what, what we're saying here is another wave of transformers eating things, 'cause right now models are pretty much single modality things. You know, you have Whisper, you have a pipeline, everything.
Uh, no, no, no. We only feed the, like the raw, the raw files. Do, do you think that will be realistic for you?
I 100% agree.
Okay.
Basically, everything that we talked about earlier with like the speaker diarization- ... and, uh, heuristics and everything, I, I completely agree, like in the, in the future that will just be put everything into a big multimodal LLM.
Okay.
And it will output, uh, everything that you want.
Yeah.
So I've also experimented with that, like just-
With, with, uh, Gemini too?
Uh, with Gemini 2.0 Flash.
Yeah, yeah.
Just for fun.
Yeah, yeah.
Because the big difference right now is still like the cost difference of doing speaker diarization this way or doing transcription this way is a huge difference to the pipeline that we've built up.
Huh, okay. I, I need to figure out what, what that cost is because in my mind 2 Flash is so cheap.
Yeah.
But maybe not cheap enough for you.
Uh, no, I mean, if you compare it to, yeah, Whisper and speaker diarization and especially self-hosting it and-
Yeah, yeah. Yeah, yeah. Okay.
But we will get there, right?
Yeah, yeah.
Like this is just a question of time. And, um, at some point, as soon as that happens, uh, we'll be the first ones to switch.
Yeah. Awesome. Anything else that you're like sort of eyeing on the horizon as like we are thinking about this feature, we're think- we're thinking about incorporating this new functionality of, of AI into our, into our app?
Yeah. We-- I mean, we-- there are so many areas that we're thinking about. Like our challenge is a bit more, uh-
Choosing.
Yeah, choosing.
Yeah.
So I mean, I think for me, like looking in-into like the next couple of years, the, the, the big areas that interest us a lot, basically four areas. Like one is content. Um, right now it's, it's podcasts. I mean, you did mention, I think you mentioned like you can also upload audiobooks and YouTube videos.
YouTube. I actually use the YouTube one a fair amount.
But in the future, we, we wanna also have audiobooks natively in, in the app and, uh, we want to enable AI-generated content. Like just think of take deep research and NotebookLM podcast generation, like put these together. That, that should be, that should be in our app.
The second area is discovery. I think in general-
Yeah, I noticed that you don't have-- So you have download counts and most snips, right? Something like that.
Yeah.
Yeah.
Yeah. On the discovery side, we wanna do much, much more. I think in general discovery as a paradigm in all apps is will undergo a change thanks to AI. You know, there has been a lot of talk-- Before Elon bought Twitter, there was a lot of talk about bring your own algorithm-
Yeah
... to Twitter.
Yeah.
And that was Ja-Jack Dorsey's big thing or like he, he talked a lot about that.
Yeah.
And I actually think this is coming but with a bit of a twist. So I, I think what actually AI will enable is not that you bring your own algorithm, but you will be able to talk. You will be able to communicate with the algorithm.
So you can just tell the algorithm like, "Hey, you keep showing me cat videos." "And I know I freaking love them, and that's why you keep showing them to me." "But please, for the next two hours I really want to like get more into AI stuff.
Do not show me cat videos." And then it will just, uh, adapt. And, um, of course the question is, you know, like big platforms like, I don't know, let's say, say TikTok, they do not have the incentive to offer that.
Exactly.
Um-
That's what I was gonna say.
But we actually like our-- we are driven by helping you learn, get the most, like achieve your goals. And so for us this is actually very much our incentive like, "Hey, no, you, you, you should be able to guide it."
Um, yeah, so that was a long way of, of, of saying that I think, um, there will happen a lot in recommendation.
Order by.
Yeah.
Um, I just have-
Most popular.
Yeah, yeah. I, I, I think collaborative filtering will be the first step, right? For, for Rexus and then, and then some LLM fancy stuff. Um-
Um, yeah. Maybe, maybe to, to go back to the question that you had before. So the other-- Like these were the first two areas.
Yep.
Like the, the other two are voice, voices and interfaces-
Yep
... and voice AI.
Well, how is this gonna exist?
Yeah. So maybe I can tell you a bit first like why I find it so interesting for us.
Yeah.
Because voice as an interface, like historically there has been so much talk about it, and it always fell flat. The reason why I'm excited about it, uh, this, this time around is with any consumer app, I, I like- To ask myself, what is the moment in my life, what is the trigger in my life that gets me to open this app and start using it?
So for example, I don't know, take Airbnb. It's, the trigger is like, ah, you want to travel, and then, and then you, you do that. Uh, then you open up the app. Apps that do not have this already existing natural trigger in your life, uh, it's very difficult for a consumer app to then get a user to-
You need a-
... open the app again.
Yep.
There's basically only one app, one super successful app that has been able to do that without this natural trigger, and that is Duolingo.
Ah.
So Duolingo, like everyone wants to learn a language, but there's, you don't have this natural moment during your day where it's like, ah, now I need to open up this app.
You have the notifications.
Exactly.
The owl memes.
Exactly.
Yeah.
So they, I mean, they gamified the shit out of it. I mean, super successful, super beautiful. They are the GOATs- ... in this, in this arena. But the much easier is actually, no, there is already this trigger, and then you don't have to do all of the streaks and leaderboards and, and everything.
Okay.
That's a bit of a, a context. Now, if you look what we are doing and our goal of, of getting people to really maximize what they get out of their listening, um, we are interested in, in... There are a couple of features where we know we can sort of 10X the value that people get out of a podcast.
Okay.
Um, but we need them to do something for that. There is friction involved because it's, it's, it's all about learning, right? It's about thinking for yourself. Like that's, these are, those are the moments when you actually start, yeah, really 10X-ing the value that you got out of the podcast instead of just consuming it.
Apply-
Yeah
... apply the knowledges. Yeah. Okay. Okay.
Yeah. Basically being forced to think about like what was actually the main takeaway for you from this episode.
Okay.
Like, uh, this is something that I like doing myself. For every episode that I listen to, I try to boil it down to, like try to decide one single takeaway.
Yeah.
Even though there might have been 10 amazing things.
You pick one.
One most important one.
Yeah.
And this is a, this is an active process that is like a forcing function in your brain to challenge all of the insights and really come up with the one thing, um, that is applicable to you and your life and what you might wanna do with it.
So it also helps you to turn it into, into action. This is, uh, basically a feature that we're interested in, but you have to get the user to, to use that, right? So when do you get the user to use that?
If this is all text-based, then we're basically playing the same game as Duolingo, where at some point you're gonna get a notification from Snipd and be like, "Hey, Swyx, come on. You know you should do this." Maybe there's a blue owl.
Um, but if you have voice, you can basically hook into the existing habits that the user already has. So you already have this habit that you listen to a podcast. You're already doing that. Once an episode ends, instead of just jumping into the next episode-
Ah
... you can now actually have your AI companion come on.
Ah.
And you can have a quick conversation. You can go through these things.
Yeah.
Um, and how that looks like in detail, like that is still... Like we need to figure that out. But just this paradigm of you're now, you're staying in the flow. Like a bit- This, this also relates to what you were saying, like AI that is invisible.
Like you're staying in the flow what you're already doing, but now we can insert a completely new experience in there that helps you get, get the most out of your listening.
Yeah. I think your framing of this is very powerful because I think this is where you are a product person more than an engineer. Because an engineer would just be like, "Oh, it's just chat with your podcast." It's like chat with PDF, chat with podcast.
Okay. Cool. But you're framing it in a different light that actually makes sense to me now, as opposed to previously, I don't chat with my podcast. Like why? I, I just listen to the podcast, right? But for you it's more about retention and learning and, and all that.
Um, and because you're very serious about it, that's why you started a company, um, so you're focused on that. Whereas yeah, I'm still me. Like I will admit, I'm still stuck in that consume, consume, consume mentality. And I know it's not good, but this is, you know, my default.
Which is why I was a little bit lost when you were saying all the things about Duolingo, and you were saying the things about the trigger. 'Cause my trigger for, for listening to the podcast is, you know, I'm by myself.
That's my trigger. But you're saying the trigger is not about listening to the podcast. The trigger is remembering and retaining and processing the podcast I just listened to.
So, no, but so, so what I meant, like you already have this trigger that gets you to start listening to a podcast.
Yes.
Like this you already have.
Yes.
And so do, I don't know, more-
Millions of people.
Yeah.
Yeah.
So there are more than half a billion monthly active podcast listeners.
Okay.
Um, so you already have this trigger that gets you to start listening. But you do not have this trigger, as you just said yourself, basically, you do not have this trigger that gets you to regularly, um-
Process. Yeah
... process this information, right? And, um, voice basically for me is, is uh, the ability to hook into your existing trigger, with the trigger that I was talking about is basically your podcast ends and you're just still listening.
So we just continue, and we can now spend... You know, this can be two minutes. Like I'm not saying now this is like a 60-minute process. I think like two minutes, three minutes, that can just come on completely naturally.
And if we manage to do that, and you start noticing as a user like, "Freaking hell, like I'm just now spending three minutes with this AI companion," but like-
Your retention is more
... I'm taking this much away.
Yeah, yeah.
And it's not... And like retention is one thing, but you like, you start to take what you've learned and apply it to what's important to you, like your thinking.
Yeah.
Um- If we get you to notice that feeling, then, yeah, then move on.
Yeah. I, I would say, like, a lot of people rely on Anki, Anki notes, like flashcards and all that to, to do that. But making the notes is also a, a chore. And, um, so I think this, I think this could be very, very interesting.
Now, I think that... I, I'm just noticing that it's, it's kind of like a different usage mode. Like, you already talked about this, you know, the, the name of Snipd. It's very Snipd-centric. And I actually originally also resisted adopting Snipd because of that.
But now you're like, you know, you observe that people are listening to long-form episodes, and they... you're talking at the end. Like, the ideal implementation of this is I browse through a bunch of Snips of the things that I'm subscribed to, I listen to the Snips, I talk with it, and then maybe it, it double-clicks on, on the, the podcast, and it, it g- goes and finds other timestamps that are relevant to the thing that I want to talk about.
Just... I was just thinking about that. May- I don't know if that's interesting.
I think these are all areas that, that we should explore.
Yeah.
Like, um, we're, we're still quite open about how this will look like in, in detail.
What are your thoughts on voice cloning? Everyone wants to continue... I have had my voice cloned, and people have talked to me, my- the AI version of me. Is that too creepy?
I, I don't think it's too creepy- ... in the future.
Okay.
With a lot of these things, you know, society is going through a change. Um, and things seem quite weird now that in the future will seem normal. Um, I think already voice cloning has become much more normalized. I remember I was at the...
I think it was 2017, uh, NIPS conference.
Mm-hmm.
Back when in-
San Diego?
No, LA.
LA.
Um-
It was the Flo Rida one? I think.
Yeah, yeah, yeah, yeah.
Yeah, Flo Rida.
Flo Rida.
Yeah. So everyone says that was peak NIPS. Yeah.
Um, I remember there was this, this, uh, talk or workshop by Lyrebird. Um, they actually got acquired by Descript later. They were doing voice cloning. And they, they were showing off their tech, and there was this huge discussion later on, like the all of the moral implications and, and ethical implications.
Um, and it really felt like this would never be accepted by society.
Mm-hmm.
And you look now, you have 11 labs, and just anyone can just clone their voice. And-
Mm
... like, no one really talks about it as like, "Oh my God, the world is going to end."
Yeah.
So I think society will get, will get used to that. In our case, I think there are some interesting applications where we'd also be super interested in working together with, uh, creators, like podcast creators, to play a bit around with this, with this concept.
I think that would be super cool if someone can, um, you know, come onto Snipd, go to the Latent Space podcast, and start chatting with, uh, AI Swix.
Yeah. Um, no, I think, I think we'd be down. We want to... Obviously, I think as an AI podcast, we should be first consumers of these things.
Yeah.
I would say that one observation I've made about podcasting, this is the general state of the market, and you can ask me your questions, uh, you know, things you wanna ask about podcasters. We are focusing a lot more on YouTube this year.
Video & Creators1:10:03
YouTube is the best podcasting platform. It is not MP3s, it is not Apple Podcasts, it is not Spotify. It's YouTube. And it's just the social layer of recommendations and the existing habit that people have of logging onto YouTube and getting, getting that.
That's my observation. I- you, you can riff on that. The only thing I would just say is, like, when you were listing your list of priorities, you said audiobooks first over YouTube, and I would switch that if I were you.
Yeah, like as in YouTube video, video podcasts. I mean, it's obvious, yeah, video podcasts are here to stay. I-
Not just here to stay. Bigger.
Yeah. What I want to do with Snipd is obviously also add video to, to the platform.
Oh, yeah?
The way I see video is I do believe it's... I like this concept of backgroundable video. I didn't come up with this concept. Uh, it was actually Gustav Södastrom.
The CPO of Spotify.
Exactly, exactly. When I speak with people, it, it remains, uh, true that they listen to podcasts when they do something else at the same time. Like, this is like 90% of their consumption. Also, if they, if they listen to on, on YouTube.
But every now and then, it's nice to have the video. It's nice if you're, for example, just watching a clip. It's nice if they sometimes mention something, like they show some slides, or they show some- something where you need to have the visual, uh, with it.
It helps you connect much more with your, uh, with the host as in like as, as, as a listener. But the biggest benefit I see with video is discovery. I think that is, that is also why YouTube has become the biggest podcast player out there because they have the discovery, and discovery in video is just so much easier and so much better and so much more engaging.
So this is the area where I'm most interested about when it comes to video and Snips, that we can provide a much better, much more engaging, and much more fun discovery experience.
For consumers?
For con- Yeah, for consumers.
Okay. I think that you almost have, like, three different audiences. The vast majority of people for you is the people listening to podcasts, right? Of course. Then there's a second layer of people who create Snips, right? Who, who add extra data, annotation, value to, to your platform.
By the way, we use the Snip count as a proxy for popularity, right? Because we have download counts, but for example, platforms like Spotify re-host our MP3 file, so we don't get any download count from Spotify. Snip count is active, like I opt in to, to listen to you, and I shared this.
Th- th- those are, those are really, really good metrics. But the third audience that you haven't really touched is the podcast creators like myself.
Yeah.
And for me, discovery from that point of view, not from your point of view, discovery for me is like I want to be discovered, and I think YouTube is still there. Twitter, obviously for me, Substack, Hacker News. Uh, I try to...
I, I really try very hard to rank on Hacker News.
Yeah.
I think when TikTok took this very seriously, they prioritized the creators of the content. And for you, the creator of the content was the Snips. But- There may be a world for you in which you prioritize the creators of the podcast.
Yeah. Yeah, interesting observation. What, what, uh, what are some of your ideas or thoughts? Do you have some specific-
Um, Riverside is the closest that has come to it. Descript is number two. Descript bought a Riverside competitor and, and, uh, as far as I can tell, it's not very, it's not been very successful. Descript just, like, has, like, has a very, very good niche, very, very good editing angle, and then just has- hasn't done anything interesting since then.
Although Underlord is good, it's not great. Like, your chapterization is better than Descript's. Again, like, they should be able to be you. They're not. And Riverside is good also. Very, very good. Very, very, very good. Like, so we, we actually recently started a second series of podcasts within LeanSpace that is YouTube only, 'cause you only find it on YouTube, and it's also shorter.
So, like, this is, like, a one-and-a-half-hour, two-hour thing. It's remote only, 30 minutes, chop, chop, s- send it all into Riverside. Riverside, pretty good for that. Not great. It doesn't do good thumbnails. It doesn't g- do... Uh, the editing's still a little bit rough.
It has, like, this auto editor where, like, you know, whoever's actively speaking, it focuses on the editor, on, on the, on the active speaker, and then sometimes it, it goes back to, like, the multi-speaker view. That kind of stuff, people like that.
Okay. But, like, the shorts are still not great. You know, like, I, I still need auto... I need to, to manually download it and then republish it to YouTube. The shorts I still need to pick. They, they mostly suck.
There's still a lot of r- rough edges there that ideally me, as a creator, like, you know what I want. You definitely know what I want. I sit down, record, press a button, done.
Yeah.
We're still not there.
Yeah.
And I think you guys could do it.
Okay. So if I can translate that for you, it's really about the simplifying the creation process of, of the podcast.
Yeah.
Yeah.
And, and I, I'll tell you what, this, this will increase the quality because the reason that most podcasts or, or YouTube videos are shit is they are made by people who don't have life experience, who are not that important in the world.
They're not doing important jobs. And so what you wanna actually enable is CEOs to each of them make their own podcasts, who are busy. They're not gonna, you know, sit there and figure out Riverside. The, a lot of the reason that people like LeanSpace is it takes an idiot like me, who could be doing a lot more with my life, ma- making a lot more money, having a real job somewhere else.
I just choose to do this 'cause I like it. But otherwise they will never get access to me, and the access to the people that I have access to.
So that's-
That's my pitch.
Cool.
Cool. Um, anything else that you normally wanna talk to podcasters about?
I think we've, we've covered everything. Uh, I guess, like, last, last message is, uh, you know, go try out Snipd.
Yeah.
It's a premium version, so you can use and try out everything for free.
And I, I-
Also happy to provide you with a, with a link that you can add to the show notes.
Yes.
Uh, try out the premium version also for free for a month.
Yeah.
If people wanna do that. Um, yeah.
Yeah.
Give it a shot.
Parting Words1:16:08
I would say, uh, yeah, thanks for coming on. Uh, I would say that the... After you demoed me, I did not convert for another four to six months because I found it very challenging to switch over, and I think that's the main thing.
Like, you, you, you basically had... You, you have import OPML, right? But there's no way to import, like, all the existing, like, half-listened-to episodes or, like, my rankings or whatever. And for that, for listeners who, who are... I have a p-a blog post where I talked about my, my, my switch.
Just treat it as a chance to clean house.
That is a good point. Yeah.
Just delete, delete things and, you know, just, uh, refocus your-
Fresh, fresh start. 2025 fresh start.
Yeah. Great. Well, thank you for working on Snipd. Thank you for coming on. You know, we usually we spend a lot of time talking to, like, big companies, like venture startups, B2B SaaS, you know, that kind of stuff.
But, uh, I think your journey as, like, you know, as a sm- small team building a B2C consumer app is the kind of stuff that we like to also feature because a lot of people wanna build what you're doing, and they don't see role models that are successful, that are confident, that are, like, having success, um, in this market, which is very challenging.
Um, so, uh, yeah, thanks for, thanks for sharing some of your, your thoughts.
Thanks. Yeah. Thanks, thanks for having me, and thank you for creating an amazing podcast-
Thank you
... and an amazing conference as well.
Thank you.






