LALatent SpaceJul 11, 2025· 1:04:10

Personalized AI Language Education — with Andrew Hsu, Speak

Andrew Hsu, CTO of Speak, explains how the company built the "third generation" of language learning by betting on AI speech and language models before they were ready. Speak focused on South Korea early, a counterintuitive move for a San Francisco startup, and now 6% of the Korean population has tried the app. The turning point came with Whisper and GPT in 2022, which enabled Speak to evolve from a listen-and-repeat tool into a full-featured AI tutor that gives real-time feedback and adapts to learners. Speak generates over $50M ARR primarily from consumer subscriptions, and Hsu details the product decisions that made it work: abandoning a free version, building custom ASR for low latency, and emphasizing functional fluency over textbook language. He also reveals that Speak is building a real-time voice platform and a knowledge graph for personalized fluency scores, with an eye toward expanding beyond language into general AI-powered education.

  1. 0:00Origins
  2. 4:35Speech
  3. 6:03Onboarding
  4. 9:47Gen 3
  5. 12:43Business
  6. 16:32Korea
  7. 19:10Early Tech
  8. 23:33LLM
  9. 30:46Fluency
  10. 39:52Future
  11. 52:11AI Tools
  12. 56:21Wrap Up

Powered by PodHood

Transcript

Origins0:00

Alessio0:06

Hey everyone, welcome to the Latent Space Podcast. This is Alessio, partner and CTO at Decibel, and I'm joined by my co-host, Wilkes, founder of Small AI.

Wilkes0:13

Hello, hello, we're back in the studio with Andrew Hsu of Speak. Welcome.

Andrew Hsu0:16

Thank you for having me.

Wilkes0:17

I have to start this off. I didn't prep you on this at all but you were a Thiel Fellow in twenty eleven.

Andrew Hsu0:23

First class. That's right.

Wilkes0:24

First class?

Andrew Hsu0:24

Yeah, yeah.

Wilkes0:24

Is that, is that the one with, um, SBF?

Andrew Hsu0:27

No, he was, I think, several years later, actually.

Wilkes0:29

Okay.

Andrew Hsu0:29

Yeah, yeah.

Wilkes0:30

What was it like? Just talk about the-

Andrew Hsu0:32

That's a good question. Haven't been asked that one in a while. It was a really crazy idea at the time and very controversial, and I think the first few years of the fellowship were definitely, "Let's just find twenty people under twenty and give them a hundred thousand dollars to drop out of college," and it could be-- It was no holds barred.

You could do anything. You could be doing some crazy research idea, a startup, anything. Um, and, uh, I actually met my current co-founder at Speak. He was in the second year of the fellowship and-

Wilkes1:06

Ooh

Andrew Hsu1:06

... made many, like, very close friends from the first, uh, few years. But, I mean, for me, it was life-changing. I, I had a very unusual path where I was actually-- I did finish college, unfortunately. I was in grad school at the time because I went to school really early.

Wilkes1:20

Yeah, I was like- ... "Aren't you, aren't you too old?" You know? Thiel likes them young.

Andrew Hsu1:23

I was

I was nineteen at the time and in grad school. It was a very accelerated path, but I think, like, I knew at the time that I was going to leave grad school and do startups anyway, and the timing lined up really well.

Wilkes1:38

Yeah, yeah.

Andrew Hsu1:39

Yeah.

Wilkes1:39

Vitalik, I think.

Andrew Hsu1:41

He was also in a later year.

Wilkes1:42

Ah.

Andrew Hsu1:43

Yeah.

Wilkes1:43

Damn. Okay, anyway. The, the-

Andrew Hsu1:45

But the first two years had-- I mean, there were some crazy s-successes. You know, Dylan from Figma. Uh, I mean, yeah, like, a lot of people.

Wilkes1:53

Awesome. Well, you know, feel free to bring in those stories as and when-

Andrew Hsu1:56

Cool, yeah

Wilkes1:56

... because obviously only, only you know, like, those kinds of people. You are now CTO co-founder of Speak.

Andrew Hsu2:01

Yes.

Wilkes2:01

Uh, I would say from a very early stage, like, one of the most successful and prominent OpenAI partners that, like, anyone would know as, like, doing, doing well and, like, teaching English to Koreans is, like, your, your, your rough, um, remit at the time.

How did that all come about?

Andrew Hsu2:18

It's funny that you say that because despite our- ... current sort of revenue scale and objectively, I think, how successful we are, we've always operated in a market, at least initially, on the other side of the world and been much, much more popular in the sort of Eastern world and a bunch of Asian markets and relatively unknown in the West.

So it hasn't really felt like we've had that sort of awareness until, you know, the past few years really. But brief story is that my co-founder and I back in twenty sixteen were fascinated by the promise of AI, and we spent a year sabbatical basically learning everything we could.

We talked to Karpathy back then actually when he was, like, just finishing grad school and did a lot of sort of self-study research, and we were just so convinced, I think fundamentally, that speech models were going like this, language models were going like this, and in the five to ten year span they would become superhuman.

And we were utterly convinced of this future, and we saw that the way people learn things and specifically learn languages, which was a very sort of human-based thing if you really care about fluency, that would completely change and we'd be able to build language tutors that were pure software, pure AI.

So that was kind of the genesis story of Speak. It took much, much longer than we expected to build a great product and find good PMF. The first few years were very painful, and I think without this really compelling vision of the future, we would've quit.

We actually, like, never pivoted. Last year, we brought the entire company to Taipei. We do this company trip every year. And we played our original YC application video on screen, and it was really funny because the things we were saying in that video were the exact same things that I still say today about the long-term vision and what we're building towards.

Um, so that was really cool to see.

Wilkes4:16

Can you summarize the long-term vision again?

Andrew Hsu4:18

It was that as speech models and language models become superhuman, that would let us create an AI language tutor that would help you become fluent faster than any human could. And I think we're like eighty to ninety percent of the tech is here now.

Speech4:35

Alessio4:35

And you have this big focus on, like, speaking. Obviously, it's in the name of the company.

Andrew Hsu4:39

Yeah, that's right.

Alessio4:39

Uh, and I think the speech models were maybe a little delayed compared to the text models. Did you ever think about, okay, maybe speech is just not gonna work for this use case? Or like, what were kinda like the valleys of, you know, discomfort, and then what were maybe some of the pivotal releases and models that you were like, "Okay, it's gonna work.

It might take a little longer, but it's gonna work"?

Andrew Hsu4:59

So we've always done custom speech stuff. The first act of the company, if you will, was before LMs, right? Before twenty twenty-two when Whisper came out, when ChatGPT came out. In the years before that, you know, like roughly two to three years is when we feel like we found PMF in South Korea and then started growing still only in that market, still only teaching English.

And we developed custom speech recognition models, and users were speaking into the app all day, so we had a ton of this non-native English speaker data, and we would use that to fine-tune models, understand our users better. We still do that today, and it's important for us for the core recording loop in many of our lessons that it's extremely fast, so we're very latency sensitive.

There's many other sort of product surfaces within the app today that are more LM-powered, where it's more open-ended real tutoring, where we actually give you feedback on what you said and the semantics and so on. So that stuff is more like Whisper-powered, more LM-powered.

But we've always had like a very fast core ASR loop that's been fully custom.

Alessio6:03

I just onboarded to the app earlier today.

Onboarding6:03

Andrew Hsu6:05

Yeah.

Alessio6:05

Unlike other apps, there's kinda like this, uh, tutor conversation that you do for onboarding.

Andrew Hsu6:09

Yeah.

Alessio6:10

I'm guessing that is mostly LLM-based, and then you're kinda judging the person respond. So I selected Spanish, and the conversation was in Spanish via text to start, and then from there started to create lessons for me.

Andrew Hsu6:20

Yeah.

Alessio6:21

Was that all unlocked from LLMs where now you can kinda have these conversations and then bring people into the speech flow?

Andrew Hsu6:26

Yeah. So before that, so we call that Magic Onboarding.

Alessio6:29

Mm-hmm.

Andrew Hsu6:29

And it was a new thing we built that was more conversational. We wanted it to feel more like you were talking with a tutor, and they were, uh, sort of learning things about you, and we would use that later to personalize the experience.

Before that, we had a, like, a much more traditional app onboarding. There's still a lot of open questions, interesting questions around what is the proper onboarding UX because a lot of people start using Speak, and they're not in a situation where they can actually speak aloud.

So we have, like, you know, fallback outlets and so on, but it's something we're, like, super actively experimenting with.

Alessio6:59

Is there a structured output behind that? You know, anything that you found implementing Magic Onboarding? I think people always wanna improve onboarding. What's the uplift? Or was there one? We still don't know yet.

Andrew Hsu7:10

The interesting thing is that, in general, because it's speaking based, which is a much higher barrier than just, like, tapping a multiple choice button, what we see is that install to sign-up rate is a decent amount lower, but trial start rate is higher.

It's still something that we have an active experiment that, um, that's running, and we're trying to be v- super agile about testing many different sort of, like, formats of this. I don't think I have, like, the final answer yet.

Alessio7:37

Yeah.

Andrew Hsu7:37

But I think the intent, the really, like, vision that we're going for here is that as soon as you download the app from the App Store, maybe you see it in an ad, the first thing that... The first interaction when you have, like, a fresh open of the app should feel pretty futuristic.

It should feel like, okay, this is, like, the new AI-native, next gen way of learning a language to fluency, and that's kind of been always our ambition. Like, we want to build something that wasn't possible before without LLM and, like, AI technology.

Alessio8:09

Yeah. I, I think, um, I wanted to go back on the onboarding soon, but there's a general idea of, like, when you replace a form with a voice bot, that you need to have some kind of state machine behind the hook.

The thing to drive, like, what else don't, don't I know about you? Let me proactively ask that. And I'm just wondering if you had any insights there, or is it literally just a state machine?

Andrew Hsu8:32

We tried both, actually.

Alessio8:33

Yeah.

Andrew Hsu8:33

Right now, I think probably what you saw is a state machine, but I think that-

Alessio8:37

Trust the AGI.

Andrew Hsu8:39

Yeah. Right. I think that things should move in a direction where it's much more of a natural conversation.

Alessio8:44

Yeah.

Andrew Hsu8:45

There is a general sense of a goal in the prompt that you can specify, and part of the hard thing here is all the guardrails, right? Like, once you start talking about what you had for breakfast yesterday, right, and trying to be, like, antagonistic to the system, then things start, like, really going off the rails.

So for a bunch of these experiences, we're pretty careful about the fallbacks, and we have a lot of evals around that. But I think where it should end up is just feeling like you have a quick three to five-minute conversation with your tutor, and then it knows a lot about you, and then you create your account, et cetera.

Alessio9:20

And you create memories? Like-

Andrew Hsu9:23

Yeah.

Alessio9:23

Yeah.

Andrew Hsu9:23

So we, we store what you're saying. We summarize. In the experience, the way it works is the tutor will ask you some sort of question like, "What are your goals around learning English or the language?" And then we will basically use a separate LM prompt to summarize, so it's not the, like, full transcript for what you said that you see.

It's more of, like, an abstracted, "Okay, here's what you care about," and we think that's a better product experience.

Gen 39:47

Alessio9:47

What were some of the other key tenants on the product? Obviously, language learning is, like, one of those consumer markets where, like, dozens of companies always try and get started, and you get these old companies like, you know, Babbel, and you got Duolingo.

Um, so speak- speaking, the, the act of speaking was, like, a big part of it. I think this memory stuff is great. I, I think if you tried some of the other apps, it's like they always try to re-ask you the same things that you got wrong before, but you're not really learning.

Is there anything else that is maybe, uh, not as obvious from the outside in the design of the app and the product that is, like, you think is really different?

Andrew Hsu10:19

I would say from a macro level, this is actually a pretty new product category, AI-powered language learning. And all these apps that you mentioned, Duolingo, Babbel, et cetera, they're more of, like, the Gen 2 of language learning. So, like, if you think about Gen 1 was Rosetta Stone if you, you know, if you remember, right?

CD-ROMs in airports. And then Gen 2 was basically mobile, so you have these very casual, massively popular mobile apps like Duolingo that I think the comp there is probably closer to a mobile game, something that feels productive, something that's very engaging, very gamified.

And-

Alessio10:54

Duolingo's very leaning into it, the gamification.

Andrew Hsu10:57

And they've done an amazing job of that, to be clear. Um-

Alessio10:59

Yeah. They, they might be the world's best people at it.

Andrew Hsu11:01

Yeah. Um, and our view is that LLMs and AI now enable Gen 3 of language learning, which is something that is very AI-native, very focused on functional fluency, which is why we do all these role plays and let you practice Spanish by talking to your Uber driver.

We don't teach vocabulary and grammar. We teach sentence patterns, and we try to get you to just repeat and drill and drill and drill, almost like you're in a gym, until it's automatic 'cause that's what speaking is, right?

Like, it has to be spontaneous and automatic. In terms of the other aspects of the design, though, we went through many, many iterations over the first few years of starting the company. This is kind of what I was mentioning about it was, like, really painful in the first four or five years, and in fact, the current version of the Speak app is not the first thing that we launched.

We had something that we call internally the Red App, which was, like, a red app icon, still a similar logo, and it was more around packs of content instead of courses where you could sort of choose any topic that you wanted to learn.

It was for many different languages for learning. It was essentially, like, not a very directed experience, and it didn't really work. It was free. It was a very basic thing. But we, in 2018, tore everything down and realized that we had to really fully change what we were doing, and that's when we decided to focus on South Korea specifically on teaching English.

We built a bunch of new lesson types, and we created our courses so that the experience was much more on rails. We realized people don't want to choose. They're already using some of their motivation on a daily basis just to open the app.

They don't wanna make another choice after that, right? Just tell me what to do, right? Like, you know, give me a big button, and then I can tap it and just start a video lesson or whatever. We also, pretty critically, I think, abandoned the free version and just went straight premium, and we kind of sidestepped the motivation question that way because we knew that there were a ton of users that really wanted to learn English and were already really motivated, so we wanted to basically filter for these users.

Business12:43

Andrew Hsu13:05

Um, so they... You know, like it was... I wouldn't say there was one silver bullet. It was kind of the combination of many learnings over three or four years, and then that started really growing in South Korea. And from there, I guess like phase two was really 2022 when LLMs came out and Whisper came out, and that allowed us to go from this more supplemental speaking practice tool to more full-featured language tutoring where we could use LLMs like 3.5 Turbo back then to give you direct feedback on your wording and on, you know, like that was kind of a weird thing to say.

A native speaker would say it this way or use a different word or whatever.

Wilkes13:46

I always do a poor job of doing this, but, uh, can we get some headline numbers, like just to get a sense of scale? Because I think maybe some audiences don't know-

Andrew Hsu13:53

Yeah

Wilkes13:53

... where you're at now in terms of your reach.

Andrew Hsu13:56

So we're now the biggest English app in South Korea.

Wilkes13:59

Yeah.

Andrew Hsu13:59

We do billboards, big celebrity campaigns, that sort of scale. Like we're, you know, very popular there. I think like 6% of the Korean population has tried us. Um, we're, you know, like well on the way in a bunch of other Asian markets like Japan, Taiwan.

Wilkes14:13

Mm-hmm.

Andrew Hsu14:13

So the Asian markets are currently our mainstay. We also teach English in 40 more countries. We're coming to the US as well, launching... I mean, we, we have Spanish, French live, and several more languages are coming this year.

That's a huge focus of the company right now. Um, in terms of revenue scale, well over 50 million ARR. It's a pretty simple business model. It's like mostly consumer.

Wilkes14:37

Subscription. Yeah.

Andrew Hsu14:38

The B2B stuff is super, super exciting, and that's also growing really fast, and I think it'll be a really meaningful part of the business.

Wilkes14:45

When did you start B2B?

Andrew Hsu14:46

About a year ago. It w-

Wilkes14:48

Okay

Andrew Hsu14:48

... it kind of... It was like very much a side bet/experiment at first- ... and then it just started working. And-

Wilkes14:54

Of course it's gonna work .

Andrew Hsu14:55

Uh, yeah. And now it's like, okay, you know, this is part of the future, right? This is a real thing. Yeah. So that's exciting.

Wilkes15:02

What's the B2B race between learning language and like real-time AI translation?

Andrew Hsu15:06

Uh-

Wilkes15:06

At Google I/O, they're like one of those like Google Beam things for like, uh, you know, for conferencing. They do real-time translation, like-

Andrew Hsu15:12

Yeah. Yeah. Yeah. So people always ask this, right? They're- they're always like, "What happens when the Babel Fish comes?" Right? When the real-time translation comes.

Wilkes15:21

And Babel Fish is the Hitchhiker's Guide.

Andrew Hsu15:23

Right. Yes, exactly. The counterexample that I always have that I think is quite illustrative is in German, the verb is at the end of the sentence, right? So if you're trying to do real-time translation from German to English, as an example, you can't actually make any progress on the English until you hear the whole German sentence and you know what the verb is at the end, right?

So like the minimum latency there is, is the full sentence. And that, that's like an example of the technical blocker for- ... like why it'll never be truly, truly perfect. But also, I think besides that, if you talk to all of our users in Asia, they don't want a translator.

The reason that they are trying to learn English is to make themselves a better person, to connect with other people. Like, they wanna be able to look you in the eye and speak English, speak the same language as you, right?

So it's actually like a very different thing. I think what will end up happening is that we will build a real-time translation feature into Speak and have it integrated into the learning experience.

Wilkes16:23

And also, like there's always that human side, right? Like I'm dating a Romanian woman.

Andrew Hsu16:27

Yeah.

Wilkes16:28

Uh, his wife is trying to learn Italian. Like, there's always that-

Andrew Hsu16:31

Yeah, absolutely

Wilkes16:32

... which is gonna keep happening. I wanna double-click on Korea.

Korea16:32

Andrew Hsu16:34

Yeah.

Wilkes16:35

I think it's like a very insightful, smart decision. Maybe people only know Korea through K-pop. But actually, I think a lot of Americans learn Korean because of K-pop. That's a, that's a side thing.

Andrew Hsu16:46

Yeah.

Wilkes16:46

But like you could have done Taiwan, you could have done China. I saw... I remember seeing a d- documentary about how China was crazy about English or mad about English. I think that was the title of the documentary.

Was it obvious? D- Were you sure when you went into Korea, or was it just a test?

Andrew Hsu17:00

We visited a bunch of Asian countries when we were thinking about how do we relaunch things, how do we focus in, and we almost chose Taiwan, actually.

Wilkes17:08

Yeah.

Andrew Hsu17:08

Um, but I think it was a little bit serendipitous. So our first employee is Korean and was my co-founder's college roommate, actually.

Wilkes17:17

Mm.

Andrew Hsu17:17

When my co-founder visited Seoul to check out the market, he asked SJ to come along as essentially a translator and to like, you know, facilitate. And I think that just went really well, and it was just very obvious from being on the ground in the market that Korea is pretty obsessed with learning English, and there is every human-based solution possible, right?

You know, like English academies, classes, skyscrapers full of classrooms, stuff like that. And our logic was basically, if we can really make headway and win this market that is chock-full of these human competitor products and all these people that fundamentally care about fluency, then we probably have something pretty real and strong PMF that we could win other markets with.

So that was the original logic and, you know, so far it's been working.

Wilkes18:14

It's retroactively obvious, which is the best kind of obvious. But, like, it's so counterintuitive that you would be the team to do this and not a Korean team, right?

Andrew Hsu18:22

Yeah, yeah.

Wilkes18:22

Where they would be... They would know because they had personal experience of, like, "I started in Korean, I learned English. Here's how you do it."

Andrew Hsu18:29

Yeah. In hindsight, super weird, right? Like, we, we were definitely s- you know, sitting in an office here in San Francisco- ... operating with users in a market all the way on the other side of the world. It would not have worked without Soonjae.

I have to give him a lot of credit here because we paid a lot of attention to the specific wording of button text in the app and locals, you know, like-

Wilkes18:51

Yeah

Andrew Hsu18:51

... like localized strings. We had a lot of reports from users pretty early on that they were shocked that it was an American company. Like, they thought, dude... Right? 'Cause you can always tell. There's always some weird wording or whatever, but there wasn't in Speak and I think that probably had, like, a large sort of non-tangible effect.

Wilkes19:08

Yeah. Focus, attention to detail.

Andrew Hsu19:10

Yeah.

Wilkes19:10

Tech stack. You... This was 2018. What were you rolling? You, you just did ASR and there was no LLM, so BERT maybe? I don't know.

Early Tech19:10

Andrew Hsu19:19

We actually had really no LLM component of it.

Wilkes19:22

Yeah.

Andrew Hsu19:22

So all the content... Oh yeah, another thing we did that I forgot to mention was we decided we needed to fully own all the content. So the way that we teach all in-house, all sort of thought from first principles, we build this thing called the Speak Method, which is basically like a pedagogical philosophy around teaching sentence patterns that you drill and then sort of combine into higher order patterns.

And all of that was in-house with, you know, our content team and our teachers. Um, and we built a lot of internal tooling to make this possible. There's just a lot of operational overhead, I would say. This is something we've struggled with to scale to many more languages, and that's, like, a big research effort within the company right now.

We're building a mobile product, right? My co-founder and I have always just loved apps and been big iPhone users, so we cared a lot about the app being native, feeling great, being high performance. The DNA of the company was always consumer.

Frankly, my co-founder and I had never worked in a real company. I dropped out of grad school, had a few failed startups, and then eventually started Speak, and he had never worked in a real company either. He, uh, just startups in the past.

So we didn't know anything about enterprise, enterprise workflows or, like, what sort of software real companies used. Um, so I think frankly, consumer was the only path. Uh, I don't think we could have done anything else. We just didn't know enough.

Um, and I think that has served us well, though, in terms of just really caring about the craft of it and wanting to build something that felt not 90 to 95%, but 95 to 100% in terms of polish.

Wilkes21:08

Was it hard to build an engineering team that did that at the time? Because ML engineering is-

Andrew Hsu21:13

That's another... Yeah

Wilkes21:13

... is very academia driven-

Andrew Hsu21:15

Yeah

Wilkes21:15

... back, back then, and then you have, like, the more consumer stuff that it's maybe more nascent.

Andrew Hsu21:19

And it's mobile. I'm now realizing that our story is very weird. So in addition-

Wilkes21:24

You only just realized?

Andrew Hsu21:26

Yeah. In addition to the market on the other side of the world, our first iOS engineer that we hired through a YC referral was in Slovenia. If you don't know where Slovenia is, look it up on Google Maps.

But it's, you know, it's like a pretty obscure little country, right?

Wilkes21:41

Middle of nowhere.

Andrew Hsu21:41

Yeah.

Wilkes21:42

Yeah.

Andrew Hsu21:42

And then we needed to hire a backend engineer, and one of his best friends was a great backend engineer and we hired him, and then this happened four more times, all in the same city. And then we were like, "Okay, we should probably just open a physical office."

So for-

Wilkes21:55

In Slovenia?

Andrew Hsu21:56

Yes. So for several years, we had an engineering office in Slovenia.

Wilkes22:01

What?

Andrew Hsu22:01

And then a few people here in San Francisco, and we still do. Now we have 90% of our core product development team in San Francisco here, office in Fidi. We're really only hiring here. But for the first, like, several years that, you know, that was, like, another very interesting sort of cultural aspect of the company, I guess.

Wilkes22:21

I think a lot of early stage founders have to do that. It's, that's the only people they can afford or whatever.

Andrew Hsu22:27

Yeah.

Wilkes22:27

Um, what are your, your tips that make that remote stage work?

Andrew Hsu22:32

For us, it wasn't really a price thing.

Wilkes22:35

Yeah.

Andrew Hsu22:35

I think legitimately thought he was the best person that we interviewed, and then it just kind of happened-

Wilkes22:41

Yeah

Andrew Hsu22:42

... that way when you roll it out.

Wilkes22:43

Yeah, it's not a price, it's more about remote work, right? Like distributed team-

Andrew Hsu22:46

Yeah

Wilkes22:47

... early stage. Like, a lot of people say, like, "No, you have to m- you have to move everyone to SF or your startup will die."

Andrew Hsu22:51

Yeah. I don't think that we were good at remote work. I don't think that my personality or my co-founder's personality is inherently very good at async, just to be perfectly frank. I actually think that almost, like, in spite of it- ...

we made it work. It was a little bit brute force. Like, I would just sync with them every single day, right?

Wilkes23:13

Yeah.

Andrew Hsu23:13

And there was pain because the time zone overlaps. Like, it was, like, exactly the most inconvenient.

Wilkes23:19

Yep.

Andrew Hsu23:20

But I think for several years we did that. We got really good at the cadence of it. I think they were excellent engineers as well, so it worked out. But if I had to do it over again, I probably wouldn't do it.

It's hard to say. Yeah.

Wilkes23:32

Shall we move to phase two on the, the LLM side? Um, that's when-

LLM23:33

Andrew Hsu23:36

Yeah

Wilkes23:36

... OpenAI started opening up, and when did they invest?

Andrew Hsu23:39

This was 2022.

Wilkes23:40

Okay.

Andrew Hsu23:41

So that was also when Whisper dropped. And-

Wilkes23:44

Yeah

Andrew Hsu23:44

... Whisper was a really exciting moment for us. It was actually since we started the company and made that prediction of, okay, in five or 10 years, speech models, language models will become superhuman level. Whisper was really that magic moment for us where we were like, "Oh, shit, I think what we predicted is here."

And I pretty distinctly remember this moment in the office when we got access to the model, and we were testing it on an audio clip of, like, a very beginner English learner in Korea saying something. And it was-- If you close your eyes as a human, you'd have no idea what they were saying.

There were four of us in the room, we all closed our eyes, and none of us had any idea, and the model got it right. So I mean, superhuman. I think that was the moment that we had been waiting on, and at the same time, LLMs were on the ascendancy.

ChatGPT would come out, I think on Thanksgiving of twenty twenty-two, and three point five Turbo came out. And I think like we kind of realized very quickly that all the pieces were clicking now, right? Like we have what we need at our fingertips now to go from something that was listen and repeat, where the user would see something on screen, hear a reference of the teacher saying the thing, and then they would just repeat the thing, right?

It was like very simple. Still a great product, by the way. You know, it still grew to like several million ARR in South Korea. Um, so clearly there was like a big market need for that.

Alessio25:12

Pre-Whisper?

Andrew Hsu25:13

Yes.

Alessio25:14

Wow.

Andrew Hsu25:14

This is from like twenty nineteen through twenty twenty-two. And then-

Alessio25:17

Yeah, that's, that's the grind.

Andrew Hsu25:18

Yeah.

Alessio25:19

You needed, you needed to hang in there.

Andrew Hsu25:20

Yeah. And again, I think that if we... There were many moments when things weren't working from twenty seventeen through twenty nineteen. We were looking in the mirror and we were like, "Why are we doing this? This is, this is crazy."

But I think we were so convinced about the vision. We just like couldn't believe that the vision would not-

Alessio25:38

Yep

Andrew Hsu25:38

... come true. So we stuck with it. So fast-forward to twenty twenty-two, the pieces started coming together. We realized that we could start building something that felt more like a language tutor, that could give you feedback, that could start explaining to you why you did something wrong.

And that was act two of Speak, a true English tutor.

Alessio26:00

This is something that a lot of founders struggle with today. It's like I'm kinda building something hoping that the models get better later.

Andrew Hsu26:07

Yeah.

Alessio26:08

How did you feel once the models got better? Did you feel like, "Okay, I am ahead of the curve because I built all this history of building product and like doing all this work?" Or did you almost feel like, "Okay, we spent all this money and time building these models, and now we're just gonna use Whisper"?

Andrew Hsu26:22

It was purely positive for us. We still kept using our custom ASR system because it was streaming real time really fast, really well fine-tuned. Whisper wasn't streaming, so it was a different use case. We used it for the more spontaneous stuff, and I think in almost every way, we were just really excited because pretty directly as the frontier of model intelligence improved, it would just unlock things on our roadmap that were locked before, if that makes sense.

And we still really operate in that mode today, where we take a model and then we try to think about, okay, how do we saturate model capability by building product on top of it? And then it happens again, right?

And then we build and saturate the model capability again. I think that's a really cool paradigm to like, you know, think about. But all the LM stuff basically allowed us to build a tutor for English, and we still didn't have like real-time voice, for example, right?

But the barriers are coming down now. Obviously, it's a really hot topic. We're actively building out a real-time voice platform that we can build a lot of more verticalized specific lesson experiences on top of that I'm super, super excited about.

I don't think they're going to replace our current lessons. They're going to be more immersive, just a different thing probably for more advanced learners.

Alessio27:42

Still language learning, though, not broadening out from language.

Andrew Hsu27:44

Yeah. So I think that language learning is interesting because it is so universal. Ninety-nine percent of people you know have certainly tried to learn a language, and it's so hard, right? Becoming fluent just has a huge failure rate, and it's something people are willing to pay for.

So I think that has been just like a pretty amazing beachhead for us, and I think we'll be doing language learning for a long time. There's a huge, huge, huge company to be built here. But our even longer-term ambition is really this idea that even beyond language, we think AI will reinvent how people learn anything, right?

It already has for me, right? I use ChatGPT to learn things every ten minutes, and I think I'm just naturally like a very curious person, so whenever I'm thinking about something, I want to know more about it, and then I'll naturally go to ChatGPT, and then I'll learn about it.

It's unlocked this like entirely new dimension of learning, and I'm spending way more time learning as well as an adult, which is really cool, and I wanna bring that in a more sort of structured, systematic way to everyone.

So I think that that's like the vision beyond language.

Alessio28:51

I'm curious to sort of double-click onto just the tech side.

Andrew Hsu28:55

Mm-hmm.

Alessio28:55

We talked a little bit about the c- the content that you, that you own and develop in-house.

Andrew Hsu28:59

Yeah.

Alessio28:59

And we talked a little bit about the onboarding memory. I assume that you have conversational memory as you, as you go, right? As y- And any other major pieces of the puzzle that really unlocked it for you?

Andrew Hsu29:09

So there's a few things I can talk about. I think one thing is, in order to go from teaching English to teaching a bunch more languages, we needed to really figure out more direct AI content generation. That was a pretty-- Right?

Because it's hard to scale, like our little studio in LA where we shoot a lot of the video lessons. Um, all of the scripts were written manually before by our content team. But we want like a hundred x more content, right?

And ten x more languages. Eventually a hundred x more language pairs, which is how we think about it. It's like, what's your native language, and then what language are you learning? And really, the only way to do that is to make it more AI-generated and, you know, very much like a AI-native company.

We wanna be on a frontier here. We want to keep a small team and to have as much leverage as possible through these types of tools. So that's a big active area where we're building out, I think, using...

You know, people overuse the word agent, but we have a tutor agent, we have a curriculum writing agent, we have a giant LM-based pipeline that creates curriculum, scaffolds it in the right way, writes the lessons themselves. That's a big active area that Will basically help us to scale to a lot more markets and a lot more languages.

So that's, like, one big thing. Another big thing is we care a lot about fluency, obviously. Specifically, we want to be able to quantify how fluent you are. So if you're learning Spanish, it's like, okay, what, what does it mean to be fluent, right?

Fluency30:46

Andrew Hsu30:46

And-

Wilkes30:47

There's no real-world test for that.

Andrew Hsu30:48

We care about real-world fluency, your ability to go to Mexico City and go to a street taco stand and actually order, right? That's very functional fluency in one aspect. You might be really good at that, but be completely unable to, like, talk about your family, right?

So the frontier of fluency is very jagged, but we're very pragmatic, and we care a lot about meeting user goals and helping them become fluent at what they care about. And we're thinking a lot about, okay, how do you quantify that?

How do you actually store a knowledge graph of everything you know about Spanish in terms of the vocabulary you know or you don't know, the s- you know, the patterns that you know or you don't know, the mistakes you made using Speak over the last month that are clustered.

Wilkes31:35

You said the magic word of knowledge graphs. Uh, is that-

Andrew Hsu31:38

Yeah

Wilkes31:38

... live? Is, is that experimental?

Andrew Hsu31:40

There are aspects of it that are live, and it's, it's a very sort of multidimensional system where we think of it as there are many aspects of fluency, right? There's many subscores. And we have a few of them that are currently live, and we're actively developing other aspects of it, and then all those will fold up into a more holistic fluency score.

The idea is that eventually, once we have a complete enough picture, everything will fold up into a number that we call the Speak score, that is a very sort of holistic measure of just, like, how good are you at Spanish, right?

And obviously, fifty-four is kinda meaningless by itself, but it does give you a general sense, right? Like, being at fifty-four versus being at five is very different, right?

Wilkes32:28

Yeah.

Andrew Hsu32:28

And I think everyone can kind of, like, intuitively understand that.

Wilkes32:31

And surprising- like, I, I would've grounded it more in real world, like, we will get you to pass this exam that is a standard, that is, like, the ESL standard or whatever.

Andrew Hsu32:40

So the way that we think about that is-

Wilkes32:42

Yeah

Andrew Hsu32:42

... we don't really teach for the test. I think it's possible in the future that we'll do a test-prep product. But in general, we care about real-world proficiency in various functional situations. So the way that we think about it is, if you're at this level, then these are the things you can do, right?

So it is exactly that.

Alessio33:00

We have that a lot in Italy. I grew up in Italy, so-

Andrew Hsu33:02

Uh-huh

Alessio33:02

... English is my second language, and there's a lot of people that pass a lot of tests and, like, get high grades in all the classes, and then they travel to the US and the UK, and it's, like, hard to speak because-

Andrew Hsu33:12

Yeah, totally

Alessio33:12

... they don't. Uh, you know-

Andrew Hsu33:13

Yeah

Alessio33:13

... I, I feel like the, the hard part is, like, being in the conversation, you know? It's, uh ... I, I think, like, when I started, my written and reading was, like, much higher than my conversation-

Andrew Hsu33:22

Yeah

Alessio33:22

... which, like, doesn't really help you if you're, like, traveling somewhere.

Andrew Hsu33:25

That, that's me for Chinese because my, my parents spoke Mandarin to me growing up, so I can understand, like, a nontrivial amount, but I, I'm very bad at speaking.

Wilkes33:33

Mm.

Andrew Hsu33:33

Yeah.

Alessio33:34

Oh.

Wilkes33:34

I heard there's a good language learning product for this.

Alessio33:38

I have one question on the course generation.

Andrew Hsu33:40

Yeah.

Alessio33:41

How do you eval that product? Like, uh, when you're asking the AI to generate courses, how do you figure out the courses are gonna be good?

Andrew Hsu33:48

Rely very heavily on our content team, and we are trying to build out an eval suite. It's really hard.

Alessio33:55

Right.

Andrew Hsu33:56

There's ... The illustrative example here is that as we try to hire and train new content writers on our content team, it's so nuanced. There's many different aspects of training them in the Speak method and how to write the right types of lessons and articulating why this form of lesson, which is subtly different from this other form of lesson, is better, right?

So we try as hard as we can to articulate that. So I think, like, forming, like, a sense of evals, using model graded evals like that, that's one piece of it. And I also think, like, in the future, a really good curriculum or lesson writer agent will probably be, like, reinforcement fine-tuned on a lot of our internal data as well.

Alessio34:44

Yeah.

Andrew Hsu34:44

That's something we're experimenting with, but it's still pretty early.

Alessio34:46

This seems like a great example of, like, you know, AI removing jobs, which is like, oh, you're creating the courses with AI, you don't have a person. But it's actually like instead of one person creating two courses, like, reviewing fifty courses-

Andrew Hsu34:58

Yeah

Alessio34:58

... that the AI generates.

Andrew Hsu34:59

Yeah.

Alessio34:59

That's kinda how you're seeing the content team.

Andrew Hsu35:00

The way that we see it, really not just for cr- for our content team members, but also I think it's perfectly applicable to engineering, is that it's, it's leverage. It just allows you to do a hundred x in the same amount of time.

We still need human review of the syllabus, the curriculum, the specific lines, et cetera. But the hope is that this will allow us to launch a hundred x more courses.

Wilkes35:23

A lot of, um, language is colloquial. I think the, the way that you put it on one of our episodes one time was the Italian that is taught in school is not the Italian Italians speak.

Andrew Hsu35:33

Yeah.

Wilkes35:34

How much of that do you adjust for informal versus formal?

Andrew Hsu35:37

Entirely. That's one of our fundamental tenets, which is that we don't teach textbook English or textbook language.

Wilkes35:45

Oh.

Andrew Hsu35:45

Like, we try very hard to-

Wilkes35:48

Teach Gen Z slang.

Andrew Hsu35:50

We don't go quite that far, but-

Wilkes35:51

Like slaking.

Andrew Hsu35:52

But we, we try to teach very casual conversational language that is actually what real people use. And like you said, that's usually very, very different. Like, if you pick up, like, a typical English textbook in Korea, it's all really traditional and weird formulations, and it's not how people actually speak.

Wilkes36:12

Yeah.

Alessio36:13

I know you're gonna release Italian soon, so I can give you a hand on that. I know in the US there's not that many dialects. There's, like, accents, but, like, most of the language is like ver-- like, the words that people use are similar.

Because I know, for example, Spanish is like, you know, Spanish spoken in Argentina is, like, very different than Spanish spoken in Mexico. How do you kinda adjust for that? Or maybe you don't, but-

Andrew Hsu36:34

So I would say that- For example, currently we teach American English, standard American English. We don't really teach other accents or other dialects. For now, given how small we are, we just have to be pragmatic and teach in the direction that most people want and most of our users know.

So we've made those decisions, like, on the content team side for American, Spanish, every language that we're teaching. But I do expect that in the future we're gonna get a lot more sharply differentiated. Like, if you wanna learn British English, then we'll teach you British English.

We'll teach you how to pronounce it, et cetera. I think all of that feels like something that Superhuman language t- you know, tutors should be able to do.

Wilkes37:12

I just think it'd be very funny if all the Koreans had, like, a very distinct Southern accent.

Andrew Hsu37:17

Yeah.

Wilkes37:18

Uh, it'd be, it'd be great. Make that happen.

Andrew Hsu37:20

Yeah.

Wilkes37:20

I do think about this because, uh, me, you know, obviously there's a moving of the goalposts. Like, now that we have this, now we want the next thing.

Andrew Hsu37:28

Yeah.

Wilkes37:28

And obviously people who are English as a second language always have an accent. Like, I haven't ... Like, a lot of people think I don't have an accent, but if you know any Singaporeans, you know I'm, I'm Singaporean.

How much accent training is important, right? Like, I think, like, actually that does help a lot with f- for people.

Andrew Hsu37:45

Yeah.

Wilkes37:45

And you cannot tokenize accents yet.

Andrew Hsu37:47

Yes, that's right. So I have two main thoughts on this. I think the first one is that communication and your ability to speak spontaneously and get a s- a concept across, an idea across, is almost fully orthogonal to pronunciation.

You can be really bad at pronunciation but still communicate effectively. So a lot of the current core product experience is about just speak as much as possible, make mistakes, don't worry about screwing something up on the accent or the pronunciation side.

Wilkes38:21

Yeah.

Andrew Hsu38:21

The important thing is that you literally move your mouth and you make the sounds, right? And it turns out there's, like, a really key psychological barrier there where people are just not willing to do this in front of a human, even if it's a human that is a teacher that you're paying, right?

So a lot of the core message of our marketing campaigns in many of our, like, biggest markets is along the lines of, like, you can make mistakes in this private space with Speak, and I think psychologically that's extremely powerful.

And then you can go and get it right more confidently in the real world after you practice with Speak. Now, having said that, people do care about their pronunciation and their accent, right? So we w- we have for English only right now a pronunciation coach that is basically like a fine-tuned version of wave2vec, which is a meta model, but we basically fine-tune it on a bunch of our own phonetic transcripts, like fine-tune data.

It works pretty well. It's currently for single words. We're going to expand it to full sentences, to more languages, et cetera. But I think that just if you look at, like, the pure market opportunity, our sense is that we really want to push people to just speak very freely as much as possible, you know, just get that volume up.

Wilkes39:39

Yeah. Yeah. In terms of immersing language learning in the real world, one of the more interesting approaches that people keep trying is to have, let's say like a Chrome extension or something on top of a page. I think Toucan was doing this.

Future39:52

Andrew Hsu39:53

There's a bunch of those, yeah.

Wilkes39:54

Yeah.

Andrew Hsu39:54

Yeah.

Wilkes39:54

And then there was another one I saw recently which was like watch a YouTube video and it'll transcribe for you but randomly mask out words.

Andrew Hsu40:02

I saw that too, yeah, yeah. Yeah.

Wilkes40:03

Uh, that was like a h- Show Hacker News.

Andrew Hsu40:05

Yeah.

Wilkes40:06

Do those work?

Andrew Hsu40:07

There's kind of the question of is, you know, is that the right product, right?

Wilkes40:10

Yeah.

Andrew Hsu40:10

I don't think-

Wilkes40:10

So, so basically the difference is your content or real world content, right? Obviously you want real world content.

Andrew Hsu40:15

I think that for work, right? So for Speak for Business, for the, for the B2B product, another part of the vision is really, like, what should a Superhuman language tutor be able to do? It should probably be able to handle kids as well as a Samsung employee that wants to transfer to the US office and wants to use Speak for work, right?

So our view there is that it's the same product. It's a different distribution mechanism, right? Consumer versus B2B. And I think that we will eventually build something like a Mac app. Maybe it'll be integrated with a browser in some way.

We're not really sure yet. But obviously in order to apply it to your day-to-day, there needs to be some way to hook into your actual sort of work documents, whatever. That's a whole can of worms. Um, we are actively thinking about it, but I think my sense is that it's not clear to me that any of these products have really taken off, and I think that there's many other approaches that are possible.

I don't have the answer, but like another example, very hypothetical future world is maybe OpenAI, you know, the, the new Jony Ive thing will come out with some hardware that-

Wilkes41:31

Ooh

Andrew Hsu41:31

... will be listening to you all day and then we can, you know, give you some sort of like very deep analysis that is integrated with the Speak app at the end of the day or like-

Wilkes41:39

Yeah

Andrew Hsu41:39

... you know, the end of the week, whatever. I don't know.

Wilkes41:40

Okay, one more time since you brought that up. I'm sure you don't, actually they haven't told you anything, but what-

Andrew Hsu41:44

I don't know anything

Wilkes41:45

... what's it gonna be?

Andrew Hsu41:46

I don't know anything.

Wilkes41:46

It's like a very, like this is the number one topic in all the parties I go to now.

Andrew Hsu41:49

Really?

Wilkes41:50

Yeah.

Andrew Hsu41:50

What's the most compelling idea you've heard?

Wilkes41:53

Okay, so there's, there's people that say Jony hates wearables.

Andrew Hsu41:55

Yeah. I've heard that too.

Wilkes41:56

And I'm like, if it's not a wearable, then you've just made a second phone. I- in that case, just make a phone.

Andrew Hsu42:03

Yeah.

Alessio42:03

I thought they said it was g- I mean, didn't Sam say that it was g- he wanted to do a phone in the past?

Andrew Hsu42:09

That was in the far, that was like in the far past.

Wilkes42:11

He says a lot of things.

Andrew Hsu42:12

He does say a lot of things.

Alessio42:13

Yes.

Wilkes42:14

Okay. Anyway, I think wearable makes sense. I think the, the, the race is to capture context.

Andrew Hsu42:19

I mean, I, I have a wearable on.

Wilkes42:20

Yeah, I, I've, we have a wearable here too.

Andrew Hsu42:22

Yeah.

Wilkes42:22

Transcribe everything.

Andrew Hsu42:23

That transcribes everything?

Wilkes42:24

Yeah.

Andrew Hsu42:25

That's cool.

Wilkes42:25

Yeah, it's a previous episode of ours with, with the-

Andrew Hsu42:27

Oh, cool

Wilkes42:27

... um, I can hook you up if you want.

Andrew Hsu42:29

Yeah.

Wilkes42:30

But yeah, I think, like, it's something that a lot of people are interested obviously 'cause it's a huge bet by them and, uh, yeah. Curious. Okay, you mentioned video. I just wanted to double-click on that a little bit.

I'm sure engagement very high for video because people love to watch video. I thought that Speak would be one of those places where, like, you just kinda leave it in your pocket, you walk, you take, you take a walk, learn to speak.

Probably that's not true?

Andrew Hsu42:53

What we've done so far is part of the course experience is a teacher video. We've tested other, more audio-forward types as well. We found that, of course, like you said, video is very engaging, but at the same time, we have a lot of users that do want to be able to walk around with the phone locked in their pocket.

So doing something that is more like voice mode with optional v- uh, you know-

Wilkes43:20

Yeah

Andrew Hsu43:20

... visuals, I think is really good. I think there's huge opportunity for a better way to learn things like listening comprehension. So I took German in grad school for two years, and I thought I was getting somewhere, but any time I listen to a native German speaker, it's so fast, it's, it's completely on a different level.

Wilkes43:43

Mm.

Andrew Hsu43:43

And I think you can imagine a plethora of really cool experiences that feel kind of like you're listening to a podcast, but it's all AI-generated, it's fully controllable, it's integrated with the app. You know, there's something there for sure.

Yeah.

Wilkes43:57

Don't wanna do AI podcast, man. We're cooked.

Andrew Hsu44:02

It's okay. We'll, we'll document, uh, the, the own ending. I, I, I mean, I think when that happens, we just end the show. Like, why not? Like... To zoom out a little bit, in the pretty near future, multimodal models will cross the threshold where they will be able to generate images a lot faster than they currently are, maybe somewhat close to real time even, right?

And audio at the same time, text at the same time, and you can imagine just like a very powerful multimodal tutor that can kind of do it all at once where there's an audio track, and then if the teacher's teaching you something, with the right timing, it chooses, "Okay, at this point, I'm about to introduce a new concept, so I'm going to show the word on screen so that the user can see how it's spelled," right?

There's a lot there. You can do generative UI. A lot of nuance there where it's easy to do it badly, but to do it well requires a fair amount of reasoning and mental modeling of what the user knows.

Wilkes44:58

Yeah.

Andrew Hsu44:59

Which feeds into what you need to show at what time. So that's probably gonna have to be like a pretty parallel set of systems.

Wilkes45:06

Have you spent any time looking at this like, uh, you know, like V03, where you do video plus audio at the same time on how you can tweak the audio part versus the video part? Because I can imagine you might work on a video part, and then you wanna change the audio generation model.

I don't actually know how the model works inside on like how much you can tweak just the language.

Andrew Hsu45:25

We haven't really looked at the video stuff much. We basically think that we're very bandwidth-constrained, right? So we're just scaling and trying to hire as fast as possible like everyone else is. And as a result, we're really focusing on just like the most in-reach, highest impact things.

I do think that the barriers are coming down very fast for all of this sort of stuff. I'm just so excited about multimodality and where things are going here because imagine if you're learning Spanish, being able to look at an image that the model generates for you and then doing Q&A on it, right?

Like a beach scene, and then the model will ask you like, "Oh, how many people are running on the beach?" And then you have to sort of respond in the target language that you're learning. Very traditional language learning exercise, but you can imagine it being fully generative, which is really cool.

Wilkes46:15

Awesome.

Andrew Hsu46:15

Lots of stuff like that.

Wilkes46:16

The engineer in me worries about inference costs, but I think you can just kinda sweep that under the rug for now.

Andrew Hsu46:22

Yeah.

Wilkes46:22

See if it works first.

Andrew Hsu46:23

Yeah.

Wilkes46:23

And then, and then you can worry about costs.

Andrew Hsu46:24

Yes.

Wilkes46:25

You mentioned a real-time voice platform. Uh, I just wanna give you the platform to, heh, platform, to talk more about that. Just like, uh, you mentioned, for example, that you're a very heavy user of the real-time API from OpenAI, and you build a bunch of tooling around it.

Andrew Hsu46:40

Yeah. So we, last year, had early access to the real-time API, and there's a very obvious sort of use case for language learning. I think one common theme that has just been pretty awesome since LLMs came out is that language learning as an application is just a really good fit for LLMs, all these model types, in almost every way, which has been just really great for Speak specifically.

For real time, I think the audio piece promises to really like infuse almost every surface in the app. You can imagine this is the primary way that you talk to your tutor, right? And an additional complication is that it needs to be multilingual, and there needs to be code switching.

So that's a pretty frontier problem right now, right? So like, I should be able, if I'm learning Spanish, to speak both English and Spanish and vice versa from the model. That's a pretty hard TTS problem today. It's actually like only a few models are able to speak two languages in the same sentence-

Wilkes47:38

You can have a router model-

Andrew Hsu47:38

And then, and pronounce them properly. Sorry?

Wilkes47:40

You can have a router model, like a tiny little router model guess which language first and then route.

Andrew Hsu47:45

Well, the problem is that there's, you can have a sub-word in- ... in a single sentence-

Wilkes47:52

Yes

Andrew Hsu47:52

... in a different language.

Wilkes47:53

Yeah.

Andrew Hsu47:54

So you can't just concatenate either-

Wilkes47:56

Yeah

Andrew Hsu47:56

... because it won't sound right.

Wilkes47:58

Yeah.

Andrew Hsu47:58

Right? It won't sound natural. That's not how humans do it. So this needs to be like a, like a very native, controllable audio, you know, function.

Wilkes48:05

Yeah.

Andrew Hsu48:05

But we are in the process of building a variety of experiences on top of the real-time API. I wanna clarify that actually nothing is in production yet, mostly for price reasons, frankly. The pricing model of the real-time API makes more sense for something like a customer support agent, where you're very directly replacing somebody that you would pay hourly otherwise, and that's how you're seeing the price model for a lot of these initial agents work out.

For us, we want our users to be able to do these real-time role plays and have these conversations for many hours a day, right, if they want. Getting cost under control is definitely a pretty key consideration right now, but we are pretty close, maybe actually a- Even by the time that this episode is released, we'll have something live.

But we have a, what I think is a really cool application of the real-time API, which is basically a new instructional lesson where it's the model actually teaching you something, like a new language concept, and it's intended to sort of augment slash play the same role as our current video lessons, which are the instructional lesson type.

And it's interactive, obviously. At certain points in the three to five minute lesson, you're interacting with the real-time API. It's semi on guardrails. There was a lot of scaffolding we needed to build to basically, number one, switch between the interactive and non-interactive portions of this lesson properly, if that makes sense, right?

There, there's, there's some portions where you're just listening or looking, and then some portions where you're actively in a short conversation, and we kind of swap back and forth. And we have, like, a bunch of sort of custom architecture and info around that.

And then there's also making the cost make sense or at least, like, semi make sense. And then there's a bunch of WebRTC infrastructure. We're at sort of, you know, not huge, but non-trivial scale either. So we definitely just-- it'll cost us millions of dollars if we do something wrong.

Wilkes50:04

Yeah.

Andrew Hsu50:05

Yeah.

Wilkes50:05

Do you do inference in Korea because of the, you know, latency and all that?

Andrew Hsu50:09

It's something that we have been increasingly paying attention to for all the real-time paths. Like, I would say two or three years ago, when real-time stuff was still quite nascent, users didn't really care as much. But I think now the standards have risen, right?

Like, latency has to be low. Everyone cares.

Wilkes50:29

Do you have a hard latency budget for responses, or do you just kinda work it out? So for example, right, like, you have a knowledge graph that you're accessing, you have content that you're retrieving. There's a lot of stuff there.

And then, like, I, you know, maybe you're using a reasoning model, probably not, but, like, that all eats into the budget.

Andrew Hsu50:46

I will say that from the, like, real-time engineering side, everyone talks about, okay, submit user request to get agent audio response, like, first bytes, right, first audio bytes. What's that latency? And then we try to get that as low as possible.

I would argue that's actually, like, a vanity metric because what you don't take into account is how the VAD works. How do you do turn detection to detect when the user is finished speaking, right? Because that can easily add, like, another second if you do it badly, and nobody talks about that for some reason, right?

Like, what you need to measure is actually when does the user stop talking to when does the model first audio come, and usually that number is much larger. That is a very domain-specific problem. You can use, like, the semantic VAD on real-time API for regular English conversation, and that will basically classify at every token how likely it is that you're done speaking as a sort of normal conversational English speaker, like in this conversation.

That's fine, but it doesn't work at all for language learners, right? If I am trying to respond in a language that I'm learning- ... I'm gonna be hesitating halfway through- ... for ten seconds or more. Right? So it needs to be fully custom probably.

This is something that we're also actively working on, but that is actually, like, the dominating factor in perceived latency.

Wilkes52:11

Coding, do you use Cursor, Windsurf, um, other autonomous agents?

AI Tools52:11

Andrew Hsu52:16

It's kind of all of the above.

Wilkes52:17

Yeah.

Andrew Hsu52:18

So I think, like, as the CTO, I view it as part of my responsibility to really set expectations, push everyone on the team, show them what's possible. We've been trying everything.

Wilkes52:32

Yeah.

Andrew Hsu52:32

And I think we try to basically set the expectation that the frontier is moving so fast, it's deeply non-intuitive.

Wilkes52:40

Mm-hmm.

Andrew Hsu52:41

If you've tried coding tools six months ago, and they weren't that great, especially if it's not TypeScript or Python, right?

Wilkes52:49

It's like most apps are the most popular languages.

Andrew Hsu52:51

Yeah.

Wilkes52:51

Like that's all it is.

Andrew Hsu52:53

We try to set a culture in the engineering team where usage of these tools as much as possible and as a default path is the expectation. And in hiring, we are now explicitly asking about this a lot, thinking about what are the types of people that are gonna be better, higher agency at trying these types of tools.

It's so important.

Wilkes53:16

Before we zoom out, anything we missed about Speak that you really wanna highlight or something that people underrate about it?

Andrew Hsu53:23

One thing that I've always been really excited about is that I feel like a lot of the foundational pieces that we're building around knowledge graph, for example, a lot of these concepts should be applicable to not just learning language, but also other things in the future.

We're already starting to see the very beginnings of this on the B2B side, where a lot of it is more like management skills and hospitality skills, communication skills, more like true L&D for enterprise, less like core pure English proficiency.

So I think you were, you know, that's like obviously immediate neighborhood. But you can imagine many academic subjects, math, biology, et cetera, you know, schools work for. Um, super excited about that.

Wilkes54:08

If I knew my employer was giving me a language tool, but then he was evaluating me on my management skills while learning the language- ... I might use it less just, you know.

Andrew Hsu54:18

Fair. Very fair.

Wilkes54:18

You know, you wanna, you wanna, you wanna separate that out.

Andrew Hsu54:21

Yeah, very fair.

Wilkes54:22

I agree overall that the knowledge graph problem is very important. We have a whole track on it for, for the conference, and I think that the amount of data can be so high. And actually, like-

Andrew Hsu54:33

Yeah

Wilkes54:33

... you want to, you wanna generate relevant triplets. I, I assume you use the normal subject, predicate, object type.

Andrew Hsu54:38

It's a bit more custom than that because it's a bit more domain specific around the way that we conceptualize the vocabulary you know-

Wilkes54:46

Yeah

Andrew Hsu54:46

... and the sentence patterns and so on.

Wilkes54:47

Yeah.

Andrew Hsu54:47

So it's, it's more specifically around, like, language learning concepts, if you will.

Wilkes54:51

Yeah.

Andrew Hsu54:51

Yeah.

Wilkes54:51

But what I think we can extract from Speak, or as it is generalized as a framework, is, um, what I've been calling sort of like the Bloom's 2 Sigma Problem type thing, like the level adjusting tutor. Like, where are you at, let me adjust my thing to where you're at, and then I'll push you up to the next level.

And I think the knowledge graph is, is a part of it, but I don't know if th- that's all of it. I've never seen a working example.

Andrew Hsu55:13

We are approaching that problem from a few different angles. I think part of it is knowledge graph. Part of it is being very careful in how we structure the curriculum so that you're placed at the right level, so that the learning path itself, which has a foundational backbone because beginner to intermediate English learners actually, like, all need to know a bunch of similar concepts.

It isn't really until you get intermediate and more advanced where that starts to, like, more sharply diverge. And A0 through B1, I would say, there's a pretty well-defined, like, sort of linear path actually. A lot of the deep thinking that we've done around how do we structure the pedagogy is also s- super useful in terms of just, like, matching people to the right level.

And then you can take this backbone and then basically modify it based on the knowledge graph, on your system's knowledge of what the user is, like, bad at versus good at.

Wilkes56:12

I think a lot of startups or ed- especially edtech, like, that is the core engine. Like, you know, once you do that, you can kind of teach anything.

Andrew Hsu56:20

Totally, yeah.

Alessio56:21

We have a few more broader fun questions.

Wrap Up56:21

Andrew Hsu56:24

Yeah.

Alessio56:24

Um, so speak.com.

Andrew Hsu56:26

Yeah.

Alessio56:26

Great domain. I looked it up. Voice.com got bought for $30 million in 2019.

Wilkes56:30

Oh, my God.

Andrew Hsu56:31

Really? When?

Alessio56:32

2019.

Wilkes56:33

Okay.

Alessio56:34

So-

Andrew Hsu56:34

Wow

Alessio56:35

... I don't know if you wanna share how much you paid for it, but-

Andrew Hsu56:37

It was a lot less. It was a lot less than that

Alessio56:38

... I, I figure it would be a lot less, but I'm curious if-

Wilkes56:40

Uh, my estimate was 100K, but-

Andrew Hsu56:42

It was, it was more than that.

Wilkes56:43

More than that. Wow.

Andrew Hsu56:44

Okay, I'm not gonna say any more about the numbers.

Alessio56:46

So what, what's the... Yeah, what was the sort? Was it easy?

Andrew Hsu56:48

Yeah.

Alessio56:49

Was it, did you use a broker? Like, we had, uh, you know, Dharmesh Shah from HubSpot, uh, who sold chat.com to OpenAI-

Andrew Hsu56:55

Yeah

Alessio56:55

... and he has a lot of very fancy domains.

Andrew Hsu56:56

That was like a $100 million deal or something, right?

Alessio56:58

That wa- that was very big.

Andrew Hsu56:59

Oh, wait, no, that was ai.com. Whatever. Um-

Alessio57:01

Uh, chat.com.

Andrew Hsu57:01

Oh, chat.com, okay.

Alessio57:02

Yeah.

Andrew Hsu57:03

We bought it several years ago. It felt very expensive for us at the time. It was a little bit of a crazy move, but I think we were very convinced that we needed a super strong consumer brand that was scalable globally, and that was just always our ambition.

Like, we want to be the way the next billion people learn languages, and we need speak.com. So we don't regret it.

Wilkes57:28

It's, it's such a, such a great word, makes for great swag. Very nice decision. Yeah, a couple other fun, fun questions. Any fun Korean celebrity stories 'cause you work with so many influencers?

Andrew Hsu57:39

We have a bunch baking right now.

Wilkes57:41

Okay.

Andrew Hsu57:41

But I think it was, you know, some- something more generally that has just been so fun on the journey. So we, we would visit Seoul every year.

Wilkes57:49

Yeah.

Andrew Hsu57:50

And seeing Speak go from nothing to the first time we saw somebody on the street using Speak. To now our main teacher in the app is, like, a mini celebrity. People come up to her on the street as she's just walking around Seoul and recognize her from the app, which is really cool.

Now we do a lot of advertising. We do billboards, TV commercials. We work with big influencers and so on. So just, like, seeing the scale of that has me kind of, like, in awe. It's, like, really cool, um, just to see something that used to be nothing.

Wilkes58:23

Yeah, I wanted you to name drop, like, Blackpink or I don't know.

Alessio58:27

I don't know.

Andrew Hsu58:27

Look, there, there's, there's some stuff baking right now.

Wilkes58:29

Yeah, okay. All right.

Alessio58:30

We talked about the Thiel Fellowship.

Andrew Hsu58:32

Yeah.

Alessio58:32

On your LinkedIn, you kinda have this hole between 2012 and 2016-

Andrew Hsu58:35

Yeah

Alessio58:35

... which you talked about. You did some startups. Any of them that you wanna share? Like, ideas that you worked on that you thought were cool.

Wilkes58:41

Maybe it was just early, but, you know-

Alessio58:43

Well, yeah, what people should revisit

Wilkes58:44

... try again.

Andrew Hsu58:45

I've always been interested in learning and education. One of the other failed startups that I did in that time, um, was... It feels silly to even talk about this because it amounted to nothing, but it was called Bloom.

You know, the Bloom's 2 Sigma Problem.

Wilkes59:02

Oh, yeah.

Andrew Hsu59:02

It was actually, like, named after that, and we were trying to, to basically build, like, a better adult learning platform and have really cool interactive JavaScript widgets for various concepts that you could learn. Didn't find PMF. I was young and didn't really know anything about business at the time either.

But I think that the common thread through actually, like, everything that I've been interested in since leaving grad school has been how do we build software, build tools that help people learn things more effectively and better and faster.

And now I feel very lucky to be in this position because obviously AI is the ultimate version of that, right? And it's been completely transformative for me personally because I just get a lot of just inherent fun and pleasure out of being able to, like, think of a concept and then, oh, now I can talk to this omniscient LLM that can tell me more about it, and I'm really good at asking the right follow-up questions that I want to know.

Um, so that, that's been completely transformative for me.

Wilkes1:00:03

Do you get a lot of, like, people using Speak for therapy? Like, you know, 'cause it's not meant to be that, but since you have inference, they will use it.

Andrew Hsu1:00:14

In 2023, when we first launched our AI role plays using GPT-4, back then, people were way more concerned about safety, right? And, you know, obviously the models now are much better at, like, refusals, and the line is sharper between what's appropriate and not.

But we did see a lot of our first users start to put in pretty questionable custom scenarios. Um- ... you probably guessed.

Wilkes1:00:39

Yeah.

Andrew Hsu1:00:39

And, you know, uh, like this was something we expected, but I think seeing the logs in person is, like, very different.

Wilkes1:00:46

Got it. Um-

Andrew Hsu1:00:47

Some shocking stuff in there.

Wilkes1:00:49

Last couple questions. One on Andre. You talked to him when, in your machine learning journey. Um, he's also-

Andrew Hsu1:00:54

That was a long time ago, yeah

Wilkes1:00:55

... he's also working on edtech now.

Andrew Hsu1:00:56

Yeah.

Wilkes1:00:57

Uh, I don't know if you've ever, ever had conversations with him on-

Andrew Hsu1:00:59

No, I haven't

Wilkes1:01:00

He's also interested in language learning, by the way, you know?

Andrew Hsu1:01:02

One thing that I think we didn't really realize early on or, like, fully internalize at least, was just, like, how deep the market is.

Wilkes1:01:10

Say more.

Andrew Hsu1:01:11

It was so universal where we really struggled to do some of the basic startup stuff around define your, like-

Wilkes1:01:18

ICP

Andrew Hsu1:01:19

... ideal customer profile-

Wilkes1:01:20

Yeah

Andrew Hsu1:01:20

... and, and like, you know, segment your users because our users were everyone. Like, we, we had parents using it with their kids, we had really old people using it, we had people using it for work. So that was kind of, like, mind-boggling.

Wilkes1:01:31

You still did customer segmentation or are you saying it doesn't matter?

Andrew Hsu1:01:34

I'm saying it was hard to do.

Wilkes1:01:35

Yeah.

Andrew Hsu1:01:35

Like, we tried, and we have a sweet spot in Korea. It's, like, 25 to 45, more professional, more white collar, but it's very... Where it's like-

Wilkes1:01:46

Yeah

Andrew Hsu1:01:46

... like, a very long tail on either side. Yeah, I think it, you know, it's a huge market, and I think it's a very special moment in time right now where it's obvious that a lot of the tech is here.

I think it's really good for humanity if we make a lot of progress here. So I'm really excited for his company, too.

Wilkes1:02:03

We started asking about the Thiel Fellowship.

Andrew Hsu1:02:06

Yeah.

Wilkes1:02:06

So maybe we can wrap with one of Thiel's favorite questions, which is, uh, what's something you believe in today that most people would not agree with you on?

Andrew Hsu1:02:13

I think that people, if you recall, expected the world to kind of explode when GPT-4 came out, and, you know, like, everything would change. And I think if you, like, go to another state outside of the Bay Area, probably even in California outside of the Bay Area, and then you ask somebody how much their life has materially changed, it's, like, pretty close to zero.

Real-world inertia is enormous. Obviously, AI is probably the most transformative technology we've ever built, but I think in a very real sense, the world hasn't changed that much either, and that's a really weird thing, right? So I think we need more builders, we need more people building applications.

It's weird to me that Speak is actually, like, not that many net new consumer AI-native applications at scale. Like, there should be way more. I would love for there to be way more. Consumer is hard.

Wilkes1:03:06

Yeah. I'm intimidated, but like you know.

Andrew Hsu1:03:09

It was just like there was never any alternative for us.

Wilkes1:03:11

Yeah.

Andrew Hsu1:03:11

Like I said before.

Wilkes1:03:12

You didn't have a choice.

Andrew Hsu1:03:13

Yeah.

Wilkes1:03:13

But also, you're very smart. But also, maybe you have some growth hack things that you can advise people on that, like, that people could, could learn. But yeah, I, I agree. I, I, think like, the general take actually is this is what we want, which is slow take-off, short timeline.

Andrew Hsu1:03:27

That's fair.

Wilkes1:03:27

Right? The, this is the two by two that everyone always talks about in AI safety. You're seeing slow takeoff, and like, maybe don't complain. We, like, we, we have a heads-up because... Or, you know, Dario's right, and like, half of us lose our jobs in the next two years.

Andrew Hsu1:03:40

Yeah. It's a, it's, it's so hard to predict. Yeah. Sometimes I get AI anxiety.

Wilkes1:03:47

Uh-huh.

Andrew Hsu1:03:48

And then I just-

Wilkes1:03:49

You get anxiety?

Andrew Hsu1:03:50

Yeah.

Wilkes1:03:50

Okay.

Andrew Hsu1:03:51

Then I just focus on our users.

Wilkes1:03:53

That's a perfect place to wrap. Thank you so much for taking the time.

Andrew Hsu1:03:55

Yeah.

Wilkes1:03:56

Thank you very much.

Andrew Hsu1:03:56

Thank you guys so much. This was great.