Quizard Origins0:00
All right, cool. Uh, we're here with Lightning Pod, with Sid Bendre of Oleve. Oleve, I screwed it up. We practiced this. But welcome. Yeah, uh, thanks for, thanks for joining us.
Yeah, thank you for having me. Uh, definitely a big fan of the podcast, so super exciting to be on it.
Awesome. I wanted to feature you because I think that you have an interesting story around doing a lot of shipping apps, uh, with a small team, but then also shipping multiple apps, right? Like, you're not just, like, doing one startup, you're kind of a startup studio.
Is that right? Maybe you could tell that story just on, like, you know, I, I think you went through Neo, um, and Z Fellows as well. Um, and, and, like, how, how, how did, how did you find your way towards this startup studio concept?
Yeah, totally. Formerly, we were actually known as Quizzer AI as a company. This is the namesake of the first product that we launched, the Quizzer AI mobile app. But we launched this at the end of January in 2023, kind of a few months after ChatGPT was sort of launched.
Um, and we launched this with a TikTok that went viral overnight. It was kind of like a POV, ChatGPT, and Photomath had a baby. Um, got a million views overnight, and that converted to about ten thousand users in under thirty hours, and then we just started scaling it from there.
A month later, we added a paywall in and basically, like, started revenue generating, tried different business models at the time. Kind of landed on weekly subscriptions initially as being a good Christian model, or just like subscriptions in general as being the best model.
Played around with tokens and stuff like that. We like-- What was interesting actually to note is, back then we had no LLM costs when we had launched, um, because we, like-- Back then, Codex was actually a model that was live.
If you remember, like, Codex had the beta launch, beta preview, and it was completely free. And we figured out if you prompt engineered it, you could actually get it to do open domain, like, language, like natural conversation about different subjects.
And so we were cycling through about ten different keys, um, using Codex to respond. And actually, when they were turning down or, like, shutting down or sunsetting Codex, we, we got an email saying we were one of the highest traffics on Codex.
Um, so it was kind of interesting to, to have that.
Like, was that a compliment or were you in trouble?
Uh, it was a, it was a compliment, but it was more like we had to now switch to three point five, which was actually gonna have a cost, and so that's where, like, we had extra, extra incentive around monetizing.
So that was, yeah, that was just an interesting piece. We then applied to, uh, the Neo Accelerator and got in. We were in their second cohort. Honestly, for us, it was super helpful and beneficial because, um, to be honest, like, all three founders, we went to University of Rochester.
Rochester isn't the most, like, I'd say, entrepreneurial community. There's a lot of interesting people there, but it doesn't-- you're not necessarily exposed to the right, like, knowledge, I'd say, to learn how to, like, really take a startup to s- and scale it.
Um, but that said, post accelerator, we moved to New York City, continued, you know, building up Quizzer. We then launched our second product, which was Unstuck AI, at the end of August in 2024, kind of taking all the playbooks for growth and product development that we'd, like, learned through Quizzer.
This was our fastest-growing product yet. We hit a million users in under nine weeks and then just started scaling it from there. And so now across the portfolio, we do all-- we do five million users. We, we hit six million, eight dollars in ARR, and we've been profitable since the first nine months of operating.
Unstuck Launch3:01
The profitability actually came pretty early, and so after Neo, we only raised like a smaller, like, angel round. Um, we raised from people like Cal Henderson, the co-founder of Slack, Amir Jang, the ex-CTO of Tinder, and Russell Kaplan, who's the, uh, current president of Cognition Labs.
Yep. I made a-
And nice suite. Yeah, I guess like, at that point, we sort of realized that like, um, we-- a lot of the play-- like in some sense, we were kind of, we were kind of going hard mode by going directly to consumer edtech, selling to a like, uh, generally broke, uh, you know, like demographic, um, you know, with more finicky spending and, and such.
And we realized that a lot of our play, playbooks for like growing and distributing and like building product and scaling product were very market agnostic. Um, in fact, you know, we started plating-- placing bets earlier that year about like, oh, which markets are ripe for popping off.
And then we'd suddenly see like people launch apps in that domain that were like quite simple, you know, technically and, and just like purely like running on the same playbooks we'd had. And we realized that, hey, we could probably do the same thing.
If we took even like maybe like fifty percent of the like amount of work we do in terms of like, you know, winning in this market, um, we could probably do really well in other markets, and that was the idea that Oleve was born of.
We're taking all the playbooks that we've learned over the last two years acro- and, and building and scaling both Quizzer and Unstuck successfully and applying it to different verticals, be it lifestyle, wellness, fitness. We have a third product that's kind of in stealth right now, and so can't necessarily talk about it, but it's our first one outside of edtech.
Um, we're super excited about it. Um, and, uh, yeah, that's kind of like why we felt that this model could work. It also meant a bit of a reorganization how the company approaches, I guess, like scaling, uh, you know, both as a company, but also our, our portfolio.
Platform Org5:01
Um, and so this is how we've kind of now structured the company under the, like, product engineering org and platform org, where the product engineers are people who individually will own like an app or two, like depending on like bandwidth that they can manage.
And the platform team, um, is actually s- um, as seen on the website, the software engineering team is quite detached from an individual product, and it's more focused on building like s-systems to scale across products. So you can think of it as like the universal architectures and components that we-- and libraries we use across our products, you know, the way we do our AI, the way we like monitor all our referral mechanisms, all, all our like notification mechanisms.
This-- The platform team also owns a lot of the like significantly technical, like engineering, like systems. The idea is that the product engineers need-- are like direct revenue lines into business, and they need to be f- solely focused on like effectively CEO-ing that project.
Um, and so they need to be focused on user experience. They're living in the metrics. They're focused on like revenue, profitability, reducing churn, you know, retention, all that sort of stuff. And this detached team can focus on like reusing like- ...
concepts across different products to make things more deterministically like successful, if that makes sense. The third piece of the platform is kind of interesting. I think you might like to hear about it, um, is effectively what I'm trying to do with the platform is start up an internal shadow org within the company that's all run by agents, effectively staffing like each business unit with agents.
You know, our growth and marketing especially, um, being the highest priority, but also like our product teams and how we discover new products and new spaces to double down in. Um, 'cause the idea eventually is to have Oleve be this like company where we hire people for, you know, their taste, um, their strategic thinking, their expertise, but a lot of the operating could be done, uh, primarily through agentic workflows or automations that we build in.
Um, and so that's kind of the like the model that we've sort of like landed on right now and we're scaling with. Um, and I think the reason why like growth and distribution matters the most is because our thesis on consumer software now is it's gonna be-- it's gonna reflect less or look less like-- or operating consumer technology company is g- is gonna look less like traditional tech companies and more like CPG companies that were like very much distribution first, always branding like at the forefront because people now have a more refined taste in how they purchase consumer software.
They've spent enough time paying for things online that they now like have a taste of when they wanna buy things or when they wanna continue subscribing to things. And so people have a more refined sense of like taste in general.
And so actually we, we are quite excited by the fact that it's a much more competitive market because it means that we can have like a lot more wins if we double down on our distribution tech or distribution like, I guess like playbooks, so to say.
Yeah. You don't have a common brand. Like, the people who use Unstuck don't necessarily know about Quizzerd, correct?
No, no. We've tried-
They don't even know it's the same.
They don't really know. Yeah, not really. Uh, some may, but generally we, we've not like crossed over between those two products. And effectively they serve two different like user persona, if I'm being honest.
Yeah. I guess on the App Store, if you, if you do a little bit of like, you know, "Oh, who made this?" And then, uh, you can ac- you can actually see on the App Store that it's the same people.
A hundred percent. Yep. Yep. Exactly.
Yeah. But like, okay-
We all-
In education, like I just, you know, I actually haven't looked at the App Store things. One hundred-- You're really good in e-tech. Um, I almost feel like-
Yeah, yeah
... okay, so, so you gave me a whole bunch of things. Sorry, I didn't wanna cut you. Was there any other things you-
No, you're good.
Okay.
No, this is good.
Yeah, I'll go through historically. Like, I just wanna double-click on like each of those things, right? And, you know, probably that'll, that'll be all the time that we have. Um-
Totally.
Quizzerd was your-- Was Quizzerd your first ever app? Like, what, like, did you have pre-previous successes before Quizzerd? Like, uh, how did you just go-
Quizard Tech8:33
No
... zero to a hundred like this?
Yeah. Quizzerd, Quizzerd was honestly like, um, Quizzerd was the first app that we'd built and put out. A lot of it came from like kind of first mover's advantage, if I'm being honest. Like, we were the first of the kind in, on the market where you could just, you know, take a picture of any sort of like, um, like, problem you're stuck on, question you're working with, concept you're working through, and sort of get like this sort of like in-depth tutoring experience around it.
The user experience was somewhat proved already though. Like if you look at companies like Photomath and Socratic, both bought by Google, um, they like had a similar idea of like a scan and solve approach where, you know, you take a picture of something that you need help with, um, and we give you, we give you support around that very specific like ask that you have.
But Codex did not have this. So are you saying this is, this is functionality that happened after Codex, right?
Well, Codex had the ability-- You mean in, in terms of, uh, like the-
The photo
... like image? Yeah. The, the first version of the app we would basically like just like OCR the, the image and then pass that OCR into a, a fine-tuned model that would be able to like extract like specific features from it that helped us better like answer those questions.
So things like what subject it is or, or like maybe like what subtopic within the subject it is. Because if we know those details, then we can best like-- we can probably respond to it in a, in more better way.
So for example, if you're using-- if you're asking a math question, um, specifically like something like arithmetic, it's probably best to go like step by step than like if you're asking a history question where like going step by step on like who was the first president of, you know, the US isn't like, uh, a step-by-step sort of response, if that makes sense.
Um, so really focused on like how do we understand what kind of question this person is asking, and how do we like deliver the best user experience to help them like understand what they like need and help them get unblocked.
Do you try to do model routing or do you not bother?
Like in terms of like-
Like, oh, it's-
I guess we like-
Let me, let me redirect to my math fine-tune or, you know, you know, that concept.
Yeah. We've thought about that. I think like we-- Initially, we spent a lot of time thinking about like this being a top like, a top way to, to go forward, i.e., like routing to different models.
A lot of people do this, yeah.
For us, we've actually like figured that the like unit, the like-- The return on it was not actually that beneficial. What makes sense for us instead is, you know, like we use base models, um, for, for like the final like response.
What makes sense-- What makes more sense for us is can we like spend more time on that fine-tuned like, uh, feature extractor up front and then like build in these like... And, and sort of model routing, it's more like prompt routing, if that makes sense.
So we like route to the right prompts, the right tools, the right examples based on the exact level of granularity we know on the-- what kind of question is coming in. And this scales in terms of like whenever you switch the base model, the base model gets better, like this builds on top of that.
So you're not always competing or in, in any way to keep up with like the best of the best in, in some sense. 'Cause like, for example, OpenAI has probably a de-dedicated team to do math, if you know what I mean.
Like, we're like very lean. So for, for us, we have to think about what is the, like, the best way to build on top of them, so that as they continue to improve that stuff, we are actually benefiting more from it.
And for us, that was more so this quote-unquote prompt routing, if that makes sense.
I mean, you know, I think OpenAI is not that interested in like high school math. They're, they're more like, you know, trying to do-
That's fair
... frontier math type things.
Uh-huh.
So, and then, um, on the distribution side, uh, mostly TikTok?
Viral Growth12:01
Yeah, TikTok and, uh, Instagram mostly-
Yeah
... is our, our bread and butter.
And you guys did it first, and then after that, you, you sort of worked with influencers to do, to do all this stuff. Is that, is that the general formula?
Yeah, yeah. We've done-- Exactly. I guess like a lot of, like, whenever we launch a product, we do it in-house a lot, and we've done a lot, a few campaigns that have done really well. So for example, like, obviously, there was the first video with, uh, Quizzer that, like, went super viral when we launched the app.
The second one that was super interesting is we ran a campaign in the fall. This being-- This is one of the videos that we've run. I, I would say go to the Quizzer account as well on TikTok-
Okay, give me one second
... 'cause we would, like, cross-post it from different things. Um, yeah.
There... Whoo. Twenty-eight million.
Yeah. Down to talk more about that. So basically, like, some of our biggest hits have been, like, in the fall of 2023, if you search TikTok for, like, you search Harvard, NYU, Boston, Columbia on TikTok, we were the first three videos, if not all fi-first five videos on TikTok.
And it's 'cause we were running a, a man on the street campaign that went super viral. If you look at the eleven point seven million view, that's from that era, where you just went around, like, asking the, the most, like, hot take, like, sort of like hot takes for college campus.
This wasn't necessarily, like, something that converted to, like, you know, product, uh, like conversions, but this was more brand building, and that effectively allowed us to, like, build a, a significant audience. People would then click into it, figure out where you're from, figure out the app, and they go and try it out.
But this is what-- This is one of our, like, most viral, like, campaigns. The second thing we did was-
How many copycats with this one?
Yep, yeah, yeah. People started copying right after. You know, we had, like, bigger companies in, in edtech as well as, like, small companies try it. But at that point, you know, we'd sort of like milked out the alpha in this concept, if that makes sense.
So it didn't... Um, the second concept that went super well was when we launched Unstuck. Actually got this concept, borrowed this concept from... Have you heard of Dupe.com?
No. Dupe?
Okay.
Looking for, uh, similar products, right?
Furniture. Yeah, yeah. Yeah, similar products. Yeah, yeah. So basically, when they launched, they had one concept with this, like, sticky note thing, um, where, like, they would basically like, uh... Like, if you go back to our, like, TikTok, I can show you the, like, uh...
Basically, all the sticky note concepts are from that. This is what-- This is how we launched Unstuck. We were able to get two hundred and fifty million views in under a month just on this concept, just ripping this concept over and over again, and this is what sort of like scaled Unstuck's like initial, like, viral growth.
Yeah.
But, like, the sticky note, is actually that important? Because I guess it makes a good-
A hundred per... It's like, it's like the, the, like... There's, there's different features of this, right? There-- Mr. Beast video in the background. I don't know if this is Mr. Beast. Sometimes it's somebody's surfer. There's the fact that the angle in that video is, like, basically pointing at a professor's ass in a very funny way.
Like, and it's like the handwriting has to be super clear. So there's a lot of detail and experimentation that's gone into a lot of our videos. We've also had, like, one, one video do so well with-- We're working with someone who done so well that, um, it brought our app all the way to number four in the education charts.
And so it was like, you know, Photomath's owned by Google, you know, bought for half a, I think, like... I, I don't really know how much, but, like, a sizable amount of money. Like, I'd assume, like, a, a fair amount.
Uh, um, you know, Duolingo. I think there was, like, this giant app studio doing this plan scanner app that was in the education category. And then there was us, you know, like, four people, like Quizzer, uh, in the number four category.
This was back when we'd launched another campaign. Yeah.
Photomath is still up there. Wow.
Yeah, Photomath is definitely up there.
What, what is-
It's one in-
Yeah.
I was gonna say about Photomath, one in every, like, four high school students, um, has used Photomath or has Photomath on them. They've got a pretty large... They've existed for, like, eight years or nine years now.
Would you say, uh, let's say Goth or, like, Solvex, these are your competition?
I would s- I would say they're definitely, like, competing with us in that Quizzer space specifically.
Oh, yeah. Okay, okay. Like, uh, do you, do you know them? Is it like a small community of, like, everyone knows e- everyone, or is it just the Wild West?
No, everyone kinda knows everyone. For example, like Goth-- But Goth is, like, a special case. Like, it's owned by TikTok.
Okay.
So they have, like, in some sense, infinite distribution-
Oh my God
... in many ways.
Kylie's gonna freak out about this.
Yeah, actually, when TikTok went down, Goth went down, too, 'cause obv- it was one of their ads. Yeah.
Okay.
Yeah. Then there's Question AI. So Question AI is interesting. They're by-- They're actually, like, the, the US brand, US, uh, a US version of a... Or let's say the US product from a, like, Chinese, like, company called Zoybank out in China.
And so they came to the US market to compete on the same thing, but they do, like, a few hundred million, I think, like, already back in, uh, back in China on their, like, products. But this is their, their jump into the market.
So we're definitely competing with some, like, a, a few players. But we've been able to, um, kinda win on the fact that whenever we go to market with, like, concept and distribution, it's stuff that really hits with our or resonates with our persona, like, user persona really well.
Yeah. I mean, you know, then the question comes to, like, how do you retain these users, right? 'Cause, uh, how do you, how are you gonna stop them jumping from app to app to app, right? Like, there's, there's gonna be one a- one after this, one after this.
Moat & Brand17:16
Like, so what's the, you know, this very VC question, what's the moat, right? Like, how do you-
Yeah
... you, you know, how do you make them loyal, sticky, whatever?
Yeah, totally. I think, like, the truth is, like, if you think about, like, how students use different products, like, a lot of them will, like, hit a paywall on one and move to another. You know what I mean?
Like, a lot of them are like... It's already a very specific case. So if your question is how do you retain users, you already have a natural churn. Every four years, a student graduates, you know? And, like, that's already, like, something you're dealing with, and you're always dealing, dealing with a new cohort.
So I think, like, um, moat in, moat in this sense is more so thinking about, like, how do we continue being the forefront of, at the forefront of distribution-
Distribution
... so that people are cons- consistently thinking about us top of mind. It's kind of like-
You can't always be on top of TikTok. You know what I mean? Like-
Exactly. But-- Oh, you mean like, you mean like we, because we can't be on top of TikTok, what do we do on that basis?
I mean, you can. It's just gonna be h-very hits driven. It's gonna be unpredictable.
Yeah.
It's gonna be... And then sometimes you're gonna not, not do well, like, just because, right?
Yeah. I think, um, I think that's where we feel that we're, we're bringing going viral down to more of a science, to the point that we actually are trying to make it a little more de-deterministic on our end.
But I will say a lot of it has come from, like, friends refer each other. They love the product. We keep it very simple. We keep it very clean. Um, and I think people sort of like-- Like, I think if you think of consumer products more as like CPG products, where it's associated to, like, some sort of brand or aesthetic or, like, aligning with some community, then the way you go to market and the way you sort of like, you know, the sort of influences you work with or the sort of like, uh, you know, UGC content you might put out there, users will align more with that and, you know, it just keeps them top of, keeps us top of mind when they use our product, i.e., like, the question would be like...
I guess like a, a counter question would be like, you know, what is-- why do people, like, buy, like, Coca-Cola instead of, like, Pepsi? It's probably because, like, people associate that brand as just being, like, much more, like, I don't know, just, like, higher value than Pepsi.
Yeah.
And I think, and I think we're playing on that-- we're pushing on that realm a little bit, where, like, sure, you know, we've got all this subscription stuff. We're building a product that's really good. We, like, try to, like, provide the best value for students, and we try to, like, make sure things are, like, high value.
But at the end of the day, like, this branding piece is probably one of the most important thing, especially in consumer tech right now, i.e., being in the front-- forefront of people's minds, being, like, on top of their feeds and being like...
And, and having a, a brand voice that resonates the most with, like, their, like, I guess, day-to-day and their lifestyle. Um, so it feels like we're, we're, like, selling to them, and, and they're, like, really, like, part of this community or part of this, like, um, I guess aesthetic that we're pitching them with our brand.
I guess the standard-- What, what I'm more looking for as well is there's, there's, there's the network effects, and then there's personalization. So the more they use you, the better you get, and the more people use you-
Personalization19:54
Yeah
... the better it gets, right?
Yeah.
Uh, so, like, you haven't really tapped those low... It's not low-hanging fruit. It's like, that's the platform.
Yeah.
We can talk about the platform. Uh, so okay.
Yeah.
You said-
Literally, that's the platform. Yeah.
So you-you're starting this platform team, or you already have started?
Yeah. We've already started it, and basically, the platform-- the goal of the platform is to be that node, right? Where we build a lot of these systems that allow us to sort of scale individual, like, brands from zero to a million as deterministically as possible, and then build a lot of the tooling around things that will let us, like, deterministically go viral on social media, um, or with other content and, like, basically be this, like a, like almost this, like, abstraction layer across different business units in the company where we just are building things to scale, like, their workflows and their playbooks.
'Cause all the playbooks that we run are very, like, I guess, like, automatable at this point. And so a lot of it, a lot of what people on the platform team do is, like, look at processes that we run in the company and then think about how can we, like, build tooling around this to make the decisions better or scale with, like, automation, i.e., like agents or stuff like that.
Yeah. Yeah. Okay, awesome. And, uh, I'm, I'm just kinda curious what you have found so far that is important for an AI engineering platform. So I see some details of the stack, which is Python, FastAPI, Flask, Mongo, and Postgres.
Engineering Hacks21:04
Uh, or I guess, I guess you, you use one, but then you're saying you're open to others, right?
Yeah, yeah.
Docker, Kubernetes, and Terraform. Any-anything else? Like any of the vector databases that you like? Any of the API, AI gateways that you like? Just literally anything. Any framework, LangChain? No?
Yeah. No, no, I c- I could talk through some of the weird, like, other hacks that we've done in the past-
Sure
... to be honest, to like, that have l-let us, like, I guess, like, scale a bit leanly. But in terms of, like, AI tooling, we've done a lot of things, like, natively, i.e., using native SDKs. I think, like, we've been-- Because it's so critical to know what goes into your prompt, I think it's been very important for us to, like, sort of like manually, like, own those prompts, i-i.e., like, build the integrations ourselves and, and build those, like, workflows ourselves.
So we haven't really, like, used a lot of, like, other, like, frameworks o-outside of that. What I will say, like, in terms of vector database stuff, I, I think a lot of, like, our decisions comes down to the fact that we're operating very lean, and so we know that, like, we can't have an engineer, like, fully own something.
And so it's like we try to outsource, like, stress, let's say, as much as possible. So actually, the, the vector DB that we, or vector search service that we work with is Azure AI Search. One, because we've got a shit ton of credits from Azure, like about, like, 2K in credits.
But on top of that, it's also because they have a lot of things handled under the hood through this, like, very seamless wrapper. The scaling is a little funky on that, but we have, like, we've built, like, solutions around it to sort of, like, keep, like, costs at bay.
I.e., like, the interesting thing with Azure AI Search is when you hit a certain tier's limit, you have to, like, migrate everything. You have to purchase a new tier and then migrate everything to that new tier, which is a little, like...
Yeah, I know. It's not, like, great. It's, like, really ass. But what we've built in on our end is like... Because something with consumer AI products is everyone does a novelty test, right? If you hear about a new product, you're gonna try something that you wouldn't have used on the daily at all.
But, like, for it to be a winning product, we have to win on the novelty as well as the actual use case. Does that make sense?
Yeah.
Um, and so what happens is, like, you hear about Unstuck, you're gonna put this, like, random video that you don't give a shit about, but, like, you're never actually gonna use again, but you wanna test the, you know, test the limits a little bit, and we have to work for that.
Um, the problem is, you know, that is gonna sit in our index, um, and it's gonna take space, and obviously, the way this pricing model of Azure AI Search works is it's storage-based. Um, so what we do have is we actually have a, a de-indexer that runs, like, um, every, like, few days, where basically we look at the index, we look at what's, like, been put in there and what hasn't been used in a while, and we take it off.
And the moment somebody, like, tries to trigger it again, we immediately reload it back in there. This has helped us, like, keep from needing to, like, scale in a crazy way because we basically, like, we basically erase all the novelty test, like, files off of the index and only keep the things that people are, like, frequently using, which grows at a much more, like-
Yeah
... manageable rate.
Uh, you, you, you take it off the index, but you don't delete the data. You just-
We don't delete the data. Yeah.
Yeah.
We only charge for what's loaded in the index. Yeah.
Okay. Okay, that's kind of a cool trick. So, like, so you're just, like, things that people use once, uh, and never use again. Yeah, I mean, that's, that's, that's, uh, and any-
Yeah, exactly
... right? This is not like a-- this is not a use-
Yeah
... specific situation. This is like a trick that-
Exactly
... could do.
Exactly. Another thing that's more interesting, though, is like, like, we actually, like, I wouldn't say the word abuse, but we don't use LaunchDarkly, like, exactly as it should be. So the story-- Let me, let me start with a little bit of the story.
So initially, we had got these, like, Azure, like, credits, about three hundred, four hundred K in them, and basically, you know, we wanted to move all our OpenAI load to Azure OpenAI, but this was, you know, about a year ago.
And back then the, the limits were super strict, i.e., we weren't gonna be h-able to, like, run our volume on Azure AI, especially when back to school was about to hit. And so, you know, I'd been on months trying to get back to them, like, trying to have them push the, like, uh, limits up for one region.
And then they were-- they came back to us and said, "Hey, it's really hard to get a region level increase, but you can stand up, uh, like, an Azure AI endpoint across multiple regions instead of a load balancer and to, like, load balance across them."
And for me, I'm like, dude, like, I'm like one engineer in the back end. I'm not managing a load balancer on top of, like, all the other things I have to do. Um, and so I was trying to figure out, like, a good way around this.
And then I realized that one night when I was setting up a LaunchDarkly, like, you know, like, staged a rollout, that effectively the stage rollout flag is kind of like a router at the end of the day. And so what we do is we use the stage roll-- like, we use LaunchDarkly flags to route between different Azure AI, uh, Azure OpenAI endpoints, um, setting up the percentages ourselves.
And so now, like, every new request is a random, like, a new, like, user effectively. And so it comes in, it, like, hits, like, you know, the, the LaunchDarkly flag, and it gets routed to one of the Azure endpoints.
And so that scales because obviously LaunchDarkly scales really well, and this is their main, like, thing. Um, and so we've done shit like that, where, like, uh, now we do a lot of our load balancing off of LaunchDarkly, where we can now manage it on one place, i.e., we can redirect traffic across products to different models and try things out.
Um, but a lot of the management had to happen without needing to redeploy, which has been super interesting for us and helped us scale really well. Yeah.
Fascinating. Um, I mean, like, do you even-- Like, I'm sure, like, most of your users are US or it actually d-does it matter region-wise?
We have, like, international users, but most of our paying users are from the US.
Yeah. Yeah. Interesting.
Yeah, we have a ton of international, though.
Yeah. Uh, I guess international is good for keeping you up in the free tier and then the paid-- it's paying con-converts. Sorry. I think, you know, I, I have, I, I have this tendency to go into business, business side, but I wanna stay on-
Go ahead
... tech side, especially, I mean, you, you know, you're a CTO, so you can answer all the tech, tech things. Um, prompts engineering, um, u-usually, like, I have been going back and forth on do I store prompts in my code base, right?
Um, and developers love storing prompts in their code base. They're like, "We can version control it. We can see it together with all the other stuff." And then you're like, "Yeah, but, you know, iteration is very slow. You have to wait for the full CI/CD process to happen."
Um, you know, it's... And you're definitely-- Like, basically, what, what is your prompt iteration process? Do you use some kind of eval framework or ops thing just for, just for-
No, that makes sense. I think, like, um, we initially when Quizard was, like, kicking off, like, I had to build an eval framework internally because there, there was not much out on the internet, like, back then about evals and stuff.
Um, we called it Excel. I don't remember what it stood for, but it was something about, like, uh, experimental something, something. It was some acronym. I'll, I'll, I'll see if I can figure it out. Uh, but basically built this benchmark out to run a lot of this, like, prompt tuning, um, for different things, specifically, like, um, tuning, like, that end result when we find a different subject, if that makes sense.
'Cause the way we have Quizard set up is effectively this decision tree where we, you know, we fine-tune this model to pa-parse out these, like, specific features, and based on that, we send it down the right, like, prompt route, if that makes sense, to help to have it give them the best user experience and the most, like, qual- like, high performance answer, if that makes sense.
And so, like, in terms of prompt management, like, for us, I think, like, the way we optimize is more on... And maybe it's not, like, typical to other tech companies, but for us, like, the experiments we've run in the past is, like, instead of trying to, like, do a lot of optimization in advance, it's kind of like we run an experiment or we run different prompts, and we see which one, like, triggers conversion better and, like, improves profitability.
And then the one that does, we'll just kind of keep that. But I think there's never been, like, a... There's never a, like, necessarily, like, a pushing need to continuously tune prompts, if that makes sense. Especially as models get better, like, you don't necessarily have to, like, keep tuning prompts.
I think, like, the better focus is, like, if you can figure out, like, I think the priority for people should be more on the orchestration of the AI workflow than the, like, actual prompts, because, um, LLM models are gonna get better at, like, understanding the shittiest prompts, to be honest, and, like, still deliver value.
But the, the best question is, like, how do you make sure you're building a non, like, putting as much determinism into a non-deterministic system? And this means, like, how do you, like, limit... It's almost like I sometimes, like, like to liken, like, prompt engineering as, like, like programming language design, where you're, like, deciding where the typing is, like, where the constraints are, and then how expressive it is, so that you can, like, both constrain what it can do, but you're allowing to be expressive in a very constrained way.
So for example, like, on Unstuck, you know, like, the agent that runs, like, the, the, like, back and forth chat around our different files, instead of having an agent that's exposed to very axiomatic tools like a search tool, a summarize tool, and, and some, like, stuff like that, we instead what we did is...
Actually, it was very, it was very interesting. The w-way we launched it was just a pure vector search, and we gave it to a beta audience. It was like, do whatever you want, and this is just vector search.
And then I would look at the data, and I'd be like, "Oh, these are all the things they were trying to do." I would categorize them, and I'd be like, "All right, great. We've now identified intents." And so for each intent, we built a deterministic, like, like, LLM-driven, of course, workflow, but it's every time that intent is hit, we know it runs the exact same process each time.
That also gives us the ability to tune that intent really well or, like, know how to fix, like, a broken user experience. Because if it was something that it was an agent that had exposure to very, like, basic tools and had to, like, stitch things together on its own, how the hell are you gonna debug that at scale?
Like, that's so hard. Um, you've kind of made the problem way harder for yourself. Um, and so instead what we have is, like- Literally all the LLM does is it classifies what is the, you know, the intent of the user, and it has some level of freedom in, like, specifying parameters to pass in to customize that intent workflow.
But at the end of the day, that intent workflow is already de- pre-determined, undetermined stick in how it'll execute. Because the idea is we wanna make sure that every user experience, whenever you hit, like, summarize some content on, on Stuck, it's always giving a very similar expected behavior, which is what, like, users want.
Like, you don't want magic one moment and, and no magic the other again, especially, like, I show a product to my friend and it, you know, I was like, "Oh my God, that was so amazing," and then when I show them the same thing it doesn't work out.
Like, that would be, like, not a great user experience, to be honest. And so, um, I think, like, less on prompt engineering and prompt versioning, but more on, like, how... Like, because this will matter less and less, I think, to be honest, and what will matter more is, like, have you built a system that will, like, allow you to scale deterministically and, like, actually manage improving your, like, workflows deterministically.
So I guess what I mean by that is I would emphasize people to focus on, like, building the most if else condition-esque, like, LLM flows, because that gives you something to actually improve, i.e., like, if I identify an intent that's doing really shittily but has high volume usage, I know the expected value of improving that intent.
And so I know how to prioritize my engineers on that, like, you know, improving the product in that way. But if you have something that's, like, if you're just prompt tuning and you don't really know why you're, you know, tuning it or what the, like, general, like, benchmark you're trying to improve with is, like, you're kind of wasting, like, I guess, like, time and energy in a, like, unknown, like, direction or Nor- North Star.
Model Picks31:45
I don't know if any of this, like, sort of answered your question, but we've kind of moved away from the whole framework of prompt tuning and more on, like, making sure the workflow is something that can be managed at scale in terms of improving it.
No, it's, it's really good. Uh, I mean, in a way it's, it's kind of like you're, you have a kind of router. You're really, uh, the classification step is still important, right? That, and that's what a router would, would do.
Exactly.
Except that the, the classification doesn't lead you to route to a different model, it just leads you to route to a different workflow that is more, uh, deterministic, as you say. Uh, I would, I would use-
Yeah
... variable rather than deterministic, because it's still not deterministic entirely.
Valid. No, valid.
Yeah. Um, and I guess you, like, people can report hallucinations and bugs, and you can see if there's any spike-
Yeah
... and, you know, you, you can go solve it.
Exactly, in a specific intent and double down that intent and fix it.
Yeah. Just kind of, actually kind of curious, have you seen interesting holes in knowledge between OpenAI versus Claude versus, uh, anyone else of, like, oh, this-
No
... is, like, really bad at history and, like, oh, this is really good at since this other thing?
I think, like, as opposed to, like, knowledge gaps in subjects, what we have found differences in is, and maybe this is kind of, like, well-known, but whenever we need to do stuff that's a little more conversational or, like, a little more, like, like, a little better at slang, like, Claude is way better at slang.
Like, whenever you ask, like, Claude to generate something that's, like, that sounds, like, human or, like, sounds like it fits in with, like, the tone of whatever, like-
Right
... you're, like, giving it to, to-
Gen Z
... place to regenerate. Yeah, it sounds... It's way better at being Gen Z versus OpenAI will just throw in hashtags and emojis. And I'm like, "Dude, like, this is not, this sounds so fake." But yeah, Claude is... Whenever we've had to do something, like run a workflow, whether it's for growth stuff or, like, in a product where we need it to sound more human and scale, like, human, like, te- like, human-sounding text, Claude has been the, the model of choice.
Yeah. No, that's, that's really interesting for... I mean, that's a win for Claude. Uh, I, I was gonna say-
Yeah
... I, when you started talking about conversation, I was like, already, you know, Claude has that reputation, but, uh, I didn't know about the writing style and, like, yeah, definitely OpenAI is very cringe. And any other models that stand out or are you just the, the, the big two labs?
I think these two are our main right now, to be honest. We're exploring Gemini as well for some of our stuff, but we haven't-
I mean, Gemini is fantastic. Yeah.
Yeah. Exactly. So we, like, I think I don't have enough reps in with Gemini to, to have, like, a, about, like, a good take on it, but so far it's just been OpenAI and Claude.
Yeah. I mean, it, I mean, what, what, what it comes down to is that you are necessarily multi-model now, and you cannot just be as your OpenAI credits, right? Like, you have-
Yeah, yep. Exactly
... some kind of model, like AP- AI gateway or whatever-
Yeah
... that normalizes these things. Yeah, cool. Okay, any other hard, uh, AI engineering problems that you wanna talk about? Like, you know, things that just you took a while to figure out.
Let me think about this. I think, yeah, I don't know if I have anything come to mind because most of it has been less on, like, like, the AI engineering, like, task itself, and more on, like, like, you, like, how to support a certain user experience, you know?
And, like, we've done little, like, a few hacks around, like, different things like that and, like, I think, like, the way we build and the way we prioritize things is how can we make sure everything is experimental, like, by nature.
So, like, you know, we launch, like, when we have, we have our transcription service, we have, like, a waterfall, you know, approach, and we have things that we can dine out. And this, uh, also is matched by LaunchDarkly.
So, like, if, you know, if one service goes down, we can reorder it in LaunchDarkly and send it, and suddenly it'll, like, prioritize a different, like, waterfall structure. And so, like, honestly, if Launch, if LaunchDarkly goes down, we're cooked.
Like, we're literally, like, operating a lot of the company off of LaunchDarkly. So definitely a big fan of, of their product.
I mean, you should, you should send them this recording once, once we're, uh, done, 'cause, like, I feel like it's important. But also, maybe, maybe you're, uh, you know, they'll be like, "What are you doing with us?" Have you, have you explored other alternatives to LaunchDarkly, or are you just a LaunchDarkly maxi, you know?
'Cause I, I, I've heard that they're expensive, but maybe, like, you're so profitable you don't care. I don't know.
Yeah. We, uh, we actually are not... We're still on their developer account, but we don't use... Yeah, yeah. Like, I think, like, the way we, like, run experiments is we don't use, we don't use them for experiments. We just use imperious feature flags.
I mean, we obviously abuse the feature flags in a way that's, like, interesting, but they don't price on that, like, the way we use it, right? They don't price on how many people are hitting your feature flag and, and that, that works for us.
The way we run experiments, though... Pardon?
I didn't know that. That's great.
Sorry, yeah. And, and the way, but the way we run experiments is... Oh, sorry. To answer your initial question, we have tried things like Statsig, um, and stuff like that, but I think, like, it- Didn't seem to provide the value that we were looking for.
The other analytics like s- like, uh, software we're, we use is Mixpanel. And then the last thing that we do is for experiments, especially for, you know, the mobile app experiments, um, okay, so, like, have you heard of RevenueCat?
So RevenueCat and their paywall has this metadata, like, tag thing, which is actually meant for, like, you to, like, render your paywall in a very specific way, not so much anything else. But we actually use that to run our experiments as well, i.e., like, when you send a paywall, you can send a metadata flag, but we attach that-- we send an experiment value that we attach to a user, and that will trigger, like, obviously some experiment function.
So for example, whether a feature's enabled or not, and then when they hit the paywall, obviously, like, we know that this user conditioned by this experiment is now seeing this paywall. So we actually do a lot of our experiments off of RevenueCat, but we've kind of like, I wouldn't say misuse, but, like, use in a different way, like, this metadata tag that they, that you can send to, like, customize a paywall on, on our end.
So it's been, like, very like, what is the best way to consolidate, like, all the things we need on the platforms we already use? And we've just been-- I think we've, like, we've been very, like, thankfully very scrappy about, like, finding ways to solve certain problems.
Um, and they've actually scaled pretty well, um, for us. And so, uh, yeah, a-actually we, like, r- we don't really pay, we don't, like, have a huge cost on LaunchDarkly. Like, to your initial point, it's been pretty good.
Um, yeah.
It's awesome. It's awesome. Okay, cool. Um, all right. Now I'll, I'll close it out with, with the, the last thing that you mentioned, which is you're also looking to have purely AI, um, automated agents running around in your, your company.
AI Agents37:49
That's-- I think a lot of people are saying that they do that, and then I, I'm, like, suspicious of, like, how much they are actually doing. Like, Klarna is doing that. It sounds like Shopify is about to do more of that.
Like, what, what are the opportunities that you see for building those things?
Yeah. I think a lot of the options we've seen, and we have already deployed some of these, like, is I guess to caveat, like, the first thing we always do is, like, try to see if we can build, like, regular automation, see if that scales, and then build the agent that, like, supports everything.
What I mean by that is, like, um, we don't wanna over-invest in trying to get an agent stood up if we can solve the problem with, like, general, like, basic automation or like a better-
Zapier
... dashboard utility. Yeah, exactly, things like that. Um, but what we've seen opportunities for are, like... Sorry, go on.
Oh, no, no. There's, there's, like, the AI-native versions of Zapier, so like Lindy-
Yeah. Yeah
... us, and Gumloop and Wordware-
Yeah
... those kinds of guys, yeah.
Hundred percent. I think for us, like, what we see opportunities are, like, for it is, is mostly on this, like, like, growth, our growth and marketing side, i.e., like, things like how we do research, like how we study concepts in the market, like, ways to like, like, like, have agents to run to, like, understand what concepts are winning and going really viral across markets that are, market segments that are very similar to the market that we're gonna launch an app within.
And so basically, like, right now we have, we literally pay, like, our, like, we literally have our, like, head of marketing who literally, like, is literally twenty-five hours a day on TikTok, you know, scrolling, like, literally, like, an ungodly amount.
Um, but it's because, like, he's, like, sort of like tuned into the algorithm itself, and honestly, like, if we could free his time by building agents to help him do all, like, replace a lot of his workflows, like, that would be a level up for us.
Why?
So honestly on our end it's like Pardon?
Like why-
Literally watch TikTok for me, you know, 'cause, like, you can learn a lot by, like, curating your, like, feeds to, like, hit a certain demographic, you know, and you could see the sort of content you get, like, shown.
Um, so a lot of the opportunities we see is in, in, like, basically, like, reproducing, like, the workflows that our marketing team, like, already does, um, internally.
Mm.
Um, and so, like, it's very much like let me just see what this dude does, let me just reproduce as an agent, and let me give him, like, a synthesis at the end of that workflow so that he can work with it, um, for the next stage of, like, what happens after researching a concept.
Um, so a lot of it is still, like, I will, like, you know, I will say a lot of it is still very nascent, so we haven't like, like, we're still in this phase of, like, trying to figure out what works and what doesn't and then scaling that.
But we see a lot more opportunities in terms of, like, research for market concept, like, uh, marketing concepts in terms of social media, as well as, like, research in terms of identifying like, you know, what's the next product market to hit, if that makes sense.
So for example, like, the way my co-founder will, like, and us will, like, the way we'd all, like, research a new market, a lot of that is very, like, systematic, but a lot of it then sometimes has some gut checks, and if you could, like, get that systematic way, like, part of it, like, out of the way through, like, agents or, like, just, like, a really powerful analytics dashboard, then we can sort of, like, scale how fast we make these decisions on, like, what's the next product line that makes sense, you know, that's highly lucrative and guaranteed to be, like, a, a profitable, like, venture to, like, spend time in.
Yeah. Awesome. Um, I mean, I, I think that's something that everybody would want. There's, uh, something that we need to figure out how to, how to do. In my company, we're, we're literally, we're trying to, we're starting to talk about, uh, sales agents that, like, reads our emails and then responds with, like, you know, information or requests or req- or, like, invoice details and all that stuff.
Um-
Nice
... I'm just, I don't know if it's over-engineering. I'm just like, I don't know, sometimes you know it'll be better if a human does it, but because you want to automate, then you have to code this agent up, and it's probably gonna suck.
Yeah.
It's probably gonna have mistakes. I don't, I don't know if it's worth it, you know?
No, I, no, a hundred percent. I think that's why it's like if you can, like, if you can first build the, like, an automation tool that scales a human to do it really well or do it faster, then you can, like, determine if you still need it to, like, build the agent to run on its own independently.
So we're not, like, we're not necessarily like gung-ho, like, "Yeah, let me put an agent to everything." Um, it's more so the first phase is to, like, can we, like, build the right automation tooling to make a human be more augm-augmented in doing their work?
And then we can sort of, like, build an agent if there's more, like, need or we need to scale that workflow, like, like, in parallel, you know, like, across, like, different, like... Like, yeah, just we need to scale it up to, to, to run many things in parallel, if that makes sense.
Yeah. Awesome. Cool. Uh, you know, taking-- went over as it is, but it was a really interesting chat. Any calls to action, I guess, if it's hiring?
Outro42:07
Um, yeah, mostly we're hiring. I believe we're building a systematic consumer portfolio firm. Um, you know, and, and we're building the system to help, you know, launch, uh, in a way and deterministically in any market, and we're coming for every market.
So if you're interested in working on, um, working on owning consumer products that go from zero to million or interested in building, uh, a company that's run purely by agents, would love to chat. Please reach out.
Yeah. Awesome. Well, thanks for joining. Uh, I'm sure this won't be the last time we chat.
Of course. Yeah. Thank you for having me.





