Intro0:00
Hey, everyone. Welcome back to our Lead in Space Lightning Pod. This is Alessio, partner and CTO at Decibel, and I'm joined by Swyx, founder of Small AI.
Hello, and today we are joined by an extra special guest, Thomas, your CEO of Raycast. Welcome.
Hey, thanks for having me. Finally making my way here. Been listening from the sidelines for a while, so happy to share a few things on our end.
Yeah. Um, I, I'll share some personal context as well. Uh, I have been using... So I think, like, people may want to know what Raycast is. Like, Raycast i- in my mind is kind of like a, a better spotlight for, for Mac, right?
Uh, it sh- there should be one universal quick access tool for anything on your, on your machine, and the default one that Apple ships isn't very smart, isn't very good. The options really suck. My- I have a new Mac setup guide that is pretty popular among other people, and the first thing I tell people to do is do things like turn off the developer settings where it lists all the source code of whatever .c files that people have.
Um, I also was an adopter of Alfred for a long time, and recently I moved to Raycast.
Nice.
So, uh, and, and actually was, was, was, I was told by an employee to do this because they, you had integrated good AI features. So I think, y- you know, uh, maybe you wanna introduce, like, your journey with, with AI, I think, you know, s- just, just to bring people up to speed.
Raycast Journey1:25
Yeah, definitely. Uh, great to hear that you adopted it, and yeah, the introduction is very accurate. So we're basically around since 2020, and started there as sort of a drop-in replacement really for macOS Spotlight, and kind of what we paid attention to is, like, giving users a, a better experience, but then also, I think what really put us on the roadmap is, like, our extensions.
So we allowed basically developers to build their own extensions. Um, it's very easy to do with TypeScript, Node, and React, and they basically can extend the functionality of Raycast. And what we did, we basically built, like, sort of this end-to-end experience.
Developers can build it. They can then publish it into a store where other users can just discover that. And there are, like, thousands of extensions now, like ranging from Linear or Notion to Zoom. And so you can connect to all those tools and, and really have, like, one interface to access all of those.
And so when two years ago, or two and a half years ago by now, I guess, AI came around, one thing that we were, which I mentioned to Alessio, is like we've been basically this global text box that always lives on your Mac and is always accessible with Command + Space or the hotkey that you use.
And so at the time when everybody was looking for a text box to put in a prompt, we basically started our AI journey as well. And initially this was, like, very easy of, like, you put in a prompt, you get an answer.
The next thing then we did was, like, integrating on the operating system level, like grabbing your selected text, then you can then run an AI command like fix spelling, and it automatically picks up the selected text that works across all the different apps you have, so you can use it to improve your email writing or your tweets or whatever.
Um, and then on top of that, we basically expanded into having, like, a fully fleshed out AI chat by now, and basically, well, we're recording this one day before a big announcement, so tomorrow we're gonna release our next big feature, which is called AI Extensions, and that basically brings together our roots of, like, having developers extending Raycast, and this time they can extend Raycast AI.
And so what it allows you to do, you basically can chat to all your apps and services in just natural language and using them in sort of an attendant way, and happy to show some of those things in a second.
Yeah, that's great, and I've, I've been following the journey along since you did YT, and you started-- For people that don't have that background, it started as a really developer-focused tool, and now I think you replaced one of my use- used to be favorite apps, Command + E.
Demo3:35
So yeah, would love... I would love to just jump into the demo. You know most people are gonna watch this on YouTube, so the visual always, always help.
Let's do it. Um, so let me share my screen real quick. So this is basically Raycast. I can launch Raycast with, like, a global hotkey. Mine is set to Option + Space. And what we're basically releasing now is a way to interact with your extensions, and these are basically apps and services in really, like, a natural language way.
And how we do this is basically we're allowing you to sort of @ mentioning your extensions. So for example, I can @ mention here Slack, and then you see a few things that I did here before, my history, and suggested prompts.
So in this case, I'm just gonna set my status to recording podcast for one hour. And so here what it does, it figures out how to do this. It sets my status and quickly comes back with a response.
So what I have done here is, like, basically just updating my Slack status without opening Slack, and that's just, like, an easy way to, like, let Raycast perform actions for you. What you also can do is you basically can bring in other extensions.
So same thing. I can, for example, use our HR system and ask who's off next week. So just if I wanna check if some of my colleagues are off. Again, it goes out, checks, like, certain tools and brings back information.
So Yugu is off 3rd March to 7th of March. So it's a really easy way to basically interact with your tools and, and getting information across all the different applications that you have installed. All those extensions are basically built by developers, so it's very easy to integrate with, with our system.
Yeah. And so all of this is through auth into the app and then calling the API of the app. You're not doing computer use or any of that AI stuff yet.
Exactly. So this is primarily, like, using, like, OAuth and then, uh, doing API calls, and then developers basically can AP- expose those information to our AI, and then we're picking it up and, and composing those information together. Um, the cool thing is also you can sort of combine those.
So what I, for example, did, I created myself a little meeting assistant. So this opens our AI chat in this case. I can quickly show that. It's basically sort of like a set of system instructions. I, for example, specify here how I wanna have my meetings formatted because it's sort of a personal taste I have.
And you see here on the bottom, it's like it brings in Google Calendar and, and the Zoom beta. So here these are two AI extensions which are automatically loaded in into my- AI chat. And so what can do now is I can sort of freely chat here and, for example, ask like, "Find me some time with Petro next week."
And then my system goes off and basically searches my contacts, checks the availability of Petro in this case, and comes back with a few different time slots. So here I can then just follow up and, like, ask, "Okay, book 11:00 AM, um, on the Monday for 30 minutes," and then basically can continue this.
And so what's smart here is basically it recognizes like, "Oh, I wanna have a meeting link here," so it creates me a meeting link, and then it suggests me this. So this is one thing that we added here is, like, for things that basically are sort of destructive or actionable, you maybe wanna gut check this, so we keep the human in the loop for something where you create an event and you maybe wanna check is this really the right attendee, and do I have a Zoom link here?
And you see a little bit of indicators where I can press Command Return. So in this case, it creates the event for me, have it 11:00 AM, 30 minutes, have it all set up. So here I don't need to jump around between different tools and figuring out how to do some of the stuff.
And so this one is, is really exciting. We feel like that can enable people to really build sort of mini automations where you have really quickly accessible for you on your desktop all the time, just a keyboard shortcut away.
Yeah. This is like the super integration of all the things. It's almost like Siri, but, you know, done well.
Fine-Tuning7:41
Yeah.
Yeah, like I, I noticed that y- you used a Ray 1 model in there. What is it? Is it... You know.
Yeah, you spotted that automatically. One thing that we did is basically we started exploring all the different models and then, and really quickly realized, like, we have such a unique use case, right? So we have quite a large set of, like, certain functions or tools we wanna call, but then also a very specific need for that.
So at some point we decided like, "Hey, let's go down the route of, like, fine-tune a model." And so we looked into all the various models we had, um, and then we picked, at the moment it's GPT-4o and GPT-4o mini, which we basically did a fine-tune to really optimize for our use case, and then basically shipping that in the app as, as Ray 1 and, and Ray 1 mini.
So they are highly optimized for our function calling and basically make it, like, one, faster, but also the most important part, more accurate because, like, you wanna make sure that those things happen as best as possible.
Yeah. Got it. Um, so you have some kind of fine-tuning data set. What do you fine-tune for? What is it not good at out of the box?
Yeah. So we, we recognized, like, it's, like, because we load in sort of different set of functions, like sometimes it gets it right, and then we started basically when it gets it wrong, um, ad- like adjusting the tool descriptions or the system problems and all those kind of tweaks you wanna do to make it better.
But we find at some point there was just a limitation, like we couldn't go further up. So then we basically started to fine-tune. We actually started to distill down first 4o to 4o mini, which was actually giving really good results for what we had basically, because you can really mimic the behavior of the bigger model, right?
And like distill that into the smaller one. Um, and so the bigger one is like sort of giving like most of the stuff, um, and basically then using that, um, with our data set to make sure that we can, like, distill it down further.
Um, we also tried various ways of, like, looking into can reasoning models get better? But I think it's like sort of this common thing which is kind of funky, like reasoning models behave very differently, and the ones that can do tool calls, we found them, like they have a very different way of using tool calls, um, so we didn't make them work yet.
And we also played around with pretty much all the models out there. Some were better, some were worse, and we sort of were going first for like accuracy of like can you call multiple tools as you saw, especially in the last example where it like needs to call a chain of tools.
Can it reason through that and figure out how to call that? And there, there was quite a few surprises there. Some models got this right on the first try, and others were completely going off. And so we find like the OpenAI models function calling-wise for our use case so far the best ones and like, yeah, basically verifying the others.
We had a beta group and testing different models, um, so basically from all providers and, yeah, find basically the best one so far. Still GPT-4o and 4o mini in a way, which is surprising almost because they're so old by now.
I see. Well, you know, it keeps changing. I was also thinking like the reasoning models would be too slow for you. You know, like-
Yeah, that's the other thing. Like one of the things-
Hard to detect it. Yeah.
Yeah, reasoning models basically became too slow as well. Like for those use cases you wanna have something rather snappy. So the first thing we did is basically make it work and then make it fast by distilling it down, which gave pretty good results for us.
How many extensions do-- does the average Raycast user have? And is that an issue? Because every single extension is an overhead for tool choice, right?
Tool Overhead11:06
Yeah. So that's why we introduced basically this @mention. We feel like very often you kinda know which service you wanna do, so you're @mentioning like a Google Calendar if you wanna chat, right? And so we're guiding basically the model of like we're just loading those amount of information into the context, and which actually like brings a pretty, pretty good UX.
Like it feels very natural, like almost like you're in Slack and you're tagging a coworker, right? Which you wanna get some work done. So this feels very good. So usually like even if a user has hundreds of extensions installed, that doesn't become a problem because we let them essentially manually filtering out.
One thing which we haven't done yet is like sort of doing the automatic one. If you just like put in something, we could probably figure out which kind of extensions you wanna use and like then bring them automatically into the context.
That's something which we'll probably explore later. But yeah, that's sort of the thing. Uh, with OpenAI what we figured out at some point is like, oh, there is a upper limit of like 128 tools you can have, um, and then the API basically gives you an error.
Uh, so we ran into some of those issues as well. I think other models allow you more tools. And then oftentimes you also have the problem, they're eating into the input, uh, token count, so can shrink basically the other information you can s- stuff in there.
And you're basically transpiling, so to speak, all of these extensions to work with every model. Uh, I think OpenAI and Tropic are pretty kinda like, you know, backward compatible, so to speak, but did you have any issues with some of the alternative models that you offer or?
Definitely. Like, I feel like every model has their little nuances. Finish reasons are different, how they're calling tools are a bit different, behavior's a bit different, formatting is a bit different. So there was, like, one thing which we basically had to go over all of the different models and patch them, and we're not using like a, a proxy, like a router or something like this, so we ha- build that internally.
But like, yeah, that's, that's one of the gnarly things you have to do. Like, even so all models say they support OpenAI compatible APIs, there are always little nuances which are slightly different and then put you off when you see it for the first time.
Uh, what else did you have to build? You mentioned, uh, you had to build a gateway. What about evals, like, uh, all, all of that fun stuff?
Evals13:28
Yeah. Yeah, so we have an internal eval system, which was actually surprisingly hard to build. Like, we felt like, "Oh, this must be a solved thing," right? Like, basically since LLMs came out, everybody talks about evals, right? And then we like, "Come here, let's do evals."
And then, like, nobody talks about how to do evals with tool calls. It's like a super unsolved thing. So the problem with tool calls is you... if you really wanna do this, you kinda need to mock your tool calls, right?
When you say like, "Hey, check the availability for this contact", you kinda need to have something that comes back. And so we basically have a system to, like, mock all those tool calls and, like, put that into the eval.
So an eval for us is like you have an input, which is to use a prompt. You have, like, a series of mocks, which is basically the tool call stuff that goes in between. And then you have sort of the expectations.
And for us, expectations are primarily two things, like, one, tool calls. Like, is the tool called with the right arguments or combination of arguments? Are the right tool called, tools called? And then also, is the output formatted properly?
Like, you might wanna... For example, when I create a meeting link, I wanna make sure there is always a Markdown link in there. So this was a surprisingly hard problem to solve. Um, the other angle that is hard for us, like, we basically wanted to give third-party developers who build those AI extensions, um, a possibility to write their own evals.
So we basically came up with a system that they can write like a DSL for those evals, and then they can run the evals themself, and so they basically can check. And I think that's, like, quite critical because, like, if we change something in our system, we wanna make sure that the AI extensions still behave.
But if a developer change something, they also wanna make sure that nothing really broke. And I think that's another thing which I don't hear much about, like, working with third-party developers is a very interesting contract because you wanna guarantee that those things basically don't break, right?
And we all know with LLMs, it's, like, always a bit tricky, like some things regress very easily, and so evals is the best bet. So that was something which we had to build basically from scratch to make sure that we have something there.
Yeah. And on the extension side, we did an episode with LindyAI, we did one with Dust. I don't think either of them allow users to build custom extension to then plug into workflows. Um, any other commentary on just, like, state of, you know, AI integration agent platforms and what you definitely wanna do differently or, um-
Yeah
Platform Design16:00
... any other fun design decisions?
Yeah. So we had sort of the luxury, quote-unquote, to have already this extension platform, right? So people can build extensions with Node, TypeScript, and React essentially, and so they can co- distribute it through the store. So we had all of that ready.
The only thing we had to do is giving, like, a new entry point for what is essentially a combination of tools and prompts. And so we basically added this, and we tried to make it, like, really natural for a developer to build that.
Like, especially people who aren't familiar with AI because I think, like, we as developers still are going through a journey, right? Like our engineers to, like, figuring out how to do this. So basically what we came up with, it's essentially you just write a TypeScript function, and you document your TypeScript function with JSDoc, and then we extract all the information from there and basically make that and pass that information to the LLMs.
And so for a developer, it feels like they're writing just like their usual TypeScript code, and that's like, I think, like, a pretty elegant way to abstract away all those other things that you usually have to think about when building function calling and tool calling, and make that much more approachable.
So all of those AI extensions are gonna be open source. We have 50 at the beginning. We worked with a bunch of developers together, and then we expect a bunch of more people building more over the next couple of weeks and months.
Automation17:22
Yep. Do you also have schedule task, kinda like things are getting run in the background, or is it mostly-
Not yet, but this was something that popped up pretty quickly, like what the team and also our, uh, beta testers basically did, they built a lot of automations, like take a browser tab here or, like, browse the internet there, combine those different AI extensions.
And so at the moment, you need to trigger them yourself. The next thing we sort of wanna look into is, like, turning it more into, like, automation/workflows that can run in the background for you. So you can think about you building one of those.
Different trigger points, it can be, like, time, it could be something completely different that triggers it. I think we have something quite unique of being on the operating system, so even you could think about some file events, maybe you download something, and then after the download, something gets triggered automatically.
That's something which we wanna explore of, like, basically bringing in even more... I guess it goes then into autonomous ways of, like, letting, letting your little AI extensions do more and more work for you in the background.
Future Vision18:24
Cool. Any other parting thoughts, like the, the future of AI for Raycast? Like, is it a major part of your business? You know, any, any other thoughts?
Yeah, definitely. Um, yeah, basically it's like a major part for us. Like, our future is basically turning macOS, and we're also now expanding to W- Windows and iOS, and really like a ni- AI native operating system. Like, I'm not sure about you guys, but for me it's like since I use a, a Mac or a computer in general, the whole experience hasn't really changed, right?
Like, I still click around and do the same things, and I feel like it's ripe for, like, a big change. But I can't really see that coming from the big operating systems. They're, like, so mature. They're, like, nothing moves anymore, right?
Um, especially, like, on computers, which are almost a bit historic in a way. And so we feel like there should be something who, like, develops that and push that forward. And, and given we're, like, sort of in between of, like, apps and, like, the operating system, we're, like, seeing ourselves as this AI layer and, like, infusing that and, and working across apps, being deeply integrated, and really think about, like, how, how a single person can basically get 10 times more things done with things like what Alessio mentioned, like maybe there is, like, multiple things running in the background, and it becomes more of like a delegation and orchestration how you work.
And really thinking about, like, how can we basically change how you use a computer going forward? 'Cause I think it has been way too long the same experience, and by seeing, like, how much AI can do already and where it goes, I feel like there's a lot of stuff up for grabs, which we're super excited about.
Excellent. Well, uh, thanks for working on Raycast, and thanks for making everyone else more productive. Thanks for jumping on, uh, and sharing about your, your updates.
Thanks for having me. It was a pleasure joining you guys.






