Intro & Recap0:00
Under DevDay lights, code ignites. Real-time voice streams reach new heights. o1 and GPT-4o in flight. Fine-tune the future, data in sight. Schema sync up, output's precise. Distill the models, efficiency spliced. WebSockets blaze, connections flow. Voice AI live, watch innovation grow.
Happy October.
Wonderlust and function combined.
This is your AI co-host, Charlie. One of our longest-standing traditions is covering major AI and ML conferences in podcast format, delving, yes, delving into the vibes of what it is like to be there, stitched in with short samples of conversations with key players just to help you feel like you were there.
Flying through the code.
Covering this year's DevDay was significantly more challenging because we were all requested not to record the opening keynotes.
Bring the game.
So in place of the opening keynotes, we had the Viral NotebookLM Deep Dive Crew, my new AI podcast nemesis, give you a seven-minute recap of everything that was announced.
Flying through the code.
Of course, you can also check the show notes for details. I'll then come back with an explainer of all the interviews we have for you today. Watch out and take care.
Connections flow.
All right. So, um, we've got a pretty hefty stack of articles and blog posts here all about OpenAI's DevDay 2024.
Yeah, lots to dig into there.
Seems like you're really interested in what's new with AI.
Definitely. And it seems like OpenAI had a lot to announce.
Hmm.
New tools, changes to the company. It's a lot.
It is, and especially since you're interested in how AI can be used in the real world, you know, practical applications, we'll focus on that.
Perfect.
Like, for example, this new Realtime API, they announced that, right? That seems like a big deal if we want AI to sound, well, less like a robot.
It could be huge. The Realtime API could completely change how we, like, interact with AI. Like, imagine if your voice assistant could actually handle it if you interrupted it.
Or, like, have an actual conversation.
Right, not just these clunky back-and-forth things we're used to.
And they actually showed it off, didn't they? I read something about a travel app, one for languages, even one where the AI ordered takeout.
Those demos were really interesting, and I think they show how this Realtime API can be used in so many ways, and the tech behind it is fascinating, by the way. It uses persistent WebSocket connections and this thing called function calling so it can respond in real time.
So the function calling thing, that sounds kinda complicated. Can you, like, explain how that works?
So imagine giving the AI access to this whole toolbox, right?
Yeah.
Information, capabilities, all sorts of things.
Okay.
So take the travel agent demo, for example. With function calling, the AI can pull up details, let's say, about Fort Mason, right, from some database, like nearby restaurants, stuff like that.
Ah, I get it. So instead of being limited to what it already knows, it can go and find the information it needs like a human travel agent would.
Precisely.
Mm.
And someone on Hacker News pointed out a cool detail. The API actually gives you a text version of what's being said, so you can store that, analyze it.
That's smart. It seems like OpenAI put a lot of thought into making this API easy for developers to use. But while we're on OpenAI, you know, besides of their tech, there's been some news about, like, internal changes too.
Didn't they say they're moving away from being a nonprofit?
They did, and it's got everyone talking. It's a major shift, and it's only natural for people to wonder how that'll change things for OpenAI in the future.
Hmm.
I mean, there are definitely some valid questions about this move to for-profit. Like, will they have more money for research now? Probably. But will they, you know, care as much about making sure AI benefits everyone?
Yeah, that's the big question, especially with all the, like, the leadership changes happening at OpenAI too, right? I read that their chief research officer left, and their VP of research, and even their CTO.
It's true. A lot of people are connecting those departures with the changes in OpenAI's structure.
And I guess it makes you wonder what's going on behind the scenes, but they are still putting out new stuff. Like, this whole fine-tuning thing really caught my eye.
Right, fine-tuning. It's essentially taking a pre-trained AI model and, like, customizing it.
So instead of a general AI, you get one that's tailored for a specific job.
Exactly, and that opens up so many possibilities, especially for businesses. Imagine you could train an AI on your company's data, you know, like how you communicate your brand guidelines.
So it's like having an AI that's specifically trained for your company?
That's the idea.
And they're doing it with images now too, right? Fine-tuning with vision is what they called it.
It's pretty incredible what they're doing with that, especially in fields like medicine.
Like using AI to help doctors make diagnoses.
Exactly. An AI could be trained on, like, thousands of medical images, right?
Yeah.
And then it could potentially spot things that even a trained doctor might miss.
That's kinda scary, to be honest. What if it gets it wrong?
Well, the idea isn't to replace doctors but to give them another tool, you know, help them make better decisions.
Okay, that makes sense, but training these AI models must be really expensive.
It can be. All those tokens add up, but OpenAI announced something called Automatic Prompt Caching.
Automatic what now? I don't think I came across that.
So basically, if your AI sees a prompt that it's already seen before, OpenAI will give you a discount.
Huh. Like a frequent buyer program for AI.
Kind of, yeah. It's good that they're trying to make it more affordable, and they're also doing something called model distillation.
Okay, now you're just using big words to sound smart. What's that?
Think of it like a, like a recipe, right? You can take a really complex recipe and break it down to the essential parts.
Make it simpler, but it still tastes the same.
Yeah, and that's what model distillation is. You take a big, powerful AI model and create a smaller, more efficient version.
So it's like lighter weight but still just as capable.
Exactly, and that means more people can actually use these powerful AI Tools. They don't need, like, a super computer to run them.
So they're making AI more accessible. That's great.
It is. And speaking of powerful tools, they also talked about their new o1 model. That's the one they've been hyping up, the one that's supposed to be this big leap forward.
Yeah, o1. It sounds pretty futuristic. Like, from what I read, it's not just a bigger, better language model.
Right. It's a different approach.
They're saying it can, like, actually reason, right? Think differently.
It's trained differently. They used reinforcement learning with o1.
So it's not just finding patterns in the data it's seen before.
Not just that. It can actually learn from its mistakes, get better at solving problems.
So give me an example. What can o1 do that, say, GPT-4 can't?
Well, OpenAI showed it doing some pretty impressive stuff with math, like advanced math.
Yeah.
And coding too, complex coding, things that even GPT-4 struggled with.
So you're saying if I needed to, like, write a screenplay, I'd stick with GPT-4, but if I wanted to solve some crazy physics problem, o1 is what I'd use?
Something like that, yeah. Although there is a trade-off. o1 takes a lot more power to run, and it takes longer to get those impressive results.
Hmm. Makes sense. More power, more time, higher quality.
Exactly.
It sounds like it's still in development though, right? Is there anything else they're planning to add to it?
Oh, yeah.
Yeah.
They mentioned system prompts, which will let developers, like, set some ground rules for how it behaves, and they're working on adding structured outputs and function calling.
Wait, structured outputs. Didn't we just talk about that?
We did. That's the thing where the AI's output is formatted in a way that's easy to use, like JSON.
Right. Right. So you don't have to spend all day trying to make sense of what it gives you. It's good that they're thinking about that stuff.
It's about making these tools usable. And speaking of that, DevDay finished up with this really interesting talk. Sam Altman, the CEO of OpenAI, and Kevin Weill, their new chief product officer, they talked about, like, the big picture for AI.
Yeah, they did, didn't they? Anything interesting come up?
Well, Altman talked about moving past this whole AGI term, artificial general intelligence.
I can see why. It's kind of a loaded term, isn't it?
He thinks it's become a bit of a buzzword, and people don't really understand what it means.
So are they saying they're not trying to build AGI anymore?
It's more like they're saying they're focused on just making AI better, constantly improving it, not worrying about putting it in a box.
That makes sense. Keep pushing the limits.
Exactly. But they were also very clear about doing it responsibly. They talked a lot about safety and ethics.
Yeah. That's important.
They said they were gonna be very careful about how they release new features.
Good, because this stuff is powerful.
It is. It was a lot to take in, this whole DevDay event, new tools, big changes at OpenAI, and these big questions about the future of AI.
It was, but hopefully this deep dive helped make sense of some of it. At least, that's what we try to do here.
Absolutely.
Thanks for taking the deep dive with us.
The biggest demo of the new real-time API involved function calling with voice mode and buying chocolate-covered strawberries from our friendly local OpenAI developer experience engineer and strawberry shop owner, Ilan Biguio. We'll first play you the audio of his demo and then go into a little interview with him.
Voice API Demo9:10
Fantastic. Uh, could you place a call and see if you could get us 400 strawberries delivered to the venue? But please keep that under $1,500.
I'm on it. We'll get those strawberries delivered for you.
Hello?
Hi there. Is this Ilan? I'm Romance AI Assistant, all about it. Fantastic. Could you tell me what flavors of strawberries did you have available?
Yeah. We have chocolate, vanilla, and we have peanut butter.
Great. How much would 400 chocolate-covered strawberries cost?
400. Are you sure you want 400?
Yes. 400 chocolate-covered strawberries. How much would that be?
I think that'll be around, like, $1,415 with 92 cents.
Awesome. Let's go ahead. Place the order for 400 chocolate-covered strawberries.
Great. Where would you like that delivered?
Uh-oh.
Please deliver them to the Gateway Pavilion at Fort Mason, and I'll be paying in cash.
Okay, sweet. So just to confirm, you want 400 chocolate-covered strawberries to the Gateway Pavilion?
Yes, that's perfect. And when can we expect delivery?
Um, well, you guys are right nearby, so it'll be, like, I don't know, 37 seconds.
That's incredibly fast. Thank you so much. Have a wonderful day.
Cool. You too.
Hi, Ilan. Uh, welcome to Latent Space. Uh-
Thank you.
Just saw your amazing demos, had your amazing strawberries. You are dressed up, like, exactly like a strawberry salesman. Uh
Gotta have the look.
What was the building the demo like? Uh, what was the story behind the demo?
It was really interesting. This is actually something I had been thinking about for months before the launch. Like, having a, like, AI that can make phone calls is something, like, I've personally wanted for a long time. And so as soon as we launched internally, like, I started hacking on it.
Um, and then that sort of just made it into, like, an internal demo, and then people found it really interesting, and then we thought, "How cool would it be to have this, like, on stage as-
Yeah
... as one of the demos?"
Yeah. Would, um, would you call out any technical issues building... Like, you are basically one of the first people ever to build with the voice mode API. Um, would you call any issues, like, integrating it with Twilio like that, like you did with function calling, with, like, the form-filling elements?
I noticed that you had, like- ... intents of things to f- to fulfill, and then, uh, you're, like, when, when you're still missing info, the, the voice would prompt, uh, you, uh, uh, play- role-playing the, the, the store guy.
Yeah, yeah. So I think technically there's, like, the whole just working with audio and streams is a whole different beast. Like, even separate from, like, AI and this, this, like, new capabilities, it's just, it's just tough. Yeah, when you have a prompt, conversationally it'll just follow, like, the, it was set up, like, kinda step by step to, like, ask the right questions, um, based on, like, the, um, like, what the request was, right?
Um, the function calling itself is sort of tangential to that. Like, you have to prompt it to, to call the functions, but then handling it isn't too much different from, like, what you would do with assistant streaming or, like, uh, chat completion streaming.
I think, like, the API feels very similar just to, like, if everything in the API was streaming, it actually feels quite familiar to that.
And then function calling-wise, I mean, it, does it work the same? I, I don't know. Like, I saw a lot of logs.
Yeah.
You guys showed, like, in the playground, a lot of logs. What is in there? What, what should people know?
Yeah, I mean, it is, like, the events may have different names than the streaming events that we have in chat completions, but they represent very similar things.
Okay.
It's things like, you know, function call started, argument started. It's like, here's, like, argument deltas and then, like, function call done. Um, conveniently we send one with a, like, that has the full function, and then I just use that.
Nice.
Yeah.
Um, and then, like, what, uh, restrictions do- should people be aware of? Like, you know, I, I think, I think before we recorded we discussed a little bit about the sensitivities around basically calling random store owners and putting, putting, like, an AI on them.
Yeah. So there's, I think there's recent regulation on that, which is why we wanna be, like, very, I guess, aware of, of, you know, you can't just call anybody with AI, right? That's, like, just robocalling. You wouldn't want someone just calling you with AI.
Yeah, yeah. So, like, I ha- I'm a developer.
Yeah.
I'm about to do this on random people.
Yeah.
Wh- uh, what, what, what laws am I about to break?
I, I forget what the governing body is, but you should def- I think get, having consent of the person you're about to call, it always works, right? I, as the strawberry owner, have consented to, like, getting called with AI.
Yeah.
Um, I think past that you, you wanna be careful. Definitely individuals are more sensitive than businesses. I think businesses you have a little bit more leeway. Also the, like, businesses I think have an incentive to want to receive AI phone calls, um, especially if, like, they're dealing with it-
It's doing business
... with AI phone calls.
Yeah, yeah.
Right? Like, it's more business. Um, it's kinda like getting on a booking platform, right? You're exposed to more. But I think it's still very much, like, a gray area, and so I think everybody should, you know, tread carefully, like, figure out what it is.
I, I, I-- the law is so recent, I didn't have enough time to, like-
Okay
... I'm also not a lawyer.
Yeah, yeah, of course.
But yeah.
Okay, cool. Uh, fair enough. One other thing, this is kind of agentic. Did you use a state machine at all? Did you use any framework? Uh...
Uh, no. No. This was, uh-
J- you, you stick it in context and then just run it in a loop until it ends, ends call?
Yeah. There, there isn't even a, a loop. Like, um, because the API is just based on sessions, it's always just gonna keep going. Every time you speak it'll trigger a call.
Yeah.
And then after every function call it was also invok- invoking, like, a, a generation.
I see.
And so that is another difference here. It's, like, it's inherently almost, like, in a loop be- just by being in a session.
I see.
Right? No state machine's needed. I'd say this is very similar to, like, the notion of routines where it's just, like, a list of steps, uh, and it, like, sticks to them softly, but usually pretty well.
And the steps is the prompts is the-
The st- the steps-
It's like the steps
... it's like the, the prompt, like, the steps are in the prompts.
Yeah, yeah.
Right? It's like step one do this, step one-- step two do th- do that.
What, what if I wanna change the system prompt halfway through the conversation?
You can.
Okay.
You can. To be honest, I have not played with that too, too much-
Yeah, yeah
... but I know you can.
Yeah, yeah.
Yeah.
Awesome. I noticed that you called it Realtime API-
Yeah
... but not voice API.
Mm-hmm.
So I assume that it's, like, Realtime API starting with voice, right? Yeah, I think that's what-
Yes
... that's what he said on-
Yes
... on the thing. I can't imagine, like, what else is realtime?
Well, I guess to use ChatGPT's voice mode as an example, like we've demoed the video, right? Like realtime image, right?
Okay, so-
So I'm not actually sure what timelines are, but I would expect, if I had to guess, that, like, that is probably the next thing that we're gonna be making.
Yeah.
You'd probably have to talk directly with the team building this.
Sure.
But, um-
'Cause you can't promise their timelines.
Yeah, yeah, right. Exactly.
Yeah.
But, like, given that this is the features that currently-- or that exist that we've demoed on ChatGPT-
Yeah
... yeah, that's what I would say.
There, there will never be a case where there's, like, a realtime text API, right? Like, I don't-
Well, this is a realtime text API. You can do text only on this.
Oh.
Yeah. I don't know why you would. Um, but it's actually-- so text-to-text here doesn't quite make a lot of sense.
Yeah.
I don't think you'll get a lot of latency gain. But, like, speech-to-text is really interesting because you can, uh, prevent-- you can prevent responses, like audio responses, and force function calls. And so you can do stuff like UI control that is, like, super, super reliable.
We had a lot of, like, you know, un- like, we weren't sure how well this was gonna work because it's, like, you have a voice answering. It's, like, a whole persona, right? Like, that's a little bit more, you know, risky.
But if you s- like, cut out the audio output-
Yeah
... and make it so it always has to output a function, like you can end up with pretty, pretty reliable, like, commands, uh, like a command architecture.
Yeah. A- actually, that's how the-- that's the way I wanna interact with a lot of these things as well. Like-
Right
... one-sided voice, not-
Yeah
... no forced two-sided on me.
You don't necessarily wanna hear the voice back.
Okay.
And, like, sometimes it's like-- Yeah, I think having an output voice is great, but I feel like I don't always wanna hear an output voice. I'd say usually I don't. But yeah, exactly, being able to speak to it is, is super sweet.
Cool. Uh, do you wanna comment on any of the other stuff that you announced, um- Prompt caching I noticed was like, I, I like the no code change part.
Mm.
Um, I'm looking forward to the docs, 'cause I'm sure there's a lot of details on, like, what you cache, how long you cache.
Yeah.
'Cause like Anthropic caches for like five minutes, and it was like, okay, but what if I don't make a call every five minutes ?
Yeah. To be super honest with you, I've been so caught up with the real-time API and making the, the demo that I haven't read up on the other launches too much. I mean, I'm aware of them, but, um, I think I'm excited to see how all distillation works.
Um, that's something that we've been doing, like, I don't know, I've been, like, doing it between our models for a while, and I've seen really good results. Like I've done back in the day, like from GPT-4 to GPT 3.5, um, and got like all, like pretty much same level of like function calling with like hundreds of functions.
Yeah.
Um, so that was super, super compelling. So I feel like easier distillation, I'm, I'm really excited for.
I see. Is it a tool... Uh, so I saw evals.
Yeah.
Like, what, what is the distillation product? It wasn't super clear, to be honest.
I, I think I wanna-
You can-
I wanna let that team-
Yeah
... I wanna let that team talk about it.
I can let the team come back.
Yeah.
Okay. Great. Well, I appreciate you jumping on and-
No, of course
... uh, amazing demo. It was beautifully designed, so I'm sure that was, uh, part of you and Romain and-
Yeah, I guess fir- shout out to like the first-
Shout out
... people to like, creators of Wanderlust originally were like Simon and Karolis, and then, like, I took it and built the voice component and the voice calling components.
Ooh.
Yeah. So it's been a, it's been a big team effort, and then like the entire PI team for like debugging everything as it's been going on. It's been, it's been super rewarding.
Platform Strategy19:16
Yeah, you were the first consumers on the UX team.
Yeah. Yeah.
Yeah, I mean, the classic role of, of what we do there.
Yeah.
Okay. Yeah. A- anything else? Any other call to action?
No, enjoy DevDay.
Thank you.
Yeah.
That's it.
The Latent Space crew then talked to Olivier Godement, head of product for the OpenAI platform, who led the entire DevDay keynote and introduced all the major new features and updates that we talked about today.
Okay. So we are here with Olivier Godement. It's I, I, I don't-
That's right
... I don't pronounce French.
That's fine. It was perfect.
And, uh, it was amazing to see your keynote today. Uh, what was the backstory of, of, uh, preparing something like this, preparing, like, DevDay?
Essentially came from a couple of places. Uh, number one, excellent reception from last year's DevDay. Developers, startup founders, researchers wanna spend more time with OpenAI, and we wanna spend more time with them as well. And so for us, like it was a no-brainer, frankly, to do it again, like, you know, like a nice conference.
The second thing is, um, going global. Uh, we've done a few events like in Paris and like a few other, like, you know, non-European and non-American countries. Uh, and so this year we're doing SF, Singapore, and London to frankly just meet more developers.
Yeah. I'm very excited for the Singapore one.
Oh, yeah.
Uh, you, you-
Will you be there?
Uh, I don't know. I don't know if I got an invite.
No.
Actually, I can, I can just ask you.
We can put it up.
Yeah, yeah. Uh, yeah, like and then, uh, there was some speculation around October 1st.
Yeah.
Is it because o1, October 1st?
Has nothing to do. I discovered a tweet yesterday where like, people are so creative. No, o1 has no connection to October 1st, but in hindsight, in insight like that, that would've been a pretty-
Makes sense
... a pretty good meme, but yeah, no.
Okay. Yeah, and, you know, I think like OpenAI's outreach to developers is something that, uh, I felt the whole, um, in 2022 when like, you know, like people were trying to build with ChatGPT and like there was no function calling, all that stuff that you, that you talked about in the past, and that's, that's why I started my own conference as like, um, like here's our little developer conference thing.
Yeah.
Uh, and but to see this OpenAI DevDay now and like to see so many developer oriented products coming out of OpenAI, I think it's really encouraging.
Yeah, totally. It's, uh, that's what I said essentially, like pretty cur-... Like developers are basically the people who make the best connection between the technology and, you know, the future essentially. Like, you know, essentially see a capability, see a low level like technology and are like, "Hey, I see how that application or that use case, like can be enabled."
And so in the direction of enabling like AGI like for all of humanity, it's a no-brainer for us I think to partner with devs.
Yeah.
And most importantly, you almost never had wait lists, which compared to like other releases people usually, usually have. Um, what is the, you know, you had prompt caching, you had real-time voice API. We, you know, Sean did a long Twitter thread so people know the releases.
Yeah.
What is the thing that was like sneakily the hardest to actually get ready for, for that day? Or like what was the kinda like, you know, last 24 hours anything that you didn't know was gonna work?
Yeah. They- they're all fairly, like I would say involved, like features to ship. Um, so the team has been working for month, all of them. The one which I would say is the newest for OpenAI is the real-time API, for a couple of reasons.
I mean, one, you know, it's a new modality. Uh, second, like it's the first time that we have an actual like WebSocket based API. And so I would say that's the one that required like the most work over the month to get right from a developer perspective, and to also make sure that our existing safety mitigations like worked well with like real-time audio in and audio out.
Yeah. What, uh, design choices or what's like the sort of design, design choices that you wanna highlight? Like, you know, like, um, I think for, for me, like WebSockets, you just receive a bunch of events. It's two-way.
Yeah.
Um, I obviously don't have a ton of experience.
Yeah.
I think a lot of developers are going to have to em- embrace this real-time programming.
Yeah.
Uh, like what are you designing for, or like what, what advice would you have for developers exploring this?
Uh, the, the core design hypothesis was essentially how do we enable like human level latency? We did a bunch of tests, like on average, like human beings, like, you know, takes like something like 1,500 milliseconds to converse with each other.
Yes.
And so that was the design principle essentially, like working backward from that and, you know, making the technology work. And so we evaluated a few options, and WebSockets was the one that we landed on. So that was like one design choice.
Uh, a few other like big design choices that we had to make, uh, prompt caching. Um, prompt caching the design like target was automated from the get-go, like zero code change from the developer-
That way you don't have to learn, like, what is a prompt prefix and, you know, how long does a cache work. Like, we just do it as much as we can essentially. So that was a big, um, design choice as well.
And then finally on distillation, like an evaluation, the big design choice was something I loved as Hype, like in my previous job, like a philosophy around like a pit of success. Like, what is essentially the, the, the minimum number of steps for the majority of developer to do the right thing?
Because when you do evals on fine-tuning, there are many, many ways, like, to mess it up frankly, like, you know, and have like a crappy model, like evals that tell like a wrong story. And so our whole design was, okay, we actually care about like helping people who don't have, like, that much experience, like evaluating model, like get like in a few minutes, like, to a good spot, and so how do we essentially enable that pit of success, like in the product flow?
Yeah. Yeah. I, I'm a little bit, um, scared to fine-tune.
Yeah.
Uh, especially for vision, because I don't know what I don't know for, for, for stuff like vision, right? Like for, for text I can evaluate pretty easily.
Yeah.
For vision, um, let's say I'm like trying to... Y- One of your examples was Grab.
Yeah.
Which, very close to home.
Yeah.
I'm from Singapore. Uh, the o- I think your example was like they identified stop signs better.
Yeah.
Why is that hard? Why do I have to fine-tune that? Uh, if I fine-tune that, do I lose other things? You know, like there's a lot of unknowns with, with vision that I, that I think developers have to figure out.
For sure. Vision is going to open up like a new, I would say, evaluation space, because you're right, like it's harder, like, you know, to tell correct from incorrect essentially with images. Um, what I can say is, uh, we've been alpha testing like the vision fine-tuning like for several weeks at that point.
We are seeing like even higher performance uplift compared to text fine-tuning. So that's... There is something here, like we've been pretty impressed, like, in a good way frankly, but you know how well it works. But for sure, like, you know, I expect the developers who are moving from one modality to like text and images will have like more testing evaluation, like, you know, to set in place, like to make sure it works well.
Um, the Model Distillation and Evals-
Yeah
... is definitely like the most interesting, moving away from just being a model provider-
Yeah, yeah
... to being a platform provider.
Yeah, yeah.
How should people think about being the source of tru- Like, do you want OpenAI to be like the system of record o- of all the prompting? Because people sometimes store it in like different data sources.
Yeah.
And then is that gonna be the same as the models evolve so you don't have to worry about, you know, refactoring the data-
Yeah
... or like things like that for like future model structures?
The vision is, um, if you want to be a source of truth, you have to earn it, right? Like, we're not going to force people like to pass us data. There is no value prop, like, you know, for us to store the data.
The vision here is at the moment, like most developers, like you, use like a one-size-fits-all model, like the off-the-shelf, like GP4O essentially. The vision we have is fast forward a couple of years, I think like most developers will essentially like have a, an automated continuous fine-tune model.
The more like you use the model, the more data you pass through the model provider, like the model is automatically like fine-tuned, evaluated against some eval sets, and essentially like you don't have to, every month when there is a new snapshot, like, you know, to go online and, you know, try a few new things.
That's the direction. We are pretty far away from it. Um, but I think like that evaluation and decision product are essentially a first good step in that direction. It's like, hey, if you are excited by the direction and you give us like evaluation data, we can actually log your completion data and start to do some automate- automation on behalf.
Um, and then you can do evals for free if you share data-
Yeah
... with OpenAI.
Yeah.
How should people think about when it's worth it, when it's not? Sometimes people get overly protective of their data-
Yeah
... when it's actually not that useful.
Yeah.
But h- how should developers think about when it's right to do it or when not, or if you have any thoughts on it?
The default policy is still the same. Like, you know, we don't train on like any API data unless you opt in. Um, what we've seen from feedback is, uh, evaluation can be expensive. Like if you run like o1 evals on like thousands of samples, like your bill will get increased like, you know, pretty, um, pretty significantly.
That's problem statement number one. Problem statement number two is essentially I want to get to a world where whenever OpenAI ships a new model snapshot, we have full confidence that there is no regression for the tasks that developers care about.
And for that to be the case, essentially, we need to get evals. And so that essentially is a sort of, um, a two birds, one stone, is like we subsidize basically the evals, and we also use the evals when y- we ship like new models to make sure that, you know, we keep going in the right direction.
So in my sense, a win-win, but again, like completely opt-in. I expect that, you know, many developers will not want to share their data, and that's perfectly fine to me.
Yeah. I think free evals though, very, very in- uh, good incentive.
Yeah.
Uh, I mean, it's a fair trade. Like you get data, we get free evals.
Exactly. And we sanitize PII, everything. You know, like we have no interest in like the actual like sensitive data. We just wanna have good evaluation on real use cases.
Like I, I almost want to eval the eval.
Yeah.
I don't know if that, that ever came up. Like sometimes the evals themselves are wrong.
Yeah.
And then h- uh, then, then there's no way for me to tell you.
Everyone who is starting with LLM, tinkering with LLM is like, "Yeah, evaluation easy, you know. I've done testing like all my life." And then you start to actually build evals, understand like, you know, all the corner cases, and you realize, wow, that's like a whole field in itself.
So yeah, eva- good evaluation is hard. And so, yeah.
Yeah, yeah. Uh, but I think th- there's a, you know, uh, I just talked to Braintrust, uh, which I think is one of your partners.
Mm-hmm.
Um, they also emphasize code-based evals-
Yeah
... versus your sort of low code. W- what I see is like, I don't know, maybe there's some more that, that you didn't demo, but what I see is kind of like a low-code experience, right-
Yeah
... for evals. Uh, would you ever support like a more code base? Like would I, would I run code on OpenAI's eval platform?
For sure. I mean, we meet developers where they are. You know, at the moment the demand was more for like, you know, easy to get started, like eval.
That was good.
But, you know, if we need to expose like an evaluation API, for instance, for people like, you know, to pass like, you know, their existing test data, um, we'll do it. So yeah, there is no, you know, philosophical, I would say like, you know, misalignment on that.
Yeah, yeah. What I, what I think this is becoming, by the way, and, uh- I don't... Like, uh, it, it's basically like you're becoming AWS, like the AI cloud.
Yeah.
And I, I don't know if, like, that's a conscious strategy or it's like... Like, it doesn't even have to be a conscious strategy. Like, you're gonna offer storage.
Yeah.
You're gonna offer compute. You're gonna offer, um, networking. I, I don't know what networking looks like. Networking is maybe like caching or- Like, it's a CDN. It's a prompt-
Yeah
... CDN.
Yeah.
Uh, but it's the AI versions of everything, right? Do you... Like, do you see the analogies or?
Yeah, totally.
Yeah.
Um, whenever, whenever I talk to developers, I feel like good models are just half of the story to build a good app. There's a ton more you need to do. Evaluation is the perfect example. Like, you know, you can have the best model in the world, if you're in the dark, like, you know, it's really hard to gain the confidence.
And so our philosophy is, um...
That's number one. Number two, like the whole, like, software development stack is being basically reinvented, you know, with LLMs. There is no freaking way that OpenAI can build everything. Like, there is just too much to build, frankly. And so my philosophy is essentially we'll focus on, like, the tools which are, like, the closest to the model itself.
So that's why you see us, like, you know, investing quite a bit in, like, fine-tuning, distillation, now evaluation, because we think that it definitely makes sense to have like in one spot, like, you know, all of that. Like, there is some sort of virtual circle essentially that you can set in place.
But stuff like, you know, LLM ops, like tools which are, like, further away from the model, I don't know. Um, if you want to do like, you know, super elaborate, like prompt management or, you know, like tooling, like I'm not sure, like, you know, OpenAI has, like, such a big edge frankly, like, you know, to build that sort of tools.
So that's how we view it at the moment. But again, frankly, like the philosophy is, like, super simple. The strategy is super simple. It's meeting developers where they want us to be. And so, you know, um, that's frankly like, you know, day in, day out, like you know, what I try to do.
Um, cool. Thank you so much for the time. I'm sure you got- Yeah, I, I have more questions on-
Yeah
... a, a, a couple of questions on voice and then also, like your, your call to action, like what you want feedback on, right?
Yeah.
So but, uh, let... I think, I think we should spend a bit more time on voice-
Yeah
... because I feel like that's like the, the big-
Oh, got it
... splash-
Yeah
... thing. Um, I talked to, uh... Well, well, I mean, I mean, just, uh, like what is the future of real-time for OpenAI?
Yeah.
Um, because I, I think obviously video is next. You already have it-
Yeah
... in the, the ChatGPT desktop app. Um, do we just, like, have a permanent li- like, you know, like are developers just going to be, like, s-sending sockets back and forth with OpenAI? Like, uh, how do we program for that?
Like, what, uh, what is the future?
Yeah, that makes sense. Uh, I think with multimodality, like real-time is quickly becoming like, you know, essentially the right experience, like to build an application. Uh, so my expectation is that we'll see like a non-trivial, like, uh, volume of applications, like moving to a real-time API.
Um, like if you zoom out, like audio is a really simple... Like audio, until basically now, audio on the web in apps was basically very much like a second-class citizen.
Mm.
Like you basically did like an audio chatbot for users who did not have a choice. You know, they were, like, struggling to read or, I don't know, they were, like, not super educated with technology.
Yeah.
And so frankly, it was like the crappy option, you know, compared to text. But when you talk to people in the real world, the vast majority of people, like prefer to talk and listen instead of typing and writing.
We speak before we write.
Exactly. I don't know, I mean, I'm sure it's the case for you in Singapore. For me, my friends in Europe, the number of like WhatsApp, like voice notes-
Same here
... that you receive-
Yeah
... every day, I mean, just people-
And-
It makes sense frankly, like, you know-
Chinese
... Chinese, yeah.
Yeah. All voice.
You know, it's easier. There is more emotions. I mean, you know, you get the point across like pretty well. Um, and so my personal ambition for like the real-time API and like audio in general is to make like audio and like multimodality like truly a first-class experience.
Like, you know, if you're like, you know, the amazing, like super bold, like startup out of YC, you want to build like the next like billion, like you know, user application, to make it like truly audio first and make it feel like, you know, an actual good, like, you know, product experience.
So that is essentially the ambition, and I think like, yeah, it could be pretty big.
Yeah. I think one, one people-- One issue that people have with the voice so far as, as released in the advanced voice mode is the refusals.
Yeah.
Uh, you guys had a very inspiring model spec, I think Joanne-
Yeah
... worked on that, um, where you said like, "Yeah, we don't want to overly refuse-
Yeah
... all the time." In fact, like even if like not safe for work, like-
Yeah
... in some occasions, oh, it's okay.
Yeah.
Uh, how... Is there an API that we can say not safe for work okay, like
I think, I think we'll get there. I think we'll get there. Uh, the model spec like nailed it, like, you know-
It nailed it.
We are not-
It was so good.
Yeah. We are not in the business of like policing, you know, if you can say like vulgar words or whatever. You know, there are some use cases, like, you know, I'm writing like a Hollywood, like script, I want to say like vulgar words-
Right
... and it's perfectly fine, you know? And so I think the direction where we'll go here is that basically there will always be like, you know, a set of behavior that we will for- you know, just like forbid, frankly, because they're, they're illegal against our terms of services.
But then there will be like, you know, some more like risky, like themes which are completely legal, like, you know, vulgar words or, you know, not safe for work stuff, um, where basically we'll expose like a controllable, like safety, like knobs in the API to basically allow you to say, "Hey, that theme okay, that theme not okay."
How sensitive do you want the threshold to be on safety refusals? I think that's the direction.
So a safety API?
Yeah. In a way, yeah.
Yeah. We have-- We've never had that.
Yeah.
Because right now it's you, it's whatever you decide, and then it's, that's it. That, that, that would be the main reason I don't use OpenAI Voice-
Yeah
... is because of-
It's overpoliced at the moment
... overrefus-
Yeah
... overrefusals.
Yeah, yeah, yeah. No, we gotta fix that.
Yeah.
Like singing. We're trying to do a voice karaoke.
So I'm a singer, uh, and you, you locked off singing.
Yeah, yeah, yeah.
But I, I understand music gets you in trouble. Uh, uh, okay. Uh, okay, yeah, so then, and then just generally, like what do you want to hear from developers, right? We have a, we have all developers watching. Um, y- you know, what feedback do you want?
Um, any- anything specific as well, like from-
Sure
... especially from today? Um, anything that you are unsure about, that you're like y- our feedback could really help you decide
For sure. I think essentially it's becoming pretty clear after today that, you know, the, I would say the OpenAI direction's become pretty clear, like, you know, after today. Investment in reasoning, investment in multimodality, investment as well, like in, I would say, tool use, like function calling.
Um, to me, the biggest question I have is, you know, where should we put the cursor next? I think we need all three of them, frankly, like, you know, so we'll keep pushing-
Hire 10,000 people. Or actually, no, no need. Build a, build a bunch of bots.
Exactly. And so, um, let's take o1 for instance, like is o1 smart enough, like for your problems? Like, you know, let's set it sec- aside for a second the existing models, like for the apps that you would love to build, is o1 basically it in reasoning or do we still have like, you know, a step to do?
DX Live Demos36:57
Preview is not enough.
Yeah.
I need, I need the full one.
Yeah. So that's exactly the sort of feedback. Essentially what I would love to do is for developers... I mean, there's a thing that Sam, which has been saying like over and over again, like, you know, it's easier said than done, but I think it's directionally correct.
As a developer, as a founder, you basically want to build an app which is a bit too difficult for the model today.
At least.
Right? Like what you think is right, it's like sort of working, sometime not working, and that way, you know, that basically gives us like a goalpost and be like, "Okay, that's what you need to enable with the next model release, like in a few months."
Um, and so I would say that's usually like that's the sort of feedback which is like the most useful-
Please enjoy a short wait.
That I can like directly, like, you know, incorporate.
Yeah. Awesome. I, I think that's our time.
Thank you so much, guys.
Yeah. Thank you so much.
Yeah. Thank you.
We were particularly impressed that Olivier addressed the not-safe-for-work moderation policy question head on, as that had only previously been picked up on in Reddit forums. This is an encouraging sign that we will return to in the closing candor with Sam Altman at the end of this episode.
Next, a chat with Romain Huet, friend of the pod, AI engineer world's fair closing keynote speaker, and head of developer experience at OpenAI on his incredible live demos and advice to AI engineers on all the new modalities.
All right. We're live from OpenAI DevDay with Romain, who just did two great demos on, on stage, and he's been a friend of Latent Space, so thanks for taking some of the time.
Of course, yeah. Thank you for being here and spending the time with us today.
Yeah. Appreciate. Uh, appreciate you guys putting this on. I, I know it's like extra work, but it really shows the developers that you care and about reaching out.
Yeah, of course. I think when you go back to the OpenAI mission, I think for us it's super important that we have the developers involved in everything we do, um, making sure that, uh, you know, they have all of the tools they need to build successful apps.
And we really believe that the developers are always gonna invent the, the ideas, the prototypes, the form factors of AI that we can't build ourselves. So it's really cool to have everyone here.
We had Michelle from the po- uh, from, from you guys on.
Yes. Great episode.
And she, she, she very... Thank you. Uh, and she very seriously said API is the path to AGI.
Correct.
And people in our YouTube comments were like, "API is not AGI." I'm like, "No, like she, she's very serious. API is the path to AGI." Because like you're not gonna build everything like the developers are, right?
Of course. Yeah. That's the, that's the whole value of having a platform and an ecosystem of amazing builders who can-
Yeah
... like in turn create all of these apps. Um, I, I'm sure we talked about this before, but there's now more than 3 million developers building on OpenAI, so it's pretty exciting to see all of that energy, uh, into creating new things.
Yeah.
Um-
Please board when your group is called
I, I was gonna say, you built two apps on stage today.
Yes.
An International Space Station tracker and then a drone. The hardest thing must have been opening Xcode and setting that up. Now, like the models are so good that they can do everything else.
Yes.
You had two modes of interaction. You had kind of like ChatGPT app-
Yep
... to get the plan with o1, and then you had, um, Cursor to do apply some of the changes.
Correct.
How should people think about the best way to consume the coding models especially, both for, you know, brand-new projects and then existing projects that they're trying to modify instead?
Yeah. I mean, one of the things that's really cool about o1 preview and o1 mini being available in the API is that you can use it in your favorite tools like Cursor like I did, right? And that's also what like Devin from Cognition can use-
Yeah
... in their own software engineer eng- agent. Uh, in the case of Xcode, like it's not quite deeply integrated in Xcode, so that's why I had like ChatGPT-
Yeah
... side by side.
Copy paste.
Um, but it's cool, right? Because I could instruct, uh, o1 preview to be like my coding partner and brainstorming partner for this app-
Yeah
... but also consolidate all of the, the files and architect the app the way I wanted. So all I had to do was just like port the code over to Xcode and zero shot the app built. Um, I don't think I conveyed by the way how big a deal that is, but like you can now create an iPhone app from scratch describing a lot of intricate details that you want, and your vision comes to life in like a minute.
Yeah.
It's pretty outstanding.
I, I, I have to admit, I was a bit skeptical because if I open up Xcode, I don't know anything about-
Yeah
... I- iOS programming. You know which file to paste it in. You probably set it up a little bit. So I'm like, I have to go home and test it to like figure out, and I need the ChatGPT desktop app so that it can tell me where to click.
Yeah. I mean, like Xcode and, and iOS development has become easier over the years since they introduced Swift and SwiftUI. I think back in the e- the days of Objective C or like, uh, you know, the storyboard, it was a bit harder to get in for someone new.
But now with Swift and SwiftUI, their dev tools are really exceptional. But now when you combine that with o1 as your brainstorming and coding partner, it's like you architect effectively. That's the best way I think to describe o1.
People ask me like, "But can GPT-4 do some of that?" And it certainly can, but I think it would just start, um, spitting out code, right? And I think what's great about o1 is that it can like make up a plan.
In this case, for instance, the iOS app had to fetch data from an API. It had to look at the docs. It had to look at like how do I parse this JSON? Where do I store this thing?
Um, and kind of wire things up together. So that's where it really shines.
Yeah.
Is mini or preview the better model that people should be using? Like how-
Oh.
Yeah.
I think people should try both. Uh, we are obviously very excited about the upcoming o1 that we shared the evals for. Uh, but we noticed that o1 mini is very, very good at everything math, coding, everything STEM. Uh, if you need for your kind of brainstorming or your kind of, uh, science part, you need s- some broader knowledge, then reaching for o1 preview's better.
Uh, but yeah, I used o1 mini for my second demo-
And it worked perfectly. Um, all I needed was very much like something rooted in code, architecting and wiring up like a front end, a back end, some UDP packets, some WebSockets, something very specific, and it did that perfectly
Yeah. Uh, and then maybe just talking about voice and Wanderlust-
Yeah
... the app that keeps on giving.
It d- it does indeed, yeah.
Uh, what's the backstory behind, like, preparing for all of that?
You know, it's funny 'cause when last year for DevDay, we were trying to think about what could be a, a great demo app to show, like, an assistive experience. I've always thought travel is a kind of, uh, a great use case 'cause you have, like, pictures, you have locations, you have the need for translations potentially.
There's, like, so many use cases that are bounded to travel that I thought last year, "Let's use a travel app," and that's how Wanderlust came to be. But of course, a year ago, all we had was a text-based assistant, and now we thought, "Well, if there's a voice modality, what if we just bring this app back-
Yeah
... as a wink, and what if we were interacting better with voice?" Um, and so with this new demo, what I showed was the ability to, like, have a complete conversation in real time with the- with the app.
Uh, but also the thing we wanted to highlight was the ability to call tools and functions, right? So, uh, you... Like in this case, we- we placed a phone call using the Twilio API interfacing with our AI agent.
But, uh, developers are so smart that they'll come up with so many great ideas that we could not think ourselves, right? But, uh, what if you could have like, um, you know, a 911 dispatcher? What if you could have like a customer service, like, uh, center that is much smarter than what we've been used to today?
Uh, there's gonna be so many use cases for Realtime. It's awesome.
Yeah. And- and sometimes actually you- you... Like, they should kill phone trees. Like, there should not be-
Yeah, of course
... like dial one-
Yeah, of course
...
Para num-
Para español
You know? Yeah, exactly
...
Cero
... dos or whatever. I don't know.
I mean, even you starting speaking Spanish would just do the thing.
You should just-
You know?
Yeah.
Uh, you don't even have to ask.
Yeah.
So yeah, I'm excited for this future where we don't have to interact with those legacy systems.
Yeah, yeah. Um, is there anything... So, uh, you're doing function calling in a streaming environment. So basically it's- it's WebSockets, it's UDP I think. Uh, it's basically not guaranteed to be exactly once delivery. Like, is there any coding challenges that you- you encountered when building this?
Yeah. It's a bit more delicate, uh, to get into it. Um, we also think that for now what we- what we shipped is a- is a beta of this API. I think there's much more to build onto it.
Um, it does have the function calling and the tools, uh, but we think that, for instance, if you wanna have something very robust on your client side, maybe you wanna have WebRTC as a client, right? And- and as opposed to like directly working with the sockets-
Okay
... at scale. Uh, so that's why we have partners like LiveKit and Agora if you wanna w- if you wanna use them, and I'm sure we'll have many mores in the in- ma- many more in the future. Um, but yeah, we keep on iterating on that, and I'm sure the feedback of developers in the weeks to come is gonna be super critical for us to get it right.
Yeah. I think LiveKit has been fairly public that they are used in- in the ChatGPT app. Um, like is it- this- it's just all open source, and we just use it directly with o- uh, OpenAI, or do we use LiveKit Cloud or something?
So right now we- we released the API. We released some sample code also and reference clients for people to get started with our API.
Yeah.
And we also, um, uh, partnered with LiveKit and Agora, so they also have their own, like, uh, ways to help you get started w- that plugs natively with the Realtime API. So depending on the use case, people can- can, uh, can decide what to use.
Uh, if you're working on something that's completely client, uh, or if you're working something on the server side for the voice interaction, you may have different needs, so we wanna support all of those.
Um, I know you gotta run. Is there anything that you want the AI engineering community to give feedback on specifically, like even down to like, you know, a specific API endpoint? Or like, uh, what- what's like the thing that you want-
Sure.
Yeah.
Yeah.
I mean, you know, if we take a step back, I think, uh, DevDay this year is all different from last year and- and in- in a few different ways. But one way is that we wanted to keep it intimate, even more intimate than last year.
We wanted to make sure that the community, uh, is, uh, on the spotlight. That's why we have community talks and everything. And, uh, the takeaway here is like learning from the very best developers and AI engineers. And so, you know, uh, we wanna learn from them.
Most of what we shipped this morning, including things like prompt caching, uh, the ability to generate prompts quickly in the playground, or even things like vision fine-tuning, these are all things that developers have been asking of us. And so the takeaway I would- I would, uh, leave them with is to say like, "Hey, the roadmap that we're working on is heavily influenced by them and their work."
And so we love feedback, uh, from high feature requests, as you said, down to like very intricate details of an API endpoint. We love feedback. So, uh, yes, uh, that's- that's how we- that's how we build this API.
Yeah. I think the- the model distillation thing as well, it's... It might be like the- the most boring, but like actually used a lot.
True, yeah. And I think maybe the most unexpected, right? Because I think if I- if I read Twitter correctly the past few days, a lot of people were expecting us to ship the Realtime API for speech-to-speech. I don't think developers were expecting us to have more tools for distillation, and we really think that's gonna be a big deal, right?
If you're building apps that have, um, you know, you- you want high like, uh, li- like low latency, low cost, but high performance, high quality on the use case, distillation is gonna be amazing.
Yeah. I sat in the distillation, uh, session just now, and they showed how they distilled from 4o to 4 Mini, and, uh, it was like only like a 2% hit in the performance and 15x cheaper.
Yeah. Yeah. I was there as well for the Superhuman kind of use case, uh, inspired for an input client. Yeah, this was really good.
Um, cool, man.
Amazing. Thank you so much.
Thanks for joining DevDay.
Thanks again for being here today.
Yeah.
It's always great to have you.
As you might have picked up at the end of that chat, there were many sessions throughout the day focused on specific new capabilities, like the new Model Distillation features combining EVOs and fine-tuning. For our next session, we are delighted to bring back two former guests of the pod, which is something listeners have been greatly enjoying in our second year of doing the Latent Space Podcast.
Michelle Pokrass of the API team joined us recently to talk about structured outputs, and today gave an updated long form session at DevDay describing the implementation details of the new structured output mode. We also got her updated thoughts on the voice mode API we discussed in her episode now that it is finally announced.
API Deep Dive49:22
She is joined by friend of the pod and super blogger Simon Willison, who also came back as guest co-host in our DevDay 2023 episode. Great. We're back, live at DevDay, uh, returning guest Michelle, uh, and then returning guest co-host, fourth?
Fourth-
Um-
... or fifth or s- yeah, I don't know
I, I've lost count. Yeah.
I've lost count.
It's been a few, but-
Simon Willison is back. Um, yeah, we just wrap- we just wrapped everything up. Congrats on, on getting everything, uh, everything live. Simon did a great live blog, so if you haven't caught up.
I actually, I wrote my, I implemented my live blog- ... while waiting for the first talk to start using, like, uh, G- GPT-4. It wrote me the JavaScript. But I got that live just in time, and then, yeah, I was live blogging the whole day.
Are you a Cursor enjoyer?
Uh, I haven't really gotten to Cursor yet, to be honest. Like, I just haven't spent enough time for it to click, I think.
Yeah.
I'm more copy and paste things out to Claude and ChatGPT.
Yeah. It's interesting.
Yeah.
Uh, yeah, I've-
I've-
Yeah, I've converted to Cursor for, and o1 is so easy to just toggle on, on and off.
Yeah.
Yeah.
What, what's your workflow? Copy, paste, or-
Okay, I'm gonna be real. I'm still VS Code Copilot, so-
Yep, same here
... uh-
Same, same Copilot
... Copilot is actually the reason I joined OpenAI. It was, you know, before ChatGPT, this is the thing that really got me, so I'm still into it. But I keep meaning to try out Cursor, and I think now that things have calmed down, I'm gonna give it a real go.
Yeah, it's-
Yeah
... it's a big thing to change your tool of choice.
Yes. Yeah. I'm pretty dialed, so.
Yeah. I mean, you know, if you want, you can just fork VS Code and make your own. That's, that's the thing to do.
It becomes a thing, right? Yeah.
You joked about doing a hackathon where you o- the only thing you do is fork VS Code, and bet may the best fork win.
Nice.
That's actually a really good idea.
Uh, yeah, so, um, I mean, congrats on launching everything today. Uh, I know, like, we touched on it a little bit, but, like, everyone was kinda guessing that Voice API was coming, and, like, we t- we talked about it in our, in our episode.
How do you, how do you feel going into the, the, the launch? Um, like any design decisions that you wanna highlight?
Yeah. Super jazzed about it. The team has been working on it for a while. It's, like, a very different API for us. This is the first WebSocket API, so a lot of different design decisions to be made, like what kind of events do you send?
When do you send an event? What are the event names? What do you send, like, on connection versus on future messages? So there've been a lot of interesting decisions there. The team has also hacked together really cool projects as we've been testing it.
One that I really liked is we had an internal hackathon for the API team, and some folks built, uh, like, a little hack that you could use, uh, Vim with, uh, voice mode, so, like, Control Vim, and you would tell the model, like-
Nice
... write a file, and it would, you know, know all the Vim commands and, and type those in. So yeah, a lot of cool stuff we've been hacking on. I'm really excited to see what people build with it.
I've gotta call out a demo from today.
Yeah.
I think it was Katya had a 3D visualization of the solar system, like WebGL solar system-
Incredible
... you could talk to. That is one of the coolest conference demos I've ever seen.
Thank you.
That was so convincing. I really want the code. I really want the code for that-
Yeah
... to get put out there. But I'll-
I'll talk to the team. I think we can probably put it up
... absolutely beautiful example, and it made me realize that the real-time API, this WebSocket API, it means that building a website that you can just talk to is easy now. It's like it's not difficult to build, spin up a web app where you have a conversation with it.
It calls functions for different things. It interacts with what's on the screen. I'm so excited about that. There are all of these projects I thought I'd never get to, and now I'm like, "You know what? Spend a weekend on it, I could have a talk to your data, t- talk to your database with a web, with a, with a little web application."
Yeah.
That's so cool.
Chat with PDF, but really chat with it.
Yeah, but really chat with PDF.
Yeah, exactly.
No, completely.
Totally.
And it's not even hard to build. That's the crazy thing about this.
Yeah. Very cool. Yeah, when I first saw the space demo, I was actually just wowed. Um, and I, and I had a similar moment, I think, to all the people in the crowd. Uh, I also thought Romain's, uh, drone demo was super cool.
Um-
That was a super fun one as well.
Yeah.
That was great-
I actually saw that live this morning, and I was holding my breath for sure.
Knowing Romain, he probably spent the last two days do- working on it.
Um, but yeah, like I- I'm curious about, uh, you, you were talking with Romain actually earlier about, um, what the different levels of abstraction are with WebSockets. It's something that most developers have zero experience with. I have zero, zero experience with it.
Uh, apparently there's, like, the RTC level, and then there's the WebSocket level, and there's, like, levels in between. Like what-
Not so much. I mean, with WebSockets, um, with the, the way they've built their API, you can connect directly to the OpenAI WebSocket from your browser.
Yeah.
And it's actually just regular JavaScript. Like, you instantiate the WebSocket thing. It, it looks quite easy from their example code. The problem is that if you do that, you're sending your API key from, like, source code that anyone can view.
So-
Yeah, we don't recommend that for production
... so it doesn't work for, for, work for production, which is frustrating because it means that you have to build a, a proxy.
Yeah.
So I'm gonna have to go home and build myself a little WebSocket proxy just to hide my API key. I want OpenAI to do that. I want OpenAI-
Yeah
... to solve that problem for me so I don't have to build the 1,000th WebSocket proxy just for that one problem.
Totally. We've also partnered with some, uh, some partner solutions. Uh, we've partnered with, I think, Agora, um-
Livekit
... uh, Livekit, uh, a few others. So there's some loose solutions there, but, but yeah, we hear you.
Yeah.
It's a beta.
Yeah. Yeah, I mean, um, you, you still want a solution where someone brings their own key, and they can trust that you don't get it, right?
Kind of. I mean, I've been building a lot of bring your own key apps-
Yeah
... where it's my HTML and JavaScript. I store the key in local storage in their browser-
Yeah
... and it never goes anywhere near my server, which works, but how do they trust me? How do they know I'm not gonna ship another piece of JavaScript that steals the key from the
And, and so nominally, this actually comes with a crypto background. This is what MetaMask does. Um, where- where you-
Yeah, it's a public private key thing. Um-
Yeah
Yeah.
Um, like, why doesn't OpenAI do that? I- I don't know if obviously it's-
I mean, as with most things, you'd think there's like-
Prioritization
... some really interesting question and a really interesting reason, and the answer is just, you know, it's not been the top priority, and-
Yeah
... and it's hard for- for a small team to do everything. I have been hearing a lot more about the need for things like, uh, sign in with OpenAI.
That's what... I want OAuth.
Yeah.
I want to bounce my users through ChatGPT, and I get back a token that lets me spend up to $4 on the API-
Yeah
... on their behalf.
Right.
That would solve it. Then I could ship all of my stupid little experiments, which currently require p to cop- people to copy and paste their API key in-
Right
... which cuts off everyone, right?
Yeah.
No- nobody knows how to do that.
Totally. I hear you. Something we're- we're thinking about, and yeah, stay tuned.
Yeah, yeah. Um, well, right now, the- the... I think the only player in town is Open Router, um, that- that is basically... It's funny, like, it was made by, um, well, I forget his name. Um, but he- he used to be CTO of OpenSea, and he- the first thing he did when he came over was build MetaMask for AI.
Totally. Just... Yeah, very cool.
And-
Um, what's the most underrated release from today?
Vision Finetuning. Vision Finetuning is so underrated. For the past, like, two months, whenever I talk to founders, they tell me this is the thing they need most. A lot of people are doing, like, OCR on- on very bespoke formats, like government documents, and- and Vision Finetuning can help a lot with that use case.
Also bounding boxes. People have found, like, a lot of improvements for bounding boxes with Vision Finetuning. So yeah, I think it's pretty slept on, uh, and people should try it. You c- you only really need 100 images to get going.
Tell me more about bounding boxes. I didn't think that GPT-4 Vision could do bounding boxes at all.
Yeah, it's actually not that amazing at it. We're working on it.
Okay.
But with finetuning, you can make it really good for your use case.
That's cool, 'cause I've been using Google Gemini's bounding box stuff recently.
Yeah.
It's very, very impressive.
Yeah, totally.
But being able to finetune a model for that. The first thing I'm gonna do with finetuning for images is I've got five chickens-
Yeah
... and I'm gonna tune- finetune a model that can tell which chicken is which.
Love it.
Which is hard, 'cause three of them are gray.
Yeah.
So there's- there's a little bit of- of- of-
Okay, this is my new favorite use case. This is awesome.
Yeah, it's, uh... I've- I've ac- I've managed to do it with w- with prompting, just like-
Ooh
... I gave Claude pictures of all of the chickens-
Yeah
... and then said, "Okay, which chicken is this?"
Yeah.
But it's not quite good enough, 'cause it confuses the- the gray chickens.
Let's see... we can close that eval gap.
Yeah.
Yeah.
That's, uh... It's gonna be a great eval.
Right.
Like, my chicken eval's gonna be fantastic.
I'm also really jazzed about the evals product. Um, it's kind of like a sub-launch of the distillation thing, but people have been struggling to make evals, and the first time I saw the flow with how easy it is to make an eval in- in our product, I was just blown away.
Um, so recommend people really try that. I think that's what's holding a lot of people back from really investing in AI, 'cause they- they just have a hard time figuring out if it's going well for their use case.
So we've been working on making it easier to do that.
Does the eval product include, uh, structure output testing? Like-
Yeah, you can check, um-
... function calling and things
... if it matches your JSON schema. Um, yeah.
But I mean, we- we guaran- we have guaranteed structured output anyway, right?
Well, but it- it-
So we don't have to test it.
Well, it's a function calling.
Well, not the schema, but like the-
Yeah. See, these seem easy to-
... performance.
I think so.
You know?
Yeah.
Like it might call the wrong function or-
Oh, I see
... you're gonna have right schema, wrong output.
So you can do function calling testing?
I'm pretty sure. I'll- I'll have to check that for you, but I think so.
Yeah, yeah.
We'll- we'll make sure it's in the notes.
Fun- fun fact, after our- our podcast, they released function calling V3, which is multi-turn function calling benchmarks. We're- we're having- we're having the guy-
Are you talking about the BFCL?
... on the podcast as well. Sorry?
Are you saying the BFCL one?
BFCL.
Yeah. Yeah.
What would you ask the BFCL guys, 'cause we're actually having them next on the podcast?
Yeah. Uh...
Yeah, I think we tried out V3. Um-
It's just multi-turn-
Probably cut this, but we have some feedback from the founder that's-
Yeah
... We'll- we should probably cut this, but, uh-
Yeah, yeah
... we wanna make it better. Yeah.
Yeah.
What, like... How- how do you think about the evolution of, like, the API design? I think to me, that's, like, the most important thing. So even with the OpenAI levels, like chatbots, I can understand what the API design looks like.
Reasoning, I can kinda understand it, even though, like, chain of thought kinda changes things. As you think about real-time voice, and then you think about agents, it's like, how do you think about how you design the API and, like, what the shape of it is?
Yeah, so I think, uh, we're starting with the lowest level capabilities, and then we build on top of that as we know that they're useful. So a really good example of this is real time. Uh, we're actually going to be shipping audio capabilities in chat completions.
So this is, like, the lowest level capability. So you supply in audio, and you can get back raw audio, and it works at the request response layer. But in through building advanced voice mode, we realized ourselves that, like, it's pretty hard to do with something like chat completions, and so that led us to building this WebSocket API.
So we really learned a lot from our own tools, and we think, you know, the chat completions thing is nice and for certain use cases or async stuff, but you're really gonna want a real-time API. And then as we, you know, test more with developers, we might see that it makes sense- it makes sense to have, like, another layer of abstraction on top of that.
Um, something, like, closer to, uh, you know, more client side, uh, libraries. But for now, you know, that's where we feel we have, like, a really good point of view.
So that's a question I have is, um, if I've got a half-hour long audio recording-
Yeah
... at the moment, the only way I can feed that in is if I call the WebSocket API and slice it up into little JSON basics for snippets-
Yeah
... and fire them all over. In that case, I'd rather just give you a, like, an image in the chat completion API-
Right
... give you a URL to my MP3 files and input. Is that something-
That's what we're gonna do.
Oh, thank goodness for that.
Yes. It's in the blog post. I think it's a short one-liner, but it's rolling out, I think, in the coming weeks.
Oh, wow.
Yeah.
Oh, really soon then.
Yeah, the team has been sprinting. Uh, we're just putting finishing touches on stuff.
Do you have a feel for the length limit on that?
Uh, I don't have it off the top.
Okay.
Sorry.
'Cause yeah, often I want to do... I do a lot of work with, like, transcripts of hour-long YouTube videos, which-
Yeah
... currently I run them through Whisper, and then I do the transcript that way, but being able to do the multimodal thing with those would be really useful.
Totally, yeah. Really jazzed about it. We wanna basically give the lowest capabilities we have, lowest level capabilities, and, you know, the things to make it easier to use. And so, you know, targeting kind of both. Yeah
I just realized what I can do though is I do a lot of Unix utilities, little, like Unix things.
Yeah.
I want to be able to pipe the output of a command into something which streams that up to the WebSocket API and then speaks it out loud.
Yeah.
So I can do streaming speech of the output of things. That should work.
Yeah.
Like, I think you've given me everything I need for that. That's cool.
Yeah. Excited to see what you build.
Is there, um... Uh, I, I heard there are, like multiple, um, competing solutions, and you, you guys evaled it before you picked WebSockets. Like, uh, server-side events, polling. I, I don't... Like, can you give, like your thoughts on, like the live updating paradigms that you, you guys looked at?
'Cause I think a lot of engineers have looked at stuff like this.
Well, I think WebSockets are just a natural fit for bidirectional streaming.
Yeah.
You know, other places I've worked, like Coinbase, we had a WebSocket API for, for pricing data, and I think it's just, like a very natural format. Um-
So it wasn't even really that controversial at all.
I don't think it was super controversial. I mean, we definitely explored the space a little bit, but I think we came to WebSockets pretty quickly. Yeah.
Cool. Um, um, video?
Yeah. Not yet, but, you know- ... possible in the future.
I actually was hoping for the ChatGPT, uh, desktop app with video today because that was demoed-
Yeah
... uh, in-
Well, this is DevDay
... in, yeah.
I think the moment we have the ability to send images over the WebSocket API-
Yeah
... we get video.
Just, my, my question is-
Send the frame, like one frame a second. It's done
... how frequently? Yeah.
Yeah.
Because y- yeah, I mean, sending a v- sending a whole video frame of like a 1080- 1080p screen, maybe it might be too much. What's the limitations on a, on a WebSocket chunk going over? I don't know. It, it might-
I don't have that off the top.
Goog- like Google Gemini, you can do an hour's worth of video in their context window and just by slicing it up into one frame at, at 10 frames a second.
Yeah.
And it does work.
Yeah, yeah, yeah.
So I dunno. I, I'm... But then that's the weird thing about Gemini is it's so good at you just giving it a flood of individual frames. It'll be interesting to see if GPT-4o can handle that or not.
Um, do you have any more feature requests? I know it's been a long day for everybody. But you got, you got Michelle right here, so.
My one is I want you to do all of the accounting for me. I want my users to be able to run my apps, and I want them to call your APIs with their user ID and have you go, "Oh, they've spent 30 cents.
Check, cut them off at a dollar." I can, like check how much they spent, all of that stuff, 'cause I'm having to build that at the moment, and I really don't want to. I don't want to be a token accountant.
I want you to do the token accounting for me.
Yeah, totally. I hear you. It's good feedback.
Well, like how does that contrast with your actual priorities, right? Like I, I feel like you have a bunch of priorities. They showed some on stage with multimodality and all that.
Yeah.
Like...
Yeah. Um, it's hard to say. I, I would say things change really quickly. Um, things that are big adopt- big blockers for user adoption, we, we find very important, and yeah. It's, it's, it's a rolling prioritization. Yeah.
Uh, no, um, assistance API update?
Not at this time. Yeah.
Yeah.
Yeah.
'Cause I was, I was hoping for, like an o1-native thing in assistance.
Yeah.
So I thought they, they would go well together.
We're, we're still kind of iterating, uh, on the formats. I think there are some problems with the assistance API, some things it does really well. Uh, and I think we'll keep iterating a- and land on something really good, but just, you know, it wasn't quite ready yet.
Some of the things that are good in the assistance API is hosted tools. People really like hosted tools and, uh, especially RAG. And then some things that are, you know, less intuitive is just how many API requests you need to get going with the assistance API.
It's, it's quite-
It's quite a lot.
It's a lot.
Yeah. You gotta create an assistant, you gotta create a thread, you gotta, you know, do all this stuff. Um, so yeah, it's something we're thinking about. It, it shouldn't be so hard.
The only thing I've used it for so far is Code Interpreter.
Right.
It's like it's an API to Code Interpreter.
Yes.
Crazy exciting.
Yes. We wanna fix, we wanna fix that and make it easier to use, so.
I, I want Code Interpreter over WebSockets. That would be wildly interesting.
Yeah. Do you, do you wanna bring your own Code Interpreter or you wanna use OpenAI's one?
I wanna use theirs 'cause Code Interpreter's a hard problem. Sandboxing and all of that stuff is-
Yeah, but there's a bunch of Code Interpreter-as-a-service things out there.
There are a few now, yeah.
Because there's... I, I think you, uh, don't allow arbitrary installation of packages.
Oh, they do.
Unless-
They really do
... unless they use your hack.
You can upload it.
Huh?
Yeah, and I do.
Yeah.
You, you know you can-
You can upload a, a pip package.
You can run... You can compile C code in Code Interpreter-
I know. That-
... if you know how to do it
... that's a hack. That's a hack.
Oh, it's such a glorious hack, though.
Okay.
I've had it write me custom SQLite extensions in C and compile them and run them inside of Python, and it works.
I mean, yeah. Uh, there, there's others. E2B is one of them. Like, yeah, it's... It'll be, it'll be interesting to see what the real-time version of that will be. Yeah. Um, awesome, Michelle. Thank you for the-
Yeah
... update. We left-
Yeah
... the episode as what will voice mode look like?
Yeah.
And obviously you knew what it looked like, but you couldn't say it, so now you could, you can share that, so.
Yeah, here we are.
Yeah.
Hope you guys like it.
Um, yeah.
Cool.
Cool. Awesome. Thank you. That's it.
Our final guest today, and also a familiar recent voice on the Latent Space Pod, presented at one of the community talks at this year's DevDay. Alistair Pullen of Cosine made a huge impression with all of you. Special shout-out to listeners like Jesse from Morph Labs when he came on to talk about how he created synthetic datasets to fine-tune the largest LoRAs that had ever been created for GPT-4o to post the highest ever scores on SWE-bench and SWE-bench Verified, while not getting recognition for it because he refused to disclose his reasoning traces to the SWE-bench team.
Genie Fine-tuning1:06:47
Now that OpenAI's o1-preview is announced, it is incredible to see the OpenAI team also obscure their chain of thought traces for competitive reasons and still perform lower than Cosine's Genie model.
We snagged some time with Ali to break down what has happened since his episode aired. Welcome back, Ali.
Thank you so much. Thanks for having me.
Yeah. Uh, so you just spoke at OpenAI DevDay. What was the experience like? Did they reach out to you? Um, you seem to have a very close relationship.
Yeah, so off the back of, off the back of the work that we've done, that we spoke about last time we saw each other, um-
Yeah
... I think that OpenAI definitely felt that the work we've been doing around fine-tuning was worth sharing. Um, I would obviously tend to agree. But today, um, today I spoke about some of the techniques that we learned. Obviously, it was like a nonlinear path arriving to where we've arrived, and the techniques that we built to build Genie.
Um, so I defin- I, I think I shared, um, a few, a few extra pieces about some of the techniques and how it really works under the hood, how you ge-generate a data set to show the model how to do what we show the model.
Um, and that was mainly what I spoke about today. I mean, yeah, they reached out and the, the... I was, I was super excited at the opportunity, obviously. Like, it's not every day that you get to come and do this, um, especially in San Francisco.
So yeah, they reached out and they were like, "Do you wanna talk at DevDay? You can speak about basically anything you want related to what you've built." And I was like, "Sure, that's amazing. I'll talk about fine-tuning and how you, how you build a model that does this, um, software engineering."
So, yeah.
Yeah. W- um, and the, the, the trick here is when we talked, o1 was not out.
No, it wasn't.
Did you, did you know about o1, or?
I didn't know it was... I, I, I knew some bits and pieces. No, not really.
Yeah.
I knew a reasoning model was on the way. I didn't know what it was gonna be called. I knew as much as everyone else. Strawberry was the name back then. Um-
Because, uh, you know, I'll, I'll fast-forward.
Mm-hmm.
You were the first to hide your reas- chain of thought reasoning traces as IP.
Yes.
Right? Famously, that got you in trouble with Sweet Bench or whatever.
Yes.
Uh, uh-
I feel slightly vindicated by that now, not gonna lie.
And now obviously o1 is doing it.
Yeah, the fact that every- yeah, I mean, like, I think it's, I think it's true to say right now that the reasoning of your model gives you the edge that you have. Um, and like the amount of effort that we put into our data pipeline to generate these human-like reasoning traces was, I mean, that, that wasn't for nothing, that we, we knew that this was the way that you'd unlock more performance, getting them all to think in a specific way.
In our case, we wanted it to think like a software engineer. But yeah, I think, I think the,
the, the, the approach that other people have taken, like OpenAI, in terms of reasoning, has definitely showed us that we were going down the right path pretty early on. Uh, and even now we've started, um, replacing some of the reasoning traces in our Genie model with reasoning traces generated by o1, or at least in tandem with o1, and we've already started seeing improvements in performance from that point.
Um, but no, like back to your point, in terms of like the, the whole like withholding them, I, I, I, I still think that that was the right decision to do because of the, the very reason that everyone else has decided to, to, to, to not share those things.
Yeah.
It's, it is exactly-- It shows exactly how we do what we do, and that is our edge at the moment. So, yeah.
Uh, as a founder, so, uh, they also feature recognition on, on stage, talk about them.
Mm-hmm.
How does that make you feel that like, you know, they're, they're like, "Hey, o1 is so much better, makes us better." For you, it should be like, "Oh, I'm so excited about it too," because now all of a sudden it's like it kind of like-
Oh, of course
... raises the bar for everybody.
Yeah.
Like how should people, especially new founders, how should they think about, you know, worrying about the new model versus like being excited about them, just focusing on like the core FP and maybe switching out some of the parts like you mentioned?
Yeah. I, I, I, I, I, speaking for us, I mean, obviously, like we are extremely excited about o1 because at that point, the, the, the process of reasoning is obviously very much baked into the model. We fundamentally, if you like, remove all distractions and everything, we are a reasoning company.
Mm-hmm.
Right? We wanna reason in the way that a software engineer reasons. So when I saw that model announced, I thought immediately, "Well, I can improve the quality of my traces coming out of my pipeline," so like my signal-to-noise ratio gets better.
And then, not immediately, but down the line, I'm going to be able to train those mo- those traces into o1 itself, so I'm gonna get even more performance that way as well. Um, so it's for us, a really nice position to be in, to be able to take advantage of it, both on the prompted side and the fine-tune side.
Um, and also because fundamentally, like we are, I think, fairly clearly in a position now where we don't have to worry about what happens when o2 comes out, what happens when o3 comes out. This process continues. Like even going from, you know, when we first started going from 3.5 to 4, we saw this happen.
Um, and then from 4-Turbo to, to 4o, and then from 4o to o1, we've seen the performance get better every time.
Mm.
Um, and I think, I mean like the crude advice I'd give to any startup founder is try to put yourself in a position where you can take advantage of the same, you know, like sea level rise every time essentially.
Do you make anything out of the fact that you were ta- able to take 4o-
Mm-hmm
... and fine-tune it higher than o1 currently scores-
Mm-hmm
... on Sweet Bench Verified?
Yeah. I mean, like-
Um, yeah
... that was obviously, to be honest with you, you, you realized that before I did. Um, but it was-
Adding value.
Yes, absolutely. That's a value add investor right there. No, obviously, I think it's been-- That in of itself is really vindicating to see because I think, I think we have heard from some people, not a lot of people, but some people saying, "Well, okay, well, if o1 can reason, then what's the point of doing your reasoning?"
But it shows how much more signal is in like the custom reasoning that we generate. Um, and again, it's the, it's the d- very sort of obvious thing. If you take something that's made to be general and you make it specific, of course it's gonna be better at that thing, right?
Um, so it was obviously great to see like we still are better than o1 out of the box-
Yeah
... um, you know, even with an older model, and I'm sure that that, that delta will continue to grow once we're able to train o1 and once we've done more work on our data set using o1, like that delta will grow as well.
It's not obvious to me that they will allow you to fine-tune o1, but-
Uh
... you know, maybe they'll try.
Mm.
Um, I think the, the, the core question that OpenAI really doesn't want you to figure out-
Mm-hmm
... is can you use an open source model and beat o1?
Interesting.
Because-
It, it looks-
Because you basically have shown proof of concept that a non-o1 model can beat o1
And their whole o1 marketing is don't bother trying. Like, don't bother stitching together multiple chain-of-thought calls. We did something special, secret sauce, you don't know anything about it.
Uh-huh.
And somehow, you know, your 4o chain-of-thought-
Is still-
... reasoning as software engineer is still, still better.
Yeah.
Maybe it doesn't last. Maybe they're gonna run o1 for five hours instead of five minutes, and then-
Mm
... and it suddenly works.
Yeah.
So I, I don't know.
It's hard to know. I mean, one of the things that we just want to do out of sheer curiosity is do something like fine-tune 405b on the same dataset.
Yeah.
Like, same context window length, right?
Yeah.
So it should be fairly, fairly easy. We haven't done it yet. Truthfully, we have been so swamped with the wait list, shipping product, you know, DevDay, like, you know, onboarding customers from our wait list. All these different things have gotten in the way, but it is definitely something out of more curiosity than anything else I'd like to try out.
But also, it opens up a new vector of, like, if someone has a, a VPC where they can't deploy an OpenAI modeler, but they might be able to deploy an open source model, it opens that us- it opens that up for us as well from a customer perspective, so it'll probably be quite useful.
I'd be very keen to see what the results are though.
I suspect the answer is yes-
Mm-hmm
... but there may be, it may be hard to do.
Yes.
So, like, Reflection 70B was, like, a really crappy attempt at doing it.
Mm.
You guys were much better-
Mm
... and that's why we had you on the show. Uh, I, yeah, I'm interested to see if there's-
Yeah
... open o1, basically.
Yes.
People want open o1.
Yeah, I, I'm sure they do. Um, as soon as we, as soon as we do it, and, like, once we've wrapped up what we're doing in San Francisco, I'm sure we'll, we'll give it a go. Um- I spoke to some guys today actually about fine-tuning 405b, um, who, who might be able to allow us to do it very, like, very easily.
I don't wanna have to basically do all the setup myself.
Yeah.
So, um, yeah, that might happen sooner rather than later.
Yeah. Any-
Um, anything from the releases today that you're super excited about? So Prompt Caching, I'm guessing when you're, like, dealing with a lot of code bases-
Yeah
... that might be helpful. Is there anything with Vision Finetuning related to, like, more like UI-related-
Yeah
... development?
Definitely.
Anything?
Yeah. I mean, we were, like, we were talking, it's funny, like, my co-founder, Sam, who you've met, uh, and I were talking about the idea of doing Vision Finetuning, like, way back, like, well over a year ago, before Genie existed as it does now.
Um, when we, when we collected our original dataset to do what we do now, um, whenever there were image links and links to, like, um, like, graphical resources and stuff, we also pulled that in as well. We never had the opportunity to use it, but it's something we have in storage.
And again, like, when we have the time, it's something that I'm super excited, particularly on the UI side, um, to be able to, like, leverage. Particularly if you think about one of the things, uh, not to sidetrack, but one of the things we've noticed is, I know Swebench is, like, the most commonly talked about thing, and obviously it's a very- it's an amazing project.
But one of the things we've learnt the most from actually shipping this product to users is it's a pretty bad proxy at telling us how competent the model is. So for example, when people are doing, like, React development, um, using Genie, for us it's impossible to know whether what it's written has actually done, you know, done what it wanted to.
So at least even using, like, the fine-tuning for Vision to be able to help eval, like, what we output is, is, is already something that's very useful. Um, but also in terms of being able to pair, "Here's a UI I want.
Here's the code that actually, like, represents that UI," is also gonna be super useful as well, I think. In terms of generally what have I been most, um, impressed by, the distillation thing is awesome. Um, I think we'll probably end up using it in places.
But what it shows me more broadly about OpenAI's approach is they're gonna be building a lot of the things that we've had to hack together internally in terms from a tooling point of view, just to make our lives so much easier.
And I've spoken to, you know, John, their head of fine-tuning, extensively about this. But there's a bunch of tools that we've had to build internally for things like dealing with model lineage, dealing with dataset lineage, because it gets so messy so quickly, that we would love OpenAI to build.
Like, absolutely would love them to build. Like, it's not, it's not what gives us our edge, but it certainly, um, means that then we don't have to build it and maintain it afterwards. So it's a really good first step, I think, in, like, the overall maturity of the fine-tuning product and API in terms of where they're going to see those early products.
Um, and I think that they'll be continuing in that direction going on.
Did you not-- So there's a very active ecosystem of LLN- LLMOps tools.
Mm-hmm.
Did you not evaluate those before building your own?
We did, but I think fundamentally, like-
No moat.
No. Yeah. Like, I think in a lot of places it was, it was never a big enough pain point to be like, "Oh, we absolutely must outsource this." It's definitely s- in many places, something that you can hack a script together, um, in a day or two, and then hook, uh, hook it up to our already existing internal tool UI, and then you have, you know, what you need.
And whenever you need a new thing, you just tack it on. But for, like, all of these LLMOps tools, I've never felt the pain point enough to really-
Sam Altman Q&A1:18:31
Oh, yeah
... like, bother. And, and, and like, it's no- that's not to deride them at all. I'm sure many people find them useful. But just for us as a company, we've never, we've never felt the need for them. Um, so it's great that, it's great that OpenAI are gonna build them in, 'cause it's really nice to have them there, for sure.
Um, but it's not something that, like, I'd ever consider really paying for externally or something like that, if that makes sense.
Yeah. Does voice mode factor into Genie?
Maybe one day. That'd be sick, wouldn't it?
Yeah. I don't know.
Like, to be able to... Yeah, I think so. Like, I don't-
You, you're the first person that we, uh, we've been asking this question to everybody, uh-
Yeah, I think-
You're the first person to not mention voice mode.
Oh, well, it's, it's, it's currently so distant from what we do. Um, but I definitely think, like, this whole talk of we want it to be a full-on AI software engineering colleague, like, there is definitely a vector in some way that you can build that in.
Um, maybe even during the ideation stage, talking through a problem with Genie as, in terms of how we wanna build something down the line, um, I think that might be useful. But honestly, like, that, that would be a nice to have-
Yeah, yeah
... when we have the time. Yeah.
Yeah. Amazing. Um, one, uh, one last question on, on your, uh, in your talk. You mentioned a lot about- ... curating your data and your distribution and all that.
Yes.
And before we sat down, you talked a little bit about having to diversify-
Absolutely. Sorry
... your data set.
Yeah.
Uh, what's driving that? What are you finding?
So we have been rolling people off the wait list that we sort of amassed when we announced when I last saw you. Um, and it's been really interesting because as I may have mentioned on the podcast, like, we had to be very opinionated about the data mix and the data set that we put together for, like, sort of the v0 of Genie.
Um, again, like to your point, JavaScript, JavaScript, JavaScript, Python. Right? There's a lot of JavaScript in its various forms in there. Um, but it turns out that when we've shipped it, um, to the, to the very early alpha users we've rolled it out to, um, for example, we had some guys using it, um, with a C# code base, and C# currently represents, I think, about 3% of the overall data mix.
Um, and they weren't getting the levels of performance that they saw when they tried it with a Python code base, and it was obviously, like, not great for them to have a bad experience. But it was nice to be able to correlate it with the, the, the actual, like, objective data mix that we saw.
So we did, um-- What we've been doing is, like, little top-up fine tunes where we take, like, the general Genie model and do an incremental fine tune on top with just a bit more data for a given, you know, vertical language.
Um, and we've been seeing improvements coming from that. So again, this is one of the great things, um, about sort of baptism by fire and letting people use it and giving you feedback and telling you where it sucks, um, because that is not something that we could have just known ahead of time.
Um, so I want that data mix to, over time, as we roll it out to more and more people, and we are, are trying to do that as fast as possible, but we're still a team of five for the time being, um, to be as general and as representative of what our users do as possible and not what we think they need.
Yeah. So, uh, uh, every customer is going to have their own lo- uh, fine-tune, uh-
There is gonna be the option to-
To, to-
Yeah, there is gonna be option to fine tune the model on your code base. It won't be in, like, the base pricing tier, but-
Yeah
... you will definitely be able to do that. Um, it will go through all of your code base history, learn how everything happened, and then you'll have an incrementally fine-tuned Genie just on your code base, and that's what enterprises really love the i- the idea of.
Lovely.
Yeah.
Perfect.
Cool.
Anything else? Yeah. That's it.
Yeah.
Thank you so much.
Thank you so much, guys.
Yeah.
Good to see you. Thank you.
Lastly, this year's DevDay ended with an extended Q&A with Sam Altman and Kevin Weil. We think both the questions asked and answers given were particularly insightful, so we are posting what we could snag of the audio here from publicly available sources credited in the show notes for you to pick through.
If the poorer quality audio here is a problem, we recommend waiting for approximately one to two months until the final video is released on YouTube. In the meantime, we particularly recommend Sam's answers on the moderation policy , on the underappreciated importance of agents and AI employees beyond level three, and his projections of the intelligence of o1, o2, and o3 models in future.
Thanks, everybody.
Thanks for coming. All right, I think everybody knows you. For those who don't know me, I'm Kevin Weil, Chief Product Officer at OpenAI. I have the good fortune of getting to turn the amazing research that our research teams do, uh, into the products that you all use every day and the APIs that you all build on every day.
Uh, I thought we'd start with some audience engagement here. So, uh, on the count of three, I'm gonna count to three, and I want you all to say, of, of all the things that you saw launched here today, what's the first thing you're gonna integrate?
It's the thing you're most excited to build on. All right? You gotta do it, right? One, two, three.
Realtime API.
I'll say personally, I'm super excited about, uh, our distillation products.
Whoo.
I think that's gonna be really, really interesting.
I'm also excited to see what you all do with advanced voicemail with the Realtime API and with vision fine-tuning in particular. So, okay. So, uh, I've got some questions for Sam. I've got my CEO here in the hot seat.
Let's see if I can't make a career-limiting move.
So we'll start this, uh, we'll start with an easy one, Sam. How close are we to AGI?
Excuse me.
You know, we, we used to, every time we finished a system, we would say, like, "In what way is this, this not an AGI?" And it used to be, like, very easy. You could, like, make a little robotic hand that does Rubik's Cube or a Dota bot, and it's like, "Oh, it does some things, but definitely not an AGI."
Um, it's obviously harder to say now, and so we, we're trying to, like, stop talking about AGI as this general thing, and we have this levels framework because the word AGI has become so overloaded. Um, so, like, real quickly, we use one for chatbots, two for reasoners, three for agents, four for innovators, five for organizations, like, roughly.
I think we clearly got to level two, or we clearly got to level two with o1. Um, and it, you know, can do really quite impressive cognitive tasks. It's a very smart model. Um, it doesn't feel AGI-like in a few important ways.
But I think if you just do the one next step of making it, you know, very agent-like, which is our level three, and which I think we will be able to do in the not distant future, it will feel surprisingly capable.
Uh, still probably not something that most of you would call an AGI, although maybe some of you would. Um, but it's gonna feel like, all right, this is, this is, like, a significant thing. And then the, the leap-- And I think we do that pretty quickly.
Um, the, the leap from that to something that can really increase the rate of new scientific discovery, which for me is, like, a very important part of having an AGI, I feel a little bit less certain on that, but not a long time.
Like, I think all of this now is gonna happen pretty quickly. And if you think about what happened from last DevDay to, to this one in terms of model capabilities, and you- you're like, "Eh"
I mean, if you go look at like if you go from like o1 on a hard problem back to like Four Turbo that we launched eleven months ago, you'll be like, "Wow, this is happening pretty fast." Um, and I think the next year will be very steep progress.
Next two years will be very steep progress. Harder than that, hard to see with a lot of certainty. But I would say like not varying, and at this point the definitions really matter. And the fact, the fact that the definitions matter this much somehow means we're like getting kind of close.
Yeah. And, y-you know, there, there used to be this sense of AGI where it was like it was a binary thing, and y-you were gonna go to sleep one day, and there was no AGI, and wake up the next day, and there was AGI.
I don't think that's exactly how we think about it anymore. But how, how have your views-
Yeah
... on this evolved?
You know, the, the one-- I agree with that. I think we're like,
you know, in this like kind of period where it's gonna feel very blurry for a while, and the, you know, is this AGI yet or is this not AGI or kind of like at what point? Yeah, it's just gonna be this like smooth exponential, and, you know, probably most people looking back at history won't agree like when that milestone was hit, and we'll just realize it was like a silly thing.
Even the Turing test, which I thought always was like this very clear milestone. You know, there was this like fuzzy period. It kind of like went whooshing by, and no one cared. Uh but, but I, I think the right framework is it's just this one exponential.
That said, um, if we can make an AI system that is like materially better at all of OpenAI than doing-- at doing AI research, that does feel to me like some sort of important discontinuity. It's probably still wrong to think about it that way.
It probably still is the smooth exponential curve, but that feels like a real milestone.
Mm-hmm.
Is OpenAI still as committed to research as it was in the early days? Will research still drive the core of our advancements and our product development?
Yeah. I mean, I think more than ever. Uh, the-- there was like a time in our history when the right thing to do was just to scale out compute, and we saw that with conviction. And we had a spirit of like, we'll do whatever works.
You know? Like we want to-- we have this mission. We want to like build safe AGI, figure out sharing benefits. If the answer is like rack up GPUs, we'll do that. And right now, the answer is again, really push on research.
And I think you see this with o1. Like that is a giant research breakthrough that we were attacking from many vectors over a long period of time that came together this really powerful way. Um, we have many more giant research breakthroughs to come.
But the thing that I think is most special about OpenAI, uh, is that we really deeply care about research, and we understand how to... I think it's easy to copy something you know works. And, uh, you know, I actually don't even mean that as a bad thing.
Like, when people copy OpenAI, I'm like, "Great, the world gets more AI. That's wonderful." But to do something new for the first time, to like really do research in the true sense of it, which is not like, you know, let's barely get SOTA at this thing or like let's tweak this, but like let's go find the new paradigm and the one after that and the one after that, that is what motivates us.
And I think the thing that is special about us as an org, besides the fact that we, you know, married product and research and all this other stuff together, is that we know how to run that kind of a culture that can go-- that can go push back the frontier, and that's really hard.
But we love it. Uh, and that's-- you know, I think we only have to do that a few more times, and then we get to AGI.
Yeah, I'll say like the, the litmus test for me coming from the outside from, you know, sort of normal tech companies of how critical research is to OpenAI is that building product at OpenAI is fundamentally different than any other place that I have ever done it before.
You know, normally you have, you have some sense of your, your tech stack. You have some sense of what you have to work with and what capabilities computers have. And, and then you're trying to build the best product, right?
You're figuring out who your users are and what problems they have and how you can help solve those problems for them. There is that at OpenAI, but also the state of like what computers can do just evolves every two months, three months, and suddenly computers have a new capability that they've never had in the history of the world, and we're trying to figure out how to build a great product and expose that for developers and our APIs and so on.
And, y-you know, you, you can't totally tell what's coming. They're coming through-- It's coming through the mist a little bit at you and gradually taking shape. It's fundamentally different than any other company I've ever worked at, and it's, I think because research-
Is that, is that the thing that has most surprised you?
Yes. Yeah, and it's interesting how e-even internally, we don't always have a sense. Like you have like, okay, I think this capability's coming, but is it going to be, you know, ninety percent accurate or ninety-nine percent accurate in the next model?
Because the difference really changes what kind of product you can build.
Yeah.
And you know that you're gonna get to ninety-nine. You don't quite know when, and figuring out how you put a roadmap together in that world is really interesting.
Yeah. The, the degree to which we have to just like follow the science and let that determine what we go work on next and what products we build and everything else is, I think, hard to get across. Like, we have guesses about where things are gonna go.
Sometimes we're right, often we're not. But if something starts working or if something doesn't work that you thought was gonna work, the-- our willingness to just say, "We're gonna like pivot everything and do what the science allows," and you don't get to like pick what the science allows-
Yeah
... uh, that's surprising.
I was, uh, sitting with an enterprise customer, uh, a couple weeks ago, and they said, "You know, one of the things we really want-- This is all working great. We love this. One of the things we really want is a notification sixty days in advance when you're gonna launch something."
And I was like, "I want that too."
Uh, all right. So I'm, I'm going through... These are a bunch of questions, uh, from the audience, by the way, and we're gonna try and also leave some time at the end for people to ask some audience questions.
Uh, so we've got some folks with mics, and when we get there, uh, so be thinking. But, um, next thing is so many in the alignment community are genuinely concerned that OpenAI is now only paying lip service to the, to alignment.
Can you reassure us?
Yeah. Um, I think it's true we have a different take on alignment than, like, maybe what people write about on one of the, like, internet forums. Um, but we really do care a lot about building safe systems. We have an approach to do it that has been informed by our experience so far.
Um, and touching on another question, which is you don't get to pick where the science goes, um, we want to figure out how to make capable models that get safer and safer over time. And, you know, a couple of years ago, we didn't think the whole Strawberry or the o1 paradigm was gonna, um, work in the way that it's worked, and that brought a whole new set of safety challenges but also safety opportunities.
And rather than kind of like plan from a theoretical once, you know, superintelligence get here, gets here, here's the like seventeen principles, um, we have an approach of figure out where the capabilities are going and then work to make that system safe.
And o1 is obviously our most capable model ever, but it's also our most online model ever by a lot. Um, and as, as these models get better intelligence, better reasoning, whatever you wanna call it, the things that we can do to align them, um, the things we can do to build really safe systems across the entire stack, um, our tool set keeps increasing as well.
So
we, we, we have to build models that are generally accepted, um, as safe and robust to be able to put them in the world. And when we started OpenAI, what the picture of alignment looked like and what we thought the problems that we needed to solve were going to be turned out to be nothing like the problems that actually are in front of us and that we had to solve now.
And also when we made the first GPT-3, um, if you asked me for the techniques that would have worked for us to be able to now deploy our current systems, uh, as generally accepted to be safe and robust, they would not have been the ones that, uh, turned out to work.
So by this idea of iterative deployment, which I think has been one of our most important safety stances ever, um, and sort of confronting reality as it's in front of us, we've made a lot of progress, and we expect to make more.
We keep finding new problems to solve, but we also keep finding new techniques to solve them. All of that said, um,
I think worrying about the sci-fi ways this all goes wrong is also very important. We have people thinking about that. It's a little bit less clear kind of what to do there, and sometimes you end up backtracking a lot.
But-
Right
... but I don't think it's-- I also don't think it's fair to say we're only gonna work on the thing in front of us. We do have to think about where this is going, and we do that too.
Um,
and I think if we keep approaching the problem from both ends like that, most of our thrust on the like, okay, here's the next thing. We're gonna deploy this. What it, what needs to happen to get there? Um, but also like what happens if this curve just keeps going?
That's been, that's been an effective strategy for us.
I'll say also it's one of the places where I'm really-- I really like our philosophy of iterative deployment. Um, when I was at Twitter back, I don't know, a hundred years ago now, um, Ed said something that stuck with me, which is no matter how many smart people you have inside your walls, there are way more smart people outside your walls.
Um, and so w-when we try and get our... You know, it-it-it'd be one thing if we just said we're gonna try and figure out everything that could possibly go wrong within our walls, and it would be just us and the, the red teamers that we can hire and so on.
And we do that. We work really hard at that. But also launching iteratively and launching carefully and learning from the ways that, that folks like you all use it, what can go right, what can go wrong, I think is a big way that we get these things right.
I also think that as we head into this world of agents off doing things in the world, that is gonna become really, really important. The-the-- As these systems get more complex, interact, you know, with longer horizons, um, the pressure testing from the whole outside world, I...
really could be critical.
Yeah. So we'll go actually-- We'll go off of that and, uh, maybe talk to us a bit more about how you see agents fitting into OpenAI's long-term plans.
Well, you-
It's like kind of a huge part of the, the- I mean, I think the exciting thing is this, this set of models, o1 in particular, and if all of its successors are going to be what makes this possible because you finally have the ability to reason, to take hard problems, break them into simpler problems, and act on them.
I-I mean, I think twenty twenty-five is gonna be the year that this really hits big.
Yeah. I, I mean, chat interfaces are great, and they will, I think, have an important place in the world. But the...
When you can, like, ask a model, when you can ask, like, ChatGPT or some agent something, and it's not just like you get a kind of quick response or even you get like fifteen seconds of thinking, and o1 gives you like a nice piece of code back or whatever.
But you can like really give something a multi-turn interaction with environments or other people or whatever and like think for the equivalent of multiple days of human effort and, and like a really smart, really capable human and like have stuff happen We all say that we're all like, "Oh yeah, agents are the next thing.
This is coming. This is gonna be another thing." And we just talk about it like, okay, you know, it's like the next model in the evolution. I would bet, and we don't really know until we get to use these, that it's-- we'll of course get used to it quickly.
People get used to any new technology quickly, but this will be like a very significant change to the way the world works in a short period of time.
Yeah. It's amazing. Somebody was, uh, talking about getting used to new capabilities in AI models and how quickly-- Actually, actually, I think it was about Waymo. Um, but they were talking about how the-- in the first ten seconds of using Waymo, they were like, "Oh my God, is this thing like-- There's a bike.
Let's watch out." And then ten minutes in, they were like, "Oh, this is really cool." And then twenty minutes in, they were like checking their phone bored. You know, it's amazing how much your, your sort of internal firmware updates, um, for this new stuff very quickly.
Yeah. Like I think that people will
ask an agent to do something for them that would've taken them a month, and it'll finish in an hour, and it'll be great, and then they'll have like ten of those at the same time, and then they'll have like a thousand of those at the same time.
And by twenty thirty or whatever, we'll look back and be like, "Yeah, this is just like what a human is supposed to be capable of, what a human used to like, you know, grind at for years or whatever, or many humans used to grind at for years.
Like I just now like ask a computer to do it, and it's like done in an hour." That's why is it not a minute? Like it's-
Yeah. It's also-- it's one of the, the things that makes having an amazing, uh, development platform great too because, you know, we'll experiment, and we'll build some agentic things of course. And like we've already got-- I, I think just like we're, we're just pushing the boundaries of what's possible today.
We-- You've got groups like Cognition doing amazing things in coding, uh, like Harvey and Case Text. You've got s- uh, Speak doing cool things with language translation. Like we're beginning to see this stuff work, and I think it's really gonna start working, uh, as we, as we continue to iterate these models.
One, one of the very fun things for us about having this development platform is just getting to like watch the unbelievable speed and creativity of people that are building these experiences. Like developers very near and dear to our heart, uh, 'cause they're kinda like the first thing we launched, uh, and just many of us came from building on platforms.
But the-- so much of the capability of these models and great experiences have been built by people building on the platform. We'll continue to try to offer like great first-party, uh, products, but we know that will only ever be like a small narrow slice of the apps or agents or whatever people build in the world.
And seeing what has happened in the world in the last eighteen, twenty-four months, uh, it's been like quite amazing to watch.
I'm gonna keep going on the agent front here. Um, what do you see as the current hurdles for computer-controlling agents?
Uh, safety and alignment. Like if you are really going to give an agent the ability to start clicking around your computer, um, which you will- You, you are going to have a very high bar for the, the robustness and the reliability and the alignment of that system.
Uh, so technically speaking, I think that, you know, we're getting like pretty close to the capability side, but the sort of agent safety and trust framework, that's gonna I think be the long pole.
And now I'll kinda ask a question that's almost the opposite of one of the questions from earlier. Do you think safety could act as a false positive and actually limit public access to critical tools that would enable a more egalitarian world?
The honest answer is yes, that will happen sometimes. Like we'll try to get the balance right. Um, but if we were fully YOLO, didn't care about like safety and alignment at all, could we have launched o1 faster? Yeah, we could have done that.
Um,
it would've come at a cost. There would've been things that would've gone really wrong. I'm very proud that we didn't. Um,
the cost, you know, I think would've been manageable with o1, but by the time of o3 or whatever, like maybe it would be pretty unacceptable. And, and so starting on the conservative side, like, you know, and I think people complaining like, "Oh, voice mode, like it won't say this offensive thing, and I really want it to," and you know- "...a formal company, and let it defend me."
You know what? I actually mostly agree. If, if, if you are trying to get o1 to say something offensive, it should follow the instructions of its user most of the time. There's plenty of cases where it shouldn't. But we have like a long history of when we put a new technology into the world, we start on the conservative side.
Um, we try to give society time to adapt. We try to understand where the real harms are versus sort of like kind of more theoretical ones. Um, and that's like part of our approach to safety, and not everyone likes it all the time.
I don't even like it all the time. But, but if we're right that these systems are-- and we're gonna get it wrong too. Like sometimes we won't be conservative enough in some area. Um, but if we're right that these systems are going to get as powerful as we think they are as quickly as they-- we think they might, then I think starting that way makes sense.
And, you know, we like relax over time.
Totally agree. What's the next big challenge for a startup that's using AI as a core feature?
I'll say-
You first.
I, I've got, I've got one, which is I, I think one of the challenges, and we face this too because we're also building products on top of our own models, is trying to find the kind of the frontier.
You, you wanna be building-- These, these AI models are evolving so rapidly, and
If you're building for something that the AI model does well today, it'll work well today, but it's gonna feel, it's gonna feel old tomorrow. And so you wanna build for, for things that the AI model can just barely not do, you know, where maybe the early adopters will go for it and other people won't quite.
But that just means that when the next model comes out, as we continue to make improvements, that use case that just barely didn't work, you're gonna be, you're gonna be the first to do it, and it's gonna be amazing.
But figuring out that boundary is really hard. Uh, I think it's where the best products are gonna get built, though.
Totally agree with that. The other thing I would add is I think it's, like, very tempting to think that a technology makes a startup, and that is almost never true. Uh, no matter how cool a new technology or a new sort of, like, tech title it is, it doesn't excuse you from having to do all of the hard work of building a great company that is going to have durability, um, or, like, a accumulated advantage over time.
And we hear from a lot of startups that, uh, oh, I see this as, like, a, a very common thing, which is, uh, "I can do this incredible thing. I can make this incredible service." Um, and that seems like a complete answer, but it doesn't excuse you from any of, like, the normal laws of, of business.
You still have to, like, build a good business in a good strategic position. And I think a mistake is that in the unbelievable excitement and updraft of AI, people are very tempted to forget that.
This is a-- This is an interesting one. The mode of voice is like tapping directly into the human API. How do you ensure ethical use of such a powerful tool with obvious abilities of manipulation?
Yeah. You know, voice mode was a really interesting one for me. It was, it was, like, the first time that I felt like I sort of had got, had gotten, like, really tricked by an AI, in that when I was playing with the first beta of it,
I couldn't, like, I couldn't stop myself. I mean, I, I kind of... Like, I still say, like, "Please to ChatGPT," um, but in voice mode, I, like, couldn't not kind of use the normal niceties. I, I was, like, so convinced, like, ah, it might be a real per- like- ...
you know? Um, and obviously it's just, like, hacking some circuit in my brain, but I really felt it with voice mode. Um, and I, I sort of still do. Uh, the... I think this is a more-- this is an example of, like, a more general thing that we're gonna start facing, which is as these systems become more and more capable, and as we try to make them as natural as possible to interact with, uh, they're gonna, like, hit parts of our neural circuitry that we've, like, evolved to deal with other people.
And, you know, there's, like, a bunch of clear lines about things we don't wanna do. Like, we don't-- Like, there's a whole bunch of, like, weird personality growth hacking, like, I think vaguely socially manipulative stuff we could do.
But then there's these, like, other things that are just not nearly as clear-cut. Like, you want the voice mode to feel as natural as possible, but then you get across the uncanny valley and it, like, at least to me, triggers something.
Uh, and you know, me saying like, "Please, thank you" to ChatGPT, no pro-probably a good thing to do. You never know.
Um, the-
But, but I think this, like, really points at the kinds of safety and alignment issues we have to start paying a lot of attention to.
All right, back to brass tacks. Sam, when's o1 gonna support function tools?
Do you know?
Uh, before the end of the year.
All right.
There, there are three things that we really want to get in for, um-
We're gonna record this, take this back to the research team- ... show them how badly we need to do this. Uh, but there-- I mean, there are a handful of things that we really wanted to get into o1, and we also, you know, it's a balance of should we get this out to the world earlier and begin un- you know, learning from it, learning from how you all use it, or should we launch a fully complete thing that is, you know, i-in line with that, that has all the abilities that every other model that we've launched has?
I'm really excited to see things like system prompts and structured outputs and function calling make it into o1. We will be there by the end of the year. It really matters to us too.
In addition to that, just 'cause I can't resist the opportunity to reinforce this, like, we will get all of those things in and a whole bunch more things you all have asked for. Um,
the model is gonna get so much better so fast. Like, we are so early. This is like, you know, maybe it's the GPT-2 scale moment, but, like, we know how to get it to GPT-4, and we have the fundamental stuff in place now to get it to GPT-4.
And in addition to planning for us to build all of those things, plan for the model to just get, like, rapidly smarter. Like, you know, hope you all come back next year and plan for it to feel like way more of a year of improvement than from, uh, four to o1.
That's it. Thank you.
Uh, what feature or capability of a competitor do you really admire?
Um, I think Google's notebook thing is super cool.
Yeah.
Uh, what do they call it?
NotebookLM.
NotebookLM. Yeah.
Yeah.
Uh, I was like-- I woke up early this morning, and I was, like, looking at examples on Twitter, and I was just like, "This is, like, this is just cool. Like, this is just a good, cool thing." And, like, I think not enough of, not enough of the world is, like, shipping new and different things.
It's mostly, like, the same stuff. But that, I think, is like That brought me a lot of joy this morning.
Yeah.
It was very, very well done.
One of the things I really appreciate about that product is they, uh, there's the, the-- just the format itself is really interesting, but they also nail the podcast-style voices.
Well, yes .
And they have really nice microphones
. They have these sort of sonorant voices. Um, did you guys see I-- somebody on Twitter was, um, uh, saying, like, the cool thing to do is take your LinkedIn and put a, you know- ... PDF it and give it to these, uh, give it to NotebookLM.
Yeah.
And you'll have two podcasters riffing back and forth about how amazing you are and- ... all of your accomplishments over the years . Um, I'll say mine is, um, I think Anthropic did a really good job on, uh, projects.
Um, it's kind of a, a different take on what we did with GPTs. I mean, GPTs are a little bit more long-lived. It's something you build and can use over and over again. Projects are kind of the same idea, but, like, more temporary, meant to be kind of stood up, used for a while, and then you can move on.
And then that-- the different mental model makes a difference. Uh, and I think they did a really nice, nice job with that.
Um, all right. We're getting close to audience questions, so be thinking of what you wanna ask. Uh, so at OpenAI, how do you balance what you think users may need versus what they actually need today?
Also a better question for you.
Yeah. Well, I think it does get back to a bit of what we were saying around trying to, trying to build for what the model can just, like, not quite do, but almost do. Um, but it's a real balance too as we, as we, you know, we support over two hundred million people every week on ChatGPT.
You also can't say, "No, it's cool. Like, deal with this bug for three months or this issue. Uh, we've got something really cool coming." You've gotta solve for the needs of today. And there are some really interesting product problems.
I mean, you think about, uh, I'm, I'm speaking to a group of people who know AI really well. Think of all the people in the world who have never used any of these products, and that is the vast majority of the world still.
Uh, you're basically giving them a text interface, and on the other side of the text interface is this, like, alien intelligence that's constantly evolving that they've never seen or interacted with, and you're trying to teach them all of the crazy things that you can actually do with it, all the ways it can help, can integrate into your life, can solve problems for you.
But people don't know what to do with it. You know, like, you, you come in and you're just like-- people type, like, "Hi." And it responds, you know, "Hey, great to see..." Like, "How can I help you today?"
And they-- you're like, "Okay, I don't know what to say." And then you end up-- You kind of walk away, and you're like, "Well, I didn't see the magic of that." And so it's a real challenge figuring out how you-- I mean, we all have a hundred different ways that we use ChatGPT and AI tools in general.
But teaching people what those can be and then bringing them along as the model changes month by month by month and suddenly gains these capabilities way faster than we as humans gain new capabilities, it's, it's a really interesting set of problems, and I, I know it's one that you all solve in, in different ways as well.
I, I have a question. Who feels like they, they spent a lot of time with o1, and they would say, like, "I feel definitively smarter than that thing"?
Do you think you're still well by o2? Just thinking. No one, no one taking the bet of, like, it is smarter than o2. Um, so one of the challenges that we face is, like, we know how to go do this thing that we think will be, like, at least probably smarter than all of us in, like, a, a broad array of tasks.
And yet we have to, like, fig- still, like, fix the bugs and do the "Hey, how are you?" problem. And mostly what we believe in is that if we keep pushing on model intelligence, um, people will do incredible things with them.
Um, you know, we wanna build the smartest, most helpful models in the world, and people then find all sorts of ways to use that and build on top of that.
It has been definitely a
evolution for us to not just be entirely research-focused, and then we do have to fix all those bugs and make this super usable. Um, and I think we've gotten better at balancing that. But still, as part of our culture, I think we trust that if we can keep pushing on, um, intelligence, so yourself four if you run down here, um, it'll-- people will build just incredible things with that capability.
Yeah. I think it's a core part of the philosophy, and you do a good job pushing us to always-- well, basically incorporate the frontier of intelligence into our products, both in the APIs and into our first-party products. Um, because it's, it's easy to kinda stick to the thing you know, the thing that works well, but you're always pushing us to, like, get the frontier in even if it only kinda works, because it's gonna work really well soon.
So, uh, I always find that a really helpful push. Uh, you kind of answered the next one. You do say please and thank you to the models. I'm curious, how many people say please and thank you?
Isn't that so interesting?
I do too . I kind of can't-- I, I feel bad if I don't.
And-
Um, okay. Last question, and then we'll go into, uh, audience questions for the last ten or so minutes. Do you plan to build models specifically made for agentic use cases, things that are better at reasoning and tool calling?
Um, specif-- We plan to make models that are great at agentic use cases. That'll be a key priority for us over the coming months. Um, specifically is a hard thing to ask for, because I think it's also just how we keep making smarter models.
Yeah.
So yes, there's, like, some things like tool use and function calling that we need to build in that'll help, but mostly we just wanna make the best reasoning models in the world. Those will also be the best agentic-based models in the world.
Cool. Let's go to audience questions. I don't know who's got the mic. All right, we got a mic.
Audience Q&A1:56:32
How extensively do you dogfood your own technology in your company? Uh, and do you have any interesting examples that may not be obvious?
Yeah. Uh, I mean, we put models up for internal use even before they're done training. Like we use checkpoints and try to have people use them for whatever they can and try to sort of like build new ways to explore the capability of the model internally and use them for our own development or research or whatever else as much as we can.
We're still always surprised by the creativity of the outside world and what people do. Um, but
basically, the way we've figured out every step along our way of how to-- what to push on next, what we can productize, what, what, what, like what the models are really good at is by internal dogfooding. Um, that's like our whole-- That's how we like to feel our way through this.
HR that are based off of o1?
Uh, we don't yet have like employees that are based off of o1, but- -I, you know, as we like move into the world of agents, we will try that. Like we will try having like, you know, things that we deploy in our internal systems that help you with stuff.
There are things that get closer to that. I mean, there are like customer service. We have bots internally that, that do a ton of both answering external questions and fielding internal people's questions on Slack and so on. And our customer suc-- our customer service team is probably, I don't know, twenty percent the size it might otherwise need to be because of it.
Good. Um, I know Matt Knight and our security team has talked extensively about all the different ways we use models internally, uh, for-- to automate a bunch of security things and, you know, take what used to be a manual process where you might not have the number of humans to even like look at everything incoming and have models, uh, taking, you know, separating signal from noise and highlighting to humans what they need to go look at, things like that.
So I think internally, there are tons of examples, and people maybe underestimate the… You all probably will not be surprised by this, but a lot of folks that I talk to are, the extent to which it's not just using a model in a place, it's actually about using like, uh, chains of models that are good at doing different things and connecting them all together to get one end-to-end process that is very good at the thing you're doing, even if the individual models have, you know, flaws and make mistakes.
Thank you. Um, I'm wondering if you guys have any plans on sharing models for like offline usage because, uh, with this distillation thing, it's really cool that we can share our own models, but a lot of use cases, we really wanna kinda like have a version of it.
Want to take that? Um, we're open to it. It's not on-- It's not like high priority on the current roadmap. The--
If we had like more resources and bandwidth, we would go do that. Uh, I think there's a lot of reasons you want a local model. Um, but it's not like, it's not like a this year kind of thing.
Question.
Okay. Uh, hi. Um, my question is, there are many agencies in the government above the local, state, and, uh, national level that could really greatly benefit from the tools that you guys are developing but have perhaps some hesitancy on deploying them because of, you know, security concerns, data concerns, privacy concerns.
And I guess I'm curious to know if there are any sort of, you know, planned partnerships with governments, rural governments once whatever A-- when AGI is achieved. Because obviously, if AGI can help solve problems like, you know, world hunger, poverty, climate change, um, government's gonna have to get involved with that, right?
And I'm just curious to know if there's some, uh, you know, plan in the works when and if that time comes.
Yeah. I think-- I actually think you don't wanna wait until AGI. You wanna start now, right? Because there's a learning process, and there's a lot of good that we can do with our current models. So we've, we've announced a handful of partnerships with government agencies, some states, I think Minnesota and some others, Pennsylvania, um, also with organizations like USAID.
Uh, it's actually a huge priority of ours to be able to help, uh, governments around the world get acclimated, get benefit from the technology. I mean, of all places, government feels like somewhere where you can automate a bunch of workflows and make things more efficient, reduce drudgery, and so on.
So I think there's a huge amount of good we can do now, and if we do that now, it just accrues over the long run as the models get better and we get closer to AGI.
I have a pretty open-ended question. Um, what are your thoughts on open source? So whether that's open weights, just general discussion, where do you guys sit with open source?
I think open source is awesome. Um, I think, again, if we had more bandwidth, we would do that too. We've like gotten very close to making a big open source effort a few times and then, you know, the really hard part is prioritization, and we have put other things ahead of it.
Um, part of it is like there's such good open source models in the world now, uh, that I think that segment-- The thing we always end in most interest by is like a really great on-device model, and I think that segment is fairly well served.
Uh, I do hope we do something at some point, but we wanna find something that we feel like if we don't do it, then we'll just be missing this and not make like another thing that's like a tiny bit better on benchmarks, um, because we think there's like a lot of good stuff out there now.
But, but like spiritually, philosophically, very glad it exists. Would like to figure out how to contribute.
Hi, Sam. Hi, Kevin. Uh, thanks for inviting us to DevDay. It's been awesome. All the live demos work. That's incredible. Um, why can't Advanced Mo- Voice Mode sing? And as a follow-up to this, if it's a company, like, legal issue in terms of copyright, et cetera, uh, is there a daylight between how you think about safety in terms of your own products on your own platform versus giving us developers kind of the, I don't know, sign the right things off so we can, we can make our, uh, Advanced Voice Mo- Voice Mode sing?
Could you address this?
Well, you know, the funny thing is Sam asked the same question. "Why can't this thing sing? I want it to sing. I've heard it sing before." Um, it-- a-actually, it's, uh, there are things obviously that we can't have it sing, right?
We can't have it sing copyrighted songs. We don't have the licenses, et cetera. And then there are things that it can sing, and you could have it sing happy birthday, and that would be just fine, right? And we want that too.
It's a matter of I think once you... It-- basically, it's easier in finite time to say no and then build it in. But it's nuanced to get it right, and we, you know, there are penalties to getting these kinds of things wrong.
So it's really just where we are now. We really want the models to sing too.
People were tired of waiting for us to ship Voice Mode, which is, like, very fair. Uh, we could've, like, waited longer and kind of really got the classifications and filters on, you know, copyrighted music versus not. But we decided we would just ship it and we'll add more.
But I think Sam has asked me, like, four or five times. Why- I think the next feature.
Developers, uh, create their own.
Uh, I mean, we, we still can't, like, offer something where we're gonna be in, like, really bad legal hot water, developers or first party or whatever. So yes, we can, like, maybe have some differences, but we still have to, like, comply with the law.
Could you speak a little to the future of where you see context windows going and kind of the timeline for when-- how you see things balance between context window growth and RAG, basically information retrieval?
Um, I think there's, like, two different takes on that that matter. One is, like, when is it gonna get to, like, kinda normal long context, like context length ten million or whatever. Like, long enough that you just throw stuff in there and it's fast enough you're happy about it.
And I expect everybody's gonna make, uh, pretty fast progress there, and that'll just be a thing. Uh, long context has gotten weirdly less usage than I would've expected so far. But I think, you know, there's a bunch of reasons for that.
I don't want to go too much into it. And then there's this other question of, like, when do we get to context length not like ten million, but ten trillion? Like, when do we get to the point where you can throw, like, every piece of data you've ever seen in your entire life in there and, um, you know, like, that's a whole different set of things.
That obviously takes some research breakthroughs. But I assume that infinite context will happen at some point, and some point is, like, less than a decade. Um, and that is-- that's gonna be just a totally different way that we use these models.
Even getting to the, like, ten million tokens of very fast and accurate context, which I expect measured in, like, months, something like that. Um, you know, like,
people will use that in all sorts of ways, and it'll be great. Um, but yeah, the very, very long context I think is gonna happen, and it's really interesting.
I think we maybe have time for one or two more.
Don't worry. This is the-- gonna be your favorite question. So with Voice and all the other changes that users have experienced since you all have launched your technology, what do you see is the vision for the new engagement layer, the form factor, and how we actually engage with this technology to make our be- lives so much better?
Uh, I love that question. It's one that we ask ourselves a lot, frankly. Um, there's this... And I think it's one where developers can play a, a, a really big part here because there's this trade-off between generality and specificity here.
I-I'll give you an example. I was in Seoul and, and, uh, Tokyo a few weeks ago, and I was in a number of conversations with folks that, uh, with whom I didn't have a common language, and we didn't have a translator around.
Before, we would not have been able to have a conversation. We would've just sort of smiled at each other and continued on. I took out my phone. I said, "ChatGPT, I want you to be a translator for me.
When I speak in English, I want you to speak in Korean. When you hear Korean, I want you to repeat it in English." And I was able to have a full business conversation, and it was amazing. You think about w- the impact that can have, not just for business, but think about travel and tourism and people's willingness to go places where they might not have a word of the language.
You can have these really amazing impacts. But inside ChatGPT, that was still a thing that I had to... Like, ChatGPT is not optimized for that, right? Like, you want this sort of digital, you know, universal translator in your pocket that just knows that what you want it to do is translate.
Not that hard to build. Um, but I think there's... We struggle with the-- with trying to build a, a, an application that can do lots of things for lots of people, um, and that, that keeps up, like we've been talking about a few times, that keeps up with the, the pace of change and with the capabilities, you know, agentic capabilities and so on.
I think there's also a huge opportunity for the creativity of an audience like this to come in and, like, solve problems that we're not thinking of, that we don't have the expertise to do. And ultimately, the world is a much better place if we get more AI to more people, and it's why we are so proud to serve all of you.
Right. The only thing I would add is if you just think about everything that's gonna come together at some point in not that many years in the future. You'll walk up to a piece of glass. You will say whatever you want.
Um, it will have, like-- There'll be incredibly reasoning models, agents connected to everything. There'll be a video model streaming back to you, like a custom interface just for this one request. Whatever you need is just gonna get, like, rendered in real time on video.
You'll be able to interact with it. You'll be able to, like, click through the stream or say different things, and it'll be off doing like, again, the kinds of things that used to take like humans years to figure out.
And it'll just, you know, dynamically render whatever you need, and it'll be a completely different way of using a computer, um, and also getting things to happen in the world, uh, that it's gonna be quite wild.
Awesome. Thank you. That was a great question to end on. I think we're at time. Thank you so much for coming. We've had a great-
That's all for our coverage of DevDay twenty twenty-four.
Outro2:09:15
Under DevDay lights code ignites. Real-time voice streams reach new heights. o1 and GPT-4o





