Introduction0:00
Hey, everyone. Welcome to the Late in Space podcast. This is Alessio, partner and CTO at Decibel Partners, and I'm joined by my co-host, Swyx, founder of Smol AI.
Good morning. Uh and, uh, today we're very excited to have Sam Colvin join us from Pydantic AI. Welcome.
Thank you so much for having me. Yeah, it's great to be here.
Sam, I heard that Pydantic is all we need. Is that true?
Origins0:24
I would say you might need Pydantic AI and Logfire as well, but, um, it gets you a long way, that's for sure.
Pydantic almost basically needs no introduction. It's, you know, it's, uh, almost three hundred million downloads in December. And obviously, uh, in the previous podcasts and discussions we've had with Jason Lu, uh, he's been a big fan and promoter of Pydantic and AI.
Yeah, it's, it's, it's weird because obviously we didn't-- I didn't create Pydantic originally for, for uses in AI. Obviously, it predates LLMs, but it's like we've been lucky that it's been picked up by that community and, and used so, so widely.
Actually, maybe, maybe we'll hear it right from you. Like what is Pydantic and maybe a little bit of the origin story.
The best name for it, which is not quite right, is a validation library. And I-- we get some tension around that name because it doesn't just do validation, it will do coercion by default. We now have strict mode, so it-- you can disable that coercion, but like by default, if you say you want, uh, an integer field and you get in a string of one, two, three, it will convert it to hundred and twenty-three and a bunch of other sensible conversions.
And as you can imagine, the like semantics around exactly when you convert and when you don't is complicated. But because of that, it's more than just validation. Back in twenty seventeen when I first started it, the like different thing it was doing was using type hints to define your schema.
That was controversial at the time. It was like genuinely disapproved of by some people. I think the success of Pydantic and libraries like FastAPI that build on top of it means that today that's no longer controversial in Python, and indeed, lots of other people have like copied that route.
But yeah, it's a data validation library that uses type hints for the, for the most part. And obviously does all the other stuff you want, like serialization on top of that. But yeah, that's the core.
Do you have any fun stories on how JSON schemas ended up being kinda like the structure output standard for LLMs? And were you involved in any of these discussions? Because I know OpenAI was, you know, one of the early adopters.
So d-did they reach out to you? Was there kinda like a structure output council that in open source that people were talking about, or was it just at random?
No, very much not. So I originally didn't implement JSON schema inside Pydantic, and then Sebastian, Sebastian Ramirez, FastAPI came along and like the first I ever heard of him was over a weekend. I got like fifty emails from him, or fifty like emails as he was committing to Pydantic, adding JSON schema long pre version one.
So the reason it was added was for OpenAPI, which is obviously closely akin to JSON schema. And then, yeah, I don't know why it was JSON sch-- w- uh, that got picked up and used by OpenAI. It was obviously very convenient for us because it meant that not only can you do the validation, but because Pydantic will generate you the JSON schema, it will-- it kind of can be one source of t-- source of truth for structured outputs and tools.
Before we dive in further on the, on the AI side of, of, of things, something I'm mildly curious about, and obviously there's, uh, Zod in, in JavaScript land. Every now and then, there's a new sort of in vogue validation library that, that takes over for quite a few years, and then maybe like some-something else comes along.
Is Pydantic done, like the core Pydantic?
I've just come off a call where we were redesigning some of the internal bits. There will be a V3 at some point, which will not break people's code half as much as V2. As in V2 was the, was the massive rewrite into Rust, but also fixing all the stuff that was broken back from like version zero point something that we didn't fix in V1 because it was a side project.
We have plans to move some of the-- basically store the data in Rust types after validation, not convert them to Python types. So then if you were doing like validation and then serialization, you would never have to go via a Python type.
We reckon that can give us somewhere between three and five times-- another three to five times speed up. That's probably like the, the biggest thing. Also, like changing how easy it is to basically extend Pydantic and define how particular types, like for example, NumPy arrays are, are validated and serialized.
But there's also stuff going on, for example, Jitter, the JSON library that does-- in Rust, that does the, the JSON parsing, has a SIMD implementation at the moment only for AMD sixty-four. So there we're like, we can add that.
We need to go and add, add SIMD for other, other instruction sets. So like there's a bunch more we can do on performance. I don't think we're gonna go and revolutionize Pydantic, but it's gonna continue to get faster, continue, hopefully to get-- allow people to do more advanced things.
We might add a binary format like CBOL for serialization, for when you'll just wanna put the data into a database and probably load it again from Pydantic. So, so there are some, some things that will come along, but like for the most part, it will just get-- it should just get faster and cleaner.
From a focus perspective, I guess like as a, as a founder too, like how did you think about the AI interest rising and then how do you kind of prioritize, okay, this is worth going into more the... And we'll talk about Pydantic AI and, and all of that.
AI Inflection5:05
What was maybe your early experience with LLMs, and when did you figure out, okay, this is like something we should like take seriously and focus more resources on it?
I'll answer that, but I'll answer which I think is like a, a kind of parallel question, which is Pydantic's weird because Pydantic existed obviously before I was starting a company. I was working on it in my spare time, and then beginning of twenty-two, I started working on the rewrite in Rust basically.
And I worked on that-- I worked on it full-time for a year and a half, and then once we started the company, people came and joined. And it's a weird-- it was a weird project because that would never get signed off inside a startup.
Like we're gonna go off and three engineers are gonna work full-on for a year in Python and Rust writing like thirty thousand lines of Rust just to release open source free Python library. The, the result of that has been excellent for us as a company, right?
As in it's made us remain entirely relevant. It's like Pydantic is not just used in the SDKs of all of the AI libraries, but I can't say which one, but one of the big foundational model companies, when they upgraded from Pydantic V1 to V2, their number one internal metric of performance is time to first token.
That went down by twenty percent. So you think about all of the actual AI going on inside, and yet at least twenty percent of the CPU or at least the latency inside requests was actually Pydantic, which shows like how widely it's used.
So we've benefited from doing that work, although it didn't-- it would have never have made financial sense in, in most companies. In answer to your question about like how do we prioritize AI, I mean, the honest truth is we spent a lot of the last year and a half building good general purpose observability inside Logfire and making Pydantic good for general purpose use cases, and the AI has kind of come to us.
Like, we just-- Not that we wanna get away from it, but, like, the appetite, uh, both in Pydantic and in Logfire to go and build with AI is enormous because it kinda makes sense, right? Like, if you're starting a new greenfield project in Python today, what's the chance that you're using GenAI?
Eighty percent, let's say, globally. Obviously, it's like a hundred percent in California, but even worldwide, it's probably eighty percent. And so everyone needs that stuff, and there's so much yet to be figured out, so much, like, space to do things better in the ecosystem in a way that, like, to go and implement a database that's better than Postgres is a, like, Sisyphean task.
Whereas building, uh, tools that are better for GenAI than some of the stuff that's about now is not very difficult, putting the actual models themselves to one side.
And then at the same time, then you release Pydantic AI recently, which is a-
Yeah
... um, you know, agent framework. And early on, I, I would say everybody, like, you know, LangChain and, like, uh, Pydantic, kinda like a first-class support, a lot of dif- these frameworks were trying to use you to be better.
What was the decision behind we should do our own framework? Were there any design decisions that you disagree with? Any workloads that you think people didn't support well?
Wasn't so much like design and workflow, although I think there are some, some things we've done differently. I think looking in general at the ecosystem of agent frameworks, the engineering quality is far below that of the rest of the Python ecosystem.
There's a bunch of stuff that we have learned how to do over the last twenty years of building Python libraries and writing Python code that seems to be abandoned by people when they build agent frameworks. Now, I can kind of respect that, particularly in the very first agent frameworks like LangChain, where they were literally figuring out how to go and do this stuff.
It's completely understandable that you would, like, basically skip some standard best practice to kinda get something built. I am shocked by the, like, quality of some of the agent frameworks that have come out recently from, like, well-respected names, which it just seems to be opportunism, and I have little time for that.
But, like, the early ones, like, I think they were just figuring out how to do stuff, and just as lots of people have learned from Pydantic, we were able to learn a bit from them. I think from, like, the, the gap we saw and the thing we were frustrated by was the production readiness, and that means things like type checking, even w- if type checking makes it hard.
Like, Pydantic AI, I will put my hand up now and say it has a lot of generics, and you need to-- it's probably easier to use it if you've written a bit of Rust and you really understand generics.
But, like-- And that is-- We're not claiming that that makes it the easiest thing to use in all cases. We think it makes it good for production applications in big systems where type checking is a no-brainer in Python.
But there are also a bunch of stuff we've learnt from maintaining Pydantic over the years that we've gone and done. So every single example in Pydantic AI's documentation is run as part of tests, and every single print output within an example is checked during tests, so it will always be up to date.
And then a bunch of things that, like I say, are standard best practice within the rest of the Python ecosystem, but are not followed, surprisingly, by some AI libraries, like coverage, linting, type checking, et cetera, et cetera, where I think these are no-brainers, but, like, weirdly, they're not followed by some of the other libraries.
And can you just give a overview of the framework itself? I think there's kinda like the LLM calling frameworks. There are the multi-agent frameworks. There's the workflow frameworks. Like, uh, what, what does Pydantic AI do?
Agent Framework10:17
I glaze over a bit when I hear all of the different sorts of frameworks, but I, I-- like, and I, I, I will tell you, when I built Pydantic, when I built Logfire, and when I built Pydantic AI, my methodology is not to go and, like, research and review all of the other things.
I kinda work out what I want, and I go and build it, and then feedback comes, and we adjust. So the fundamental building block of Pydantic AI is agents. The exact definition of agents and how you wanna define them is obviously ambiguous, and our things are probably sort of agent-lets, not that we would wanna go and rename them to agent-let, but, like, the point is, you probably build them together to build something, and most people will call an agent.
So an agent in our case has, you know, things like a prompt, like system prompt and some tools and a structured return type if you want it. That covers the vast majority of, of cases. There are situations where you wanna go further in the most complex workflows where you want graphs, and I resisted graphs for quite a long time.
I was s-sort of, of the opinion you didn't need them, and you could use standard, like, Python flow control to do all of that stuff. I had a few arguments with people, but I, I basically came around. Yeah, I can totally see why graphs are useful.
But then we have the problem that by default they're not type-safe. Because if you have a, like, add edge method where you give the names of two different edges, there's no type checking, right? Even if you go and do some-- And not, not all the graph libraries are AI-specific.
So there's a, there's a graph library called-- But it allow-- It does like, basically, it does runtime type checking, ironically using Pydantic to try and make up for the fact that, like, fundamentally, their graphs are not typed-- type-safe.
Well, I like Pydantic, but it... that's not a real solution to have to go and run the code to see if it's safe. There's a reason that static type checking is so powerful. And so we kind of, from a lot of iteration, eventually came up with a system of using normally data classes to define nodes, where you return the next node you want to call and where we're able to go and introspect the return type of a node to basically build the graph.
And so the graph is inherently type-safe. And once we got that right, I, I, uh, was and am incredibly excited about graphs. I think there's, like- Masses of use cases for them, both in GenAI and other development. But also software's all gonna have interact with GenAI, right?
It's gonna be like web. There's no gonna no longer be a, like a, a web department in a company, is there? There's just like all the developers are building for web, building with databases. The same is gonna be true for GenAI.
Graphs12:33
Yeah. I think on your docs, you call an agent a container that contains a system prompt, function tools, structure result, dependency type, model, and then model settings. Are the graphs, in your mind, different agents? Are they, uh, different prompts for the same agent?
What are, like, the structures in your mind?
So we were compelled enough by graphs once we got them right, that we actually merged a PR this morning. That means our agent implementation, without changing its API at all, is now actually a, a, a graph under the hood, as in it is, it is built using our graph library.
So graphs are basically a lower level tool that allow you to build these complex workflows. Our agents are technically one of the many graphs you could go and build, and we just happened to build that one for you because it's a very common, commonplace one.
But obviously, there are cases where you need more complex workflows where the current agent assumptions don't work, and that's where you can then go and use graphs to build more complex things.
You said you were cynical about graphs. Uh, what changed your mind specifically?
I guess people kept giving me examples of things that they wanted to use graphs for, and my like, "Yeah, but you could do that in standard flow control in Python," became a, like, less and less compelling argument to me because I've maintained those systems that end up with, like, spaghetti code.
Yeah.
And I could see the appeal of this, like, structured way of defining the workflow of my code. And it's really neat that, like, just from your code, just from your type hints, you can get out a Mermaid diagram that defines exactly what can go and happen.
Right. Yeah. Y- uh, you do have very neat, uh, implementation of, um, sort of inferring the graph from type hints, I guess is-
Yeah
... is what I would call it. Yeah, I think the question always is I, I have gone back and forth. Um, I, you know, used to work at Temporal, where we would actually spend a lot of time complaining about graph-based workflow solutions like AWS Step Functions, and we would actually say that we were better because y- you could use normal control flow that you already knew and worked with.
Yours, I guess, is like a little bit of a nice compromise. Like, um, it looks like normal Pythonic code-
Mm-hmm
... but you just have to keep in mind what the type hints actually do, uh, with the, quote-unquote, magic that, uh, the graph construction does.
Yeah, exactly. And if you look at the internal logic of actually running a graph, it's incredibly simple. It's basically call a node, get a node back, call that node, get a node back, call that node. If you get an end, you're done.
We will add in soon support for, well, basically storage, so that you can store the state between each node that's run, and then you'd be, uh, the idea is you can then distribute a graph and run it across compute.
And also, I mean, the other weird-- the other bit that's really valuable is across time, because it's all very well if you, if you look at, like, lots of the graph examples that, like, Claude will give you if it, if it gives you an example.
It gives you this lovely enormous Mermaid chart of, like, the workflow, for example, managing returns if you're an e-commerce company. But what you realize is some of those lines are literally one function calls another function, and some of those lines are wait six days for the customer to print their, like, piece of paper and put it in the post.
And if you're writing, like, your demo project or your, your, like, proof of concept, that's fine because you can just say, "And now we call this function." But when you're building-- When you're in real, in real life, that doesn't work.
And now how do we manage that concept of basically being able to start somewhere else in the, in our code? Well, this graph implementation makes it incredibly easy because you just pass the node that is the start point for carrying on the graph, and it continues to run.
So it's things like that where I was like, "Yeah, I can just imagine how things I've done in the past would be fundamentally easier to understand if, if we had done them with graphs."
You say im- imagine, but, like, right now, does Pydantic AI actually resume, you know, six days later, like you said? Or is this, this is, like, a theoretical thing we can go someday?
I think it's basically Q&A. So there's an AI that's asking the user a question, and effectively you then call the CLI again to continue the conversation, and it basically instantiates the node and calls the graph with that node again.
Now, we don't have the logic yet for effectively storing state in the database between individual nodes. That, uh, we're gonna add soon, but, like, the rest of it is basically there.
It does make me think that not only are you competing with LangChain now and obviously Instructor, uh, and now you're going into sort of the more, like, uh, orchestrate-y things like Airflow, Prefect, Daxter, those guys.
Yeah, I mean, we're, we're good friends with the Prefect guys, and Temporal have the same investors as us, and I'm sure that my investor Bogumil would not be too happy if I was like, "Oh yeah, by the way- ...
as well as trying to take on Datadog, we're also going off and trying to take on Temporal and everyone else doing that." Obviously, we're not doing all of the, the infrastructure of deploying that, right? Yet, at least. We're, you know, we're just building a Python library.
And, like, what's crazy about our graph implementation is sure, there's a bit of magic in, like, introspecting the return type, you know, extracting things from a union, stuff like that. But, like, the actual calls, as I say, is literally call a function and get back a thing and call that.
It's, like, incredibly simple and therefore easy to maintain. The question is, how useful is it? Well, I don't know yet. I think we have to go and find out. We have a whole-- We've had a, like, slew of people joining our Slack over the last few days and saying, "Tell me how good Pydantic AI is.
How good is Pydantic AI versus LangChain?" And I refuse to answer. That's your job to go and find that out, not mine. We've built a thing. I'm compelled by it, but I'm obviously biased. The ecosystem will work out what the useful tools are.
Bogumil was my board member when I was at Temporal, and, uh-
Yeah
... I, I think, I think just generally also having been a workflow engine investor and participant in this space, uh, it's a big space. Like, everyone needs different flavors of orchestration.
Yeah.
The one thing that I would say, like, y- yours, you know, as a library, you don't have that much control over, over the infrastructure. I do like the idea that- Each new agents or whatever unit of work, whatever you call that, should spin up in its sort of isolated boundaries.
Yeah.
Whereas y-yours, I think, around everything runs in the same process, but you id-ideally wants to sort of spin out its own little container of things.
I agree with you a hundred percent. And we will... It would work now, right? As in, in theory, you're just like, as long as you can serialize-
Yeah
... the calls to the next node, you just have to-- all of the different containers basically have to have the same, the same code. I mean, I'm super excited about Cloudflare Workers running Python and being able to install dependencies.
And if Cloudflare could only give me my invitation to the, the private beta of that, we would be exploring that right now, because I'm super excited about that as a, like, compute level for some of this stuff, where exactly what you're saying, basically.
You can run everything as an individual like worker function and distribute it, and it's resilient to failure, et cetera, et cetera.
And it spins up like a thousand instances simultaneously. You know, you, you want it to be sort of truly serverless, uh, at once. Uh, actually, I, I know we have some Cloudflare friends who are listening, so, uh, hopefully we'll-
In London especially
... they'll get in front of the line.
I was in Cloudflare's office last week shouting at them about other things that frustrate me. I have a love-hate relationship with Cloudflare. Their, their tech is awesome, but because I use it the whole time, I then get frustrated.
So, um, yeah. I'm sure I will, I will, I will get there soon.
Just a side tangent on Cloudflare. Is Python supported, uh, full? I actually wasn't fully aware that of what the status of that thing is.
Yeah. So, so Pyodide, which is Python running inside the browser, in scripting, is supported now by Cloudflare. They just, they basically, they're having some struggles working out how to manage, ironically, dependencies that have binaries, in particular Pydantic, because these workers where you can have thousands of them on a given metal machine, you don't wanna have a diff-- you, you basically wanna be able to have a shared, shared memory for all the different Pydantic installations effectively.
That's the thing they, they work out. They're, they're working out. But Hood, who's my friend, who is the primary maintainer of Pyodide, works for Cloudflare, and that's basically what he's doing, is working out how to get Python running on, running on Cloudflare's, like, network.
I mean, the, the nice thing is that your binary is really written in Rust, right?
Yeah.
Which also compiles the WebAssembly.
Yeah.
So maybe there's a way that you build-- you have just a different build of, uh, Pydantic and, uh, that ships with, uh, with whatever your distro for Cloudflare Workers is.
Yes, that's exactly... So Pyodide has builds for, for Pydantic Core and for things like NumPy and basically l- all of the popular, like, binary libraries. Yeah, it's just basic. And you're doing exactly that, right? You're using Rust to compile to WebAssembly, and then you're calling that shared library from Python, and it's unbelievably complicated, but it works.
Okay. Staying on graphs a little bit more, and then I wanted to go to some of the other features that you have in Pydantic AI. I see in, in your docs there are sort of four levels of agents.
There's single agents, there's agent delegation, programmatic agent handoff, uh, that seems to be what OpenAI Swarms would be like, and then the last one, graph-based control flow. Would you say, like, that those are sort of the mental hierarchy of, of how these things go?
Yeah, roughly.
Okay. Uh, you, you, you had some, uh, expression around OpenAI Swarms.
Well, and, and indeed, OpenAI have got in touch with me and basically, maybe I'm not supposed to say this, but basically said- ... uh, that Pydantic AI looks like what Swarms would become if it was production ready. So-
Yeah
... yeah, I mean, like, yeah, which makes sense.
Awesome.
Yeah, I mean, in fact, it was specifically saying, how can we give people the same feeling that they were getting from Swarms that led us to go and implement graphs? Because my like, "Just call the next agent with Python code" was not a satisfactory answer to people, so it was like, okay, we gotta go and have a better answer for that, that like led us to get to graphs.
Yeah. I mean, it's a minimal viable, uh, graph in, in, in some sense. So what are the s- the shapes of graphs that people would... should know? Uh, so the way that I would phrase this is, I think Anthropic did a very good public service and also kind of surprisingly influential blog post, I would say, when they wrote Building Effective Agents.
Uh, we actually have the authors coming to speak, uh, at my conference in New York, uh, which I think you're giving a workshop at.
Yeah. I'm trying to work it out-
Yeah
... but yes, I think so.
Tell me if you're not. But, but yeah, I mean, like I... That, that was the first, I think, authoritative view of like what kinds of graphs exist in agents, and let's give each of them a name so that everyone is on the same page.
So I'm just kinda curious if, if you have community names or top five patterns of graphs.
I don't have top five pu-patterns of graphs. I, I would love to see what people are building with them, but, but like it's been, it's only been a couple of weeks. And of course, the, the point is that like because they're relatively unopinionated about what you can go and do with them, they don't suit...
Like, you can go and do lots of, lots of things with them, but they don't have the structure to, to go and have like specific names as, as much as perhaps like some other systems do. I think what our agents are, which have a name, and I, I can't remember what it is, but this basically system of like decide what tool to call, go back to the center, decide what tool to call, go back to the center, and then exit.
One form of graph, which as I say, like our agents are effectively one implementation of a graph, which is why under the hood they are now using graphs. And it'll be interesting to see over the next few years whether we end up with these like predefined graph names or graph structures or whether it's just like, "Yep, I built a graph."
Or whether graphs just turn out not to match people's mental image of what they want and die away. We'll see.
I think there is always a appeal, um. Every developer eventually gets graph religion and goes, "Oh, yeah, everything's a graph." And then they, they probably over rotate and go, go too far into graphs.
Yeah.
And then they have to learn a whole bunch of DSLs, and then they're like, "Actually, I didn't need this," and they scale back a little bit.
Yeah. I'm at the, I'm at the beginning of that process. I'm, I'm currently a like graph maximalist. Although I haven't... can't say I've put actually any into production yet, but, um, yeah.
This has a lot of philosophical connections with other work coming out of UC Berkeley on compound AI systems. I don't know if you know of O'Hare. This is the Gartner ... world of, of things where they, they need some kind of industry terminology to, to sell it to enterprises.
I don't know if you, you, you know about any of that stuff.
I haven't. I probably should. I should probably do it 'cause I should probably get better at selling to enterprises. But no, no I don't. Not right now.
This is really the, the argument is that instead of putting everything in one model, you have more control and more maybe observability to if you break everything out into composing little models-
Yeah
... and chaining them together. Uh, and obviously then you'll need an orchestration framework to do that.
Yeah. And, and it makes complete sense. And one of the things we've seen with agents is they work well when they work well, but when they-- even if you have the observability through Logfire that you can see what was going on, if you don't have a nice hook point to say, "Hang on, this has all gone wrong," you have a, a relatively blunt instrument of basically erroring when you exceed some kind of limit.
But, like, what you need to be able to do is effectively iterate through these runs so that you can have your own control flow where you're like, "Okay, we've gone too far." And that's where one of the neat things about our graph implementation is you can basically call next in a loop rather than just running the full graph, and therefore you have this opportunity to, to break out of it.
But, but yeah, basically it's the same point, which is like if you have too big a unit of work, to some extent whether or not it involves gen AI, but obviously it's particularly problematic in gen AI, you only find out afterwards when you've spent quite a lot of time and/or money- ...
when it's gone off and done, done the wrong thing.
I'll drop on this. We're not gonna resolve this here, but I'll drop this and then we can move on to the next thing.
Yeah.
This is the common way that we, we developers talk about this, and then the machine learning researchers look at us and laugh and say, "That's cute," and then they just train a bigger model, and they wipe us out in the next training run.
So I think there's a certain amount of we are fighting the bitter lesson here. We're fighting AGI and, you know, when AGI arrives, this will all go away. Obviously, on latent space we don't really discuss that because I think AGI is kind of this hand-wavy concept that isn't-
Yeah
... super relevant. Uh, but I think we have to respect that. For example, you could do chain of thoughts with graphs, and you could manually orchestrate a nice little graph that, that does, like, reflect, think about if you need more, more inference time computes, you know, that's the hot term now.
Yeah.
And then think again and, you know, scale that up. Or you could train Strawberry and, and DeepSeek or one.
Right. I saw someone saying recently, oh, they were, they were really optimistic about, uh, agents because models are getting faster exponentially. And I, like, took a certain amount of self-control not to describe that it wasn't exponential, but I-- my, my main point was if, if models are getting faster as quickly as you say they are, then we don't need agents and we don't really need any of these abstraction layers.
We can just give our model an, a, you know, access to the internet, cross our fingers and hope for the best. Agents, agent frameworks, graphs, all of this stuff is basically making up for the fact that right now the models are not that clever.
In the same way that if you're running a customer service business and you have loads of people sitting answering telephones, the less well-trained they are, the less that you trust them, the more that you need to give them a script to go through.
Whereas, you know, so if you're running a bank and you have the-- lots of, uh, customer service people who you don't trust that much, then you tell them exactly what to say. If you have-- If you're doing high net worth banking, you just employ people who you think are gonna be charming to other rich people and set them off to go and have coffee with people, right?
And the same is true of, of models, right? The more intelligent they are, the less we need to tell them, like, structure what they go and do and constrain the, the routes in which they take.
Yeah. Yeah. Yeah. Agree with that. Uh, so I'm happy to move on. There's sort of other, other parts of Pydantic AI that are, uh, worth commenting on, and this is, like, my last rant. I, I, I promise. So obviously every framework needs to do its sort of model adapter layer, which is: Oh, you can-
Yeah
... easily swap from OpenAI to Claude to Groq. Uh, you also have, which I didn't know about, Google GLA, which I, I didn't really s- know about until I saw this in your, in your docs, which is Generative Language API.
Model Gateway27:58
I assume that's AI Studio.
Yes. Google don't have good names for it. So Vertex is very clear. That seems to be the API that, that, like, some of the things use, although it returns five oh three about twenty percent of the time, so.
Vertex?
No, Vertex fine, but the-
Oh, oh, GLA. Yeah
... Generative Language API.
Yeah, yeah. I agree with that.
Yeah.
Yeah.
So, so we have, again, another example of, like, where I think we go the extra mile in terms of engineering is we run on every commit, at least commit to main, we run tests against the live models.
Oh, okay.
Not, not lots of tests, but like a handful of them. And we had a point last week where, yeah, GLA one was failing every single run, one of their tests would fail, and we-- I think we might even have commented out, out that one at the moment.
So, like, all of the models fail more often than you might expect, but, like, that one seems to be particularly likely to fail. But, but Vertex-
Yeah
... seems to... is the same API, but much more reliable.
My rant here is that, you know, versions of this appear in LangChain and all-- every single framework has to have its own-
Yep
... little thing or, uh, version of that. I would put to you, and then, you know, this is, this can be agree to disagree, that y- this is not needed in Pydantic AI. I would much rather you adopt a layer like LiteLLM or what's the other one in, in, in JavaScript, Portkey.
And that's their job. They focus on that one thing, and they, they normalize APIs for you. All new models are automatically added, and you don't have to duplicate this in- inside of your framework. So for example, if I wanted to use DeepSeek, I'm out of luck 'cause Pydantic AI doesn't have DeepSeek yet.
Yeah, it does.
Oh, it does. Okay. I'm sorry. Yeah, but you know what I mean. Does-- Should this live in your code, or should it live in a, in a layer that's kind of like your API gateway that's a, that's a defined piece of infrastructure that people have?
And I think if, if a company who are well-known, who are, like, respected by everyone had come along and done this at the right time, maybe we should have done it a year and a half ago and said, "We're gonna be the, like, universal AI layer," that would have been a, like, credible thing to do.
I've heard varying reports of LiteLLM, uh, is the truth, and it didn't seem to have exactly the type of safety that we needed. Also, as I understand it, and again, I haven't looked into it in great detail, part of their business model is proxying the request through their, through their own system to do the generalization- That would be an enormous put off to, to an awful lot of people.
Honestly, the truth is I, I don't think it is that much work unifying the model. I get where you're coming from. I kind of see your point. I think the truth is that everyone is centralizing around OpenAI's SD-- uh, API as the, as the one to do.
So DeepSeek, uh, support that. Groq with a K support that. Uh, Ollama also does it. Well, for-- I mean, if there is that library right now, it's more or less the OpenAI SDK.
Right. Okay.
And it's, like, very high quality. It's w-well type checked. It uses Pydantic, so I'm biased, but, I mean, I think it's, I think it's pretty well respected anyway.
There's different ways to do, do this because also, like, not-- it's not just about normalizing APIs, you have to do secret management and, and all that stuff.
Yeah. And there's also... because there's, like, Vertex will-- there's Vertex and Bedrock, which to one extent or another, effectively they host multiple models, but they don't unify the API, but they do, uh, unify the auth, as I understand it.
Although we are-- we're halfway through doing Bedrock, so I, I don't know about it that well. But, like, they're kind of weird hybrids because they support multiple models, but, but like, like I say, the auth is centralized.
Yeah. I'm surprised they don't unify the API. That s- that seems like, uh, something that I would do. You know, we can-
Yeah
... we can discuss all this all day. There's a lot of APIs.
I agree it would be nice if there was a universal one that, that we didn't have to go and build.
Test & Eval31:39
And I guess the other side of, you know, writing model and picking models like evals, how do you actually figure out which one you should be using? I know you have one. First of all, you have very good support for mocking in unit tests, which is something that a lot of other frameworks don't do.
So, you know, my favorite Ruby library is VCR, um, because it just, you know, it just lets me store the HTTP request and replay them. That part I'll kind of skip. I think you have basically like this test model where like just through Python, you try and figure out, uh, what the model might respond without actually calling the model, and then you have the function model where people can kinda customize outputs.
Yeah.
Any other fun stories maybe from there, or is it just what you see is what you get, so to speak?
On those two, I think what you see is what you get. On the evals, I think watch this space. I think it's something that, like, again, I was s-somewhat cynical about for some time. Still have my cynicisms about some of the...
Well, it's unfortunate that so many different things are called evals, and it would be nice if we could agree what they are and what they're not. But look, I think it's a really important space. I think it's something that we're gonna be working on soon, both in Pydantic AI and in Logfire to try and support better because, like, it's an unsolved problem.
Yeah. You do say in your doc that anyone who claims to know for sure exactly how your eval should be defined can safely be ignored.
We'll delete that sentence when we tell people how to do their evals.
Exactly. I was like, "We need, we need a snapshot of this today." And so let, let's talk about eval. So there's kinda like the vibe evals, which is what you do when you're building, right? Because you cannot really, like, test it that many times to get statistical significance.
And then there's the production eval. So you also have Logfire, which is kinda like your observability product, which I tried before. It's very nice. What are some of the learnings you've had from building an observability tool for LLMs?
Um, and yeah, a-as people think about evals even, like, what are the right things to measure? What are, like, the right number of samples that you need to actually start making decisions?
I'm not the best person to answer that is the truth, so I'm not gonna come in here and tell you that I think I know the answer, uh, on the exact number. I mean, we can do some back of the envelope statistics calculations to work out that, like, having thirty probably gets you most of the statistical value of g- having two hundred for, you know, by definition fifteen percent of the work.
But the exact, like how many examples do you need, for example, that's a much harder question to answer because it's, you know, it's deep within the how models, uh, operate. In terms of Logfire, one of the reasons we built Logfire the way we have, and we allow you to write SQL directly against your data, and we're trying to build the, like, powerful fundamentals of observability, is precisely because we know we don't know the answers.
And so allowing people to go and innovate on how they're gonna consume that stuff and what-- how they're gonna process it is in... we think that's valuable. Because even if we come along and offer you an evals framework on top of Logfire, it won't be right in all regards, and we want people to be able to go and innovate and being able to write their own SQL, connect to the API, and effectively query the data like it's a database with SQL allows people to innovate on that stuff.
And then that's what, you know, allows us to do it as well. I mean, we do a bunch of, like, testing what's possible by basically writing SQL directly against Logfire a-as any user could. I think the other, the other really interesting bit that's going on in observability is OpenTelemetry is centralizing around semantic attributes for GenAI.
So it's a relatively new project. A lot of it's still being added at the moment, but basically the idea that, like, they unify how, uh, both SDKs and/or agent frameworks send observability data to, to any OpenTelemetry en-endpoint. And so again, we, we can go and having that unification allows us to go and, like, basically compare different libraries, compare different models much better.
That stuff's in, uh, in a very like early stage of development. It's one of the things we're gonna be working on pretty soon is basically, I suspect Pydantic AI will be the first agent framework that implements those semantic attributes properly because, again, we control Pydantic AI, and we can say this is important for observability.
Whereas most of the other agent frameworks are not maintained by people who are trying to do observability, with the exception of LangChain, where they have the observability platform, but they chose not to go down the OpenTelemetry route, so they're like plowing their own furrow, and there's, you know, they're a lot...
They're even further away from standardization.
Observability35:51
Can you maybe just give a quick overview of how OTEL ties into the AI workflows? There's kinda like the question of is, you know, a trace and a span like a LLM call? Is it the agent is kinda like the broader thing you're, you're tracking?
How should people think about it?
Yeah. So they have a... There's a PR that I think may have now been merged from someone at IBM talking about remote agents and trying to support this concept of remote agents within GenAI. I'm not particularly compelled by that because I don't think that, like, that's actually by any means the common use case, but, like, I suppose it's fine for it to be there.
The m-majority of the stuff in OTEL is basically defining how you would instrument- A given call to an LLM. So basically the actual LLM call, what, what data you would, uh, send to your telemetry provider, how you would structure that.
Apart from this slightly odd stuff on, uh, remote agents, most of the, like, agent level consideration is not yet implemented in-- is not yet decided effectively, and so there's a bit of ambiguity. Obviously, what's good about OTel is you can in the end send whatever attributes you like.
But yeah, there's quite a lot of churn in that space in exactly how we store the data. I think the-- one of the most interesting things, though, is that if you think about observability traditionally, it was, sure, everyone would say, "Our observability data is very important.
We must keep it safe." But actually, companies work very hard to basically not have anything that sensitive in their observability data. So if you're a doctor in a hospital and you search for a drug for an STI, the sequel might be sent to the observability provider, but none of the parameters would.
So it wouldn't have the, like, patient number or their name or the drug. With GenAI, that, that distinction doesn't exist because it's all just, like, messed up in the text. If you have that same patient asking an LLM how to-- what drug they should take or how to stop smoking, you can't extract the PII and not send it to the observability platform.
So the sensitivity of the data that's gonna end up in observability platforms is gonna be at a, like, basically different order of magnitude to what's in a-- what, what you would normally send to Datadog. Of course, you could make a mistake and send someone's password or their card number to Datadog, but that would be seen as a, as, as a, like, mistake, whereas in GenAI, a lot of data is gonna be sent.
And I think that's why companies like LangSmith and are trying hard to offer observability on-prem, because there's a bunch of companies who are happy for Datadog to be cloud-hosted, but want self-hosted, self-hosting for this observability stuff w-with GenAI.
And are you doing any of that today? Because I know in each of the spans you have, like, the number of tokens, you have the context. You're just storing everything and then you're gonna offer kinda like a self-hosting for the platform, basically is the idea.
Yeah. We'll offer self-host. So, so we have scrubbing roughly equivalent to what the other observability platforms have. So if we, you know, if we see password as the key, we won't send the value. But like, like I say, that doesn't really work in GenAI.
So we're, we're accepting we're gonna have to store a lot of data, and then we'll offer self-hosting for those people who can afford it and who need it.
And then this is, I think, the first time that most of the workload's performance is depending on a third party. You know? Like, if you're looking at Datadog data, usually it's your app that is driving the latency and, like, the memory usage and all of that.
Here you're gonna have spans that maybe take a long time to perform because the GLA API is not working or because, like, OpenAI is kinda like, uh, overwhelmed. Do you do anything there since, like, the provider is almost, like, the same across customers, you know?
Like, are you trying to surface these things for people and say, "Hey, this was, like, a very slow span, but, like, actually all customers using OpenAI right now are seeing the same thing, so maybe don't worry about it," or...?
Not yet. We do a few things that, uh, people don't generally do in OTel. So we send, we send information at the beginning of a trace as well as, uh, sorry, at the beginning of a span, as well as when it finishes.
By default, OTel only sends you data when the span finishes. So if you think about a request which might take, like, twenty seconds, even if some of the intermediate spans finished earlier, you can't basically place them on the page until you get the, the top level span.
And so if you're using standard OTel, you can't show anything until those requests are finished. When those requests are taking a few hundred milliseconds, it doesn't really matter. But when you're doing GenAI calls or when you're, like, running a batch job that might take thirty minutes, that, like, latency of not being able to see the span is, like, crippling to understanding your application.
And so we've-- we do a bunch of slightly complex stuff to basically send data about a span as it starts, which is closely related.
Yeah. Any thoughts on all the other people trying to build on top of OpenTelemetry in different languages too? There's, like, the OpenLLMmetry project, which doesn't really roll off the tongue. But how do you see the, yeah, the, the future of this kind of tool set?
Is everybody gonna have to build... Why does everybody want to build their own open source observability thing to then sell?
I mean, we, we are not going off and trying to instrument the likes of the OpenAI SDK with the new semantic attributes, because at some point that's gonna happen, and it's gonna live inside OTel, and we might help with it.
But we're a tiny team, we don't have time to go and do all of that work. So OpenLLMmetry, like, interesting project, but I suspect eventually most of those semantic, like, that instrumentation of the big-- of the SDKs will live, like I say, inside the main OpenTelemetry repos.
What happens to the agent frameworks, what data you basically need at the framework level to get the context is kind of unclear. Uh, I don't think we know the answer yet. But I mean, I was on the... I guess this is kinda semi-public, because it was on-- I was on the call with the OpenTelemetry call last week talking about GenAI, and there was someone from Arise talking about the challenges they have trying to get OpenTelemetry data out of LangChain, where it's not, uh, like, natively implemented.
And obviously they're having quite a tough time, and I was realizing, hadn't really realized this before, but how lucky we are to primarily be talking about our own agent framework where we have the control, rather than trying to go and instrument other people's.
Sorry, I actually didn't know about this semantic conventions thing. It looks like, yeah, it's, it's merged into main OTel. What should people know about this? I had n- I never heard of it before.
Yeah. I, I, I think it looks like a, looks like a great start. I think there's some, there's some unknowns around how you send the messages that go back and forth, which is kind of the m-most important thing of all, and that has moved out of attributes and into OTel events.
OTel events in turn are moving from being on a span to being their own top level API where you send data. So there's a, like, there's a bunch of churn still going on. I'm impressed by how fast the OTel community is moving, moving on this project.
I guess they, like everyone else, get that, that, you know, this is important and it's something that, like, people are crying out to get instrumentation of. So, so I'm, I'm kind of pleasantly surprised at how fast they're moving, but it makes sense.
I, I'm just kinda browsing through the specification. Like, I can already see that this basically bakes in whatever the previous paradigm was. So now they, they have, like, you know, genai.usage.prompt tokens and genai.usage.completion tokens, and obviously now we have reasoning tokens as well.
And then only one form of sampling, which is top P. But you're basically baking in, or sort of reifying- ... things that you think are important today, but like it, it's not a super foolproof way of-
Yeah
... doing this for a future.
Yeah. I mean, that's what's neat about OTel is you can always go and send another attribute, and that's fine.
Sure.
It's just there are a bunch that are agreed on.
Okay.
But I would say, you know, to come back to your previous point about whether or not we should be relying on one centralized abstraction layer, this stuff is moving so fast that if you start relying on someone else's standard, you risk basically falling behind because you're relying on someone else to keep things up to date.
Or you fall behind because you got, you got other things going on.
Yeah, yeah. That's fair. That's fair.
Database43:19
Any other observations just about building Logfire, actually? Uh, let's just talk about this. You had recently announced, uh, Logfire. I was kind of only familiar with Logfire because of your Series A announcement. I actually thought you were making a separate company.
I remember some amount of confusion with you when, when that, when that came out. So, uh, yeah, so, so to be clear, it's Pydantic Logfire, and the company is, is one company that has kind of two products, an open source thing and a, and a observability thing, correct?
Yeah.
I- I was just kinda curious, like any learnings building Logfire? Like, uh, so classic question is, do you use ClickHouse? Is this like the, the, the standard persistence layer?
Uh-
Any, any learnings doing that?
We don't use ClickHouse. We started building our database with ClickHouse, moved off ClickHouse onto Timescale, which is a Postgres extension to do analytical databases.
Wow.
And then moved off Timescale onto DataFusion, and we're basically now building-- Well, it's, it's DataFusion, but it's, uh, kind of our own database. Bogumil is not entirely happy that we went through three databases before we chose one, I'll say that, but, like, we've got to the right one in the end.
I think we could have realized that Timescale wasn't right. I think ClickHouse-- They both taught us a lot, and we- we're in a great place now, but like, yeah, it's been a, it's been a real journey on the database in particular.
Okay. So, you know, as a database nerd, I have to like double-click on this, right? Like, so ClickHouse is supposed to be the ideal-
Yeah
... back end for anything like this. And then moving from ClickHouse to Timescale is another counterintuitive move that I didn't expect because, you know, Timescale is like an extension on top of Postgres, not super meant to-- for like high volume logging.
But like, yeah, tell us, tell us that those, uh, decisions.
So at the time, ClickHouse did not have good support for Post-- for JSON. I was speaking to someone yesterday and said ClickHouse doesn't have good support for JSON and got roundly stepped on because apparently it does now. So they've obviously gone and built their proper JSON support.
But like back when we were trying to use it, uh, I guess a, a year ago or a bit more than a year ago, everything had to be a map, and maps are a, are a pain to try and do like looking up JSON type data.
And obviously all these attributes, everything you're talking about there in terms of the gen AI stuff, you can choose to make them top-level columns if you want, but the simplest thing is just to put them all into a big JSON pile.
Uh, and that was a problem with ClickHouse. Also, ClickHouse had some really ugly edge cases, like by default or at least until I complained about it a lot, ClickHouse thought that two nanoseconds was longer than one second 'cause they compared intervals just by the number, not the unit.
And I complained about that a lot, and then they caused it to raise an error and just say you have to have the same unit. Then I complained a bit more and event- I think, as I understand it now, they have some-- they convert between units.
But like stuff like that when all you're looking at is, when a lot of what you're doing is comparing the duration of spans, was really painful. Also, things like you can't subtract two date times to get an interval.
You have to use the date sub function. But like the fundamental thing is because we want our end users to write SQL, the like quality of the SQL, how easy it is to write matters way more to us than if you're building like a platform on top where your developers are gonna write the SQL and once it's written and it's working, you don't mind too much.
So I think that's like one of the fundamental differences. The other problem that I have with ClickHouse and in fact Timescale is that like the ultimate architecture, the like Snowflake architecture of binary data in object store queried with some kind of cache from nearby, they both have it, but it's closed source, and you only get it if you go and use their hosted versions.
And so even if we had got through all the problems with Timescale or ClickHouse, we would end up like, you know, they would wanna be taking their eighty percent margin, and then we would be wanting to take... That would basically leave us less space for margin.
Whereas DataFusion properly open source, all of that same tooling is open source. And, and for us as a team of people with a lot of Rust expertise, DataFusion, which is implemented in Rust, we can literally dive into it and go and change it.
So for example, I found that there were some slowdowns in time-- in DataFusion's string comparison kernel for doing like string contains, and it's just Rust code, and I could go and rewrite the, the string comparison kernel to be faster.
Or for example, DataFusion when we started using it didn't have JSON support. Obviously, as I've said, it's something we needed. We were a-- I was able to go and implement that in a weekend using our JSON parser that we built for Pydantic Core.
So it's the fact that like DataF- DataFusion is like for us the perfect mixture of a toolbox to build a database with, not a database. And we can go and implement stuff on top of it in a way that like if you were trying to do that in Postgres or in ClickHouse-- I mean, ClickHouse would be easier 'cause it's C++, relatively modern C++, but like as a team of people who are not C++ experts, that's much scarier than DataFusion for us.
Yeah. That was a beautiful rant.
That's funny. Most people don't think they have agency on these projects. They're kinda like, "Oh, I should use this," or, "I should use that." They're not really like, "What should I pick so that I contribute the most back to it?"
You know? So, uh, but I think like you obviously have a open source first mindset, so that makes a lot of sense.
I think if we were probably better as a, a startup, uh, were a better startup and faster moving and just like, like headlong determined to get in front of customers as fast as possible, we should have just stuck with ClickHouse.
I hope that long term we're in a better place for having worked with DataFusion. Um, and we like, we're quite engaged now with the DataFusion community. Andrew Lamb, who maintains DataFusion, is an advisor to us. We're in a really good place now, but yeah, it's definitely slowed us down relative to just like building on ClickHouse and, and moving as fast as we can.
Competition48:34
Okay. We're about to zoom out and do Pydantic Run and all the other stuff, but, you know, my last question on Logfire is really, you know, at some point you run out of sort of community goodwill just because like, "Oh, I use Pydantic.
I love Pydantic. I'm gonna use Logfire." Okay. Then you start entering the territory of the Datadogs, the Sentries, and the Honeycombs.
Yep.
So where are you going to really spike here? What differentiator here?
I wasn't writing code in two thousand and one, but I'm assuming that there were people talking about like web observability. And then web observability stopped being a thing, not because the web stopped being a thing, but because all observability had to do web.
If you were talking to people in two thousand and ten or two thousand and twelve, they would have talked about cloud observability. Now that's not a term because all observability is cloud first. The same is gonna happen to GenAI.
And so whether or not you're trying to compete with Datadog or with a rise in LangSmith, you've got to do first class-- you've got to do general purpose observability with first-class support for AI. And as far as I know, we're the only people really trying to do that.
I mean, I think Datadog are starting in, in that direction. And to be honest, I think Datadog is a much like scarier company to compete with than the AI-specific observability platforms because in my opinion, and I've also heard this from lots of, lots of customers, AI-specific observability where you don't see everything else going on in your app is not actually that useful.
Our hope is that we can build the first general purpose observability platform with first-class support for AI, and that we have this open source heritage of making-- of putting developer experience first that other companies haven't done. For all I'm a like fan of Datadog and what they've done, if you search Datadog logging Python, and you just try as a, like a non-observability expert to get something up and running with Datadog and Python, it's not trivial, right?
That's something Sentry have done amazingly well, but like there's enormous space in most of observability to do DX better.
Since you mentioned Sentry, I'm curious how you thought about licensing and all of that. Obviously, your MIT license, you don't have any rolling license like Sentry has, where like you can only use an open source, like the one-year-old version of it.
Was that a hard decision?
So to be clear, Logfire is closed source, so Pydantic and Pydantic AI are MIT licensed and like, you know, properly open source and then, and then Logfire for now is, is completely closed source. And in fact, the, the struggles that Sentry have had with licensing and the like weird pushback the community gives when they take something that's closed source and make it source available just meant that we just avoided that whole, whole subject matter.
I think the other way to look at it is like in terms of either headcount or revenue or dollars in the bank, the amount of open source we do as a company is we've gotta be up there with the, with the most prolific open source companies and like, like I say, per head.
And so we didn't feel like we were morally obligated to make Logfire open source. We have Pydantic. Pydantic is the, the, you know, a foundational library in, uh, Python. That and now Pydantic AI are our contribution to open source, and then Logfire is like openly for, for profit, right?
As in we're not claiming otherwise, we're not sort of trying to walk a line if it's open source, but we're really, we wanna make it hard to deploy, so you probably wanna pay us. We're trying to be straight that it's to pay for.
We could change that at some point in the future, but it's not a, an immediate plan.
Playgrounds51:48
All right, so the first one I, I saw this new, I don't know if it's like a product you're building, the Pydantic.run, which is a Python browser sandbox. What was the inspiration behind that? We talk a lot about code interpreter for, for lamps.
I'm an investor in a company called E2B, which is a code sandboxes as a service for remote execution.
Yep.
What's the Pydantic.run story?
So Pydantic.run is again, completely open source. I, I have no interest in making it into a product. We just needed a sandbox to be able to demo Logfire in particular, but also Pydantic AI. Um, so it doesn't have it yet, but I'm, I'm gonna add a basically a proxy to OpenAI and the other models so that you can run Pydantic AI in the browser, see how it works, tweak the prompt, et cetera, et cetera, and we'll have some kind of limit per day of what you can spend on it or like what, what the spend is.
The other thing we wanted to be able to do was to be able to, when you log into Logfire, we have quite a lot of drop off of like a lot of people sign up, find it interesting, and then don't go and create a project.
And I-- my intuition is that they're like, "Oh, okay, cool, but now I have to go and like open up my development environment, create a new project, do something with the right token. Ugh, I can't be bothered," and then they, they drop off and they forget to come back.
And so we wanted a really nice way of being able to say, "Click here," and you can run it in the browser and, and see what it does. As I think happens to all of us, I sort of started, started seeing if I could do it a week and a half ago, got something to run and then ended up, you know, improving it and suddenly I've spent, spent a week on it.
But I think it's useful.
Yeah, I remember maybe a couple, two, three years ago, there were a couple companies trying to build in-the-browser terminals exactly for this. It's like, you know, you go on GitHub, you see a project that is interesting, but now you gotta like clone it, run it on your machine.
Sometimes it can be sketchy. This is cool, especially since you already make all the, the docs runnable in your docs. Like you said, you kinda test them. It sounds like you might just have.
So yeah, the plan is that on every example in Pydantic AI-
Yeah
... there's a button that basically says Run, which takes you into Pydantic.run, has that code there, and depending on how hard we wanna push, we could also have it like hooked up to Logfire automatically. So there's a like, "Hey, just come and join the project," and you can see what that looks like in, in Logfire.
That's super cool. Super cool.
Yeah, I think that's one of the biggest, personally for me, one of the biggest drop-offs from open source projects is kinda like do this and then as long as something-- as soon as something doesn't work, I just drop off, so.
It takes some discipline. You know, like, um, there's been very many versions of this that I've been through in my career where, uh, you had to extract this code and run it, and it always falls out of date.
Often we would have these, this concept of transclusion where we have a separate code examples repo that we want with that, and they'll be pulled into our docs and-
Yeah
... it never wor- never really works. It takes a lot of discipline. So kudos to you on this.
And it was, it was years of maintaining Pydantic and people complaining, "Hey, that example's out of date now," that eventually we went and built, uh, pytest examples, which is another... The hardest to search for open source project we ever built because obviously, as you can imagine, if you search pytest examples, you get examples of, uh, how to use pytest.
But the pytest examples will basically go through both your code inside your docstrings to look for Python code and through markdown in your docs and extract that code and then run it for you and run linting over it and soon run type checking over it.
So and that's how we keep our examples up to date. But now, now we have these like hundreds of examples, all of which are runnable and, and self-contained or if they, if they refer to the previous example, it's already structured that they have to be able to import the code from the previous example.
So Why don't we give someone a nice place to just be able to actually run that using OpenAI and see what the output is?
Lovely.
All right, so that's kind of up at ending. In the notes here, I, I, I just like going through people's X, uh, account, not Twitter. Um, so for four years, you've been saying we need a plain text successor to Jupyter Notebooks.
Yep.
I think people maybe have gone the other way, which may get even more opinionated, like with, uh, X and, like, uh, all these kinda like notebook companies.
Well, yes. So in reply to that, uh, someone replied and said Marimo is that, and sure enough, Marimo is, is really impressive, and I've subsequently spoken to, spoken to the Marimo guys and got to angel invest in their, I can't re- I think it's seed round.
So, like, Marimo is very cool. It's, it's doing that, and they... Marimo also notebooks also run in the browser, again, using Pyodide. In fact, I nearly didn't build Pydantic.run because we were just gonna use Marimo. But my concern was that people would think Logfire was only to be used in notebooks, and I wanted something that, like, ironically felt more basic, felt more like a terminal to- so that no one thought it was, like, just for notebooks.
Yeah. There's a lot of, uh, notebookators out there.
And indeed, my-- I have very strong o- opinions about, you know, proper no- like Jupyter Notebooks, this idea that, like, you have to run the cells in the right order. I mean, a whole bunch of things. It's basically, like, worse than Excel or similarly bad to Excel.
Oh, so you, you are a notebookator then invested in a notebook.
I have this rant called Notebook, which was, like, my, my attempt to build an alternative that is but mostly just a rant about the ten reasons why no- notebooks are just as bad as Excel. But Marimo, et al, the, the new ones that are text-based at least solve a whole bunch of those problems.
Agree with that. Yes. I, I was kind of wishing for something like a better notebook, and then I saw Marimo. I was like, "Oh, yeah, these guys have... are ahead of me on, on this." Um-
Yep
... I don't know if I would do the sort of annotation-based thing. Like, you know, a lot of people love the, "Oh, annotate this function," and it, it just adds magic. I think similarly to what Jeremy Howard does with his stuff.
It seems a little bit too magical still, but hey, it's a big improvement on notebooks.
Yep. Yeah, agreed.
Just as on the LLM usage, like the IPYNB file, it's just not, it's just not good to put in, uh, in LLMs. So, um-
Yeah
... just that alone, I think should be a-
It's just not good to put in LLMs. True.
It's really not. They freak out. They're like-
True.
It's not good to put in GIF either. I mean, I freak out.
London57:44
Okay. Well, we will, we will kill IPYNB at some point. And yeah, any other takes? Uh, I was gonna ask you just, like, broaden out just about the, the London scene. You know, what's it like building out there, you know, over, over the pond?
I'm a, I'm an evening person, and the good thing is, like, I can get up late and then work late because I'm speaking to people in the US a lot of the time. So I got invited just earlier today to, uh, some drinks reception about AI at, uh, 10 Downing Street with the Prime Minister.
So I'm, I'm feeling positive about the UK right now on AI, but I think, look, w- like, everywhere that isn't the US and China knows that we're, like, way behind on AI. I think it's good that the UK is, like, beginning to say, "This is an opportunity, uh, not just a risk."
I keep being told, "You should be at more events. You should be, like, you know, hanging out with AI people more." My instinct is like I'd rather sit at my computer and write code. I think that, like, is probably a more effective way of, uh, getting people's attention.
I'm like, a bit of me thinks I should be sitting on Twitter, not on, not in San Francisco chatting to people. I think it's probably a bit of, a bit of a mixture, and I, I could probably do with being in the, in the States a bit more, and I think I'm gonna be over there a bit more this year.
But, like, uh, there's definitely the risk if you're in somewhere where everyone wants to chat to you about, about code, where you don't write any code, and that's, that's a failure mode.
I, I would say, yeah, uh, uh, definitely for sure there is a scene, and, um, you know, one, one way to really fail at this is to, is to just be involved in that scene and, and have that eat up your time.
But be at the right events and, uh, the ones that I'm running are, are, are, are, are good events hopefully.
Okay. Well, come to AI 工程 , of course.
What I say is, like, use those things to produce high-quality content that travels in a different medium than you normally would be able to. Because there's some selectivity, because there's a broad-- there's a focused community on that thing, uh, they will discover your work more.
It will be highly produced. You know, that, that's the pitch over there on, on why at least I do conferences. And then in terms of talking to people, I, I always think about this a, a three strikes rule.
So after a while it gets repetitive, but maybe like the first ten, 20 conversations you have with people, if the same stuff keeps coming up, that is an indication to you that people, like, want a thing, and it helps you prioritize in a more long form way than you can get in shallow interactions online, right?
Yeah.
So that in-person eye to eye, like, this is my, my pain at work, and you see the pain, and you're like, "Oh, okay. Like, if I do this for you, you will love, uh, our tool." And, like, you can't really replace that.
It's customer interviews, really.
Yeah, I agree entirely with that. I think that, I think there's a... You're, you're right on a lot of that. And I think that, like, it's very easy to get distracted by what people are saying on, on Twitter and LinkedIn.
That's another thing.
Very hard to correct for which of those people are, are actually building this stuff in production in, like, serious companies and which of them are on day four of learning to code, 'cause they have equally strident opinions. And in their, like, few characters, they, they seem equally valid.
But which one's real and which one's not, or, or which one is from someone who really knows their stuff is, is hard to know.
Anything else, Sam? What do you wanna get off your chest?
Nothing in particular. I think we... I've really enjoyed our conversation. I, I would say I think if anyone who has, like, looked at, at Pydantic AI, we know it's not complete yet. We know there's a bunch of things that are missing.
Embeddings, like, uh, storage, MCP and, and tool sets and stuff like that. We're trying to be deliberate and do stuff well, and that involves not being feature complete yet, but, like, keep coming back and looking in a few months because we're, we're pretty determined to get there.
We know that this stuff is like, whether or not you think that AI is gonna be the next Excel, the next internet, or the next industrial revolution, it's gonna affect all of us enormously. And so as a company, we, we get that, like, making Pydantic AI the best agent framework is existential for us.
You're also the first series A company I see that has no open roles for now. Every founder that comes on our podcast, the, the call to action is like, "Please come work with us."
We are not hiring right now. W- I wanna... I would love, uh, bluntly for Logfire to have a bit more commercial traction and a bit more revenue before I, before I hire some more people. It's quite nice having a few years of runway, not a few months of runway.
So I'm not in any, any great appetite to go and, like, destroy that runway overnight by hiring another, another 10 people. Even if, like, we h- the whole team is, like, rushed off their feet, kind of doing, as you said, like three to four startups at the same time.
Awesome, man. Thank you for joining us.
Thank you very much.






