Introductions0:00
Hey everyone, welcome to the Latent Space Podcast. This is Alessio, partner and CTO in residence at Decibel Partners, and I'm joined by my co-host Swyx, founder of Smol AI.
Hey, and today in the studio we have Kanjun from Imbue. Welcome.
Thank you.
So, uh, we, you and I have, I guess, crossed paths a number of times. Uh, and you're f-former- you're formerly named Generally, Generally Intelligent, and you've just announced your rename, rebrand in a huge, humongous way, so congrats on all of that.
Thank you.
And, uh, we're here to dive in, into deeper detail on Imbue. Uh, we'd like to introduce you, um, just, uh, uh, on a, on a high level basis, but then have you go into a little bit more of your personal side.
Um, so you graduated, uh, your BS and MS in MIT at, uh, at MIT, and you also spent some time at the MIT Media Lab, one of the most, most famous, I guess, computer hacking-
Ah, yeah
... labs in the world.
True.
Um, any, any fun stories from that time?
Yeah, I built a, uh, electronic textiles, so like boards that, uh, make it possible to make like soft clothing, uh-
Yeah
... like you can sew circuit boards into clothing and then makes clothing electronic. It's not that useful.
You wrote a book about that?
I wrote a book about it, yeah.
Yeah, yeah.
Teach-- Basically, the idea was to like teach young women computer science in this route, because what we found was that, uh, young girls, they would be like really excited about math until about sixth grade, and then they're like, "Oh, um, math is not, not good anymore, uh, because I, I don't feel like the type of person who does math or does programming.
But I do feel like the type of person who does crafting." So it's like, okay, what if you combine the two?
Yeah. Yeah. Awesome. Awesome. Um, always more detail to dive into on that. Um, but then you graduated MIT and you went to, went straight into biz ops at Dropbox, where you're, uh, eventually chief of staff, which is a pretty interesting role we can dive into later.
And then it seems like the founder bug hit you. You were basically a three times founder at Embark, Sorceress, and now Generally Intelligent/Imbue.
Mm-hmm.
Um, what should people know about you on the personal side that's not on your LinkedIn, that- ... um, you're something you're very passionate about outside of work?
Yeah, I think, um, if you ask any of my friends, they would tell you that I'm obsessed with agency, like human agency and human potential.
That's work. Come on.
That's not work. What are you talking about? Um...
Okay. Uh, so like, uh, what's, what's, what's an example of human agency that you try to promote?
Yeah, like, uh, with all of my friends, I have a lot of conversations with them that's like helping figure out what's blocking them.
Yeah.
Um, and I, I guess I do this with a team kinda automatically too. And I think about it for myself often, like building systems. I have a lot of systems to like help myself be more effective. At Dropbox, I used to give this onboarding talk called "How to Be Effective," um, which people liked.
I think like 1,000 people heard this onboarding talk, and I think maybe Dropbox was more effective. Um, and I think I just really believe that like, as humans, we can be a lot more than we are, um, and it's what drives everything.
I guess completely outside of work, I do dance. I do partner dance.
Nice.
Yeah.
Yeah. Lots of interest in, uh, that stuff, especially in like the sort of group living, um, houses in, in San Francisco, which I've been a little bit part of, and you've... I was also around one of those.
That's right, yeah. I started The Archive with two friends. Oh, with Josh, my co-founder-
Exactly
... and a couple of folks-
Some history there. Yeah
... in 2015. That's right. And GPT-3, our housemates built, so
Was that the, I guess, the precursor to Generally Intelligent that, that you started, um, doing more things with Josh? Is that how that relationship started?
Yeah, so Josh and I-
You kinda go into that story. Yeah
... yeah, Josh and I are f- uh, this is our third company together. Our first company, Josh, uh, poached me from Dropbox-
Oh
... for Embark. And, uh, there we built a really interesting technology, uh, laser raster projector VR headset. And then we were like, VR is not the thing we're most passionate about. And actually, it was, you know, kinda early days when we were...
both realized like we really do believe that in our lifetimes, like computers that are intelligent are going to be able to allow us to do much more than we can do today as people, and be much more as people than, than we can be today.
And, um, at that time, we actually, after Embark, we were like, should we s- like work on AI research or start an AI lab? Um, a bunch of our housemates were joining OpenAI, and we actually decided to do something more pragmatic, to apply AI to recruiting, and to try to understand like, okay, if we are actually trying to deploy these systems in the real world, what's required?
And that was Sorceress. That taught us so much about what, uh... That was maybe an AI agent in a lot of ways.
Mm.
Um, like what does it actually take to make a product that people can trust and rely on? Um, I think we never really fully got there, um, and it's taught me a lot about what's required. Um, and it's kind of like, I think informed some of our approach and some of the way that we think about how, uh, these systems will actually get used by people in the real world.
Just to go one step deeper on that. So you're build- you're building AI agents in 2016, before it was cool. Um, what... So you got some milestone. You raised $30 million. Um, something was working. So what, what do you think like you succeeded in doing, and then what did you, uh, try to do that did not pan out?
Yeah. So the product worked quite well. Uh, so Sorceress was an AI system, uh, that basically kind of looked for candidates that could be a good fit, and then helped you reach out to them. And this was, you know, a little bit early.
We didn't have language models to help you reach out, so we actually had a team of writers that like, you know, customized emails. Um, and we automated a lot of the customization. Uh, but the product was pretty magical.
Like candidates would just be interested and land in your inbox, and then you can talk to them. As a hiring manager, that's such a good experience. Um, I think there were a lot of learnings, both on the product and market side.
On the market side, recruiting is a market that is endogenously high churn, which means... Because people start hiring, and then we hire the role for them, and they stop hiring.
Mm.
So the more we succeed-
Oh
... the more they-
It's like the whole dating business.
It's the dating business.
Yeah.
Exactly. Exactly. It's exactly the same problem as the dating business.
Mm-hmm. Mm-hmm.
And I was really passionate about like, can we help people, you know, find work that is more exciting for them? A lot of people are not excited about their jobs, and a lot of companies are doing exciting things, and the matching could be a lot better.
Um, but the dating business kind of, um, phenomenon like put a damper on that.
Yeah.
So we, we had a good... Like it's actually pretty good business. Um, but- As with any business with, like, relatively high churn, the bigger it gets, the more revenue we have, the slower growth becomes, because Nick percent, if 30% of that revenue you lose year over year, then it becomes a, a worse business.
Yeah, so that was the dynamic we noticed quite early on, um, after our Series A. I think the other really interesting thing about it is we realized what was required for people to trust that these candidates were, like, well-vetted and had been selected for a reason.
Um, and it's what actually led us, you know, a lot of what we do at Imbue is working on interfaces to figure out, how do we get to a situation where when you're building and using agents, these agents are trustworthy to the end user?
That's actually one of the biggest issues with agents that, you know, go off and do longer range goals, is that I have to trust, like, did they actually think through this situation? Um, and that really informed a lot of our work today.
Origin Story7:13
Yeah. Let's jump into GI now, Imbue. Um, when did you decide recruiting was done for you, and you were ready for the, the next challenge? And how did you pick the agent space? I feel like in 2021 it wasn't as mainstream as it is today.
Yeah. So the LinkedIn says that it started in 2021, but actually we started thinking very seriously about it in early 2020, late 2019, early 2020. Um, not exactly this idea, but, uh, in late 2019, so I mentioned our housemates, Tom Brown and Ben Mann, they're the first two authors on GPT-3.
So what we were seeing is that scale is, scale is starting to work, um, and language models probably will actually get to a point where, like, with hacks, they're actually going to be quite powerful. And it was hard to see that at the time, actually, because, uh, like, GPT-3, the early versions of it, you know, there are all sorts of issues.
We're like, "Oh, that's not that useful." But we could kind of see, like, okay, you keep improving it in all of these different ways, um, and it'll get better. And so what Josh and I were really interested in is, how can we get computers that help us do bigger things?
Like, you know, there is this kind of future where I think a lot about, uh, you know, if I were born in 1900 as a woman, like, my life would not be that fun. Uh, I'd spend most of my time, like, carrying water and literally, like, getting wood to put in the stove to cook food, and, like, cleaning and scrubbing the dishes and, you know, uh, getting food every day because there's no refrigerator.
Like, all of these things, very physical labor. And what's happened over the last 150 years since the Industrial Revolution is we've kind of gotten free energy. Like, energy is way more free than it, it was 150 years ago.
And so as a result, we've built all these technologies, like the stove and the dishwasher and the refrigerator, and we have electricity, and we have, uh, infrastructure, running water, all these things that have totally freed me up to do what I can do now.
And I think the same thing is true for intellectual energy. We don't really see it today, but, like, because we're so in it, but our computers have to be micromanaged. You know, part of why people are like, "Oh, you're stuck to your screen all day."
Well, we're stuck to our screen all day because literally nothing happens unless I'm doing something in front of my screen. I don't, you know, I can't send my computer off to do a bunch of stuff for me. There is a future where that's not the case, where, you know, I can actually go off and do stuff and trust that my computer will pay my bills and figure out my travel plans and do the detailed work that I am not that excited to do so that I can, like, be much more creative and able to do things that I, as a human, am very excited about and collaborate with other people, and there are things that people are uniquely suited for.
So that's kind of always been the thing that is really exciting, uh, has been really exciting to me. Like, Josh and I have known for a long time, I think, that AI, uh, you know, whatever AI is, it would happen in our lifetimes.
And, um, and the personal computer kind of started giving us a bit of free intellectual energy, and this is, like, really the explosion of free intellectual energy. So in early 2020, we were thinking about this, and, uh, what happened was self-supervised learning basically started working across everything.
So it worked in language. Uh, SimCLR came out. I think MoCo had come out, Momentum Contrast had come out earlier in 2019. SimCLR came out in early 2020, and we were like, okay, for the first time, self-supervised learning's working really well across images and text, and suspect that, like, okay, actually, it's the case that machines can learn things the way that humans do.
Um, and if that's true, if they can learn things in a fully self-supervised way, uh, because, like, as people, we are not supervised. We, like, go Google things and try to figure things out. So if that's true, then, like, what the computer could be is much different, you know, is much bigger than what it is today.
And so we started exploring ideas around, like, how do we actually go-- We didn't think about the i- the fact that we could actually just build a research lab. So we were like, okay, what kind of startup could we build to, like, leverage self-supervised learning so that it eventually becomes something that allows computers to become much more, uh, kind of able to do bigger things for us?
But that became Generally Intelligent-
Reasoning Models11:26
Mm-hmm
... which started as a research lab.
Yep. And so your mission is, uh, you aim to rekindle the dream of the personal computer. So when did it go wrong? And what are, like, your first, um, products and kind of like, uh, user-facing things that you're building to rekindle it?
Yeah. So what we do at Imbue is we, uh, train large foundation models optimized for reasoning, and the reason for that is because reasoning is actually, we believe, the biggest blocker to agents or systems that can do these larger goals.
Um, if we think about, you know, something that writes an essay, like when we write an essay, we, like, write it. We don't just output it and then we're done. We, like, write it, and then we look at it, and we're like, "Oh, I need to do more research on that area.
I'm gonna go do some research and figure out." And then come back and, "Oh, actually, it's not quite right, the structure of the outline, so I'm gonna re- rearrange the outline, rewrite it." It's this very iterative process, and it requires thinking through like, "Okay, uh, what am I trying to do?
Is the goal correct?" Also, like, has the goal changed as I've learned more? Also, you know, as a tool, like when should I ask the user questions? I shouldn't ask them questions all the time, but I should ask them questions in higher risk situations.
Um, how certain am I about the, like, flight am I, I'm about to book? Um, there are all of these notions of like risk certainty, playing out scenarios, figuring out how to make a plan that makes sense, how to change the plan, what the goal should be, that are, uh, things, you know, that we lump under the bucket of reasoning.
And models today, they're not optimized for reasoning. It turns out that there's not actually that much explicit reasoning data on the internet, um, as you would expect. And so we get a lot of mileage out of optimizing our models for reasoning in pre-training.
And then on top of that, we build agents ourselves, and we, I can get into, we really believe in serious use, like really seriously using the systems and trying to get to an agent that we can use every single day, tons of agents that we can use every single day.
And then we experiment with interfaces, uh, that help us better interact with the agents. So those are some set of things that we do on the kind of model training and agent side. And then, uh, the initial agents that we build, a lot of them are trying to help us write code better, because code is most of what we do every day.
And then on the infrastructure and theory side, we actually do a fair amount of theory work to understand, like how do these systems learn, and then also, like what are the right abstractions for us to build good agents with, um, which we can get more into.
And, uh, if you look at our website, we have a lot of tools. Um, we build a lot of tools internally. We have a, like, really nice automated hyperparameter optimizer. Um, we have a lot of really nice infrastructure, and it's all part of the belief of like, okay, let's try to make it so that the humans are doing the things humans are good at, as much as possible.
So out of our very small team, we get a lot of leverage.
And so would you still categorize yourself as a research lab now, or are you now in startup mode? Is that-
Um-
... a transition that is conscious at all?
That's a really interesting question. I think we've always intended to build, you know, to try to build the next version of the computer, enable the next version of the computer. Um, the way I think about it is there's a right time to bring a technology to market.
So Apple does this really well, actually. iPhone was under development for 10 years, AirPods for five years. Um, and Apple has a story where, you know, iPhone, uh, the first multi-touch screen was created. They actually were like, "Oh, wow, this is cool.
Uh, let's like productionize iPhone." They actually brought, uh, they like did some work prod- trying to productionize it and realized this is not good enough, and they put it back into research to try to figure out, like how do we make it better?
What are the interface pieces that are needed? And then they brought it back into production. So I, I think of production and, and research as kind of like these two separate phases, and internally, we have that concept as well, um, where like things need to be done in order to get to something that's usable, and then when it's usable, like eventually we figure out how to productize it.
What's the culture like to make that happen? To have both like, kinda like product-oriented, research-oriented, and as you think about building the team, I mean, you just raised 200 million. I'm sure you wanna hire more people. Uh, what, what are like the, the right archetypes of people that work at Imbue?
Team Culture15:25
Hmm. Yeah, I would say we have a very unique culture in a lot of ways. Um, I think a lot about social process design, so how do you design social processes that enable people to be, you know, effective?
Um, I like to think about team members as creative agents. So- Because most companies, they think of their people as assets, and they're very proud of this. And I think about like, okay, what is an asset? It's something you own, uh, that provides you value that you can discard at any time.
This is a very low bar for people.
Yeah.
This is not what people are. Um, and so we try to enable everyone to be a creative agent and to really unlock their superpower. So a lot of the work I do, you know, I was mentioning ear- earlier, I'm like obsessed with agency.
A lot of the work I do with, with team members is try to figure out like, you know, what are you really good at? What really gives you energy, and where can we put you such that, um, and how can I help you unlock that and grow that?
Um, so much of our work, you know, in terms of team structure, like much of our work actually comes from people. Carbs, our hyperparameter optimizer, came from Abe trying to automate hi- his own research process, uh, doing hyperparameter optimization, and he actually pulled some ideas from plasma physics, he's, he's a plasma physicist, to make the local search work.
A lot of our work on evaluations comes from a couple members of our team who are like obsessed with evaluations. We do a lot of work trying to figure out like, how do you actually evaluate if the model is getting better?
Is the model making better agents? Is the agent actually reliable? Um, and so a lot of things kind of like, I think of people as making their, like them-shaped blob inside Imbue. And I think, you know, yeah, that's the kind of person, uh, kind of person that we're, we're hiring for.
We're hiring product engineers and data engineers and, uh, research engineers and all these roles. Um, you know, we have a project. We have projects, not teams. Um, we have a project around data, data collection and data engineering. That's actually one of the key things that improve the model performance.
We have a pre-training kind of project, uh, some, and with some fine-tuning as part of that. And then we have an agents project that's like trying to build on top of our models, as well as use other models, um, in the outside world to try to make agents that then we actually use as programmers every day, so all sorts of different, different projects.
As a founder, you're, you're now sort of a capital allocator among all these different investments, effectively, and different projects. Um, and I was interested in how you mentioned that you're, you're optimizing for, uh, im- improving reasoning, specifically inside of your pre-training, which I, I assume is just a lot of data collection.
We are optimizing reasoning inside of our, uh, pre-trained models, and a lot of that is about data, and I can talk more about like what, you know, what d- exactly does it involve. Um, but, uh, actually, big, maybe 50% plus of the work is figuring out, even if you do have models that reason well, like the models are still stochastic.
The way you prompt them still Makes, is kind of random. Like, makes them do random things. And so how do we get to something that is actually robust and reliable as a user? How can I, as a user, trust it?
You know, I was mentioning earlier, um, when I talk to other people building agents, they have to do so much work, like, to try to get to something that they can actually productize, and, um, it takes a long time, and agents haven't been productized yet for, partly for this reason, is that, like, the abstractions are very leaky.
Um, you know, we can get, like, 80% of the way there, but like self-driving cars, like, the remaining 20% is actually really difficult. We believe that, and we have internally, I think, um, some things that, like an interface, for example, that, um, lets me really easily, like, see what the agent execution is, fork it, try out different things, modify the prompt, um, modify, like, the plan that it, it is making.
Uh, this type of interface, it makes it so that I feel more like I'm collaborating with the agent as it's executing, as opposed to it's just, like, doing something as a black box. Um, that's an example of a, of a type of thing that's, like, beyond just the model pre-training.
But on the reasoning, yeah, on the model pre-training side, like, reasoning is a thing that we optimize for, and a lot of that is about, yeah, what data do we put in.
Code & Data19:47
Yeah. It, it's interesting just 'cause I, I always th- think like, you know, out of the levers that you have, the resources that you have, I think a lot of people think that running a foundation model company or a research lab is gonna be primarily compute.
And I think the share of compute has gone down a lot o- over the past three years.
Yeah.
Uh, it used to be the main story, like, the, the main way you scale is you just throw more compute at it. Uh, and now it's, like, FLOPs is not all you need. You need better data, you need better algorithms, and, uh, I wonder where that shift has gone.
I don't, uh, this is a very vague question, but is it, like, 30, 30, 30 now? Is it, like, maybe even higher? So, uh, one way I'll put this is, um, people estimate that Llama 2 maybe took about $3 to $4 million of compute, but probably $20 to $25 million worth of labeling data.
Um, and I'm like, "Okay. Well, that, that's a very different story than all these other foundation model labs raising hundreds of millions of dollars and spending it on GPUs."
Yeah. Uh, data is really expensive.
Um, we generate a lot of data, and so that does help. Um, the generated data is close to actually good, as good as human-labeled data. Um-
S- so generated data from other models?
From our own models-
From your own models
... or other models, yeah.
Do you feel like, uh, and this is, th- there's, there's s- there's certain variations of this. Uh, there's the sort of the constitutional, um, AI approach from Anthropic, and, um, basically models sampling, training on data from other models.
I feel like there's a little bit of, like, contamination in there, or, uh, to put it in a statistical form, you're resampling a distribution that you already have that you already know doesn't match human distributions.
Yeah.
Yeah. How do you feel about that, basically? Just philosophically.
S- so when we're optimizing models for reasoning, we are actually, like, trying to, like, uh, make a part of the distribution really spiky.
Avalon Lessons21:43
Yes.
So in a sense, like, that's actually what we want. We, we want to, because the internet is a sample of the human distribution that's also skewed in all sorts of ways.
Yes.
Um, you know, that is not the data, uh, that we necessarily want these models to be trained on. And so, uh, I don't worry about it that much. Like, what we've seen so far is that it seems to help.
When we're generating data, we're not, we're not really randomly generating data. We generate very specific things, uh, that are, like, reasoning traces and that help optimize reasoning. Code also is a big piece of improving reasoning, so yeah. Uh, generated code is not that much worse than, like, regular human-written code.
You might even say it, it can be better in a lot of ways.
Yeah.
So yeah. So we are trying to already do that.
What are some of the tools that you saw that you thought were not a good fit? So you built Avalon, which is, um, your own simulated world, and when you first started, the kinda like meta game was, like, using games to simulate things, uh, using, you know, uh, Minecraft, and then OpenAI's, like, the gym thing, and all these things.
And y- your thing, I think in one of your other podcasts, you mentioned, like, Minecraft is, like, way too slow to actually do any serious work. Uh, what led you to like-
Is this true? Yeah.
I didn't, I didn't say it.
Yeah, this, this I didn't say.
I don't know. That's above my pay grade. Uh, but Avalon is, like, 100 times faster than Minecraft for, for simulation.
Oh.
Um, when did you figure that out, that you needed to just, like, build your own thing? Was it, um, kinda like your engineering team was like, "Hey, this is too slow?" Was it more a long-term investment?
At that time, we built Avalon as a research environment to help us learn particular things. And one thing we were trying to learn is, like, how do you get an agent that is able to do many different tasks?
Uh, like, RL agents at that time and environments at that time, what we heard from other RL researchers was the, like, biggest thing keeping, holding the field back is lack of benchmarks that let us, uh, kind of explore things like planning and curiosity and things like that, and have the agent actually perform better if the agent has curiosity.
And so we were trying to figure out, like, okay, how can we have agents that are, uh, like, able to handle lots of different types of tasks in a, uh, without the, the reward being pretty handcrafted? Um, that's a lot of what we had seen, is that, like, these very handcrafted rewards.
And so Avalon has, like, a single reward, um, it's, you know, across all tasks. And what it taught us, and it also allowed us to kind of create a curriculum, so we could make the level more or less difficult.
And it taught us a lot. Uh, maybe two primary things. One is, with no curriculum, RL algorithms don't work at all.
Mm-hmm.
So that's actually really interesting. Um-
For, for the non-RL, RL specialist, what is a curriculum in your terminology?
Uh, so a curriculum in this particular case is, uh, basically the ... environment, Avalon lets us generate simpler environments and harder environments for a given task. What's interesting is that the simpler environments, you know, what you'd expect is the agent succeeds more often, so it gets more reward.
Uh, and so, you know, kind of my intuitive way of thinking about it is, okay, the reason why it learns much faster with a curriculum is it's just getting a lot more signal. And, uh, that's actually an interesting kind of like general intuition to have about training these things.
Uh, it's like what kind of signal are they getting and like, and what, and like how can, how can you help it get a lot more signal? Um, the second thing we learned is that, uh, reinforcement learning is not a good vehicle.
Like pure reinforcement learning is not a good vehicle for planning and reasoning. So these agents were not able to... They were able to learn all sorts of crazy things. They could learn to climb, like hand over hand in VR climbing.
They could learn to open doors, like very complicated, like multiple switches and a lever, uh, open, open the door. But, uh, they couldn't do any higher level things, and they couldn't do those lower level things consistently necessarily. Um, and as a user, we were like, "Okay, as a user, I do not want to interact with a pure reinforcement learning end-to-end RL agent.
As a user, like I need much more control over what that agent is doing." And so that actually started to get us on the track of thinking about, okay, how do we do the reasoning part in language? And we were pretty inspired by our friend Chelsea Finn at Stanford was, I think, working on SayCan at the time, um, where it's basically a, um, uh, an experiment where they have robots kind of, you know, trying to do different tasks, and they actually do the reasoning for the robot in natural language.
And it worked quite well, um, and that led us to start experimenting very seriously with reasoning.
How important is the language part for the agent versus for you to inspect the agent, you know? Like, is it the, the interface to kind of the human in- on the loop really important, or?
Yeah. I personally think of it as it's much more important for us, the human user.
Ooh.
So I think you probably could get end-to-end agents that work and are fairly general, um, at some point in the future, uh, but I think you don't want that. Like, we actually want agents that we can, like perturb while they're trying to figure out what to do.
Because, you know, even a very simple example, um, internally we have like a type error fixing agent, and we have like a test generation agent. Test generation agent goes off the rails all the time. And I wanna know, like, why did it generate this particular test?
What was it thinking? Did it consider, you know, the fact that, uh, this is calling out to this other function? Like formatter agent, if it ever comes up with anything weird, I wanna be able to debug, like what happened?
With RL end-to-end stuff, like we couldn't do that.
Uh, so it, it sounds like you have a bunch of agents operating internally within, uh, the company. Um, what's your most, I guess, successful agent, and what's your least successful one?
Yeah. A type of agent that works moderately well is like fix this w- uh, the color of this button on the website, or like, uh, like change the color of this button.
Which is now Sweep.dev is doing that. Exactly that.
Oh, perfect. Okay.
But like-
Well, we should just use Sweep.dev.
Well, I mean, uh, okay, I, I don't know how m- how often do you have to fix the color of the, of the button.
Right.
Right? Because all of them-
Yeah
... raise money on the idea that they can go further.
Yeah.
And my fear when encountering something like that is that there's some kind of unknown asymptote ceiling that's going to prevent them. That they're gonna run head on into that you've already run into.
Oh, we've definitely run into such a ceiling. Um-
But what is the ceiling? Is there a name for it? Like what, what, like-
Uh, I mean, for us, uh, we think of it as reasoning plus these tools.
Yeah.
So, um, reasoning plus abstractions, basically.
Yeah.
I think actually you can get really far with current models, um, and that's why it's so compelling. Like, we can pile debugging tools on top of these current models, have them critique each other, and, and critique themselves, and do all of these like, uh, you know, spend more compute and inference time, context hack, um, you know, retrieval augmented generation, uh, et cetera, et cetera, et cetera.
Like the pile of hacks actually does get us really far. And y- kind of like trying to get more signal out of the channel. Um, we don't like to think about it that way. It's what, it's what the default approach is, is like trying to get more signal out of this noising cha- noisy channel.
But the issue with agents is, as a user, I want it to be mostly reliable.
Mm-hmm.
Um, it's kinda like self-driving in that way. Like, it's not as bad as self-driving. Like in self-driving, you know, you're like hurtling at 70 miles an hour. It's like the hardest agent problem. But I think one thing we learned from Sorceress, and one thing we're learn- we've learned inter- like by using these things internally, is we actually have a pretty high bar for these agents to work.
Mm-hmm. Mm-hmm.
Um, you know, it is actually really annoying if they only work 50% of the time. And, uh, we can make interfaces to make it slightly less annoying. But yeah, there, uh, there is a ceiling that we've, we've en- encountered so far, and we need to make the models better, and we also need to make the kind of like interface to the user better, and also a lot of the like, you know, critiquing, uh, we have a lot of like generation methods, um, kind of like spending compute and inference time generation methods that help, uh, things be more robust and reliable.
But it's still not 100% of the way there. So to your question of like what agents work well and what doesn't work well, like most of the agents don't work well.
Mm-hmm.
And we're slowly making them work better by improving the underlying-
That's good
... model and improving these parameters.
I, I think that that's comforting for a lot of people who are feeling a lot of imposter syndrome not being able to make it work. And I think, uh, y- the, the fact that you share their struggles, I think, uh, also, uh, helps people understand how early this is.
Yeah, definitely. It's very early, and I hope what we can do is help people who are building agents actually, like be able to deploy them. Um, I think, you know, that's the gap that we see a lot of today, is everyone who's trying to build agents, to get to the point where it's robust enough to be deployable, it just, it's like an unknown amount of time.
Okay.
Yeah.
Well, so this goes back into what Imbue is gonna offer as a product or a platform. How are you going to actually help people deploy those agents?
Yeah. So our current hypothesis, I don't know if this is actually going to end up being the case, um, we've built a lot of tools for ourselves internally around, like, debugging, around, like, abstractions or techniques after the model generation happens, like after the language model generates, uh, the text, uh, like interfaces for the user, uh, and the underlying model itself, uh, like models talking to each other.
Maybe some set of those things, kind of like an operating system, some set of those things will be helpful for other people. Um, and we'll figure out what set of those things is helpful for us to make our agents...
Like, what we wanna do is get to a point where we can, like, start making an agent, deploy it, it's reliable, like very quickly. And there's a similar analog to software engineering, like in the early days, in the '70s, in the '60s, like to program a computer, like you have to go all the way down to the registers.
But, um, and write things in assem- uh, eventually we had assembly. That was like an improvement. Then we wrote programming languages with these higher levels of abstraction, and that allowed a lot more people to do this and much faster, and the software created is much less expensive.
And I think it's basically a similar route here, where we're like in the like bare metal phase of agent building, and we will eventually get to something with much nicer abstractions.
Robust Agents32:12
So you touched a little bit on the data before. We had this conversation with George Hotz, and we were like, "There's not a lot of reasoning data out there, and can the models really understand?" And his take was like, "Look, with enough compute, you're not that complicated as a human."
Like, the model can figure out eventually why certain decisions are made. What's been your experience? Like, as you think about reasoning data, like do you have to do a lot of like manual work? Or like is there a way to prompt models to kinda like extract the reasoning from actions that they see?
We don't think of it as, "Oh, throw enough data at it, and then it will figure out what, like what the plan should be." Uh, I think we're much more explicit. So we have a lot of thoughts internally, like many documents about what reasoning is.
You know, a way to think about it is, as humans, we've learned a lot of reasoning strategies over time. We are better at reasoning now than we were 3,000 years ago. Um, an example of a reasoning strategy is noticing you're confused.
Uh, and like then when I notice I'm confused, I should ask like, "Huh, what was the original claim that was made? What evidence is there for this claim?" Uh, et cetera, et cetera. Does the evidence support the claim?
Is the claim correct? This is like a reasoning strategy that was developed in like the 1600s, you know, f- with like-
Wow
... the advent of science, of science. That's an example of a reasoning strategy. There are tons of them we employ all the time, lots of heuristics that help us be better at reasoning. And, um, we didn't always have them, and because they're invented, like we can generate data that's much more specific to them.
So I think internally, yeah, we have a lot of thoughts on what reasoning is, and we generate a lot more specific data. We're not just like, "Oh, it'll figure out reasoning from this black box," um-
Mm
... or like, "It'll figure out reasoning from the data that, that exists."
Yeah. I mean, the scientific method is like a good example. And if you think about hallucination, right? And people are thinking, "How do we use these models to do net new, like scientific research?" And, you know, if you go back in time and the model was like, "Well, the Earth revolves around the sun," and people are like, "Man, this model is crap."
It's like, "What are you talking about?" Like, "The sun revolves around the Earth." It's like, how, uh, how do you see the future? Or like, do you think we can actually... Like, if the models are actually good enough, but we don't believe them, it's like, h-how, how do we, how do we make the two live together?
Say you're like, you use Imbue as a scientist to do a lot of your research, and Imbue tells you, "Hey, I think this is like a serious path you should go down," and you're like, "No, that sounds impossible," like how is that trust gonna be built, and like what are some of the tools that maybe are gonna be there to, to inspect it?
Yeah. So like one, one element of it is, uh, like as a person, like I need to basically get information out of the model such that I can try to understand what's going on with the model. So then the second question is like, "Okay, how do you do that?"
Uh, and that's kind of... Some of our debugging tools, they're not necessarily just for debugging, they're also for like interfacing with and interacting with the model. So like, if I go back in this reasoning trace and like change a bunch of things, what's gonna happen?
Like, what does it conclude instead? Um, so that kinda helps me understand, like what are its assumptions? Um, and it, you know, we think of these things as tools. Um, and, and so it's really about like, as a user, how do I use this tool effectively?
Like, I need to be willing to be convinced as well. Um, it's like, yeah, how do I use this tool effectively, and what can it help me with, and what can it tell me?
So there's a lot of mention of code in your, in your process, um, and I was hoping to dive in even deeper. I think we might run the risk of giving people the impression that you, you view code or you use code, um, just as like a, a, a tool within, within yourself, uh, with- within Imbue just, just for coding assistance.
And, and I think there's a lot of informal understanding about how adding code to language models improves their reasoning capabilities. I wonder if there's any research or findings that you have to share that, um, uh, talks about the intersection of code and reasoning.
Hmm. Yeah. So the way I think about it intuitively is like code is the most explicit example of reasoning data-
Structured
... on the internet.
Yeah.
Yeah. And it's not only structured, it's actually very explicit, which is nice. You know, it says, "This variable means this."
Yeah.
"And then it uses this variable, and then the function does this." Like, as people, when we talk in language, it takes a lot more to kind of like extract that, like explicit structure out of, like our, our language.
And so that's one thing that's really nice about code is it, I see it as almost like a curriculum for reasoning. I think we use code in all sorts of ways, like, uh, the code, the coding agents are really helpful for us to understand, like what are the limitations of the agents?
Uh, the code is really helpful for the reasoning itself. But also code is a way for models to act. So by generating code, it can act on my computer. And, you know, when we talk about rekindling the dream of the personal computer, kind of where I see computers going is, you know, like computers will eventually become these much more malleable things, where I, as a user, today, I have to know how to write software code, like, in order to make my computer do exactly what I want it to do.
But in the future, if the computer is able to generate its own code, then I can actually interface with it in natural language. Um, and so we, you know, one way we think about agents is, is kind of like a natural language programming language.
Uh, it's a way to program my computer in natural language that's much more intuitive to me as a user. And these interfaces that we're building are essentially IDEs for users to program our computers in natural language.
What do you think about the other, the different approaches people have, kind of like text first, browser first, like multi-ON? Um, the, what do you think the in- the best interface will be? Or like, what is your, you know, thinking today?
Uh, I think chat is very limited as an interface. It is sequential, um, where these agents don't have to be sequential. So with a chat interface, if the agent does something wrong, I have to, like, figure out how to, like, how do I get it to go back and start from the place I wanted it to start from?
So in a lot of ways, like, chat as an interface, I think Linus, Linus Lee you had on, on this. I really like how he put it. Chat as an interface is skeuomorphic. So in the early days, when we made word processors on our computers, they had notepad lines, because that's what we understood, uh, you know, these, like, objects to be.
Chat, like texting someone, is something we understand. So texting our AI is something that we understand. But today's word documents don't have notepad lines. Um, and similarly, the way we want to interact with agents, like chat, is a very primitive way of interacting with agents.
Uh, what we want is to be able to inspect their state and to be able to modify them and fork them and all of these other things. And we internally have kind of, like, think about what are the right representations for that?
Like, architecturally, uh, like, what are the right representations? What kind of abstractions do we need to build? And how do we build abstractions that are not leaky? Because if the abstractions are leaky, which they are today, like, you know, this stochastic generation of text is like a leaky abstraction.
I cannot depend on it. And that means it's actually really hard to build on top of. But our experience and belief is actually by building better abstractions and better tooling, we can actually make these, make these things non-leaky, and now you can build, like, whole things on top of them.
So these other, other interfaces, because of where we are, we don't think that much about them.
Cool. Yeah, I mean, you mentioned this is kind of like the Xerox PARC moment for AI. Um, and we had a lot of stuff come out of PARC, like the, yeah, the what you see is what you get editors and, like, MVC and all this stuff.
Yes.
But yeah. But then we didn't have the iPhone at PARC. We didn't have all these, like, higher things. What do you think it's reasonable to expect in, like, this era of AI? You know, call it, like, five years or so.
Like, what are, like, the things we'll build today, and what are things that maybe we'll see in kind of like the second wave of, of products?
I think the waves will be much faster than before. Like, what we're seeing right now is basically like a continuous wave. Let me zoom a little bit earlier.
Mm-hmm.
So people like the Xerox PARC analogy I give, but I, I think there are many different analogies. Like, one is, uh, the, like, analog to digital computer is another analogy to where we are today. The analog computer Vannevar Bush bu- built in the 1930s, I think, and it's like a system of pulleys.
Mm-hmm.
And it can only calculate one function. Like, it can calculate, like, an integral, and that was so magical at the time, 'cause you actually did need to calculate this integral a bunch. But it had a bunch of is- issues, like in analog, errors compound.
And so there was actually a set of breakthroughs necessary, uh, in order to get to the digital computer, like, uh, Turing's decidability, Shannon. I think the, like, whole, like, relay circuits are, are, um... can be thought of as, can be mapped to Boolean operators, and a set of other, like, theoretical breakthroughs, which essentially they were creating abstractions for these, like, very analog circuits.
Mm-hmm.
Uh, and digital had this ni- nice property of, like, being error correcting. And, and so when I talk about, like, less leaky abstractions, that's what I mean. That's what I'm kind of pointing a little bit to. It's not gonna look exactly the same way.
And then the Xerox PARC piece, a lot of that is about, like, how do we get to computers that, as a person, I can actually use well? And the interface actually helps it unlock so much more power. So the sets of things we're working on, like the sets of abstractions and the, the interfaces, like, hopefully that, like, help us unlock a lot more power in these systems.
Like, hopefully that'll come not too far in the future. Um, I could see a next version, uh, like maybe a little bit farther out. It's like an agent protocol, so a way for different agents to talk to each other and call each other.
Ooh.
Um, kind of like HTTP. Um-
Ooh. Do you know it exists already?
Yeah, there is a nonprofit that's working on one. I think it's a bit early, but-
Yeah
... it's interesting to think about right now. Uh, part of why I think it's early is because the issue with agents is it's not quite like the internet, where you could, like, make a website, and the website would appear.
The issue with agents is that they don't work. Um, and so it may be a bit early to figure out what the protocol is before we really understand how could these agents get constructed. Um, but, you know, I think that's, I think it's a really interesting question.
While we're talking on this agent-to-agent thing, there's been a bit of research recently on some of these approaches. Um, I tend to just call them extremely complicated chain of thoughting. But um, any, any perspectives on kind of, uh, Meta-GPT, I think is the name of the paper.
I don't know if you care about, uh, ind- the, at the level of individual papers coming out. Um, but I, I did read that recently, and it, it, it, TLDR, it beat GPT-4 in human eval by role-playing a software agent, development agency.
Instead of having a sort of single shot, a single role, you have multiple roles, and how, uh, having all of them criticize each other as agents communicating with other agents.
Yeah, I think this is an example of an interesting abstraction-
Yeah
... of like, okay, can I just plop in this, like, multi-role critiquing and see how it improves my agent? Um, and can I just plop in chain of thought, tree of thought, plop in these other things and see how they improve my agent?
Um,
one issue with this kind of prompting is that it's still not very reliable. It like... There's one lens which is like, okay, if you do enough of these techniques, you'll get to high reliability. And I think actually that's not an, that's a pretty reasonable lens.
We take that lens often.
Mm-hmm.
Um, and then there's another lens that's like, okay, n- but it's starting to get really messy, what's in the prompt, and, like, how do we deal with that messiness? Um, and so maybe you need, like, cleaner ways of thinking about and constructing these systems, and we also take that, we also take that lens.
So yeah, I think both are necessary.
Side question, because, uh, I, I feel like this also brought up another question I had for you. Like, uh, I, I f- I noticed that you work a lot with your own benchmarks, your own evaluations of what is valuable.
And, uh, I would say, I would contrast your approach with OpenAI, as OpenAI tends to just lean on, "Hey, we played StarCraft," um, or, um, "Hey, we ran it on the SAT or the, uh, you know, the AP Bio test and, and that did results."
Um, b- basically, um, is benchmark culture ruining AI?
Um, or is, is that actually a good thing? Because everyone knows what an SAT is, and that's fine.
I think it's important to use both public and internal benchmarks.
Yeah.
Part of why we build our own benchmarks is that there are not very many good benchmarks for agents, actually.
Yeah.
And to evaluate these things, uh, we actually need to think about it in a slightly different way. Um, but we also do use a lot of public benchmarks-
Okay
... for, like, is the reasoning capability in this particular way improving? Um, so yeah.
Yeah.
It's good to use both.
Like, like, so for example, uh, like the Voyager paper coming out of, uh, NVIDIA, um, played "Minecraft" and set, set their own benchmarks on, uh, getting the diamond axe or whatever, and, uh, and exploring as much of the territory as possible.
And I don't know how that's received among... Like that's, that's obviously fun and novel for the rest of the AI engineer, like, enthus- like, the people who are new to the scene. But for, for people like yourself, who's, who you build your own Ev- uh, you build Avalon just because you already found deficiencies with, with using "Minecraft," like, is that valuable as an approach?
Oh, yeah. I love Voyager. I mean, Jim, I think, is awesome.
Yeah.
And I really like the Voyager paper, and I think it has a lot of really interesting ideas, which is, like, the agent can create tools for itself-
Yeah
... and then use those tools.
A- and, and he had the idea of the curriculum as well, which, which is something that we talked about earlier, yeah.
Exactly. Exactly, and, and that's, like, a lot of what we do. We built Avalon mostly because we couldn't use "Minecraft" very well to, like, learn the things we wanted.
Yeah.
And so it's, like, not that much work to build our own, uh... It took us, I don't know, uh, we had, like, eight engineers at the time, took about eight weeks. So six weeks.
Nice.
Yeah.
Yeah. And OpenAI built their own as well, right?
Yeah, exactly.
In eight days.
Uh, it's just nice to have control over our environment.
Yeah, build a sandbox.
But if you're doing... Yeah, our own sandbox to really try to inspect our research, our own research questions. But if you're doing something like experimenting with agents and trying to get them to do things, like, "Minecraft" is a really interesting environment.
Um, and so Voyager has a lot of really interesting ideas in it.
Yeah, cool. Uh, one more element that w- we had on this list, which, which was context and memory. Um, I think that's, that's kind of like the, the foundational, quote-unquote, "RAM" of, of our era. Um, I think, I think Andrej Karpathy has al- has already made this comparison, so there's nothing new here.
Um, but it, it's just the, the amount of working knowledge that we can fit into one of these agents, and it's not a lot, right? Like, especially if you need to get them to do long-running tasks, if they, they need to self-correct from errors that they observe while, while operating in their environment.
Um, do you see this as a problem? Do you think we're gonna just trend to infinite context and that'll go away? Or how do you think we're gonna deal with it?
Memory47:30
When you talked about what are, what's gonna happen in the first wave and then in the second wave, I think what we'll see is we'll get, like, relatively simplistic agents pretty soon, and they will get more and more complex.
Um, and there's, like, a future wave in which they are able to do these, like, really difficult, really long-running tasks. And, uh, the blocker to that future, one of the blockers is memory. And, and so, and that was true of computers too, you know.
Uh, I think when Von Neumann made the Von Neumann architecture, he was like, "The biggest blocker will be mem-- uh, like, we need this amount of memory," which is like, I don't, I don't remember exactly, like 32 kilobytes or something, "to store programs, and that will allow us to write software."
Um, he didn't say it this way 'cause he didn't have these terms, but... And then that only really was, uh, like happened in the '70s with the microch- microchip revolution. And so, like, it may be the case that we're waiting for some research breakthroughs or some other breakthroughs in order for us to have, like, really good long-running memory.
Um, but in the meantim- and then in the meantime, agents will be able to do all sorts of things that are a little bit smaller than that. I do think with the pace of the field, we'll probably come up with all sorts of interesting things, like, you know, RAG is already very helpful.
Good enough, you think?
It may be.
Yeah.
Uh, good enough for some things.
How is it not good enough?
I don't know. I just think about a situation where you want something that's like an AI scientist.
Mm.
As a scientist, I have learned so much, um, about my field, and a lot of that data is n- you know- Maybe hard to fine-tune or on, or maybe hard to like put into pre-training. A lot of that data, I don't have a lot of like repeats of the data that I'm seeing.
You know, my, my understanding is kind of like so at the edge that, yeah, like if I'm a scientist, I've like accumulated so many little data points, and ideally I'd wanna store those somehow or like use those to fine-tune myself as a model somehow, um, and, or like have better memory somehow.
I don't think RAG is enough for that kind of thing. But RAG is certainly enough for like user preferences and, um, and things like that. Like what should I do in this situation? What should I do in that situation?
That's a lot of tasks. We don't have to be a scientist right away.
Uh, I have a hard question if I can, if you don't mind me being bold.
Yeah.
I think the most comparable lab to Imbue is Adept.
Hmm.
Whatever. I mean, it, it, you know, research labby with like some amount of productization on, on the horizon, but not just yet, right? Um, why should people work for Imbue over Adept?
The way I think about it is I believe in our approach.
Mm-hmm.
Um, so maybe this is a general question of like competitors.
Yes.
And the way I think about it is, like we're in a historic moment. Uh, like this is 1978 or something.
Love it.
Apple's about to start. Lots of things are starting at that time, and IBM also exists, and all of these other big companies exist. And like we know what we're doing. We're building reasoning foundation models, trying to make agents that actually work reliably, that are inspectable, that we can modify, that we have a lot of control over.
And I think we have like a really special team and culture, and like that's like what we are. Like, I have a sense of where we wanna go, of really trying to help the computer be a much more powerful tool for us.
And kind of like the type of thing that we're doing is we're trying to like build something that enables other people to build agents. Uh, and, you know, build something that really can be maybe something like an operating system for agents.
I know that that's what we're doing. I don't really know what everyone else is doing, you know? Like and-
Yeah
... I like talk to people and, and have some sense of what they're doing. Um, and I think it's a mistake to focus too much on what other people are doing, because like extremely focused execution on the right thing is what matters.
And so, uh, you know, to the question of like why us, I think like strong focus on, uh, reasoning, which we believe is the biggest blocker, on, uh, like inspectability, which we believe is really important for user, user experience, um, and also for the power and capability of these systems.
Uh, building non-leaky good abstractions so that, which we believe is solving the core issue of agents, which is around reliability and kind of be- being able to make them deployable. And then, uh, like really seriously trying to use these things ourselves, so like every single day, and getting to something that like we can actually ship to other people that becomes something that is a platform.
Like feels like it could be, could be Mac or, or Windows. So yeah.
I love the dogfooding approach. Uh, that's extremely important. And, uh, you will not be surprised how many agent companies I talk to that don't use their own agent.
Oh, no. That's not good.
It's a big surprise. Yeah.
Yeah. I think, uh, if we didn't use our own agents, then we would have all of these beliefs about how good they are.
The only other follow-up that you had, based on the answer you just gave, was, um, uh, do you see yourself releasing m- models, or do you see yourself-- Like what is the artifacts that you want to produce that lead up to the general, uh, operating system that you wanna have, have people use, right?
Mm-hmm.
Um, and so a lot of people just as a byproduct of their work, just to say like, "Hey, I'm still shipping," you know, is like, "Here's, here's a model along the way." Adept took, I don't know, three years, but th- they released Persimmon, uh, which, which, uh, recently, right?
Like do you think that kind of approach is something on your horizon, or do you think there's something else that you can release that pe- that can show people, um, here's kind of the idea, not the end product, but here's the byproduct of what we're doing?
Yeah. I don't really believe in releasing things to show people, like, "Oh, here's what we're doing"-
Yeah
... that much. I'm, I think as a philosophy, we believe in releasing things that will be helpful to other people.
Yeah.
And so I think we may release models, or we may release tools that we think will help agent builders. Um, ideally, we would be able to do something like that, but I'm not sure exactly what they look like yet.
I think more companies should get into the releasing evals and benchmarks game.
Yeah. Something that we have been talking to agent builders about is co-building evals. So we build a lot of our own evals. Um, and every agent builder tells me basically evals are their biggest issue.
Of course. Right.
And so, yeah, we're exploring right now. And if you are building agents, you know, this is like a callout. If you are building agents, please reach out to me.
Yeah.
Because I would love to like figure out how we can be helpful, um, and like wha- based on what we've seen.
Cool. Well, that's a good call to action, and I know a bunch of people that I can send your way.
Cool. Great.
Awesome. Um, yeah, we can zoom out to other interests now.
We got a lot of stuff. Um, so we had Sharoukh from Lexica on the podcast. He had a lot of interesting questions on his website.
Big Questions54:17
Okay.
Uh, you similarly have a lot of them. Um-
Yeah. I need to do this. I, I'm very jealous of people with personal websites where they are like, "Here's the high-level questions of goals of humanity that I wanna, you know, set people on," and I don't have that.
This is great.
Yeah. It's never-
This is good. You should
... it's never too late, Sean.
Yeah.
It's never too late.
Exactly. Um, there were a few that stuck out as related to your work that maybe you're kinda learning more about it. So one is, why are curiosity and goal orientation often at odds? And from a human perspective, I get it.
It's like, you know, would you wanna like go explore things or kinda like focus on your career? How do you think about that from like an agent perspective, where it's like, should you just stick to the task and try and solve it as in the guardrails as possible?
Or like should you look for alternative solutions?
Yeah. This is a great question. So the problem with these questions is that I'm still confused about them.
Mm-hmm.
So our discussion, in our discussion, I will not have good answers. I will be still confused. Um, why are curiosity and goal orientation so at odds? I think one thing that's really interesting about agents actually is that they can be forked.
So, uh, like, you know, we can take an agent that's, uh, executed to a certain place and said, "Okay, here, like fork this and do a bunch of different things."
Mm-hmm.
"Uh, try a bunch of different things." Some of those agent can be goal-oriented, and some of them can be curio- like more curiosity driven. Uh, you can prompt them in slightly different ways. That's something I'm really curious about, like what would happen if in the future, you know, we were able to actually go down both paths.
As a person, why I have this question on my website is I really find that, like, I really can only take one mode at a time.
Mm-hmm.
Um, and I don't understand why. And, like, is it inherent in, like, the kind of context that needs to be held? That's why I think from an agent perspective, like forking it is really interesting. Like I can't fork myself to do both, um, but I maybe could fork an agent to like at a certain point in a task, yeah-
Mm-hmm
... to explore both.
Yeah.
Yeah.
How has the thinking changed for you as the funding of the company changed?
Ooh.
So that's, that's one thing that I think a lot of people in this space think is like, "Oh, should I raise venture capital?" Like, "How should I get money?" Um, how do you feel your options to be curious versus like goal-or-goal oriented has changed as you raise more money and kind of like the company has grown?
Oh, that's really funny. Actually, things have not changed that much. Um, so we raised our series A 20 million in late 2021, and our entire philosophy at that time was, and still kind of is, is like how do we figure out the stepping stones, like collect stepping stones that eventually let us build agents, the kind of these new computers that help us do bigger things.
And there was a lot of curiosity in that, and there was a lot of goal orientation in that. Um, like the curiosity led us to, uh, build Carbs, for example, uh, this hyperparameter optimizer, where Abe-
Great name by the way.
Thank you.
Is there a story behind that name?
Yeah, Abe loves carbs. Um- It's also, uh, cost aware. So as soon as he came up with cost aware, he was like, "I need to figure out how to make this work." Um, but the cost awareness of it was really important.
So that curiosity led us to this really cool hyperparameter optimizer. That's actually a big part of how we do our research. It lets us experiment on smaller models, and for those, uh, experiment results to carry to larger ones.
Which you also published a scaling laws, uh, thing for, which is great. Um, I think the scaling laws paper from OpenAI was like the biggest, and from Google I think, was, uh, gr- greatest public service to machine learning that, uh, that any research lab can do.
Yeah, totally. Um, yeah, and I think, yeah, what was nice about Carbs is like gave us scaling laws for all sorts of hyperparameters.
Yeah.
So, and then there's some goal-oriented parts, like Avalon, you know, it was like a six to eight week sprint, uh, for all of us, and we like got this thing out.
Mm-hmm.
Uh, and then now different projects do like more curiosity or more goal orientation at different times.
Another one of your questions that we highlighted was, how can we enable artificial agents to permanently new, learn new abstractions and processes? I think this is might be called online learning.
Yeah. So I struggle with this because, um, you know, that scientist example I gave, as a scientist, I've like permanently learned a lot of new things, and I've updated and created new abstractions and learned them pretty reliably. Um, and you were talking about like, okay, we have this RAM, um, that we can store learnings in, but, you know, how well does online learning actually work?
And, and the answer right now seems to be like as models get bigger, they fine-tune faster, so they're more sample efficient as they get bigger.
Because they already had that knowledge in there. They're, they're just, you're just kind of unlocking it.
Maybe partly, maybe because they already have-
Yeah
... like some subset of the representation.
Yeah.
Partly they just memorize things-
Mm-hmm
... uh, more, which is good. So maybe this question is going to be solved, but I still don't know what the answer is.
Um, just, I don't know, have an interf- has, have a platform that continually fine-tunes for you as you work on that domain, which is something I'm working on. Um, 'cause I'm interested in it.
Well, that's great. We would love to use that.
We'll talk more. Um, okay, so, uh, two more questions just about your general activities, and, uh, I think you've just been very active in the San Francisco tech scene. Um, you're a founding member of Software Commons?
Community59:54
Oh, yeah, that's true.
Uh, tell me more, 'cause I, I, you know, I, by the time I knew about SBC, it was already a very established thing, but what was it like in the early days? Like what, what's the story there?
Yeah, the story is Ruchi, uh, who started it, uh, was the VP of operations at Dropbox, and-
I see
... I was the chief of staff, and we worked together very closely. She's actually one of the investors in Sorceress, and SBC is an investor in, uh, in, Imbue. And at that time, Ruchi was like, "You know, I would like to start a space for people who are figuring out what's next."
Lightning Round1:00:22
Um, and we were figuring out what's next post Ember, uh, those three months, and she was like, "Do you wanna just like hang out in this space?" And we're like, "Sure." And it was a really good group, I think, uh, Waseem from Pilot, Waseem and Jeff from Pilot, the folks from Zulip, and a bunch of other people at that time.
It was a really good group. We just hung out. There was no programming. It's much more official than it was at that time.
Yeah. Now it's like a Y- like YC before YC type of thing.
That's right, yeah. At that time, we literally it was a bunch of friends hanging out in the space together.
And was this concurrent with The Archive?
Uh, oh, yeah. Actually, I think we started-
Just-
... The Archive around the same time
... you're just like really big into community. But also, like, so, you know, I run a hacker house, right? And I, I'm also part of, uh, hopefully what becomes like the next, uh, Software Commons or whatever.
Mm-hmm.
Um, but like w- what are the principles in organizing ... communities like that with really exceptional people that are, that go on to do great things. Um, do you have to be really picky about who joins? Are you j- like, all, did all your friends just magically turn out super successful like that?
Yeah, I think... So I think we-
You know it's not normal, right? Like, this is, this is very special. It's... And a lot of people want to do that and fail.
Mm.
And you are one, like, you had the co-authors of GPT-3 in your house.
That's true, and a lot of other really cool people that you'll eventually hear about.
And Pi- co-founders of Pilot and anyone else you wanna sh- uh, I, I don't want, I don't want you to pick your friends, but b-
I mean-
There's some magic special sauce in getting people together in, in one workspace, living space, whatever, right? And that's part of why I'm here in San Francisco.
Mm-hmm.
And I, I, I would love for more people to learn about it and also maybe get inspired to build their own.
One adage we had when we started The Archive was you become the, the average of the five people closest to you.
Yes.
And I think that's roughly true, and good people draw good people. So, um, there are really two things. Like, one, we were quite picky, um, and we... It mattered a lot to us, like, is this someone where if they're hanging out in the living room, we'd be really excited to come hang out?
Yeah.
Um, two is I think we did a really good job of creating a high-growth environment and an environment where people felt really safe. Um, we actually apply these things to our team, and it works remarkably well as well.
Oh.
So I do a lot of basically how do I create safe spaces for people, uh, where it's not just, like, safe, blah, um, but, like, it's, like, a safe space where people really feel inspired by each other. And I think at The Archive, we really made each other better.
Um, my friend Mike O'Neillson called it a self-actualization machine.
My goodness. Okay.
And I think, yeah, people came in and-
Was he a part of The Archive?
He was not, but he hung out a lot. Um-
Oh, honorary member. Friend of The Archive.
A friend of The Archive, yeah.
Yes.
Like, the culture was that we learned a lot of things from each other about, like, how to make better life systems and how to think about ourselves and psychological debugging, and, um, a lot of us were founders, so having other founders going through similar things was really helpful.
And a lot of us worked in AI, and so having other people to talk about AI with was really helpful. And so I think all of those things led to a form of idea flux and also kind of, like...
So I think a lot about, like, the idea flux and the kind of, like, uh, default habits or default impulses. It led to a set of idea flux and default impulses that led to some really interesting things, um, and led to us doing much bigger things, I think, than we otherwise would have decided to do because it felt like taking risks was less risky.
So that's something we do a lot of on the team is, like, how do we make it so that taking risks is less risky? And there's a term called scenius.
Yes. I was think- Kevin Kelly.
Kevin Kelly, scenius.
I was, I was gonna feed you that word, but I didn't wanna, like, bias you.
Yes. Yes. I think maybe, like, a lot of what I'm interested in, in, is constructing a s- a kind of scenius.
Yes.
Um, and The Archive was definitely a scenius in a particular, or, like, getting toward a scenius in a particular way. Um, and, uh, Jason Benn, my Archive housemate and who now runs the neighborhood, uh, has a good way of putting it.
If genius is from your genes, scenius is from your scene.
Oh.
And yeah, I think, like, maybe a lot of the community-building impulse is from this, like, interest in what kind of idea flux can be created. Um, you know, there's a question of, like, why did Xerox PARC come out with all of this interesting stuff?
It's their scenius. Why did, uh, Bell Labs come out with all this interesting stuff? Maybe it's their scenius. Why didn't, you know, why didn't the transistor come out of Princeton?
Yeah.
Um, and the other people working on it at the time.
I always think it's remarkable how you hear a lot about Alan Kay, and I just read a bit, and apparently Alan Kay was, like, the most junior guy at Xerox PARC.
Yeah, definitely. He's just the one who talks about it.
He talks the most. Yeah, exactly. Um, yeah, so I, you know, hopefully I'm also working towards contributing that scenius. Uh, I call mine the more provocative name of The Arena.
Oh, interesting.
So you can be-
That's quite provocative
... in the arena.
So are you fighting other people in the arena?
No, no.
You never know.
Which is-
Any day, any day in the mission is, uh, is an adventure.
You're in the arena fi- uh, you're in the arena trying stuff, as they say.
Oh.
Um, uh, you are also a GP at Outset Capital, um, where you also co-organize the Thursday Nights in AI, uh, where hopefully, uh, someday I'll eventually speak. But w-
You're on the roster.
I'm on the roster. Thank you so much. Um, w- why, so why spend time being a VC, uh, and, and organizing all these events? You're also a very busy CEO, and, you know, why, why spend time with that?
Why, why is that an important part of your life?
Yeah. For me personally, I really like helping founders. So, um, Ali, my, uh, investing partner, is fortunately amazing, and she does everything for the fund. Um, so she, like, hosts the Thursday night events, and she finds, uh, folks who we could invest in, and she does basically everything.
Josh and I, um, are her co-partners. So Ali was our former chief of staff at Sorceress.
Ah.
We just thought she was amazing.
Yeah.
Um, and she wanted to be an investor, and Josh and I also, like, care about helping founders and kind of, like, giving back to the community. What we didn't realize at the time when we started the fund is that it would actually be incredibly helpful for Imbue.
So, uh, talking to AI founders who are building agents and working on, you know, similar things is really helpful. They could potentially be our customers, and they're trying out all sorts of interesting things. And I think being an investor looking at the space from the other side of the table, it's just a different hat that I routinely put on, and it's helpful to see the space from the investor lens as opposed to from the founder lens.
Um, so I find that kind of, like, hat switching valuable. It maybe would lead us to do slightly different things.
Let's just wrap with the lightning round.
Okay.
So we have three questions: acceleration, exploration, and then a takeaway. So the acceleration question is, what's something that already happened in AI that you thought would take much longer to be here?
I think the rate at which we discover new capabilities of existing models and kind of like build hacks on top of them to make them work better is something that has been surprising and awesome. And the rate of kind of like the community, the AI, the research community building on its own ideas.
Cool. Um, exploration/request for startups. If you weren't building Imbue, what AI company would you build?
Hmm.
Every founder has, like, their, like, number two.
Really?
Yeah. I don't know.
Wow. I cannot imagine building any other thing than Imbue. Um-
Wow. Well, that's a great answer too. That's an interesting, that's like-
It's like obviously the thing to build.
Okay.
It's like obviously work on the fundamental platform.
Yeah. Um, okay. So, so the previous, I think that, that was my attempt at innovating this question, but the previous one was, uh, what was the most interesting unsolved question in AI?
Yeah. I think probably the most interesting unsolved question, and my answer is kind of boring, but the most interesting unsolved questions are these questions of how do we make these stochastic systems into things that we can, like, reliably use and build on top of?
Yep.
Uh, and yeah, take away, what's one message you want everyone to remember?
Maybe two things. Like, um, one is I didn't think in my lifetime I would necessarily be in, like able to work on the things I'm excited to work on in this moment, but we're in a historic moment and that where we'll look back and be like, "Oh my God, the future was invented in these years."
There is maybe a set of messages to take away from that. One is like, um, AI is a tool, uh, like any technology, and, you know, when it comes to things like what might the future look like? Like, we like to think about it as it's like just a better computer.
It's like a much better, much more powerful computer that gives us a lot of free intellectual energy that we can now, like, solve so many problems with. You know, there are so many problems in the world where we're like, "Ugh, it's not worth a person thinking about that," and so things get worse and things get worse.
No one wants to work on maintenance. Um, and like this technology gives us the potential to actually be able to s- like allocate intellectual energy to all of those problems, and the world could be much better, like could be much more thoughtful because of that.
I'm so excited about that. And there are definitely risks and, and dangers. And, um, we actually do a fair... Something I didn't talk about is we do a fair amount of work on the policy side. On the safety side, like we think about safety and policy in terms of, uh, engineering, uh, theory, and also regulation.
And, um, kind of comparing to like the automobile or the airplane or any new technology, there's like a set of new possible, like capabilities and a new set of new possible dangers that are unlocked with every new technology.
And so on the engineering side, like we think a lot about engineering safety, like how do we actually engineer these systems so that they are inspectable and, you know, why we reason in natural language so that the systems are very inspectable, so that we can like stop things if anything weird is happening.
That's why we don't think end-to-end black boxes are a good idea. On the theoretical side, we like really believe in like deeply understanding, like what are they learning? Like when we actually fine-tune on individual examples, like what's going on?
When we're pre-training, what's going on? Um, like debugging tools for these agents to understand, like what's going on? Um, and then on the regulation side, I think there's actually a lot of regulation that already covers many of the dangers, uh, like that might, that people are talking about.
And there are areas where there's not much regulation, and so we focus on those areas where there's not much regulation. So some of our work is actually we built an agent that helped us analyze, uh, the like 20,000 pages of policy proposals submitted to the Department of Commerce request for AI policy proposals.
And we like looked at what were the problems people brought up and what were the solutions they presented, and then like did a summary analysis and k- like, you know, built, built agents to do that. Um, and now the Department of Commerce is like interested in using that as a tool to like analyze proposals.
And so, um, a lot of what we're trying to do on the regulation side is like actually figure out where is there regulation missing and, uh, and, and how do we actually, in a very targeted way, try to solve those, uh, missing areas.
So I guess if I were to say like what are the takeaways, it's like the future could be really exciting, uh, if we can actually get agents that are able to do these bigger things. Reasoning is the biggest blocker, uh, plus like these sets of abstractions to make things more robust and reliable.
And, uh, and there are, you know, things where we have to be quite careful and thoughtful about how do we deploy these and what kind of regulation should go along with it, so that this is actually a technology that when we deploy it, it is protective to people and not harmful.
Awesome.
Wonderful.
Uh, yeah. Thank you so much for your time, Kanjun.
Cool. Thank you.
That's it.
Thank you so much.
Thank you. Any questions?
Uh.






