LALatent SpaceApr 11, 2024· 1:05:28

Supervise the Process of AI Research — with Jungwon Byun and Andreas Stuhlmüller of Elicit

Andreas Stuhlmüller and Jungwon Byun, co-founders of Elicit (formerly the nonprofit Ought), have built an AI research assistant that automates literature review and reasoning by breaking complex tasks into transparent, step-by-step processes. Their philosophy—"supervise the process, not just the outcome"—led them to start with human simulations before GPT-3 enabled a product pivot. Elicit now uses both open-source and closed models (e.g., GPT-4, Claude Haiku) for summarization, data extraction, and uncertainty flags, and recently launched computational notebooks for scalable, reusable workflows. The company transitioned from nonprofit to a Public Benefit Corporation, reached $1M revenue in four months, and now employs 12 people, focusing on senior software engineers to build reliable orchestration from unreliable components.

  1. 0:00Introductions
  2. 9:03Co-founders
  3. 10:47Product Vision
  4. 13:57Product History
  5. 19:46Early Prototypes
  6. 26:10GPT Impact
  7. 29:16Cost Optimization
  8. 36:44Notebooks
  9. 45:03Budget & Uncertainty
  10. 48:04Long Context
  11. 54:34Underrated Features
  12. 58:23Future & Mission

Powered by PodHood

Transcript

Introductions0:00

Alessio0:00

Hey everyone, welcome to the Latent Space Podcast. This is Alessio, partner and CTO in residence at Visible Partners, and I'm joined by my co-host Swyx, founder of Smol AI.

Swyx0:09

Hey, and today we are back in the studio with Andreas and Jungwon, uh, from Elicit. Welcome.

Jungwon Byun0:15

Thanks, guys.

Andreas Stuhlmüller0:16

It's great to be here.

Jungwon Byun0:17

Yeah.

Swyx0:17

So I'll introduce you separately, but also, um, you know, we'd love to learn a little bit more about you personally. Um, so Andreas, it looks like you started Elicit or Ought first-

Andreas Stuhlmüller0:28

That's-

Swyx0:29

... and, and, and, uh, Jungwon start- joined later.

Andreas Stuhlmüller0:32

That's right. Although she did, I guess, f- for all intents and purposes, the Elicit and also the Ought that existed before then, um, were very different from what, what I started. So, uh, I think, uh, it's like fair to say that she co-founded it.

Swyx0:46

And Jungwon, you're co-founder and COO of, uh, Elicit now.

Jungwon Byun0:50

Yeah, that's right.

Swyx0:50

So there's, there's a little bit of a, a history to this. Um, I'm not super aware of like the, the sort of journey. Um, I, I was aware of Ought and Elicit as, um, sort of a nonprofit-type situation, and recently you turned into, uh, sort of like a B corp?

Andreas Stuhlmüller1:04

Public benefit corporation.

Jungwon Byun1:05

Yeah.

Swyx1:06

So yeah, may- maybe if you, if you want, you could take us through that journey of, um, you know, finding the problem. Um, you know, obviously you're working together now, so like how do you get together to, to decide to leave your, uh, startup career to, to join him?

Jungwon Byun1:20

Mm-hmm.

Andreas Stuhlmüller1:21

Yeah, it's truly a very long journey. I guess truly it, it kind of started in Germany when I was born. Um, so e- e- even as a kid, I, I was always interested in AI. Like I kind of went to the library, there were books about how to write programs in QBasic, and like some of them talked about how, how to implement chatbots.

Um, I guess Eliza-

Jungwon Byun1:40

To, to be clear, he grew up in like a tiny village on the outskirts of Munich called Dinkelscherben, where it's like a very, very idyllic German village.

Swyx1:48

Oh.

Andreas Stuhlmüller1:49

Yeah.

Jungwon Byun1:49

Yeah.

Andreas Stuhlmüller1:49

Important to the, to the story, so- Um, but basically the, the main thing is, uh, I've kind of always been thinking about AI my entire life, and been thinking about, well, at some point this is gonna be a huge deal.

It's gonna be transformative. How can I, how can I work on it? And, um, was thinking about it from when I was a teenager. Um, did a-- after high school, did a year where I, um, started a startup with the intention to like become rich, and then once I'm rich I can, uh, affect the trajectory of AI.

Uh, did not become rich. Um, decided to go back to college and study cognitive science there, which was like the closest thing I could find at the time to AI. Um, in the last year of, of college, moved to the US to d- do a PhD, um, at MIT, working on broadly kind of new programming languages for AI, because it kind of seemed like the existing languages were not great at expressing world models and learning world models, doing Bayesian inference.

Um, was always thinking about, well, ultimately the goal is to actually build tools that help people reason more clearly, um, ask better... k- ask and answer, uh, better questions and make better decisions. But for a long time, it had seemed like the technology to put reasoning in machines just wasn't there.

And, uh, then... And so, so, so initially at the end of my, at, at the end of my postdoc at Stanford was thinking about, well, what, what to do? I think the, the standard path is you become an academic and, uh, do research.

But, um, it's really hard to actually build interesting tools as an academic. You can't really hire great engineers. Uh, everything is kind of on a paper to paper timeline. And so I was like, well, maybe I should start a startup, a- and pursued that for a little bit, but it seemed like it was too early, because you, you, you could have tried to do an AI startup, but, uh, probably would not have been the kind of AI startup we're, we're seeing now.

So then decided to just start a nonprofit research lab that's gonna do research for a while until we better figure out how to do thinking in machines, and that was Ought. Um, and then over time, it became clear how to actually...

how to do actual kind of, build actual tools for reasoning. And, uh, then only kind of over time, we developed a better way to, uh... Well, I, I'll let you fill in some of the details here.

Jungwon Byun4:10

Yeah, so I guess m- my story, um, maybe starts around 2015. I'd-- I kind of wanted to be a s- founder for a long time, um, and I wanted to, um, work on an idea that like really tested, t- uh, that stood the test of time for me, like an idea that stuck with me for a long time.

And then starting in 2015, actually originally I became interested in AI-based tools from the perspective of mental health. So there are a bunch of people around me who are really struggling. One, one really close friend in particular is really struggling with mental health and didn't have any support, and it didn't feel like there was anything before kind of like getting hospitalized that could just help her.

And so luckily she came and stayed with me for a while, and we were just able to talk through some things, but it seemed like, you know, lots of people might not have that resource, and something maybe AI-enabled could be much more scalable.

Um, I didn't feel ready to start a company then. That's, you know, 2015. Um, and I also didn't feel like the technology was ready, so then I went into fintech and like kind of learned how to do the tech thing.

Um, and then in 2019, I felt, you know, like it was time for me to just jump in and, and build something on my own. I really wanted to create. Um, and at the time, there were kind of two interesting...

I looked around at tech and felt like not super inspired by the options. I just, I didn't want to have a tech career ladder, or like I didn't wanna like climb the career ladder. Um, there were two kind of interesting technologies at the time.

There was AI and there was crypto. Uh, and I was like, well, the AI people seem like a little bit more nice. Or maybe like slightly more trustworthy. Um, both super exciting, but, but yeah, I kind of like, um, threw my bet in with the, on the AI side.

Um, and then I got connected to Andreas, and actually the way he was thinking about pursuing the research agenda at Ought was really compatible with what I had envisioned for an ideal AI product, something that helps kind of take down really complex thinking, overwhelming thoughts, and breaks it down into small pieces.

And then this, this kind of mission that we need AI to help us figure out what we ought to do, um-

Swyx6:08

Ooh

Jungwon Byun6:08

... was really inspiring, right? Yeah, 'cause I think it was clear that we were building the most powerful optimizer of our time. Um, but as a society, we hadn't figured out how to direct that optimization potential, and if you kind of direct tremendous amounts of optimization potential at the wrong thing, that's really disastrous.

So the goal of Ought was make sure that if we build the most transformative technology of our lifetime, it can be used for something really impactful, like good reasoning, like, not just generating ads. My background was in marketing, but, like, still, I was like, "I wanna do more than generate ads with this."

And also, if this, these AI systems get to be super intelligent enough that they are doing this really complex reasoning, that we can trust them, that, that they are aligned with us and we have ways of evaluating that they're doing the right thing.

So that's what Ought did. We did a lot of experiments. Um, this was like, you know, like Andreas said, before foundation models really, like, took off. Um, a lot of our kind- a lot of the issues we were seeing were more in reinforcement learning, but we, we saw a future where AI would be able to do more kind of logical reasoning, um, not just kind of extrapolate from numerical trends.

Um, so we orche- we actually kind of, um, set up experiments with people, where kind of people stood in as super intelligent systems, and we effectively gave them context windows, so they would have to, like, read a, read a bunch of text, and some, like, one person would get less text and one person would get all the text, and the person with less text would have to evaluate the work of the person who could read much more.

So, like, in a world-- We were basically simulating, like, in, you know, 2018, 2019, a world where an AI system could read significantly more than you, and you as the person who couldn't read that much had to evaluate the work of the AI system.

Swyx7:50

Huh.

Jungwon Byun7:51

Yeah, so it was a lot of, um, a lot of the, what, the work we did, and from that, we kind of iterated on this idea that, you know, uh, the idea of breaking complex tasks down into smaller tasks, like complex tasks, like open-ended reasoning, logical reasoning, into smaller tasks so that it's easier to train AI systems on them, and also so that it's easier to evaluate the work of the AI system when it's done.

Um, and then also kind of, you know, really pioneered this idea, uh, this, uh, the importance of supervising the process of AI systems, not just the outcomes. And so a big part of how then, like, how Elicit is built is we're very intentional about not just throwing a ton of data into a model and training it and then saying, "Cool, here's, like, scientific output."

Swyx8:31

Mm-hmm.

Jungwon Byun8:31

Like, that's not at all what we do. Um, our approach is very much like what are the steps that an expert human does? Or what is, like, an ideal process? As granularly as possible, let's break that down and then train AI systems to perform each of those steps very robustly.

When you train like that from the start, after the fact, it's much easier to evaluate. You can, like, it's much easier to troubleshoot at each point, like, where did something break down? Um, so yeah, we were working on those experiments for a while, and then at the start of 2021, decided to build a product, because when you do research, um, I think maybe I'm curious-

Co-founders9:03

Swyx9:03

Do you mind, do you mind if I, uh... 'Cause y- I think you're about to go into more modern-

Jungwon Byun9:07

Yeah, yeah

Swyx9:07

... Ought and Elicit, and I, I just wanted to, because I think a lot of pe- people are in where you were, like, sort of 2018, '19-

Jungwon Byun9:15

Uh-huh

Swyx9:15

... where you, you chose a partner to work with.

Jungwon Byun9:18

Yeah.

Swyx9:18

Right? And you didn't know him.

Jungwon Byun9:19

Yeah, yeah.

Swyx9:20

You were just kind of cold introduced.

Jungwon Byun9:21

Yep.

Swyx9:22

A lot of people are cold introduced.

Jungwon Byun9:23

Mm-hmm.

Swyx9:23

I've been cold introduced to tons of people, and I never work with them. Uh, I'm, if y- I assume you had a lot, a lot of other options, right? Like, how do you advise people to make those, make those choices?

Jungwon Byun9:32

Yeah, we were not totally cold introduced, so we had one of our closest friends introduced us. Um, and then Andreas had written a lot on, on the Ought website, a lot of blog posts, a lot of publications, and I, I just read it, and I was like, "Wow, this is, this sounds like my writing."

And even other people, my, some of my closest friends I asked for advice from, they were like, "Oh, this sounds like your writing." But I think I also had some kind of, like, things I was looking for. I wanted someone with a complementary skill set.

I want someone who was very values aligned, and yeah, I think that was, that was all a good fit.

Andreas Stuhlmüller10:04

We also did a pretty lengthy mutual evaluation process where we had a, a Google Doc-

Jungwon Byun10:09

Mm-hmm

Andreas Stuhlmüller10:09

... where we had all kinds of questions for each other, and, uh, I, I think it ended up being around 50 pages or so of, like, various, like, questions and back and forth.

Swyx10:17

Was it the YC list? Uh, there's some lists going around for co-founder questions.

Andreas Stuhlmüller10:20

No, we just made our own questions.

Swyx10:23

You made your own.

Jungwon Byun10:23

Yeah.

Andreas Stuhlmüller10:23

But I presume, I'd, I'd guess it's probably related in that you ask yourself, well, well, what are the values you care about? How would, would you approach various decisions and things like that.

Jungwon Byun10:31

I shared, like, all of my past performance reviews.

Swyx10:34

Yeah?

Jungwon Byun10:34

Yeah.

Swyx10:35

Yeah. And he had never had any, so-

Jungwon Byun10:37

No.

Swyx10:41

Yeah, sorry, I just had to-

Jungwon Byun10:42

Yeah, no, go ahead

Swyx10:42

... a lot of people are going through that phase, and you kind of skipped over it, and I was like, "No, no, no, no, there's, like, an interesting story here."

Jungwon Byun10:46

Yeah.

Swyx10:47

Yeah.

Product Vision10:47

Andreas Stuhlmüller10:47

So before we jump into what Elicit is today, uh, the, the history is a bit, uh, counterintuitive. So you start with figuring out, oh, if we had a super powerful model, how would we align it, how would we use it?

Jungwon Byun11:00

Mm-hmm.

Andreas Stuhlmüller11:00

But then you were actually like, "Well, let's just build the product so that people can actually leverage it." And I, I think there are a lot of folks today that are now back to where you were maybe five years ago that are like, "Oh, what if this happens?"

Rather than focusing on actually building-

Jungwon Byun11:13

Mm

Andreas Stuhlmüller11:13

... something, uh, useful with it. What, what clicked, uh, for you to, like, move into Elicit, and then we can cover that story too. I think in many ways, the approach is still the same because the way we are building Elicit is not let's train a foundation model to do more stuff.

It's like, let's build a scaffolding such that we can deploy powerful models to good ends. So I, I think it is different now that, in that we're, we actually have, like, some of the models to plug in, but if in 2018, '17, we had had the models, we could have run the same experiments, um, we did run with humans back then just with models.

Swyx11:47

Mm-hmm.

Andreas Stuhlmüller11:48

And so in many ways, our philosophy is always like, let's think ahead to the future. What models are gonna exist in one, two years, uh, or longer, and how can we make it so that they can actually be t- deployed in kind of transparent, controllable ways?

Jungwon Byun12:03

Yeah, I think motivationally, we both are kind of product people at heart, and we just want to... The research was really important, and it didn't make sense to build a product at that time. But at the end of the day, the thing that always motivated us is imagining a world where high-quality reasoning is really abundant.

Um- And AI was just kind of the most-- is, is the technology that's gonna get us there. And there's a way to guide that technology with research, but it's also really exciting to have... You can have a more direct effect through product, because r-with research, you ha-kind of, you publish the research and someone else has to implement that into the product, and the product felt like a more direct path, and we wanted to concretely have an impact on people's lives.

So, um, I think, yeah, I think it, the kinda personally the motivation was we, we want to build for people.

Alessio12:49

Yep. Um, and then just, just to recap, uh, as well, like the models you were using back then were like, I don't know, were they like BERT-type stuff or T5 or I don't know, I don't know what timeframe we're talking about here.

Andreas Stuhlmüller13:02

So the-- I guess to be clear, at the very beginning, we had humans, um, do the work.

Alessio13:08

Yeah.

Andreas Stuhlmüller13:08

And then the initial-- I think the first models that kind of make sense were GPT-2 and TNLG and like the, the early, early, um, yeah, early generative models. We, we do also use like T5-based models e-even, even now.

Um, but, uh, started, yeah, started with GPT-2.

Alessio13:26

Yeah, cool. I'm just kind of curious about like how do you start so early, you know? Like now it's obvious you s-where to start, but back then it wasn't.

Jungwon Byun13:33

Mm-hmm. Yeah, I used to nag Andreas a lot. I was like, "Why are you talking to this..." I don't know. I was like, "GPT-2s like clearly can't do anything." And I was like, "Andreas, you're wasting your time like playing with this toy."

But yeah, he was right.

Alessio13:47

So what's the history of what Elicit actually does as a product? I think today, um, you recently announced that after four months you got to a million of revenue. Obviously, a lot of people use it, get a lot of value.

Product History13:57

Alessio13:57

But, um, it would-- it initially l- kind of like structured data extraction from papers. Uh, then you had, um, yeah, kind of like concept grouping, and today it's maybe like a more full stack research enabler, kind of like paper understander platform.

What's, wha-what's the definitive, uh, definition of what Elicit is, and how did you get here?

Jungwon Byun14:18

Yeah. We, we say Elicit is an AI research assistant. I think it will continue to evolve. It has evolved a lot, and it will continue to research. And that's why we-- part of why, you know, we're so excited about building and research, 'cause there's just so much space.

I think the current phase we're in right now, um, we talk about it as like really trying to make Elicit, um, the best place to understand what is known. So it's all, it's all-

Alessio14:37

Oof

Jungwon Byun14:37

... a lot about like literature summarization. There's a ton of information that the world already knows. It's really hard to navigate, um, hard to make it relevant. So a lot of it is around document discovery and processing and analysis.

Um, I really want to make... I w-kind of wanna import some of the incredible productivity improvements we've, we've seen in software engineering and data science in, into research. So it's like how can we make researchers like data scientists of text?

Um, that's why we're launching this, um, new set of features called Notebooks. It's very much inspired by computational notebooks like Jupyter Notebooks, you know, DeepNote or Colab, um, because they're so powerful and so flexible, and ultimately, when people are trying to get to an answer or understand insight, they're kind of like manipulating evidence and information.

Today, that's all packaged in PDFs, which are super brittle. But with language models, we can decompose these PDFs into their underlying claims and evidence and insights, and then let researchers mash them up together, remix them, and analyze them together.

So, um, so yeah. I would say quite simply, E-overall, Elicit is an AI research assistant. Right now we're focused on, um, and text-based workflows, but long term really wanna kind of go further and further into reasoning and decision-making.

Alessio15:51

And when you say AI research assistant, this is kind of meta research. So researchers use Elicit as a research assistant. It's not a, a generic you can research anything, um, type of tool, or it could be, but like what are people using it for today?

Andreas Stuhlmüller16:06

Yeah. So specifically, uh, I guess in, in science, a lot of people use human research assistants to do things. Uh, like, uh, you, you tell your k-kind of grad student, "Hey, here are a couple of papers. Can you look at all of these, see which of these have kind of sufficiently large populations, and actually study the disease that I'm interested in, and then write out like what, what are the experiments they did, what are the interventions they did, what are the outcomes, and kind of organize that for me?"

And, uh, the first phase of understanding what is known really focuses on automating that workflow, because a lot of that work is pretty rote work. I think it's not the kind of thing that we need humans to do.

Language models can do it. And then if language models can do it, then you can obviously scale it up much more than, than a grad student or undergrad research assistant would be able to do.

Jungwon Byun16:52

Yeah, the use cases are pretty broad. So we do have people who just come... A very large percent of our users are, are just using it personally or for p-a mix of personal and professional things. Um, people who care a lot about, uh, like health or biohacking or people, you know, parents who have children with a kind of rare disease and want to understand the literature directly.

So there is a, an individual kind of consumer use case. We're most focused on the power users, so that's where we're really excited to build. Um, so Elicit was very much inspired by this workflow in literature called systematic reviews or meta-analysis, which is basically the human state-of-the-art for summarizing scientific literature.

It typically involves like five people working together for over a year, and they kind of first start by trying to find the maximally comprehensive set of papers possible, so it's like 10,000 papers, and they kind of systematically narrow that down to like hundreds or 50, extract key details from every single paper, usually have two people doing it and like a third person reviewing it.

So it's like an incredibly, um, laborious, time-consuming process, but you see it in every single domain, so in science, in machine learning, um, in policy. And so if you can-- And it's very... Because it's so structured and designed to be reproducible, it's really amenable to automation.

So that's kind of the one, the workflow that we wanna automate first, and then, uh, you can-- you make that accessible for any question and make, you know, kind of these really robust living summaries of science. Uh, so yeah, that's one of the workflows that we're starting with.

Alessio18:18

Our previous guest, Mike Conover, he's building a new company called Brightwave, which is an AI research assistant for financial research.

Jungwon Byun18:24

Mm.

Alessio18:24

Um, how, how do you see the future of these tools? Like, does everything converge to like a, a god researcher, uh, assistant, or is every domain going to have its own thing?

Andreas Stuhlmüller18:34

I think that's a good and mostly open question. Um- I do think there are some differences across domains. For example, some research is more quantitative data analysis, and other research is more kind of high-level cross-domain thinking. And, uh, we definitely want to contribute to the broad generalist reasoning type space.

Like, if, if researchers are making discoveries, often it's like, "Hey, this thing in biology is actually analogous to, like, these equations in economics or something." And that, that's just fundamentally a thing that, where you need to reason across domains.

Um, so I think there will be, at least within research, I think there will be, like, one best platform more or less, um, for, for this type of generalist research. I think there may still be, like, some particular tools, like, for genomics, like if particular types of modules of genes and proteins and whatnot.

But for a lot of the kind of high-level reasoning that humans do, I think, I think that is a more of open or type all thing.

Jungwon Byun19:29

Mm-hmm.

Swyx19:30

I wanted to ask a little bit deeper about, I guess, the, the workflow that, that you mentioned. I, I like that phrase. Um, I see that in your UI now.

Jungwon Byun19:37

Mm-hmm.

Swyx19:38

But that's as it is today. Um, and I think you were about to tell us about how it was in 2021, and how it maybe progressed. Like, what, how has this workflow evolved over time?

Early Prototypes19:46

Jungwon Byun19:46

Yeah. So the very first version of Elicit actually wasn't even a research assistant. It was a f- it was like a forecasting assistant. So, um, we sat down and we were thinking about what is, you know, what are some of the most impactful reasonings that if, uh, types of reasoning that if we could scale up AI would really transform the world.

And at the t- the first thing we started, we actually start- started with literature review, but we're like, "Ugh, so many people are gonna build literature review tools. Let's, let's not start there." Uh, and so then we, we focused on geopolitical forecasting.

So I don't know if you're familiar with, like, Manifold or-

Swyx20:16

Manifold markets, yeah.

Jungwon Byun20:18

Yeah.

Swyx20:18

Yeah.

Jungwon Byun20:18

That kind of stuff.

Swyx20:18

And Manifold.ai.

Jungwon Byun20:18

Before Manifold. Yeah. Yeah. Um, so not, we're not, not predicting relationships. We're predicting, like, is China gonna invade Taiwan? Um-

Swyx20:26

Markets for everything.

Jungwon Byun20:27

Yeah.

Andreas Stuhlmüller20:27

That's in a relationship, aren't they?

Jungwon Byun20:29

Yeah, that's fair. Um-

Swyx20:29

Yeah. Yeah, it's true.

Jungwon Byun20:31

And then we worked on that for a while, and then after GPT-3 came out, I think by that time we, um, kind of realized that the... Originally, we were trying to help people convert their beliefs into probability distributions, and so take fuzzy beliefs, but, like, model them more concretely.

And then after a few months of iterating on that, just realized, oh, the thing that's blocking people from making interesting predictions about important events in the world, um, is less kind of on the probabilistic side and much more on the research side.

And so that, that kind of combined with the very generalist capabilities of GPT-3 prompted us to make a more general research assistant.

Swyx21:07

Mm.

Jungwon Byun21:07

Then we spent a few months iterating on what a, what even is a research assistant. So we would embed with different researchers. We built data labeling, um, workflows in the beginning, kind of right off the bat. We built, um, ways to find, uh, uh, like, experts in a field, and, like, ways to ask good research questions.

So we just kind of iterated through a lot of workflows, and it was, it was... Yeah, no one else was really building at this time, and it was, like, very quick to just do some prompt engineering and see, like, what is a task that is at the intersection of what's good, cap- technologically capable and, like, im- important for researchers.

And we had, like, a very nondescript landing page. It said nothing. But somehow people were signing up, and, and we had the sign-up form that were like, it was like, "Why are you here?" And everyone was like, "I need help with literature review."

And we're like, "Uh, literature review, that sounds so hard. I don't even know what that means." We're like, "We don't wanna work on it." Um, but then eventually we're like, "Okay, everyone is saying literature review. It's overwhelmingly people want-"

Swyx21:59

In all domains, not like medicine-

Jungwon Byun22:00

Yeah

Swyx22:00

... or physics or, just all domains.

Jungwon Byun22:02

Yeah. And we also kind of personally knew literature review was hard, and if you look at the graphs for academic literature being published every single month, you guys know this in machine learning, it's like up into the right-

Swyx22:11

Ugh

Jungwon Byun22:11

... like superhuman amounts of papers. Um, so we're like, "All right, let's just try it." I was really nervous, but Andreas was like, "This is kind of, like, the right problem space to jump into, even if we don't know what we're doing."

So my take was like, "Fine. This feels really scary, but let's just launch a feature every single week and double our user numbers every month. And if we can do that, we'll, like, we'll fail fast, and we will find something."

I was worried about, like, getting lost in the kind of academic white space. Um, so the very first version was actually a weekend prototype that Andreas made. Do you wanna explain how that worked?

Andreas Stuhlmüller22:43

I mostly remember theirs was really bad. So, so the, the, the thing I remember is, uh, you enter a question, and it would give you back a list of claims. So, so your question could be, I don't know, how does creatine affect cognition?

It would give you back some claims, um, that are ba- to some extent based on papers. But they were often irrelevant. The papers were often irrelevant. And so we ended up soon, uh, just printing out a bunch of examples of results and putting them up on the wall so that we would kind of feel the constant shame of having such a bad product, and, uh, would be incentivized to make it better.

And I think over time it has gotten a lot better, but I think, I think, uh, the, the initial version was, like, really very bad.

Jungwon Byun23:23

Yeah. But it was basically, like, a natural language summary of an abstract, like kind of a one-sentence summary, and which we still have. Um, and then, and then as we learned kind of more about this systematic review workflow, we started expanding the capability so that you could extract a lot more data from the papers and do more with that.

Swyx23:38

A- and were you using, like, embeddings and cosine similarity, that kind of stuff for retrieval, or was it keyword-based or...?

Andreas Stuhlmüller23:45

I think the very first version didn't even have its own search engine. I think the very first version probably used the Semantic Scholar API or something, something similar.

Swyx23:55

Yeah.

Andreas Stuhlmüller23:55

And, uh, only later when we discovered that, uh, that API is not very semantic- ... uh, we, uh, then built our own search ver- s- search, and that has helped a lot.

Swyx24:04

No, and then w- we're gonna go into, like, more recent product stuff. But like, um, you know, I think you seem the more sort of startup-oriented business person-

Andreas Stuhlmüller24:13

Mm-hmm

Swyx24:13

... and, and you seem the sort of more ideologically, like, interested in research obviously 'cause of your, your PhD. Um, w- what kind of market sizing were you guys thinking, right? Like, 'cause y- you're, you're here saying, like, "We have to double every month."

Jungwon Byun24:24

Mm-hmm.

Swyx24:25

And I'm like, "Well, I don't know how you make that conclusion from- ... from, from this," right? Especially also as a nonprofit at the time.

Jungwon Byun24:32

Mm-hmm. Um, yeah, I think, I mean, market size-wise, I think we... I felt like in this space where so much was changing, and it was very unclear Which of, what of today was actually gonna be true tomorrow. We just like really rested a lot on very simple fundamental principles, which is like research is r- if you can understand the truth, that is very economically beneficial, like valuable.

If you-

Swyx24:56

Okay

Jungwon Byun24:56

... like know the truth.

Swyx24:56

Just on principle-

Jungwon Byun24:57

Yeah

Swyx24:57

... that, that's enough for you.

Jungwon Byun24:58

Yeah. Research is obvi- like the key to many breakthroughs that are very commercially valuable.

Swyx25:03

'Cause my, my version of it is students are poor and they don't pay for anything, right? But that's obviously not true-

Jungwon Byun25:08

No, that's not what we found out

Swyx25:08

... as you guys have found out. But I, I, you know, uh, you had to have some market insight for me to have believed that, but I think you, you, you skipped that.

Jungwon Byun25:15

Yeah.

Andreas Stuhlmüller25:16

Yeah. I mean, we did encounter, I guess, uh, talking to VCs for our seed round. A lot of VCs were like, "You know, researchers, they don't have any money. Uh, why don't you build a legal assistant?" And, uh, I think in some short-sighted way, maybe that, that's, that's true.

Um, but I think in the long run, R&D is such a big space of the economy. I think if you can substantially improve how quickly pe- people find new discoveries or avoid, um, kind of controlled trials that don't go anywhere, I think that's just huge amounts of money.

Jungwon Byun25:49

Mm.

Andreas Stuhlmüller25:50

And there are a lot of questions obviously about between here and there, but I think as long as the fundamental principle-

Swyx25:55

Mm

Andreas Stuhlmüller25:55

... is there, we were, we were okay with that, and I guess we found some investors who also were.

Swyx26:00

Yeah. Yeah. C- congrats. Uh, I mean, I'm sure, I'm sure we can talk, cover the sort of flip, uh, later. Uh, yeah, I think, I think you were- you were about to start us on like GPT-3 and how like that changed things for you.

GPT Impact26:10

Swyx26:10

It, it's, it's funny, like I guess every major GPT version, you, you have like some big insight.

Jungwon Byun26:14

Mm-hmm.

Andreas Stuhlmüller26:16

I think it's a little bit less true for us than for others-

Swyx26:18

Okay

Andreas Stuhlmüller26:18

... because we always believed that there will basically be human-level machine work. And so, uh, it is definitely true that in practice for your product, as new models come out, your product starts working better. You can add some features that you couldn't add before.

But I don't think we really ever had the moment where we were like, "Oh wow, that is super unanticipated. We need to do something entirely different now-

Swyx26:44

Mm

Andreas Stuhlmüller26:44

... from what was on the roadmap."

Jungwon Byun26:47

I think, I think GPT-3 was a big change 'cause it kind of said, oh, now is the time to build, to u- that we can use AI to build these tools, and then GPT-4 was maybe a little bit more of an extension of GPT-3.

It felt less like a level sh... GPT-3 over GPT-2 was like qualitative level shift, and then GPT-4 was like, okay, great, now it's like, you know, much less ex- more accurate. We're more accurate on these things. We can answer harder questions.

But the shape of the product had already taken place by that time.

Swyx27:13

I kinda wanna ask you about this like sort of pivot that you made.

Jungwon Byun27:16

Mm-hmm.

Swyx27:16

But I, I guess that was just a way to sell what you were doing, which is you're adding, uh, extra features on, on grouping by concepts.

Jungwon Byun27:23

Uh, when GPT-3-

Swyx27:24

The GPT-4 piv- pivot, quote unquote pivot that you-

Jungwon Byun27:26

Oh yeah, yeah, exactly

Swyx27:27

... you, you worked on.

Jungwon Byun27:27

Right, right, right. Yeah. Yeah. When we launched this workflow, now that GPT-4 was available, um, basically we had... Elicit was at a place where given a table of papers, we... Like we have very tabular interfaces, so given a table of papers, you can extract data across all the tables.

Um, but that's still... Like you kind of wanna take the analysis a step further. Um, and sometimes what you'd care about is not having a list of papers, but a list of arguments, a list of effects, a list of interventions, a list of techniques, and so, um, that's when we-- That's one of the things we're working on is now, now that you've extracted this information in a more structured way, can you pivot it or group by whatever the information that you, um, extracted to have more insight first information still supported by the academic literature.

Swyx28:13

Yeah. That was a big revelation when I saw it. Yeah, uh, b- b- basically, I, I think I'm very just impressed by how, uh, first principles, your ideas around w- what the workflow is, and I think that's why you're not as, um, re- reliant on like the LLM improving, because, uh, it's actually just about improving the workflow that you would recommend to people.

Jungwon Byun28:32

Mm-hmm.

Swyx28:33

Um, today we might call it an agent, I don't know. But it's, um, it's not as-- You're not reliant on the LLM to drive it. You- it's reliant on your sort of this is the way that Elicit does research-

Jungwon Byun28:44

Mm-hmm

Swyx28:44

... and this is what we think is most effective based on talking to our users.

Jungwon Byun28:47

Yep, that's right. Yeah, I think we're very... The problem space is still huge. Like if it's like this big, we are, we are all still operating at this tiny part, bit of it. So, you know, I think, I think about this a lot in the context of moats.

People are like, "Oh, what's your moat? What happens if GPT-5 comes out?" It's like, if GPT-5 comes out, there's still like all of this other space that we can go into.

Swyx29:07

Yeah.

Jungwon Byun29:07

And so I think being really obsessed with the problem, which is very, very big, has, has helped us like stay robust and just kind of directly incorporate model improvements and then keep going.

Cost Optimization29:16

Swyx29:16

And then I first encountered you guys with Charlie.

Jungwon Byun29:19

Mm-hmm.

Swyx29:20

Uh, you can tell us about that, that project. Um, basically, yeah, like how much did cost become a concern as you're working more and more with OpenAI? Um, how do you manage that relationship?

Jungwon Byun29:30

Let me talk about who Charlie is.

Swyx29:32

Oh, sure, sure, sure.

Jungwon Byun29:32

And then you can talk about the tech-

Swyx29:33

That's all right. Yeah, yeah

Jungwon Byun29:33

... 'cause Char- Charlie is a, is a special character. So Charlie, um, when we found him was, had just finished his freshman year at the University of Warwick. I think he had heard about us on some Discord, and then he applied, and we saw it, we were like, "Wow, who is this freshman?"

And then we just saw that he had done so many incredible side projects. Um, and we were actually on a team retreat in Barcelona visiting our head of engineering at that time, and everyone was talking about this wunderkind.

They're like, "This kid." And then on our take-home project, he had done like the best of anyone th- to that point. Um, and so people were just like so excited to, to hire him. So we hired him as an intern, and then we were like, "Charlie, what if you just dropped out of school?"

And so then we convinced him to take a year off, um, and he is just incredibly productive, and I think the thing you're referring to is like at the start of 2023, Anthropic kind of launched their constitutional AI paper.

Um, and within a few days, I think four days, he had basically implemented that in production, and then we had it in app like a, a week or so after that, and, and he has since kind of contributed to major improvements, like cutting, cutting costs down to like a 10th of what they were.

Um, it's really large scale. But yeah, you can, you can talk about the technical stuff.

Andreas Stuhlmüller30:40

Yeah. On the constitutional AI project, this was for abstract summarization, where, um- In, in Elicit, if you run a query, it'll return papers to you, and then it will summarize each paper with respect to your query for you on the fly.

And that's a, like, really important part of Elicit because you... Like, Elicit does it so much. Like, if you, if you run a few searches, it'll have done it a few hundred times for you. And so we cared a lot about this both being, like, fast, cheap, and also, um, very low on hallucination.

I, I think if, if Elicit hallucinates something about the abstract, that's, that's really not good. And so what Charlie did in that project was create a constitution that, like, expressed what kind... What, what are the attributes of a good summary.

It's like, uh, everything in summary is reflected, uh, in the, in the, the actual abstract and, uh, is, like, very concise, et cetera, et cetera. And, uh, then used RLHF with a model that was trained on the Constitution to, um, basically fine-tune a better summarizer.

And as for-

Jungwon Byun31:45

An open source model, I think

Andreas Stuhlmüller31:46

... Yeah, on, on an open source model.

Jungwon Byun31:48

Yeah.

Andreas Stuhlmüller31:48

Yeah. Um, and I think that might still be in use.

Jungwon Byun31:51

Yeah. Yeah, definitely. Yeah, I think at the time, the models hadn't been trained at all to be faithful to a text, so they were just generating.

Andreas Stuhlmüller31:58

Mm-hmm.

Jungwon Byun31:58

So then when you asked them a question, they tried too hard to ask- answer the question and didn't try hard enough to answer the question given the text or answer what the text said about the question. So we had to basically teach the models to do that specific task.

Swyx32:12

How, um, do you monitor the ongoing performance of your models? Not to get too LLM, LLM ops-y, but, um, you, you, like, you are one of the larger, more well-known operations doing NLP at scale. I, I guess effectively, like, um, you have to monitor these things, and, um, nobody has a good answer that I talk to.

Andreas Stuhlmüller32:32

Yeah, I don't think we have a good answer yet. Um, I think the answers are actually a little bit clearer on the just kind of, uh, ro- basic robustness side of, I think, uh, where you can import ideas from normal, uh, software engineering and s- normal kind of DevOps.

You're like, "Well, you need to monitor kind of latencies and response times and uptime and whatnot."

Swyx32:52

Yeah. I think when, when we say performance, it's more about-

Andreas Stuhlmüller32:54

Yeah. Yeah. Yeah

Swyx32:54

... hallucination rate and-

Andreas Stuhlmüller32:55

And, and then things like hallucination rate, where I think there, the really important thing is, um, training time. So we care a lot about having our own internal benchmarks for, uh, model development that are, that reflect the distribution of user queries, so that we can know ahead of time how well is the model gonna perform on different types of tasks.

So the tasks being summarization, question answering, given a paper ranking. And for each of those, we wanna know what's the, what's the distribution of things the model is gonna see so that we can have, um, well-calibrated predictions on how well the model is gonna do in production.

And I think, yeah, there, there's, like, some chance that there's distribution shift, and actually the things users enter are gonna be different, but I think that's much less important than getting the kind of training right and having very high qual- quality, well-vetted data sets at training time.

Jungwon Byun33:49

I think we also end up effectively monitoring by trying to evaluate new models as they come out, and so that, that, like, kind of prompts us to, like, check th- go through our eval suite every couple of months and then...

Yeah. And, and so every time a new model comes out, we have to see, like, okay, which one is, is... How is this performing relative to production and, and what we currently have.

Swyx34:07

Yeah. I mean, since we're, we're on this topic, any new models that have really caught your eye this year? Like, Claude came out-

Jungwon Byun34:12

Claude, yeah

Swyx34:12

... a bunch.

Jungwon Byun34:12

I think Claude is pretty, pretty... I think the team's pretty excited about Claude.

Swyx34:16

Uh, excellent.

Andreas Stuhlmüller34:16

Yeah.

Swyx34:16

Yeah.

Andreas Stuhlmüller34:16

Specifically, I think Claude Haiku is, like, a good point on the kind of Pareto frontier. So I think it's like, it's not the m- it's neither the cheapest model, nor is it the most accurate, most high quality model, but it's just, like, a really good trade-off between cost and accuracy.

Swyx34:35

You apparently have to 10-shot it to make it good. Um, I tried using Haiku for summarization, but, uh, zero-shot was not great.

Jungwon Byun34:42

Hmm.

Swyx34:42

Um, and yeah, then they were like, "You know, it's a skill issue. You have to try harder."

Jungwon Byun34:47

Hmm. Interesting. Yeah, we also used, um, I think GPT-4 unlocked pro-

Andreas Stuhlmüller34:51

Do you mean turbo, or?

Jungwon Byun34:53

Yeah. Uh, yeah. I think it unlocked, um, tables for us-

Andreas Stuhlmüller34:58

Hmm

Jungwon Byun34:58

... processing data from tables, which was huge.

Andreas Stuhlmüller35:00

GPT-4 Vision.

Jungwon Byun35:01

Yeah.

Swyx35:02

Yeah. Did you try, like, Phu? I guess you can't try Phu 'cause it's non-commercial. Um, that's the Adept model-

Jungwon Byun35:08

Yeah, we haven't tried that one

Swyx35:08

... from the multi-model. Yeah, yeah.

Jungwon Byun35:09

Yeah.

Swyx35:09

Yeah, but Claude is multi-modal as well.

Jungwon Byun35:11

Yeah.

Swyx35:12

Yeah. I think the, the interesting, uh, insight that we got from talking to David Luan, who is CEO of Adept, uh, was that multimodality has, like, effectively two different flavors. Like, one is the, we recognize images from a camera in the outside natural world, and actually, the, the more important multimodality for knowledge work is screenshots.

Jungwon Byun35:31

Mm-hmm.

Swyx35:31

And, you know-

Jungwon Byun35:32

Yeah

Swyx35:32

... PDFs and charts and graphs.

Jungwon Byun35:34

Yeah. Yeah. Mm-hmm.

Swyx35:36

So we need a, a new term for that kind of multimodality.

Jungwon Byun35:38

Yeah.

Andreas Stuhlmüller35:39

But is, is the claim that current models are good at one or the other?

Swyx35:42

They're, they're over-indexed 'cause of the history of computer vision is Coco.

Jungwon Byun35:46

Mm-hmm.

Swyx35:46

Right? So now we are like, "Oh, actually, you know, s- screens are more important."

Jungwon Byun35:51

Yeah, processing weird handwriting and-

Swyx35:52

OCR

Jungwon Byun35:53

... yeah.

Swyx35:53

Yeah, handwriting.

Jungwon Byun35:54

Yeah.

Swyx35:54

Yeah. You mentioned a lot of, like, uh, closed model lab stuff, and then you also have, like, this, uh, uh, open source model fine-tuning stuff. Like, what is your workload now between closed and open?

Andreas Stuhlmüller36:04

It's a good question. I think-

Swyx36:05

Is it half and half?

Andreas Stuhlmüller36:06

It's, uh-

Swyx36:07

Is that even a relevant question or not? Is this, this a nonsensical question?

Andreas Stuhlmüller36:11

It depends a little bit on, like, how you index, whether you index by, like, computer cost or number of queries. Um, I'd say, like, in terms of number of queries, it's maybe similar. In terms of, like, cost and compute, I think the closed models make, make up more of the budget since the main cases where you wanna use closed models are cases where-

Swyx36:30

Intelligence

Andreas Stuhlmüller36:30

... they're just smarter, uh, where there are no, where no existing, um, open source models are quite smart enough.

Swyx36:36

We have a lot of interesting technical questions to go in, but just to wrap the

Alessio36:41

It kind of like UX evolution. Now you have the notebooks.

Jungwon Byun36:44

Mm-hmm.

Notebooks36:44

Alessio36:44

Um, we talk a lot about how chatbots are not-

Jungwon Byun36:47

Mm

Alessio36:48

... the final frontier, you know. Uh, how did you decide to get into notebooks, um, which is a very iterative, kinda like interactive interface and yeah, maybe learnings from that?

Jungwon Byun36:58

Yeah. This is actually our fourth time trying to make this work.

Alessio37:02

Okay.

Jungwon Byun37:02

Um, I think the first time was probably in early 2021. Um, at the time we built something... I think we, I think because we've always been obsessed with this idea of task decomposition and like branching, we always want, we always wanted a way, a, a tool that could be kind of unbounded, where you could keep going, um, where you could do a lot of branching, where you could kind of apply language model operations or computations on other tasks.

Alessio37:26

Mm.

Jungwon Byun37:26

So in 2021, we had this thing called composite tasks, where you could use GPT-3 to brainstorm a bunch of research questions, and then take each research question and decompose those further into sub-questions. And this kind of, again, that like task decomposition tree type thing was, was always very exciting to us.

But that was like, it didn't work, and it was kind of overwhelming. Um, then at the end of '22, I think we tried again, and at that point we were thinking, "Okay, we've done a lot with this literature review thing.

We also wanna start helping with kind of adjacent domains and different workflows. Like, we wanna help more with machine learning. What does that look like?" And as we were thinking about it, we were like, "Well, there are so many research workflows.

Like, how do we not just build kind of three new workflows into Elicit, but make Elicit really generic to lots of workflows? What is like a generic composable system with nice abstractions that can like, scale to all these workflows?"

So we like iterated on that a bunch, and like didn't quite narrow the problem space enough, or, or like quite get to what we wanted. And then I think it was at the a- beginning of 2023, where we're like, "Wow, computational notebooks kind of enable this," where they have a lot of flexibility, um, but you know, kind of robust primitives such that you can, you can extend the workflow and it's, it's not limited.

It's not like you ask a query, you get an answer, you're done. You can just constantly keep building on top of that. Um, and each little step seems like a really good, um, kind of unit of work for the language model.

So that's, that's a... And, and also it was just like really helpful to have a bit more kind of preexisting work to, to emulate. Um, so that's, that was, yeah, that's kind of how we ended up at comp- at computational notebooks for Elicit.

Andreas Stuhlmüller39:03

Maybe one thing t- that's worth making explicit is the difference between computational notebooks and chat, because on the surface they seem pretty similar. It's kind of this iterative interaction-

Alessio39:12

Mm

Andreas Stuhlmüller39:12

... where you add stuff. Um, and you, it's almost like i- i- in both cases you have a back and forth between you enter stuff, and then you get some output, and then you enter stuff. But the important difference in our minds is with notebooks you can define a process.

So, um, in, in data science you can be like, "Here's like my data analysis process that takes in a CSV and then does some extraction, and then generates a figure at the end." And, uh, you can prototype it using a small CSV, and then you can run it over a much larger CSV later.

And similarly, the vision for notebooks in our case is to not make it this like one-off chat interaction, but to allow you to then say kind of... If, if you, if you start, and first you're like, "Okay, let me just analyze a few papers and see, do I get to the correct, like, conclusions for those few papers?"

Can I then later go back and say, "Now let me run this over 10,000 papers, um, now that I've debugged the process using a few ch- papers?" And that's an interaction that doesn't fit quite as well into the chat framework because that's more for kind of quick back and forth interaction.

Alessio40:15

Do, do you think in notebooks as kinda like structure editable chain of thought, m- basically-

Andreas Stuhlmüller40:21

Mm

Alessio40:21

... step-by-step? Like i- is that kind of where you, where you see this going, and then are people gonna reuse notebooks as like templates? In, maybe in traditional notebooks it's like cookbooks, right? You share a cookbook-

Andreas Stuhlmüller40:32

Mm-hmm

Alessio40:32

... you can start from there. Um, is this similar in, in Elicit?

Andreas Stuhlmüller40:35

Yeah, that's exactly right. Uh, so we, that's our hope, that people will build templates, share them with other people. Um, I think chain of thought is maybe still like kind of one level lower on the abstraction hierarchy than we would think of notebooks.

I think, I think we'll probably want to think about more semantic pieces, like a, a building block is more like a, a paper search or an extraction-

Alessio40:56

Mm-hmm

Andreas Stuhlmüller40:57

... or, um, a list of concepts. Um, and then the model's detailed reasoning will probably often be one level down. You always want to be able to see it, but you don't always want it to be front and center.

Alessio41:10

Yeah. What's the difference between a notebook and an agent, since everybody always asks me, "What's an agent?" Like, how do you think about where the, where the line is?

Andreas Stuhlmüller41:17

In the notebook world, I would generally think of the human as the agent in the first iteration. So you have the notebook and the human kind of adds little action steps. And then the next s- point on this kind of progress gradient is, okay, now you can use language models to predict which action would you take as a human.

And at some point you're probably gonna be very good at this. You'll be like, "Okay, I can like, in some cases I can with 100%, 99.9% accuracy predict what you do." And then you might as well just execute it, like why wait for the human?

And eventually, as you get better at this, you can, that will just look more and more like agents taking actions as opposed to do, you doing the thing. Um, and, uh, l- I think templates are a specific case of this, where you're like, "Okay, well, there's just particular sequences of actions that you often wanna chunk and, uh, have available as primitives, just like in normal programming."

And, uh, those are, y- you can view them as action sequences of agents, or you can view them as more like the normal programming language abstraction thing. And I think those are two, two valid views.

Alessio42:20

Mm-hmm. Yeah. How do you see this change as, like you said, the models get better and, and you need less and less human actual interfacing with the model, you just get the results? Like, how does the UX and, um, the way people perceive it change?

Jungwon Byun42:35

Yeah, I think this, um, kind of interaction paradigms for evaluation is not really something the internet has encountered yet, because right now, up to now, the internet has all been about like get- getting data and work from people.

Um, but so increasingly, I-- So I-- Yeah, I really want kind of evaluation both from an interface perspective and from, like, a technical perspective and operation perspective to be a power, superpower for Elicit, 'cause I think over time, models will do more and more of the work, and people will have to do more and more of the evaluation.

Um, so, so I think, yeah, in terms of the interface, some of the things we have today are, um, you know, for every kind of language model generation, there's some citation back, and we kind of directly-- we try to highlight it, highlight the ground truth in the paper that is most relevant to whatever Elicit said and make it super easy so that you can click on it and quickly see in context, um, and, and validate whether the text actually supports the answer that Elicit gave.

So I think we'd probably want to scale things up like that, like the ability to kind of spot-check, um, the models work super quickly, scale up interfaces like that, and, um-

Swyx43:37

Who, who would spot check? The, the user?

Jungwon Byun43:39

Yeah, to start it would be the user.

Swyx43:41

Uh-huh.

Jungwon Byun43:41

Um, one of the other things we do is also kind of flag the model's uncertainty. So we have models report out, "How confident are you that this was the sample size of this study?" The model's not sure, we throw a flag, and so the user knows to prioritize checking that.

Um, so again, we can kind of scale that up, so when the model's like, "Well, you know, I went and searched for Google... Uh, searched this on Google, I'm not sure if that was the right thing," I have an uncertainty flag, and the user can go and be like, "Oh, okay, that was actually the right thing to do or not."

Swyx44:07

So, uh, I've tried to do uncertainty readings from models. Uh, I, I don't know if you have this live.

Jungwon Byun44:13

Mm-hmm.

Swyx44:13

You do?

Jungwon Byun44:13

Yeah. Mm-hmm.

Swyx44:14

Okay. Um, 'cause I just didn't find them reliable because they just hallucinated their own uncertainty. Uh, I would love to base it on log probs or something more native within the model rather than generated, but i- okay, it sounds like they scale properly for you.

Jungwon Byun44:29

Yeah.

Swyx44:29

Yeah, okay.

Jungwon Byun44:29

We found it to be pretty calibrated. They're-- It varies on the model.

Swyx44:32

Okay.

Jungwon Byun44:33

Yeah.

Andreas Stuhlmüller44:33

Yeah. I, I think in some cases, we also use the different models for the uncertainty estimates than for the question answering. So one model would say, "Here's my chain of thought, here's my answer," and then a different type of model...

Let, let's say the first model is Llama, and let's say the second model is GPT-3.5, could be, could be different. Uh, and then the, the other-- the second model just looks over the results and is like, "Okay, what's, uh...

How, how confident are you in this?" And I think sometimes using a different model can be better than using the same model.

Swyx45:03

Yeah, you know, on top of your models, evaluating models, obviously you can do that all day long. Uh, like, what's your budget? Like, what... Um, because y- your queries fan out a lot.

Budget & Uncertainty45:03

Jungwon Byun45:13

Mm-hmm.

Swyx45:14

And then you have models evaluating models. Um, uh, you know, one person typing in a sen- a question can lead to a thousand calls.

Andreas Stuhlmüller45:23

It depends on the project. So, um, if the project is, um, basically a, a systematic review that otherwise human research assistants would do, then, um, the project is basically the human equivalent spend, and the sp- the spend can get quite large for those projects.

Uh, certainly, uh, I, I don't know, let's say $100,000.

Jungwon Byun45:43

For the project, yeah.

Andreas Stuhlmüller45:44

Yeah. Um, so i- in those cases, you're happier to spend compute than in the kind of shallow search case where someone just enters a question because, I don't know, maybe-

Swyx45:54

Feel like it. Yeah

Andreas Stuhlmüller45:55

... I, I heard about creatine, what's it about? Uh, probably don't wanna spend, to spend a lot of compute o- on that. And, uh, just sort of being able to invest more or less compute into getting more or less accurate answers is, I think, one of the core things we care about, and that I think is c- currently undervalued in the AI space.

I think currently you can choose which model you want, and you can sometimes tell it to, uh, I don't know, you'll tip it and, uh-

Swyx46:20

Mm-hmm. Yeah

Andreas Stuhlmüller46:20

... or it'll try harder, or you can, like, try various things and to get it to work harder. But you don't have great ways of con- converting, uh, willingness to spend into better answers-

Swyx46:29

Mm-hmm

Andreas Stuhlmüller46:30

... and we really want to build a product that has this sort of unbounded flavor where, like, I mean, as much as you care about... Like, if you care about it a lot, you should be able to get-

Jungwon Byun46:38

Yeah

Andreas Stuhlmüller46:38

... really high-quality answers, uh, really double-checked in every way.

Swyx46:42

Yeah.

Jungwon Byun46:42

And, and you have a credits-based pricing.

Andreas Stuhlmüller46:45

Mm-hmm. Yeah.

Jungwon Byun46:45

So unlike most products, it's not a fixed monthly fee.

Andreas Stuhlmüller46:47

Right, exactly.

Jungwon Byun46:48

Uh, yeah.

Andreas Stuhlmüller46:48

So some of the higher costs are tiered. So, like, for most casual users, they'll just get the abstract summary, which is kind of an open source model. Then you can, you know, add more columns which have more extractions and these uncertainty features, and then you can also add those same columns in high accuracy mode, which also par- parses the table, so we kind of stack the complexity and the cost.

Swyx47:10

You know the fun thing you can do with a credit system, which is data for data, um, or, uh, I, I don't know what I'm, what I mean by that. Basically, you can give people more credits if they give data back to you.

Jungwon Byun47:20

Yeah.

Swyx47:21

I don't know if you've already done that.

Jungwon Byun47:22

I've, I've thought about... I've-- We've thought about something like this. It's like, if you don't have money but you have time-

Swyx47:26

Yes

Jungwon Byun47:27

... how do you exchange that?

Swyx47:28

Trade. Yeah.

Jungwon Byun47:28

Yeah.

Swyx47:28

It's a fair trade.

Jungwon Byun47:29

Yeah. I think, I think it's interesting. We haven't quite operationalized it, and then there's, you know, there's been some kind of, like, adverse selection. Like, you know, for example, it would be really valuable to get feedback on our models, so maybe if you were willing to give more ro- robust feedback on our results, we could give you credits or something like that.

Swyx47:43

Yeah.

Jungwon Byun47:43

But then there's kind of this-

Swyx47:45

Adverse selection

Jungwon Byun47:45

... will people take it seriously?

Swyx47:46

Yeah, you want the good people.

Jungwon Byun47:46

Exactly.

Swyx47:47

Can you tell who are the good people?

Jungwon Byun47:49

Not right now, but yeah, maybe at the point where we can, we can offer it. We can offer it up then.

Swyx47:53

The in- the perplexity of questions asked, you know. The high-- If it's higher perplexity, these are smarter people.

Jungwon Byun47:57

Yeah. Yeah. Maybe. Yeah.

Andreas Stuhlmüller47:58

If you make a lot of typos in your queries, you're not gonna get credits, it's strange.

Long Context48:04

Swyx48:04

Negative social credit.

Jungwon Byun48:05

Mm-hmm.

Swyx48:06

Uh, it's very topical right now to think about the threat of long context windows. Um, you know, all these models that we are d- we're talking about these days are all, like, a billion token plus. Is that relevant for you?

Do you... Can you make use of that? Is that just prohibitively expensive 'cause you're just paying for all those tokens, or are you just doing RAG?

Andreas Stuhlmüller48:23

It's definitely relevant. And, uh, um, when we think about search, I think as, as many people do, we think about kind of a staged pipeline of retrieval, where first you use a kind of s- s- semantic search database, uh, with embeddings, get, like, the l- in our case, maybe 400 or so most relevant papers.

And then, then you still need to rank those. And I think at that point it becomes pretty interesting to use larger models. So, um, specifically in the past, I think a lot of ranking was kind of Per item ranking where you would score each individual item, maybe using inc- increasingly expensive scoring methods, and then rank based on the scores.

But I think list-wise reranking where you have a model that can see all the elements is a, a lot more powerful. Because often you can only really tell how good a thing is in comparison to other things, and, uh, what things should come first, it really depends on like, well, what other things are available.

Maybe you even care about diversity in your results. You don't wanna show like t- 10 very similar papers as the first 10 results. So I think along context models are quite interesting there. And especially for, for our case where, um, we care more about power users who are perhaps a little bit more willing to wait a little bit longer to get higher quality results relative to, um, people who just quickly check out things because why not?

Um, I think, uh, being able to spend more on longer context is quite valuable.

Jungwon Byun49:47

Yeah, I think one thing the longer context model's changed for us is maybe a focus from breaking down tasks to breaking down the evaluation. So before, um, you know, if we wanted to answer a question from the full text of a paper, we had to figure out how to chunk it and like find the relevant chunk, and then answer based on that chunk.

And the nice thing was then you know kind of which chunk the model used to answer the question. So if you want to help the user check it, yeah, you can be like, "Well, this was the chunk that the model got."

And now if you put the whole text in the paper, you have to go back. You have to like kind of find the chunk like more retroactively, basically. And so you need kind of like a different set of abilities and, and obviously like different technology to figure out.

You still want to sh- point the user to the supporting quotes in the text, but then like the, the interaction is a little different.

Alessio50:33

You like scan through and find some ROUGE score-

Jungwon Byun50:36

Yeah

Alessio50:36

... ceiling or floor.

Andreas Stuhlmüller50:39

Yeah, I think, I think, I think there's an interesting space of, uh, almost research problems here. Because y- you would ideally make causal claims, like, uh, if this hadn't been in the text, the model wouldn't have said this thing.

And, uh, maybe you can do expensive approximations to that where like, I don't know, you just throw a chunk off the paper and re-answer and see what happens. But hopefully there are better ways of, of doing that, um, where you just get that inf- that kind of counterfactual information for free from the model.

Alessio51:07

Do you think at all about the cost of maintaining RAG versus just putting more tokens in the window? I think in software development, a lot of times people buy developer productivity things so that we don't have to worry about it.

Um, context window's kind of the same, right? You have to maintain chunking and like RAG retrieval and like reranking and all of this versus I just shove everything into the context, and like it costs a little more, but at least I don't have to do all of that.

Uh, is that something you, you thought about at all?

Jungwon Byun51:33

I think we still like hit up against context window, context limits enough that like it's not really do we still wanna keep this RAG around? It's like we do still need it-

Alessio51:44

Mm-hmm

Jungwon Byun51:44

... for the scale of the work that we're doing. Yeah.

Andreas Stuhlmüller51:46

And I think there are different kinds of maintainability. Uh, g- in one sense, I think you're right that the t- throw everything into the context window thing is easier to maintain because s- you just can swap out a model.

Um, in another sense, it's if things go wrong, it's harder to debug. Where like if you know here's the process that we go through to, uh, go from 200 million papers to an answer, and there are like little steps and you understand, okay, this is, this is the step that finds the relevant paragraph or what- whatever it may be, um, you'll know which step breaks if the, the answers are bad.

Whereas if, uh, it's just like, uh, a new model version c- came out, and now it suddenly doesn't find your needle in a haystack anymore, then you're like, "Okay, what can you do?" Uh, you're kind of at a loss.

Alessio52:31

Mm-hmm. Yeah. Let, let's talk a bit about, um, yeah, needle in a haystack and like, uh, maybe the opposite of it, which is like hard grounding. I don't know if that's like the best name to, to think about it.

But I was using one of these chat with your documents features, and I put the AMD MI300 specs and the, um, new, you know, Blackwell chips from NVIDIA, and I was asking questions. And, uh, asked it, "Does the AMD chip support NVLink?"

And the response was like, "Oh, it doesn't say in the specs."

Jungwon Byun52:59

Mm.

Alessio53:00

But if you ask GPT-4 without the docs, it would tell you-

Jungwon Byun53:03

Mm

Alessio53:03

... no, because NVLink, it's a NVIDIA technology.

Andreas Stuhlmüller53:06

It's supposed to, it's supposed to NV.

Alessio53:07

Yeah.

Andreas Stuhlmüller53:07

Come on.

Alessio53:08

It, it, it just says in the thing. Uh, how, h- how do you think about that where like having the context sometimes suppress the knowledge that the model has?

Andreas Stuhlmüller53:16

It really depends on the task because I think sometimes that is exactly what you want. So, uh, imagine you're a researcher and you're writing the background section of your paper, and you're trying to s- describe what these other papers say.

Um, you really don't want extra information to be introduced there.

Alessio53:30

Mm.

Andreas Stuhlmüller53:30

In other cases where you're just trying to figure out the truth and, uh, you're giving the documents because you think they will help the model figure out what the truth is, um, I think you do want ... If the, if the model has a hunch that there might be something that's not in the papers, you do want to surface that.

I think ideally you still don't want the model to just tell you. I think probably the ideal thing looks a bit more like, um, agent control where the model can issue a query that, uh, then is intended to surface documents-

Alessio54:01

Mm

Andreas Stuhlmüller54:01

... that substantiate its hunch. Um, so I would ... That's maybe a, a reasonable middle ground between model just telling you and model being fully limited to the papers you give it.

Alessio54:10

Mm-hmm.

Jungwon Byun54:11

Yeah, I would say it's, they're just kind of different tasks right now, and the task that Elicit is mostly focused on is what do these papers say? Um, but there's another task which is like just- Give me the best possible answer, and that, give me the best possible answer sometimes depends on what do these papers say, but it can also depend on other stuff that's not in the papers.

So ideally, we can do both, and then kind of h- you know, do this overall task for you more going forward.

Underrated Features54:34

Alessio54:34

All right. This was, uh, we see a lot of, a lot of details, but just to zoom back out a little bit, what are maybe the most underrated features o- of Elicit, um, and what is one thing that maybe the users surprise you the most by, by using it?

Jungwon Byun54:48

I think the most powerful feature of Elicit is the ability to, um, extract, add columns to this table, which effectively extracts data from all of your papers at once. It's still-- It's well-used, but there are kind of many different extensions of that, that I think users are still discovering.

So one is we let you give a description of the column, we let you give instructions of the, of co- of a column, we let you create custom columns. So we have, like, thirty-plus predefined fields that users can extract, like what were the methods?

What were the main findings? How many people were studied? Um, and then, uh, and we can-- we actually show you basically the prompts that we're using to extract that from our predefined fields, and then you can fork this, and you can say, "Oh, actually, I don't care about the population of people, I only care about the population of rats."

Like, you can change the instructions. So, um, I think users are still kind of discovering that there's both this predefined, easy-to-use default, but that they can extend it to be much more specific to them, and then they can also ask custom questions.

One use case of that is you can, you can start to create different column types that you might not expect. So rather than just creating generative answers, like a description of the methodology, you can say, "Classify the methodology into a prospective study, a retrospective study, um, or, you know, a case study," and then you can filter based on that.

It's like all using the same kind of technology and, and the interface, but it, it unlocks different workflows. Uh, so I think that, that like, the ability to ask custom questions, give instructions, and specifically use that to create different types of columns, like classification columns, is still pretty underrated.

Um, in terms of use case, uh, I spoke to someone who works in medical affairs at a genomic sequencing company recently. So they, you know, they... doctors kind of d- you know, order these genomic tests, um, these sequencing tests to kind of identify if a patient has a particular disease.

This company helps them process it, and this person basically interacts with all the doctors and if the doctors have any questions. My understanding is that medical affairs is kind of like customer support-

Alessio56:48

Mm-hmm

Jungwon Byun56:48

... or customer success in, in pharma. Um, so this person, like, talks to doctors all the lo- all day long, and, um, uh, one of the things they, they started using Elicit for is, like, putting the results of their tests as the query.

Like, this test showed, you know, this percentage popu- you know, presence of this, and forty percent that, and whatever. What, what do we think is kind of the, you know, what, what, like, genes are present here or something?

Or, or what's, what's in this sample? Um, and getting kind of a list of academic papers that would support their findings, and using this to help doctors interpret their tests. Uh, and so we talked about, okay, cool, like if we built, you know, it would-- he, he's pretty interested in kind of doing a survey of infectious disease specialists and getting them to evaluate a l- you know, having them write up their answers, comparing it to Elicit's answers, trying to see, can, can Elicit start being used to interpret the results of these diagnostic tests?

Because the way they ship these tests to doctors is they report on a really wide array of things. Um, and he was saying that at a large, well-resourced hospital, like a city hospital, there might be a team of infectious disease specialists who can help interpret these results.

But at under-resourced hospitals or more rural hospitals, the physician, the primary care physician can't interpret this, the, the t- the test results. So then they can't order it, they can't use it, they can't help their patients with it.

Um, so thinking about, you know, kind of a, an evidence-backed way of interpreting these tests is definitely kind of an extension of the product that I hadn't considered before. But yeah, the idea of, like, using that to g- bring more access to physicians in all different parts of the country and helping them interpret complicated science is pretty cool.

Future & Mission58:23

Alessio58:23

Yeah. Um, we had Ken-Jun from Imbue on, on the podcast, and we talked about better allocating scientific resources. How do you think about these use cases and maybe how Elicit can help drive more research? And do you see a world in which, you know, maybe the models actually do some of the research before, uh, suggesting us?

Andreas Stuhlmüller58:42

Yeah, I think, um, that's, like, very close to what we care about. So our, our product values are systematic, transparent, and unbounded. And I think, uh, to make research more, especially more systematic and unbounded, I, I think is, like, basically the thing that's at stake here.

Alessio58:58

Mm-hmm.

Andreas Stuhlmüller58:59

So ideally, uh, people would think, well, what are... For, for example, um, I recent- was recently talking to people in longevity, and, uh, I think there isn't really one field of longevity. There are kind of different scientific sub-domains that, um, are surfacing various things that are related to longevity.

And I think if you could more systematically say, "Look, here, here, um, here are all the different inter- interventions we could do, and, uh, here's the expected ROI of these experiments. Here are, here's, like, the evidence so far that supports those being, uh, either, like, likely to surface new information or not.

Uh, here's the cost of these experiments." I, I think you could be in so much more systematic than, uh, science is today. Uh, probably, uh, yeah, I, I, I'd guess in like ten, twenty years, we'll look back and it will be, uh, incredible how unsystematic science was back in the day.

Jungwon Byun59:49

Yeah, and, um, I think this is... As we start to-- So, like, our, our view is kind of have models catch up to expert humans today, you know, or whatever. Start with kind of novice humans, and then increasingly expert humans, and then at some point...

And, and but we really want the models to kind of, like, earn their right to the expertise. Um, so that's why we do things in this very step-by-step way. That's why we don't just, like, throw a bunch of data and apply a bunch of compute and hope we get good results.

But obviously, at some point, you hope that once it's kind of earned its stripes, it can surpass human researchers. Um, but I think that's where making sure that the model's processes are really explicit and transparent and that it's really easy to evaluate is important.

Because if it does surpass human understanding, people will still need to be able to audit its work somehow or spot-check its work somehow to be able to reliably trust it and use it. Um, so yeah, that's kind of why the process-based approach is, is really important.

Andreas Stuhlmüller1:00:43

And on the, on the question of, um- Will models do their own research? I think one feature that models currently don't have that will need to be better there is, uh, better world models. I, I think currently models are just not great at representing kind of what's going on in a particular situation or domain in a way that allows them to, mm, come to interesting, surprising conclusions.

Um, I think they're very good at, like, coming to, I don't know, conclusions that are nearby-

Swyx1:01:12

Mm-hmm

Andreas Stuhlmüller1:01:12

... um, to conclusions that people have come to, but, uh, not as good at kind of reasoning and being-- making surprising connections maybe. Um, and so having deeper models of how... Let's see. What are the underlying structures of different domains?

How are they related or not related, I think will be an important ingredient for models actually being able to make novel contributions.

Swyx1:01:33

On the topic of hiring more expert humans- ... uh, you've hired some very expert humans. Uh, my friend Maggie Appleton joined you guys, I think, maybe a year ago-ish. Um, in fact, we-- Um, I think you got-- You're doing an offsite, and we're actually organizing our big sort of AI, AI UX meetup, uh, around whenever she's in town in San Francisco.

Jungwon Byun1:01:51

Oh, amazing.

Swyx1:01:52

How big is the team? Um, you know, how, how have you sort of transitioned your company into, into this sort of PBC and, and sort of w- the, the plan for the future?

Jungwon Byun1:02:01

Yeah. We're 12 people now. Um, mostly about half of us are in the Bay Area and then distributed across US and Europe. Um, a mix of m- mostly kind of in roles in engineering and product. Um, yeah, and I think that the transition to PBC was really not that eventful because I think we were already-- even as a nonprofit, we were already shipping every week, so very much operating as a product.

Swyx1:02:23

Very much like a startup, yeah.

Jungwon Byun1:02:24

Yeah. And then I would say the, the kind of PBC component was to, you know, very explicitly say that we have a mission that we care a lot about. There are a lot of ways to make money. We think our mission will make us a lot of money, but we are going to be opinionated about how we make money.

We're gonna take the, the version of making a lot of money that's in line with our mission. But it's, like, all very-- It's very convergent. Like, Elicit is not going to make any money if it's a bad product, if it doesn't actually help you discover truth and, and, and do research, research more rigorously.

So, um, I think, I think for us, the kind of mission and the success of the company are very, very intertwined. Um, so a big part of... Yeah, we're hoping to grow the team quite a lot this year.

Um, probably some of our highest, uh, priority roles are in engineering, um, but also, um, opening up roles in m- more in design and product marketing, go-to-market. Um, yeah. Do you want to talk about the roles?

Andreas Stuhlmüller1:03:14

Yeah. Broadly, we are just looking for senior software engineers and, uh, don't need any particular AI expertise. A lot of it is just, uh, I, I guess, how do you, uh, build good orchestration for complex tasks?

Swyx1:03:27

Mm-hmm.

Andreas Stuhlmüller1:03:27

So we, we talked earlier about, uh, kind of these are sort of notebooks, scaling up, task orchestration, and, uh, the-- I think a lot of this looks more like traditional software engineering than it does look like machine learning research, and I think the people who are, like, really good at, uh, building good abstractions, uh, building applications that can, uh, kind of survive even if some of their pieces break, like making reliable components out of unreliable pieces, I think those, those are the people we are looking for.

Swyx1:03:57

You know, that's exactly what I used to do. Have you explored any of the-

Jungwon Byun1:04:00

Do you want to come work with us?

Swyx1:04:00

... existing- I mean, I can talk about this all day. Have you explored the existing orchestration frameworks, Temporal, uh, Airflow, Daxter, Prefect?

Andreas Stuhlmüller1:04:09

We've looked into them a little bit. I think we have some specific requirements around, um, kind of being able to stream work back very quickly to, um, our users. Uh, those, uh, those could definitely be relevant.

Swyx1:04:20

Okay. Well, you're hiring. I'm sure we'll plug all the links. Thank you so much for coming. Any parting words? Any words of wisdom, mottos you live by?

Jungwon Byun1:04:29

No, I, I think it's a-- I think it's a really important time for humanity, so I hope everyone listening to this podcast can think hard about exactly which part of the-- w- how they want to participate in this story.

There's so much to build, and we can be really intentional about what we align ourselves with. Um, I think there are just-- there are a lot of applications that are going to be really good for the world and a lot of applications that are not.

And so, yeah, I hope people can take that seriously and, and kind of seize the moment.

Swyx1:04:57

Yeah. I love how intentional you guys have been. Thank you for sharing that story.

Jungwon Byun1:05:00

Thank you.

Andreas Stuhlmüller1:05:00

Yeah. Thank you for coming on.