# How AI is Eating Finance - with Mike Conover of Brightwave

Latent Space · 2024-06-11

<https://addtry.com/45829207-c73e-4d99-83b9-b9a1f9172f30>

Brightwave founder Mike Conover explains how his vertical AI startup builds a 'partner in thought' for finance professionals, using a systems-of-systems approach that decomposes problems into specialized subsystems rather than relying on large context windows or monolithic agents. He argues that the evolution from DALL-E's 1,024-token context to today's million-token models has not solved synthesis; smaller, focused reasoning units still outperform. Conover emphasizes that the real competitive edge lies in custom training data and human annotation, not pre-training, and that models are converging in capability—so value shifts to fine-tuning data that elicits specific behaviors. He details why Brightwave avoids spreadsheets in favor of conversational analysis, and predicts AI will automate idea generation and second-order derivative bets, but humans must remain the final synthesizers and deciders. The episode also covers hiring for vertical AI, the role of knowledge graphs, and the diminishing economic incentive for companies to train their own foundation models.

## Questions this episode answers

### What does Brightwave do and how is it used by finance professionals?

Mike Conover describes Brightwave as a 'partner in thought' for finance professionals that generates tight, actionable financial analysis from any question. For example, it can analyze Nvidia's GPU market position in light of rare earth metal shortages, tracing supply chain impacts. It is used by a $20 billion hedge fund's equities team for deep thesis development, by wealth managers to prepare for client conversations, and by investor relations teams to rapidly cover market-moving documents.

[9:36](https://addtry.com/45829207-c73e-4d99-83b9-b9a1f9172f30?t=576000)

### Why doesn't Brightwave simply use large context windows to analyze financial documents?

Mike Conover explains that large context windows are empirically insufficient for deep synthesis. Models have a characteristic output length of about 1,200 tokens and produce shallow summaries when given vast documents. Instead, Brightwave decomposes the problem into subsystems that understand document structure and user intent, uses multiple evaluation passes, and avoids just 'throwing all SEC filings into a million-token window' which fails to yield useful, non-consensus insight.

[13:34](https://addtry.com/45829207-c73e-4d99-83b9-b9a1f9172f30?t=814000)

### Why doesn't Brightwave generate Excel spreadsheets like other financial AI tools?

Mike Conover states that spreadsheets are fault-intolerant—a subtle formula error can invalidate all assumptions, wasting effort. Instead, Brightwave acts as a dialogue partner that allows users to pivot away from uninteresting findings. Financial modeling is also highly personal, with each fund having unique key drivers, so a one-click model does not solve the real problem of developing a thesis. The goal is to enhance qualitative reasoning, not automate spreadsheet creation.

[45:56](https://addtry.com/45829207-c73e-4d99-83b9-b9a1f9172f30?t=2756000)

## Key moments

- **[0:00] Intro**
  - [1:35] Mike Conover's Nature Communications paper demonstrated that 500 million job transitions predict next quarter S&P 500 market cap changes.
- **[4:52] Product**
  - [4:52] Brightwave raised a $6 million seed round led by Decibel Partners, with participation from Point72 and Moonfire Ventures.
  - [6:08] Brightwave co-founder Brandon Qatar, former CTO of a federally regulated derivatives exchange, filed his first deep learning patent in 2018.
  - [10:19] A $20 billion crossover hedge fund's equities team uses Brightwave to deep-dive into investment theses across multiple steps of the supply chain.
- **[12:49] Context**
  - [13:06] Mike Conover finds that large context windows do not empirically produce deep insight across vast documents; sub-passage reasoning gives more specificity.
  - [14:32] Mike Conover's book report analogy shows why summarizing large contexts loses depth: a synopsis of each chapter retains more specificity.
  - [18:20] Mike Conover asserts: 'Language models are not tools for creating new knowledge; they're tools for helping me create new knowledge.'
  - [20:00] Brightwave checks factuality by comparing outputs from multiple models, using entailment checks, and letting users double-check surprising findings with extra compute.
- **[22:03] Feedback**
  - [22:44] In finance, where there's no ground truth, Brightwave captures user feedback through revealed preferences: the queries users ask next reveal their beliefs.
- **[25:10] Evals**
  - [25:38] Brightwave combines LLM supervision with human annotator benchmarks; for nuanced outputs like transcendent poetry, only human evaluation suffices.
  - [29:31] Most companies keep custom training datasets private as competitive advantage, unlike Databricks which open-sourced Dolly 15K.
- **[30:40] Retrieval**
  - [31:48] Temporal-aware retrieval at Brightwave prevents convincingly wrong answers by aligning search results with the time intent of a user's query.
  - [34:00] Brightwave's privacy architecture, informed by co-founder Brandon Qatar's experience operating a regulated exchange, ensures sensitive data control from the start.
  - [35:36] Brightwave uses composable, context-aware prompts that instruct the model to treat data sources differently based on their provenance and importance.
  - [38:01] Brightwave extracts structured knowledge graphs from documents in a single pass, creating pivotable data artifacts that enhance generative reasoning.
- **[38:16] Knowledge Graphs**
  - [40:16] Mike Conover views fine-tuning as differentiating model behaviors like stem cells, not as a way to imbue new facts; that is RAG's role.
- **[40:17] Fine-tuning**
  - [42:51] Mike Conover finds that high-quality reasoning from agentic systems does not require anthropomorphizing language models with cute prompts.
- **[43:14] Philosophy**
  - [44:35] Classical machine learning and deep learning provide statistical guarantees that LLMs lack, so Brightwave uses them where guarantees are needed.
  - [45:42] Brightwave deliberately avoids Excel spreadsheets because formula bugs are not fault-tolerant; instead, it offers a dialogue-based investigation modality.
  - [49:20] Mike Conover predicts AI will increase financial analysis sophistication just as VisiCalc commoditized arithmetic, though model production may stay human.
- **[50:28] AI Funds**
  - [50:28] Mike Conover says many hedge funds with systematic trading groups already operate like large machine learning teams, suggesting AI hedge funds already exist.
  - [55:12] Brightwave surfaced long-duration energy storage as a second-order derivative bet on AI, bridging renewables and data center power needs.
- **[56:49] Open Source**
  - [1:00:33] Mike Conover argues the next AI innovation wave will come from creating instruction tuning and fine-tuning data that elicits specific new model behaviors.
  - [1:02:55] Alessio asserts the half-life of a dataset is much longer than that of a model; the Pile dataset outlasts year-old models.
- **[1:03:43] Hiring**
  - [1:03:54] Brightwave is hiring across AI engineering, classical ML, systems engineering, frontend, and design, fitting jobs to people.

## Speakers

- **Alessio** (host)
- **Mike Conover** (guest)

## Topics

Industry, Startups, Reasoning

## Mentioned

Bloomberg (company), Brightwave (company), Databricks (company), Decibel (company), LinkedIn (company), Moonfire Ventures (company), Point72 (company), Scale (company), Skip Flag (company), Snorkel (company), Workday (company), Bard (product), DALL-E (product), Excel (product), Llama (product), Seeking Alpha (product)

## Transcript

### Intro

**Alessio** [0:00]
Hey everyone, welcome to the Later in Space podcast. This is Alessio, partner and CTO in residence at Decibel Partners, and I have no co-host today, as you can see. Um, Swyx is in Vienna at ICLR, having fun in, in Europe.

And, uh, we're in the brand-new studio. Um, as you might see if you're on YouTube, there's still no sound panels on the wall. Mike tried really hard to put them up, but the, the glue, the glue, uh, i- is a little too old for that.

Um, so if you s- hear any echo or anything like that, um, sorry. But- ... we're doing, we're doing the best that we can. Um, and today we have our first repeat guest, Mike Conover.

**Mike Conover** [0:35]
Hey.

**Alessio** [0:36]
Um, welcome Mike, who is now the founder of Brightwave.

**Mike Conover** [0:38]
That's right.

**Alessio** [0:38]
Not, not at Databricks anymore.

**Mike Conover** [0:40]
That's right, yeah.

**Alessio** [0:42]
Um-

**Mike Conover** [0:42]
Pleased to be back.

**Alessio** [0:43]
Yeah, no, no. The... Our last episode was one of the fan favorites, and I think this will be, uh, just as good. So for, for those that have not listened to the first episode, which might be many because the podcast has grown a lot since then, thanks to people like Mike who have interesting conversations on it, um, you spent, uh, a bunch of years doing ML at some of the best companies on the internet.

Um, things like Workday, you know, Skip Flag, LinkedIn, um, most recently at Databricks, where you were leading the, the open source large language models, um, team working on DALL-E.

**Mike Conover** [1:15]
Mm-hmm.

**Alessio** [1:16]
Um, and now you're doing Brightwave, which is in the financial services, uh, space, but this is not something new, you know. I think when you and I first talked, um, about Brightwave, I was like, "Why is this guy doing a financial services company?"

**Mike Conover** [1:28]
Sure.

**Alessio** [1:29]
And then you look at your background and you were doing papers on Nature about, uh, on the Nature magazine about-

**Mike Conover** [1:35]
Yeah

**Alessio** [1:35]
... uh, LinkedIn data predicting S&P 500 stock movement-

**Mike Conover** [1:38]
Yep

**Alessio** [1:38]
... like many, many years ago. So what's, uh, what, what's kind of like the, uh, some of the tying elements in your background that maybe people are overlooking that brought you to do this?

**Mike Conover** [1:48]
Yeah, sure.

**Alessio** [1:49]
Um, yeah.

**Mike Conover** [1:50]
So I would say

my... So my PhD research was funded by DARPA, and it ha- we had access to the Twitter data set early in the, early in the natural history of the availability of that data set, and it was focused on the large scale structure of, uh, propaganda and misinformation campaigns.

And LinkedIn, we had planet scale descriptions of the structure of the global economy. And so primarily my work was, was homepage newsfeed relevant, so when you go to linkedin.com, you would see updates from one of our machine learning models.

Um, but additionally, I was a research liaison as part of the economic graph challenge and, um, had this Nature Communications paper where we demonstrated that 500 million jobs transitions can be hierarchically clustered as a network of labor flows and are predictive of next quarter S&P 500 market cap changes.

And at Workday, I was director of financials machine learning. Um, and you start to see how organizations are organisms. And I, I think of the way that like, uh, an accountant or the market, uh, encodes information in databases similar to how social insects, for example, organize their work and make collective decisions about where to allocate resources or time and attention, and, um, that especially with the, the work on Twitter, we would see, um, network structures relating to polarization emerge organically out of the interactions of many individual components.

And so, like much of my professional work has been focused on this idea that our lives are governed by systems that we're unable to see from our locally constrained perspective. And that we... When we, when humans interact with technology, they create digital trace data that allows us to observe the structure of those systems as though through a microscope or a telescope.

And particularly with regards finance, I think this is the ultimate, the markets are the ultimate manifestation and record of that collective decision-making process that humans engage in.

**Alessio** [3:56]
Just to start going off script right away.

**Mike Conover** [3:58]
Sure.

**Alessio** [3:58]
Uh, how do you think about some of these interactions creating the polarization and how that reflects in the language models today because they're trained on this data? Like, do you think the models pick up on these things-

**Mike Conover** [4:09]
Absolutely

**Alessio** [4:09]
... on their own as well?

**Mike Conover** [4:09]
Absolutely. Yeah, I, I think they are a compression of the world as it existed-

**Alessio** [4:13]
Mm-hmm

**Mike Conover** [4:13]
... at the point in time when they were pre-trained. And so I, I think absolutely the... And you see this in Word2vec too.

**Alessio** [4:20]
Mm-hmm.

**Mike Conover** [4:20]
I mean, just the, the semantics of how we, uh, think about gender as rela- relates to professions are, um, encoded in the structure of these models and like language models I think are, you know, much more sort of, uh, complete representation of, of human sort of beliefs.

**Alessio** [4:42]
Mm-hmm.

**Mike Conover** [4:42]
Yeah.

**Alessio** [4:43]
That's awesome. So we left you at Databricks last time.

**Mike Conover** [4:45]
Yeah.

**Alessio** [4:45]
You were building DALL-E. Tell us a bit more about Brightwave. This is the first time you're really talking about it publicly, so-

**Mike Conover** [4:51]
Yeah. Yeah. It's a, it's a pleasure. I mean, so we, we've raised a $6 million seed round, uh, including participate led by Decibel, uh-

### Product

**Alessio** [4:59]
Woo

**Mike Conover** [4:59]
... who we love working with, and including participation from Point72, one of the largest hedge funds in the world, and, and Moonfire Ventures. And, um, we are, uh, focused on... Like if you think of the job of an active asset manager, the, the work to be done is to understand something about the market that nobody else has seen in order to identify a mispriced asset.

And it's our view that that is not a task that is well-suited to human intellectual attention span. And so much as I was gesturing towards the ability of these models to perceive more than a human is able to, um, we think that there's a, a unique, historically unique opportunity to expand individual's ability to reason about the structure of the economy and the markets.

Um, it's not clear that you get superhuman reasoning capabilities from human level demonstrations of skill, and by that I mean the pre-training corpus, but then additionally the fine-tuning corpuses.

**Alessio** [5:51]
Mm-hmm.

**Mike Conover** [5:51]
I think you largely mimic the demonstrations that are present, uh, at, at model training time. Um, but from a working memory standpoint, these models outclass Humans and their ability to, to reason about these systems.

**Alessio** [6:05]
Yep. And you started Brightwave with Brandon, um-

**Mike Conover** [6:08]
Yeah. Yeah

**Alessio** [6:08]
... what's the, what's the story? You two worked together at Workday, uh, but he also has a really relevant background.

**Mike Conover** [6:14]
Yes. So Brandon Qatar is my co-founder, uh, the CTO, and he's a, a very special human. So he's, has a deep background in finance. So he was the former CTO of a federally regulated derivatives exchange. Um, but his first deep learning patent was filed in 2018.

And so he ha- he spans worlds. He has experience building mission-critical infrastructure in highly regulated environments for finance use cases, um, but also, uh, was, was very early to the deep learning, uh, party and, and understand. He led, at Workday, um, was the tech lead for semantic search over hundreds of millions of resumes and, um, job listings.

Um, and so just has been working with, uh, information retrieval f- you know, and neural information retrieval methods, um, for a very long time. And so, uh, was a, an exceptional person, and I'm, I'm glad to count him among, um, the people that we're doing this with.

**Alessio** [7:14]
Yeah. And a great fisherman, so-

**Mike Conover** [7:16]
Yeah.

**Alessio** [7:17]
Um, that's good.

**Mike Conover** [7:17]
Very talented.

**Alessio** [7:18]
That's always important.

**Mike Conover** [7:19]
Very talented. Very enthusiastic fisherman. Yeah, exactly.

**Alessio** [7:21]
Um, and then you have a bunch of amazing engineers. Then you have folks like JP, who used to work at Goldman Sachs.

**Mike Conover** [7:27]
Yeah.

**Alessio** [7:27]
How should people think about team building in this more vertical domain? Obviously, you come from a, a deep ML background-

**Mike Conover** [7:33]
Yeah

**Alessio** [7:33]
... but you also need some of the industry sides. What's, what's the right balance?

**Mike Conover** [7:36]
Yeah, I mean, this, it's... So I think one of the things that's interesting about building verticalized solutions in AI in 2024 is that historically, um,

you need the AI capability. You need to understand both how the models behave and then how to get them to interact with other kinds of, uh, machine learning subsystems that together perform the work-

**Alessio** [8:02]
Mm-hmm

**Mike Conover** [8:02]
... of a system that can reason on behalf of a human. There are also material sy- systems engineering problems in there. So I saw, I forget who this is attributed to, but a, a tweet that, uh, sort of made reference to all the AI comp- all of the traditional software companies are trying to hire AI talent-

**Alessio** [8:17]
Mm-hmm

**Mike Conover** [8:17]
... and all the AI companies are trying to hire systems engineers, and that is 100% the case. Getting these systems to behave, um, in a predictable and repeatable and observable way is, um, equally challenging to a lot of the methodological challenges.

But then you bring in, whether it's law or medicine or public policy or in our case, finance, um, I think a lot of the most valuable, like Grammarly is a good example of a company-

**Alessio** [8:45]
Mm

**Mike Conover** [8:45]
... that has generative work product that is, um, valuable by most humans. Whereas in finance, the q- the character of the insight, the depth of insight, and the non-consensusness of the insight, um, really requires, uh, fairly deep-

**Alessio** [9:03]
Mm-hmm

**Mike Conover** [9:03]
... domain expertise. And like even operating in exchange, uh, I mean, when we went to, to raise a round, a lot of people said, "Why, why don't you start a hedge fund?"

**Alessio** [9:10]
Mm-hmm.

**Mike Conover** [9:11]
Um, and it's like that is a totally sep- there are many, many separate-

**Alessio** [9:15]
Right

**Mike Conover** [9:15]
... skills that are unrelated to, um, AI in, in that problem. And so like we've, um, we've brought toget- brought into the fold domain experts, um, in finance who, who can help us, um, evaluate the character and sort of steer the system.

**Alessio** [9:30]
Mm-hmm. Yep.

**Mike Conover** [9:31]
Yeah.

**Alessio** [9:31]
So that's the team. What does the system actually do? What's the, the Brightwave product?

**Mike Conover** [9:36]
Yeah, I mean, it's, it's in- it does many, many things. Um, but it will, it, it acts as a partner in th- in th- a partner in thought to finance professionals. So, um, you can ask Brightwave, um, a question like, "How is Nvidia's position in the GPU market impacted by rare earth metal shortages?"

And it will identify, um, as thematic contributors to an investment decision, um, or, or developing your thesis that, um, in response to export controls on A100 cards, uh, China has put in place licensors on the transfer of germanium and gallium, which are not rare earth metals, but they're semiconductor production inputs, and has expanded its control of African and South American mining operations.

And so we see, uh, if, if you think about, you know, we, we have a, a $20 billion crossover hedge fund. Uh, their, their equities team uses this tool to go deep on a thesis. So I was describing this like, you know, multi-

**Alessio** [10:34]
Yeah

**Mike Conover** [10:34]
... multiple steps into the value chain or supply chain for, for companies. Um, we see, uh, you know, wealth management professionals using Brightwave to get up to speed extremely quickly as they step into nine conversations tomorrow, uh, with clients who are assessing, like sh- "Do you know something that I don't?

Can I trust you to, to be a steward of my, um, my financial wellbeing?" Um, we see investor relations teams using Brightwave to, um... You just think about the universe of coverage that a person working in finance needs to be aware of.

**Alessio** [11:10]
Mm-hmm.

**Mike Conover** [11:10]
The ability to rip through filings and transcripts and have a very comprehensive view of the market. Um, it's, it's extremely rate limited by how quickly a person is able to read, and not just read, but like, um, solve the blank pa- page problem-

**Alessio** [11:26]
Mm-hmm

**Mike Conover** [11:26]
... knowing what to say about a fact or a finding.

**Alessio** [11:29]
What else can you share about customers that you're working with?

**Mike Conover** [11:33]
Yeah. So we, we have seen, uh, traction that far exceeded our expectations, uh, from the market. This, the...

You sit somebody down with a system that can take any question and generate tight, actionable financial analysis on that subject, and the product kinda sells itself.

**Alessio** [11:52]
Mm-hmm.

**Mike Conover** [11:53]
And so we see, um, many, many different funds, firms, and strategies that are making use of Brightwave. So you've got, um, a- 10-person owner-operated registered investment advisor, so the classical wealth manager, you know, $500 million in AUM. Um, we have crossover hedge funds that have tens and tens of billions of dollars in assets under management.

Very different use case, so that's more investment research. Um, whereas a wealth manager's gonna use this to step into client interactions-

**Alessio** [12:19]
Mm-hmm

**Mike Conover** [12:19]
... just exceptionally well prepared. We see investor relation te- or investor relations teams. Um, we see corporate strategy types that are, um, needing to understand very quickly, um, new markets, new themes. Um, and just the ability to very quickly develop a view on any investment theme or sort of strategic, um, consideration is, is broadly applicable to many, many different kinds of personas.

### Context

**Alessio** [12:50]
Yep. Um, yeah, I, I, I can attest to the product selling itself given that I'm a user.

**Mike Conover** [12:55]
Yeah.

**Alessio** [12:56]
Um, let, let's jump into some of the technical challenges-

**Mike Conover** [12:59]
Yeah

**Alessio** [12:59]
... and, and work behind it because, um, there are a lot of things. So as I mentioned, you were on the podcast about a year ago.

**Mike Conover** [13:05]
Yep.

**Alessio** [13:06]
You had released DALL-E from Databricks-

**Mike Conover** [13:08]
Mm-hmm

**Alessio** [13:08]
... which was one of the, the first open source LLMs.

**Mike Conover** [13:10]
Yep.

**Alessio** [13:11]
Uh, DALL-E had a whopping 1,024 tokens, uh, of context size. And today, you know, I think 1,000 tokens, a model would be unusable-

**Mike Conover** [13:21]
You, you lose that bet out

**Alessio** [13:22]
... basically. Yeah, exactly. Um, how, how did you think about the evolution of context sizes as you built the company? Um, and where we are today, what are things that people get wrong? Any, any commentary there?

**Mike Conover** [13:34]
Sure. So I-- we, we take a very, very much take a systems of systems approach. So I think that this is...

When I started the company, I think I had more faith in the ability of large context windows to generally solve problems relating to synthesis. And actually, like, if you think about the attention mechanism and, um, the way that it computes similarities between tokens at a distance, um, I on some level believed that as you would scale that up, you would have the ability to simultaneously perceive-

**Alessio** [14:07]
Mm-hmm

**Mike Conover** [14:07]
... um, and draw conclusions across vast disparate bodies of content. And I think, um, that is, does not empirically seem to be the case. So when, when, for example, you... and this is something anybody can try, um, take a very long document, like needle in a haystack, I think sure, we can do information retrieval on like specific, uh, fact-finding activities pretty-

**Alessio** [14:32]
Mm

**Mike Conover** [14:32]
... pretty easily. But if you-- I kind of think about it like summarizing. If you write a book report on an entire book versus a synopsis of each individual chapter-

**Alessio** [14:42]
Mm-hmm

**Mike Conover** [14:43]
... there, there is a characteristic output length for these models. Let's say it's about 1,200 tokens. Like, it is very difficult to get any of the commercial LLMs or Llama to write 5,000 tokens. And you think about it as like, what is the conditional probability that I generate an end token?

It just gets higher-

**Alessio** [15:02]
Yeah

**Mike Conover** [15:02]
... the more tokens are in the context window prior to that, um-

**Alessio** [15:06]
Mm-hmm

**Mike Conover** [15:07]
... sort of next inference step. And so, um, if I have 1,000 words in which to say something, the level of specificity and the level of depth when I am assessing a very large body of content-

**Alessio** [15:23]
Mm-hmm

**Mike Conover** [15:23]
... is going to necessarily be less than if I am saying something specific about a sub passage. And so we-- and if you think about, um, drawing a parallel to consumer internet companies like LinkedIn or Facebook, um, there are many different subsystems with it.

So l- let's take the Facebook example. Facebook, uh, almost certainly has, uh, I mean, you can see this in your profile, your inferred interests.

**Alessio** [15:49]
Mm-hmm.

**Mike Conover** [15:50]
What are the things that it believes that you care about? Those assessments almost certainly feed into the feed relevance algorithms that would judge what you are... You know, am I, am I gonna see you snow- or am I gonna show you snowboarding content?

I'm gonna show you, um, aviation content.

**Alessio** [16:06]
Mm-hmm.

**Mike Conover** [16:06]
It, it's the outputs of one machine learning system feeding into another machine learning system. And I think with modern RAG and sort of agent-based reasoning, it is really about creating subsystems that do specific tasks well. And I think the problem of, um, deciding how to decompose large documents into more kind of atomic reasoning units is-

**Alessio** [16:30]
Mm-hmm

**Mike Conover** [16:31]
... is still very important. Now, it, there, it's an open question whether you have, um, in the-- whether that is a model that is addressable by pre-training or, um, instruction tuning. Like, can you, uh, have synthesis-oriented, uh, demonstrations in-

**Alessio** [16:50]
Mm-hmm

**Mike Conover** [16:51]
... the, at, at training time? And, and now this problem is more robustly solved, um, because synthesis is quite different from, uh, complete the next word in The Great Gatsby.

**Alessio** [17:01]
Yep.

**Mike Conover** [17:02]
Um, but it, I think empirically is not the case that you can just throw all of the SEC filings in, you know, a million token context window and get-

**Alessio** [17:12]
Mm-hmm

**Mike Conover** [17:12]
... um, deep insight that is useful out the other end.

**Alessio** [17:17]
Yeah. Yeah. And I think that's the main difference about what you're doing. It's not about summarizing. It's about coming up with different ideas and kinda like thought threads to-

**Mike Conover** [17:27]
Yes. Precisely

**Alessio** [17:28]
... to follow.

**Mike Conover** [17:28]
Yeah. And I think this specifically, like helping a person know, um... You know, if I think that GLP-1s are gonna blow up the diet industry, um, identifying and putting in context a negative result from a human clinical trial that's-- or for example, that adherence rates to Ozempic after a year are just 35%, what are the implications of this?

**Alessio** [17:52]
Mm-hmm.

**Mike Conover** [17:53]
So it, there's an information retrieval component, and then there's a not just presenting me with a summary of like, "Here's, here are the facts," but like, what does this entail?

**Alessio** [18:02]
Yep.

**Mike Conover** [18:02]
And how does this fit into my worldview, my fund strategy? Um-

Broadly, I think that, you know, so I mean, this, this idea I think is very, um, eloquently puts it, which is-- And this is, this is not my insight, but that language models-- And help, help me know who said this.

You may, you may be familiar. But language models are not tools for creating new knowledge. They're help- they're tools for helping me create new knowledge.

**Alessio** [18:27]
Mm.

**Mike Conover** [18:27]
Like they themselves do not do that work. Um, I think that-that's the f- presently the right way to think about it.

**Alessio** [18:33]
Mm-hmm. Yeah. I, I read a tweet about needle in the haystack actually being harmful to some of this work.

**Mike Conover** [18:41]
Sure.

**Alessio** [18:41]
Because now the model is like too focused on recalling everything-

**Mike Conover** [18:44]
Yeah. Yeah, yeah

**Alessio** [18:44]
... versus saying, "Oh, that doesn't matter," or like ignoring certain things.

**Mike Conover** [18:47]
Sure, sure, sure.

**Alessio** [18:47]
If you think about a S1 filing, like-

**Mike Conover** [18:49]
Yeah

**Alessio** [18:50]
... eighty-five percent is like boilerplate. It's like-

**Mike Conover** [18:53]
Yes

**Alessio** [18:53]
... you know, previous performance doesn't guarantee future performance.

**Mike Conover** [18:57]
Yeah.

**Alessio** [18:57]
Like the company might not be able to turn a profit in the future, blah, blah, blah.

**Mike Conover** [19:00]
Yeah. Yeah, yeah.

**Alessio** [19:00]
All these things, they always come up again.

**Mike Conover** [19:02]
COVID and currency fluctuations.

**Alessio** [19:03]
Yeah, yeah, yeah.

**Mike Conover** [19:03]
Yeah, sure.

**Alessio** [19:03]
Yada, yada, yada.

**Mike Conover** [19:04]
Yes.

**Alessio** [19:04]
We have a large workforce a- a- and all of that. Um, have you had to do any work at the model level to kinda like make it okay to forget these things? So like have you found that like kinda like making it a smaller problem, then cutting-- putting them back together kinda solves for that?

Um-

**Mike Conover** [19:21]
Absol- absolutely, and I think this is where, um, having domain expertise around the structure of these documents. So if you look at the different chun-chunking strategies that you can employ to, um, understand like what is the intent of this clause or phrase, and then really, um, be selective at retrieval time-

**Alessio** [19:41]
Mm-hmm

**Mike Conover** [19:41]
... in order to get the information that is most relevant to a user query based on the semantics of that unique document, and I, I think it is not, um-- it's certainly not just a sliding window over that, that-

**Alessio** [19:53]
Yeah

**Mike Conover** [19:54]
... that corpus.

**Alessio** [19:55]
Yeah. Um, and then the flip side of it is obviously factuality.

**Mike Conover** [19:59]
Yeah.

**Alessio** [19:59]
You don't-

**Mike Conover** [20:00]
Yep

**Alessio** [20:00]
... want to forget things that that were there. Um, how do you tackle that?

**Mike Conover** [20:04]
Yeah. I mean, of, of course, that's-- it's a very deep problem, and I think, you know, I'll be s- a little circumspect about the specific-

**Alessio** [20:11]
Mm-hmm

**Mike Conover** [20:11]
... kinds of methods we use. But, um, this sort of multiple passes over the material and saying, "How convicted are you that what you're saying is in fact true?" And you-- I mean, you can take generations from multiple different models and compare and contrast and say like, "Do these both reach the same conclusion?"

Um, we-- You can treat it like a voting problem. Um, you-- we train our own models to assess, um, you know-- You can think of this like entailment. Like is, is this supported by the underlying primary sources? And I think that you have, um, methodological approaches to this problem-

**Alessio** [20:52]
Mm-hmm

**Mike Conover** [20:52]
... but then you also have product affordances. So like there's a, a great blog post on Bard from the Bard team describing, um, an abil- it was sort of a design-led, um, product innovation that allows you to ask the model to double-check the work.

**Alessio** [21:07]
Yeah.

**Mike Conover** [21:07]
So if you have a surprising finding, we can let the user, uh, discretionarily spend more compute to double-check the work. And I think that you want to build product experiences that are fault-tolerant, and it is unclear that you ever-- That the difference between like hallucination and creativity is-

**Alessio** [21:27]
Right

**Mike Conover** [21:27]
... fuzzy.

**Alessio** [21:27]
Yeah, yeah.

**Mike Conover** [21:28]
And so do you, do you ever get language models with next token prediction as the loss function that are, um, guaranteed to not contain factual misstatements? That is not clear. Now, maybe being able to invoke code interpreter-

**Alessio** [21:44]
Mm-hmm

**Mike Conover** [21:45]
... you do like code, uh, generation-

**Alessio** [21:47]
Yeah

**Mike Conover** [21:47]
... and then execution in a secure way, um, helps to solve some of these problems, especially for quantitative reasoning. That, that may be the case. But, um, for right now, I think you need to have product affordances that allow you to live with the reality-

### Feedback

**Alessio** [22:04]
Mm-hmm

**Mike Conover** [22:04]
... that these things are, um,

are fallible.

**Alessio** [22:08]
Yep. Yeah. We did, um, our LHF 201 episode just talking about-

**Mike Conover** [22:13]
Yeah

**Alessio** [22:13]
... different methods and, and whatnot. How do you think about something like this where it's maybe unclear in the short term, even if the product is right, you know? It might give, it might give an insight that might be right, but it might not prove until later.

So it's kinda hard for the user to say, "That's wrong."

**Mike Conover** [22:30]
Mm.

**Alessio** [22:31]
Because actually it might be-- Like you think it's wrong.

**Mike Conover** [22:33]
Sure, sure.

**Alessio** [22:33]
You know, like an investment, that's kinda what it comes down to.

**Mike Conover** [22:35]
Yeah.

**Alessio** [22:35]
You know, some people are wrong.

**Mike Conover** [22:36]
Yep.

**Alessio** [22:36]
Some people are right. Um, how, how do you think about some of the product features that you need in something like this to bring user feedback into the mix-

**Mike Conover** [22:44]
Yeah

**Alessio** [22:44]
... and maybe how you approach it today and how you think about it long term?

**Mike Conover** [22:48]
Yeah. Well, I mean, I think that to your point about the model may h- the model may make a statement which is not actually verifiable. It's like this, this may be the case. I think that is where

the reason we think of this as a partner in thought is that humans are always going to have access to information that has not, not been digitized.

**Alessio** [23:08]
Mm-hmm.

**Mike Conover** [23:08]
And so in finance, you see that especially so with regards to expert call networks, um, the sort of like uns- like your in- the unstated investment theses that a portfolio ma-manager may have. Like we just don't do biotech.

**Alessio** [23:24]
Yep.

**Mike Conover** [23:24]
Um, or we do not believe that, um-- We think that Eli- Eli Lilly is actually very exposed because of how unpleasant it is to take exemp- Ozempic, right?

**Alessio** [23:35]
Mm-hmm.

**Mike Conover** [23:35]
Those, those are things that, um, are beliefs about the world, but that may not be, um, like falsifiable right now. And so I think you have to-- You can again take pl- pages from consumer web playbook and think about, um, personalization.

**Alessio** [23:53]
Mm-hmm.

**Mike Conover** [23:54]
So it is getting a person to articulate everything that they believe is not A realistic task.

**Alessio** [23:59]
Yep.

**Mike Conover** [24:00]
Um, Netflix doesn't ask you to describe-

**Alessio** [24:03]
Mm-hmm

**Mike Conover** [24:03]
... what kinds of movies you like, and they give you the option to vote, but nobody does this. Um, and so what I think you do is you observe people's revealed preferences.

**Alessio** [24:13]
Mm-hmm.

**Mike Conover** [24:13]
Like, what... So one of the capabilities that our system exposes is, um, given everything that Brightwave has read and assessed, and, like, the, the sort of synthesized financial analysis, what are the natural next questions that this, that a person investigating-

**Alessio** [24:29]
Mm-hmm

**Mike Conover** [24:29]
... this subject should ask? And you can think of this chain of thought and this deepening, um, kind of investigative process. And the

direction in which the user steers the attention of this system, uh, reveals information about what do they care about-

**Alessio** [24:47]
Mm-hmm

**Mike Conover** [24:47]
... what do they believe, what kinds of things, um, are important. And so, uh, at the individual level, but then also at the, the fund and firm level, you can develop, like, an implicit, um, representation-

**Alessio** [25:03]
Yeah

**Mike Conover** [25:03]
... of your beliefs about the world in a way that, um,

you just, you're never gonna get every- somebody to write everything down like that.

**Alessio** [25:10]
Mm-hmm.

### Evals

**Mike Conover** [25:10]
Yeah.

**Alessio** [25:11]
Yeah. Um, how does that tie into one of our other favorite topics, evals?

**Mike Conover** [25:16]
Yeah.

**Alessio** [25:16]
Uh, we had David Luan from Adaptin.

**Mike Conover** [25:18]
Yeah.

**Alessio** [25:18]
He mentioned they don't care about benchmarks because their customers-

**Mike Conover** [25:21]
Yeah

**Alessio** [25:21]
... don't work on benchmarks.

**Mike Conover** [25:22]
Yeah, yeah.

**Alessio** [25:22]
They work on cus- uh, you know, business results. Uh, how do you think about that for you and maybe, um, as you build a, a new company-

**Mike Conover** [25:30]
Yeah

**Alessio** [25:30]
... when is the time to, like, still focus on the benchmark versus when it's time to, like, move on to your own evaluation using maybe, yeah, uh, labelers or-

**Mike Conover** [25:37]
Yeah

**Alessio** [25:37]
... or whatnot?

**Mike Conover** [25:38]
So I mean, we, we use a fair bit of LLM supervision, uh, to evaluate multiple different subsystems. And I think that one of the reasons that... I mean, we, we, we pay human annotators to evaluate the quality of, of the generative outputs, and I think that that is always the, the reference standard.

But we frequently first turn to LLM supervision-

**Alessio** [25:59]
Mm-hmm

**Mike Conover** [25:59]
... as a way to, um, have, uh, whether it's at fine-tuning time or even for subsystems that are not generative, like what is the quality of the system? And I think we will generate a small corpus of high-quality domain expert annotations and then always compare that against how well is-

**Alessio** [26:21]
Mm-hmm

**Mike Conover** [26:21]
... either LLM supervision or even just a heuristic, right? Like w- a, a simple thing you can do. This is, this is a technique that we do not use, but as an example, um, do not generate any integers or any-

**Alessio** [26:33]
Mm-hmm

**Mike Conover** [26:33]
... numbers that are not present in the underlying source data.

**Alessio** [26:37]
Mm.

**Mike Conover** [26:37]
Right? You know-

**Alessio** [26:38]
Sure

**Mike Conover** [26:38]
... if you're doing RAG, you can just say you can't name numbers that are not pr-

**Alessio** [26:41]
Mm-hmm

**Mike Conover** [26:42]
... you know, it's very sort of heavy-handed. But you can take the annotations of a human evaluator and then compare that. I mean, Sn- Snorkel kind of takes a similar perspective.

**Alessio** [26:51]
Yeah.

**Mike Conover** [26:51]
Like multiple different weak, um,

sort of supervision data sets-

**Alessio** [26:57]
Mm-hmm

**Mike Conover** [26:57]
... can give you substantially more than any one of them does on their own. And so I think you want to compare the quality of any evaluation against human-generated, the, the sort of like benchmark. Uh, but at, at the end of the day, like eventually you...

Especially for things that are nuanced.

**Alessio** [27:14]
Right.

**Mike Conover** [27:14]
Like, is this transcendent poetry?

**Alessio** [27:17]
Mm-hmm.

**Mike Conover** [27:17]
There's just no way to multiple choice your way out of that.

**Alessio** [27:21]
Right. Yeah.

**Mike Conover** [27:21]
You know? And so, um, really where I think a lot of the flywheels for some of the large LLM companies are, are it's methodological obviously, but it's also just data generation.

**Alessio** [27:35]
Mm-hmm.

**Mike Conover** [27:35]
And you think about like, you know, for anybody who's done, um, crowdsource work, and this I think applies to the high-skilled human annotators as well. Like you look at the Google Search Quality Evaluator Guidelines, it's like a, a 90 or 120-page rubric-

**Alessio** [27:50]
Right

**Mike Conover** [27:50]
... describing like what is a high-quality-

**Alessio** [27:52]
Mm-hmm

**Mike Conover** [27:52]
... search result. And it's like very difficult to get, on a human level, people to re- reproducibly follow a rubric.

**Alessio** [28:00]
Mm-hmm.

**Mike Conover** [28:01]
And so what is your process for orchestrating that motion? Like how do you, um, articulate what is high quality insight? I think that's where a lot of the, um, work actually happens. Uh, and that it's sort of the last resort.

Like I- everything... Like ideally you wanna automate everything, but-

**Alessio** [28:22]
Mm-hmm

**Mike Conover** [28:22]
... ultimately the most interesting problems right now are those that are not especially automatable.

**Alessio** [28:28]
One thing you did at Databricks was the, um... Well, that, not that you did specifically, but the team there was like the DALI 15K-

**Mike Conover** [28:35]
Yeah, yeah

**Alessio** [28:35]
... dataset.

**Mike Conover** [28:36]
Yeah.

**Alessio** [28:36]
Um, you mentioned, um, people misvalue the, the, the, the value of this data. Why has no other company done anything similar with like creating this like employee-led, uh, dataset? You could imagine, you know-

**Mike Conover** [28:49]
Sure, yeah

**Alessio** [28:49]
... some of these like, uh, Goldman Sachs.

**Mike Conover** [28:51]
Yeah.

**Alessio** [28:51]
They got like thousands and thousands of people in there.

**Mike Conover** [28:53]
Yeah, yeah.

**Alessio** [28:53]
Um, obviously they have different privacy and whatnot requirements.

**Mike Conover** [28:56]
Sure.

**Alessio** [28:56]
Uh, do you think more companies should do it? Like, uh, do you think there's like a misunderstanding of how valuable that is or, uh, yeah.

**Mike Conover** [29:05]
So I, I think Databricks is a very special company and led by people who are, um, very sort of, uh, cour- courageous I guess-

**Alessio** [29:15]
Mm-hmm

**Mike Conover** [29:15]
... is one word for it. Just like, "Let's just ship it." And I think it's, um, it's unusual and it's also because, uh, I think in tr- like m- most companies will recogni- like if they go to the effort to produce something like that, they recognize that it is competitive advantage to have it-

**Alessio** [29:31]
Mm-hmm

**Mike Conover** [29:31]
... and to be the only-

**Alessio** [29:32]
Right

**Mike Conover** [29:32]
... company that has it. And I think Databricks is in unu- is in an unusual position in that they, um, benefit from more people having access to these kinds of-

**Alessio** [29:40]
Mm-hmm

**Mike Conover** [29:41]
... um, sources. But you also saw Scale.

**Alessio** [29:43]
Mm-hmm.

**Mike Conover** [29:43]
I guess they, they haven't released it.

**Alessio** [29:44]
Well, yeah.

**Mike Conover** [29:45]
But they have, yeah.

**Alessio** [29:45]
I'm sure they have it because they charge people a lot of money, so.

**Mike Conover** [29:48]
But they, they created that alternative to the, uh, GSMK 8K-

**Alessio** [29:53]
Mm-hmm

**Mike Conover** [29:53]
... I believe was, was the, is how that's said. And, um, I th- I guess they're, you know, they, they too are not releasing that.

**Alessio** [30:02]
Yeah.

**Mike Conover** [30:02]
Um.

**Alessio** [30:03]
It's, it's interesting because I talk to a lot of enterprises, and a lot of them are like, "Man, I spend so much money on scale." And I'm like, "Why don't you just do it?"

**Mike Conover** [30:12]
Yeah.

**Alessio** [30:12]
And they're like, "What?"

**Mike Conover** [30:13]
But this-- So I think this again, again gets to the, like, the human process orchestration.

**Alessio** [30:19]
Mm-hmm.

**Mike Conover** [30:19]
Like ha- like it's one thing to do, like, a single monolithic push to create a training data set like that or an evaluation corpus, but I think it's another to have a repeatable process. And a lot of that, I think realistically is, um, pretty unsexy, like people management-

**Alessio** [30:37]
Yeah

**Mike Conover** [30:37]
... work.

**Alessio** [30:38]
Yeah.

**Mike Conover** [30:38]
Yeah.

**Alessio** [30:38]
Um-

**Mike Conover** [30:38]
So that's probably a big part of it.

### Retrieval

**Alessio** [30:40]
We have these, uh, four wars of AI framework.

**Mike Conover** [30:43]
Yeah.

**Alessio** [30:43]
Uh, the data quality wa- war, we, we kinda touched on a little bit now. About RAG.

**Mike Conover** [30:48]
Yeah.

**Alessio** [30:48]
That's, like, the other battlefield.

**Mike Conover** [30:49]
Yeah.

**Alessio** [30:49]
Uh, RAG and context sizes and kinda like all these, all these different things.

**Mike Conover** [30:53]
Sure.

**Alessio** [30:53]
Uh, you work in a space that has, um, a couple different things. One, um, temporality of data is important because every quarter there's new data.

**Mike Conover** [31:02]
Yep.

**Alessio** [31:02]
And, like, the new data usually overrides the previous one.

**Mike Conover** [31:05]
Yep.

**Alessio** [31:05]
So you cannot just, like, do semantic search and hope-

**Mike Conover** [31:07]
Right, right, right

**Alessio** [31:07]
... that you get the latest one.

**Mike Conover** [31:09]
Yes.

**Alessio** [31:09]
Um, and then you have obviously very structure numbers thing that are-

**Mike Conover** [31:13]
Yep

**Alessio** [31:14]
... very important to the token level.

**Mike Conover** [31:15]
Yeah.

**Alessio** [31:15]
Like, you know, 50% gross margins and 30% gross margins are very different.

**Mike Conover** [31:20]
Yes.

**Alessio** [31:20]
But, you know, the tokenization's not that different.

**Mike Conover** [31:22]
Yep.

**Alessio** [31:23]
Um, any thoughts on, like, how to build a system to, to handle all of that? As much as you can share, of course.

**Mike Conover** [31:28]
Yeah. Ab- absolutely. So I think this, again, rather than having open-ended retrieval, open-ended reasoning, you-- our approach is to decompose the problem into multiple different subsystems that have specific goals. And so y- I mean, temporality is a great example.

Um, the-- When you think about time, I mean, just look at all of the libraries for managing calendars.

**Alessio** [31:55]
Mm-hmm.

**Mike Conover** [31:56]
Time is kind of at the intersection of language and math, and this is one of the places where without, um, without taking specific technical m-measures to ensure that you get high quality narrative overlays of statistics that are changing over time and have a description of how a P/E multiple is-

**Alessio** [32:17]
Mm-hmm

**Mike Conover** [32:17]
... increasing or decreasing and a RAG, like a retrieval system that is aware of the time, sort of the time intent of the user query, right? If I'm asking something about breaking news, like, that's going to be very different than if I'm looking for a, a thematic-

**Alessio** [32:34]
Yep

**Mike Conover** [32:34]
... account of the past, you know, 18 months in Fed interest rate policy. Um, you have to have retrieval systems that are, to your point, like, if I just look for something that is a nearest neighbor without any of that, uh, temporal or other qualitative metadata overlay, you're just gonna get a, a kind of a bag of facts.

**Alessio** [32:57]
Mm-hmm.

**Mike Conover** [32:58]
And that, that is, like, explicitly not helpful, um, because the, the worst failure state for these systems is that they are, um, wrong in a convincing way.

**Alessio** [33:08]
Mm-hmm.

**Mike Conover** [33:09]
And so I think, at least presently, you have to have, um, subsystems that are aware of the semantics of the documents or a-aware of the semantics of the intent behind the question, um, and then have multiple evaluat-- We, we have multiple evaluation steps.

**Alessio** [33:26]
Mm-hmm.

**Mike Conover** [33:26]
Once you have, um, the generated outputs, we, we assess it multiple different ways to know, um, is, is this a, a factual statement given-

**Alessio** [33:35]
Yeah

**Mike Conover** [33:35]
... the, the sort of content that's been retrieved.

**Alessio** [33:37]
Mm-hmm. Yep. And what about-- I think people think of financial services, they think of privacy, confidentiality. Um-

**Mike Conover** [33:46]
Yeah

**Alessio** [33:46]
... w-what's kinda like customers' interests in that as far as, like, sharing documents and, like, how much of a deal breaker is that if you don't have them? Um, I don't know if you wanna share any on, uh, any about that and how you think about architecting the product.

**Mike Conover** [34:00]
Yeah. So one of the things that gives our customers a high degree of confidence, um, is the fact that Brandon operated a federally regulated derivatives exchange. Um, that experience in these highly regulated environments-- I mean, I additionally, at, at Workday, I worked with the financials product and, you know, without going into specifics, like, um, it's exceptionally sensitive data.

**Alessio** [34:29]
Mm-hmm.

**Mike Conover** [34:29]
Um, and you have multiple tenants, and it's, uh, just important that you take the right approach to being a steward of that material. And so from the start, we've built in a way that anticipates, um, the need for controls on how that data is managed and, and who has access to it and how it is, um, treated-

**Alessio** [34:49]
Mm-hmm

**Mike Conover** [34:49]
... throughout the life cycle. And so that for, um, our customer base, where frequently the most interesting and alpha-generating material is not publicly available-

**Alessio** [35:01]
Mm-hmm

**Mike Conover** [35:01]
... um, has given them a, a great degree of confidence in sharing, um, some of this, the most sensitive and interesting material, um, with systems that are able to combine it with content that is, um, either publicly or semi-publicly available-

**Alessio** [35:17]
Mm-hmm

**Mike Conover** [35:18]
... um, to create non-consensus insight into some of the most interesting and challenging problems in finance.

**Alessio** [35:24]
Yeah. We always say RAGs are recommendation systems for LLMs. Uh, how, how do you think about that when you have private versus public data, where sometimes you have public data, so that's one thing.

**Mike Conover** [35:34]
Mm.

**Alessio** [35:34]
But then the private is like, well, actually, you know, we got this, like, inside model, like, with this inside scoop that we gotta, like, figure out.

**Mike Conover** [35:41]
Yeah, yeah.

**Alessio** [35:42]
Um, how do you think in the RAG system a-about a value of-

**Mike Conover** [35:46]
Yeah

**Alessio** [35:46]
... these different documents? You know, I know a lot of it is secret sauce, but, uh-

**Mike Conover** [35:49]
No, no, it's, it's fine. I mean, I think that there is, um...

So I, I will gesture towards this by way of saying context where prompting.

**Alessio** [36:00]
Mm-hmm.

**Mike Conover** [36:00]
So you can have prompts that are composable and that have different, um- Sort of command units that like-

**Alessio** [36:11]
Mm-hmm

**Mike Conover** [36:11]
... may or may not be present based on the semantics of the content that is being populated into the, the RAG context window. And so that, that's something we make great use of, which is, um, where, where is this being retrieved from?

**Alessio** [36:24]
Mm-hmm.

**Mike Conover** [36:24]
What does it represent? And what should be in the instruction set in order to treat and respect the underlying contents, not just as like, here's a bunch of text-

**Alessio** [36:35]
Mm-hmm

**Mike Conover** [36:36]
... like, you figure it out.

**Alessio** [36:37]
Yep.

**Mike Conover** [36:37]
But, um, this is, this is important in the following way, or, you know, these, this aspect of the SEC filings are just categorically uninteresting.

**Alessio** [36:46]
Mm-hmm.

**Mike Conover** [36:46]
Or this, um, this is sell side analysis from a favored source.

**Alessio** [36:52]
Right.

**Mike Conover** [36:52]
And so it's that, um, creating it much like you have with the, the qualitative- the, the problem of organizing the work of humans, you have the problem of organizing the work of all-

**Alessio** [37:03]
Mm-hmm

**Mike Conover** [37:03]
... of these different AI subsystems and getting them to, um, propagate what they know through-

**Alessio** [37:11]
Mm-hmm

**Mike Conover** [37:11]
... the rest of the stack so that, you know, if you have multiple, you know, seven, 10 sequence, uh, inference calls, that all of the relevant metadata is propagated through that system and that you are, um, aware of where did this come from?

How convicted am I that it is-

**Alessio** [37:28]
Mm-hmm

**Mike Conover** [37:28]
... a source that should be trusted? I mean, you see this also just, um, in analysis, right? So different, like Seeking Alpha-

**Alessio** [37:36]
Right

**Mike Conover** [37:36]
... is a good example of just a lot of people with opinions.

**Alessio** [37:40]
Mm-hmm.

**Mike Conover** [37:40]
And some of them are great, some of them are really mid. And how do you, um, how do you build a system that is aware of, uh, the user's preferences for different sources?

**Alessio** [37:52]
Mm-hmm.

**Mike Conover** [37:53]
Um, I think it, this is all, uh, related to how you... And we talked about systems engineering. It's all-

**Alessio** [37:59]
Yeah

**Mike Conover** [37:59]
... related to how you actually build the systems.

**Alessio** [38:01]
Mm-hmm. And then just to kind of wrap on the RAG side, how should people think about knowledge graphs and, um, kinda like extraction-

**Mike Conover** [38:10]
Yeah

**Alessio** [38:10]
... from documents versus just, like, semantic search over the documents?

**Mike Conover** [38:13]
Yeah, so this knowledge graph extraction is an area where we're making a pretty substantial investment. And so this--

### Knowledge Graphs

**Mike Conover** [38:21]
I think that it is underappreciated how powerful, like there's the generative capabilities of language models, but there's also the ability to, um, program them to function as arbitrary machine learning systems-

**Alessio** [38:35]
Mm-hmm

**Mike Conover** [38:35]
... basically for marginally zero cost. And so the ability to extract structured information from huge, like, sort of unfathomably large bodies of content in a way that is single pass.

**Alessio** [38:51]
Mm-hmm.

**Mike Conover** [38:51]
So rather than having to, um, reanalyze a document every time that you perform inference or respond to a user query, we, uh, believe quite firmly that you can also, in an additive way, perform single pass extraction over, over this body of text and then bring that into the RAG context window.

And, um,

this, this really sort of levers off of my experience at LinkedIn, where you had this structured graph representation of the global economy where you said person A works at company B.

**Alessio** [39:30]
Mm-hmm.

**Mike Conover** [39:30]
We believe that there's an opportunity to create a, a knowledge graph that has resolution that greatly exceeds what any, you know, whether it's Bloomberg or LinkedIn currently has access to, where we're getting as granular as person X submitted congressional testimony that was critical of organization Y, and this is the language that is attached-

**Alessio** [39:49]
Mm-hmm

**Mike Conover** [39:50]
... to that testimony. And then you have a structured data artifact that you can pivot through and reason over that is complementary to the generative capabilities that language models expose. And so it's the same technology being applied to multiple different ends, and this is manifest in the product surface where it's a highly facetable, pivotable product, uh, but it also enhances the reasoning capability-

**Alessio** [40:14]
Mm-hmm

**Mike Conover** [40:15]
... of the system.

**Alessio** [40:16]
Yeah. Uh, you know when you mention you don't wanna re-query, like, the same thing over and over, a lot of people may say, "Well, I'll just fine-tune this information in the model," you know?

### Fine-tuning

**Mike Conover** [40:26]
Mm-hmm.

**Alessio** [40:26]
Um, how do you think about that? That, that was one thing when we started working together, you were like, "We're not building foundation models." A lot of other startups were like, "Oh, we're building the finance financial model."

**Mike Conover** [40:37]
Sure.

**Alessio** [40:37]
"Like the finance foundation model or whatever."

**Mike Conover** [40:40]
Yeah.

**Alessio** [40:40]
Um, when, when is the right time for people to do, um, fine-tuning versus RAG? It's like any heuristics that you can share that you use to, to think about it?

**Mike Conover** [40:49]
So we, in general, I do not... I'll just say, like, I don't have, um, a strong opinion about how much information you can imbue into a model that is not present in pre-training-

**Alessio** [41:04]
Mm-hmm

**Mike Conover** [41:04]
... through large scale fine-tuning. Um, the benefit of RAG is the groun- the capability around grounded reasoning. So the, you know, forcing it to attend to a collection of facts that are, are known and available at inference time and sort of like materially, like only using these facts.

**Alessio** [41:24]
Mm-hmm.

**Mike Conover** [41:24]
Um, the, at least in my view, the, the role of fine-tuning is really more around, I think of like language models kind of like a stem cell. And then under fine-tuning they differentiate into different kinds of-

**Alessio** [41:37]
Mm

**Mike Conover** [41:38]
... um, specific cells, so kidney or an eye cell. And, um, we-- If you think about, like specifically, like I don't think that unbounded agentic behaviors are useful.

**Alessio** [41:52]
Mm-hmm.

**Mike Conover** [41:52]
And that instead a s- a useful LM system is more like, um, a finite state machine, where the behavior of the system is occupying one of many different behavioral regimes and making decisions about: What state should I occupy next-

**Alessio** [42:08]
Mm-hmm

**Mike Conover** [42:08]
... in order to satisfy th-the goal? Um, as you think about the graph of those states that the language mo- that your system is moving through- Once you develop conviction that one behavior is useful and repeatable and im- like worthwhile to, um, differentiate down into a specific kind of subsystem-

**Alessio** [42:30]
Mm-hmm

**Mike Conover** [42:30]
... that's where like fine-tuning and like specifically generating the training data, like having h- having human annotators-

**Alessio** [42:38]
Yeah

**Mike Conover** [42:38]
... produce a corpus that is useful enough to get a specific class of behaviors. That, that's kind of how we use fine-tuning rather than trying to imbue new, net new information into these systems.

**Alessio** [42:51]
Mm-hmm. Yep. Um, and how... But, you know, people are always trying to turn LLMs into humans. It's like, "Oh, this is my reviewer."

**Mike Conover** [43:01]
Sure, sure, sure.

**Alessio** [43:01]
"This is my editor."

**Mike Conover** [43:03]
Yeah, yeah.

**Alessio** [43:03]
Uh, I know you're not in that camp.

**Mike Conover** [43:04]
Yeah.

**Alessio** [43:05]
Uh, so any, any thoughts you have on like how people should think about, um, yeah, how to refer to models-

**Mike Conover** [43:11]
Yeah

**Alessio** [43:11]
... like, and yeah.

**Mike Conover** [43:12]
So I mean, we, we've talked a little bit about this, and it's, it is an... It's notable that I think there's a lot of anthropomorphizing going on, and that it reflects the difficulty of evaluating the systems. Is it like...

### Philosophy

**Mike Conover** [43:26]
Does the saying that you are, uh, that you're the, a journal editor for Nature-

**Alessio** [43:33]
Mm

**Mike Conover** [43:33]
... does that help? Like, you're, you know, you've got the editor, and then you've got the reviewer, and you've got the, you know, the... You're the private investigator.

**Alessio** [43:40]
Mm-hmm.

**Mike Conover** [43:40]
You know? It's like this is, I think, um... Literally, we wave our hands and we say, "Maybe if I tell you that I'm gonna tip you, that's gonna help."

**Alessio** [43:48]
Mm-hmm.

**Mike Conover** [43:48]
And it sort of seems to. And like maybe it's just like the more cycles, the more compute that is attached to the prompt and, and then the sort of like chain of thought, uh, at inference time, it's like maybe that's all that we're really doing and that it's kinda like hidden, um-

**Alessio** [44:05]
Hmm

**Mike Conover** [44:05]
... hidden compute. But I-- Our experience has been that you can get really, really high quality reasoning from roughly an agentic system without, um, needing to be too cute about it.

**Alessio** [44:21]
Mm-hmm. Right.

**Mike Conover** [44:22]
You can describe the task and, um, you know, within well-defined bounds, uh, you, you don't need to, um, treat the LLM like a person-

**Alessio** [44:33]
Mm-hmm

**Mike Conover** [44:33]
... in order to get it to generate high quality outputs.

**Alessio** [44:35]
Yeah. And the other thing is like all these agent frameworks are, uh, assuming everything is an LLM, you know, versus-

**Mike Conover** [44:43]
Yeah, for sure. And I think this is one of the places where, um, traditional machine learning has a real material role to play in producing a system that hangs together.

**Alessio** [44:55]
Mm-hmm.

**Mike Conover** [44:56]
And there are, you know, guaranteeable, like statistical promises that classical machine learning systems to include traditional deep learning can make about, you know, what is the set of outputs and like what is the characteristic distribution of those outputs that LLMs cannot afford.

And so like one of the things that we do is we, as a philosophy, try to choose the right tool for the job, and so sometimes that is, um, a, a de novo model that has nothing to do with-

**Alessio** [45:25]
Mm-hmm

**Mike Conover** [45:26]
... LLMs that does one thing exceptionally well, and whether that's, uh, retrieval or, uh, critique or multi-class classification. And, um, I think having many, many different tools in your toolbox is, is always valuable.

**Alessio** [45:42]
This is great. Um, so there's kind of the, the missing piece that maybe people are wondering about. Uh, you do a financial services company and you don't do anything in Excel. Uh, what's the what's the story behind why you're doing partner in thought-

**Mike Conover** [45:56]
Sure

**Alessio** [45:56]
... versus, hey, this is like a AI-enabled model that understands any stock-

**Mike Conover** [46:00]
Yep

**Alessio** [46:00]
... and, and all that?

**Mike Conover** [46:01]
Yeah. And, and to be clear, we, we do... Brightwave does a fair amount of quantitative reasoning. I think what is an explicit non-goal for the company is to, uh, create Excel spreadsheets.

**Alessio** [46:12]
Mm-hmm. Yeah.

**Mike Conover** [46:13]
And I think when you look at the, uh, products that work in that way, you can spend hours with an Excel spreadsheet and not notice a subtle bug. And as a... Like that is a highly non-fault tolerant-

**Alessio** [46:30]
Mm-hmm

**Mike Conover** [46:30]
... product experience where you encounter a misstatement in a financial model in terms of how a formula is composed and all of your assumptions are suddenly violated, and now it's effectively wasted effort. So as opposed to the partner in thought modality, which is yes and.

Like if, if the model says something that you don't agree with-

**Alessio** [46:50]
Mm-hmm

**Mike Conover** [46:51]
... you can say, "Taken under consideration. I, you know, this is not interesting to me. I'm gonna pivot to the next finding or claim." And it's more like a dialogue. Um, the other piece of this is that the financial modeling is often very-- When we, when we talk to our users, it's very personal.

So they have a specific view of how a company is structured. They have the, you know, one key driver of asset performance-

**Alessio** [47:16]
Mm-hmm

**Mike Conover** [47:16]
... that they think is really, really important. Um, it's kind of like the difference between writing an essay and having an essay-

**Alessio** [47:23]
Hmm

**Mike Conover** [47:23]
... I guess. Like the pro- the purpose of homework is to actually develop, what do I think about this?

**Alessio** [47:30]
Right.

**Mike Conover** [47:30]
And so it's not clear to me that like push a button, have a financial model is, is solving the actual problem that-

**Alessio** [47:39]
Mm-hmm

**Mike Conover** [47:39]
... the financial model affords. Um, and so we... That said, we take great efforts to have exceptionally high quality quantitative reasoning. So we, um... You know, you think about, um... And I, I won't get into too many specifics about this, but, uh, we deal with a fair number of documents that have tabular data that is, uh, really important to making informed decisions.

**Alessio** [48:03]
Mm-hmm.

**Mike Conover** [48:03]
And so the way that, uh, our RAG systems operate over and retrieve from tabular data sources is, um, it's something that we, um, place a, a great degree of emphasis on. It's just I think the, the f- the medium of- Like Excel spreadsheets is just-

**Alessio** [48:21]
Yeah

**Mike Conover** [48:21]
... I think not the right play for this class of technologies as they exist in 2024.

**Alessio** [48:28]
Mm-hmm. Yeah, what about, uh, 2034?

**Mike Conover** [48:31]
2034?

**Alessio** [48:31]
Like are people still gonna be making Excel models or... I, I, I like, yeah-

**Mike Conover** [48:35]
Yeah

**Alessio** [48:35]
... I think to me the most interesting thing is like how are the models abstracting people away from some of these more-

**Mike Conover** [48:43]
Mm-hmm, mm-hmm

**Alessio** [48:43]
... syntax-driven thing and making-

**Mike Conover** [48:45]
Yeah

**Alessio** [48:45]
... them focus on what matters to them.

**Mike Conover** [48:48]
Yeah, I wouldn't be able to tell you what the future l- uh, you know, 10 years from now looks like. I think anybody who could convince you of that, um, is, is not necessarily somebody to be trusted. Um, I do think that-- So draw- let's draw the parallel to, um, accountants in the, the '70s.

**Alessio** [49:08]
Mm-hmm.

**Mike Conover** [49:08]
So VisiCalc, I believe, came out in 1979, and historically the core idea-- You know, you would have, as an accountant, as a finance professional in the '70s, like I'm the one who runs the-- I'm, I run the numbers.

I do the arithmetic.

**Alessio** [49:20]
Mm-hmm.

**Mike Conover** [49:20]
That's like my main job. And we think that, I mean, you just look, the-- now that's not a job anybody wants.

**Alessio** [49:29]
Mm-hmm.

**Mike Conover** [49:29]
And the sophistication of the analysis that a person is able to perform as a function of having access to powerful tools like computational spreadsheets, um, is, is just much greater. And so I think that with regards to language models, it, it is probably the case that there is a, a play in the workflow where it, um, is commenting on your analysis within that, you know, s-spreadsheet-based context, or it is taking information from those models and sucking this into a system that does qualitative reasoning-

**Alessio** [50:05]
Mm-hmm

**Mike Conover** [50:05]
... on top of that. Um, but I think the, um, the, uh... It is an open question as to whether the actual production of those models maintain-- is, is still a human task.

**Alessio** [50:17]
Mm.

**Mike Conover** [50:17]
But I think the sophistication of the analysis that is available to us, um, and the completeness of that analysis, um, just necessarily in-increases over time.

### AI Funds

**Alessio** [50:28]
Mm-hmm. Yeah. Uh, what about AI hedge funds? Obviously, I mean, we have quants today, right? But those are more-

**Mike Conover** [50:34]
Yeah

**Alessio** [50:34]
... uh, kinda like momentum driven, kinda like signal driven-

**Mike Conover** [50:37]
For sure

**Alessio** [50:37]
... and less about long thesis driven. Uh, do you think that's a possibility? Uh...

**Mike Conover** [50:43]
It's-- This is an interesting question. I would, I would put it back to you and say like how, um, how different is that from what hedge funds do now? I think there's-- The, the more that I have learned about

how teams at hedge funds actually behave, and you look at like systematics desks or sy-

**Alessio** [51:02]
Mm-hmm

**Mike Conover** [51:02]
... semi-systematic trading groups, man, it's a lot like a big machine learning team.

**Alessio** [51:07]
Mm-hmm.

**Mike Conover** [51:07]
And it's, you-- I sort of think it's interesting, right? So like if you look at video games and traditional like Bay Area tech, there's not a ton of like, uh, talent mobility between-

**Alessio** [51:20]
Mm-hmm

**Mike Conover** [51:20]
... those two communities. You have people that work-

**Alessio** [51:22]
Yeah

**Mike Conover** [51:22]
... in video games and people that work in like SaaS software. And it's not that like cognitively they would not be able to work together. It's just like a different set of skill sets, a different set of relationships, and it's kinda like network clusters that-

**Alessio** [51:33]
Mm-hmm

**Mike Conover** [51:33]
... don't interact. And I think there's probably a similar phenomenon happening with regards to, uh, machine learning within the active-

**Alessio** [51:44]
Mm-hmm

**Mike Conover** [51:44]
... asset allocation community. And so like

I'm-- It's actually not clear to me that we don't have AI hedge funds now.

**Alessio** [51:53]
Mm-hmm.

**Mike Conover** [51:53]
The question of whether you have an AI that is operating at a trading desk-

**Alessio** [51:57]
Right. Yeah

**Mike Conover** [51:57]
... like that, that seems a little, um...

Maybe. Like the-- I don't have line of sight to something-

**Alessio** [52:06]
Oh, no

**Mike Conover** [52:06]
... like that existing yet.

**Alessio** [52:07]
I mean, I'm-

**Mike Conover** [52:07]
Yeah

**Alessio** [52:07]
... always curious.

**Mike Conover** [52:08]
Yeah.

**Alessio** [52:08]
You know, I, I think about, uh, asset management on a few different ways, but uh, venture capital is like extremely power law driven.

**Mike Conover** [52:17]
Yeah.

**Alessio** [52:17]
It's really hard to do machine learning in power law businesses-

**Mike Conover** [52:20]
Yes

**Alessio** [52:20]
... because, you know, the distribution of outcomes-

**Mike Conover** [52:22]
Yeah, yeah

**Alessio** [52:22]
... is like so small. Versus public equities, most high-frequency trading is like very, you know, bell curve-

**Mike Conover** [52:29]
Yeah

**Alessio** [52:30]
... uh, normal distribution. It's like even if you just get 50.5%-

**Mike Conover** [52:34]
Yep

**Alessio** [52:34]
... uh, at the right scale, you're gonna make a lot of money.

**Mike Conover** [52:36]
Yeah.

**Alessio** [52:36]
Uh, and I think AI starts there, right? And today, most high-frequency trading is already AI driven. You know, Renaissance, uh, started-

**Mike Conover** [52:44]
Yeah, precisely, right

**Alessio** [52:44]
... a long time ago using, using these models. Uh, but I'm curious how it's gonna move closer and closer to like power law businesses, right? I would say some boutique hedge funds, their, uh, their pitch is like, "Hey, we're differentiated because we only do kinda like these long-only strategies that are like thesis driven versus-

**Mike Conover** [53:03]
Yep

**Alessio** [53:03]
... uh, you know, movement driven."

**Mike Conover** [53:05]
Yep.

**Alessio** [53:05]
Um, and most venture capitalists will tell you, "Well, our fund is different because we have this unique thesis on-

**Mike Conover** [53:10]
Sure, sure, sure

**Alessio** [53:11]
... uh, on this market." Um, a-and I think like five years ago, I read this blog post about why machine learning would never work in venture because the things that you're investing in today, there are just like no precedent that should tell you this will work, you know?

**Mike Conover** [53:25]
Right. Right, right.

**Alessio** [53:25]
Most new companies, Amato will tell you, "This is not gonna work."

**Mike Conover** [53:28]
Yes.

**Alessio** [53:29]
You know? Versus the closer you get to the public companies, the more any innovation is like, okay, this is kinda like this thing that happened.

**Mike Conover** [53:36]
Yeah.

**Alessio** [53:37]
Um, and I feel like these models are quite good at generalizing and thinking. Again, going back to the partner and thought, like thinking about second order-

**Mike Conover** [53:44]
Yeah

**Alessio** [53:45]
... things. But-

**Mike Conover** [53:45]
And that's, that's maybe where-- So concrete exampl- I think it certainly is the case that we tell retrospective-- To your point about venture, we tell retrospective stories where it's like, well, here, here was the set of observable facts.

This was knowable at the time-

**Alessio** [54:01]
Mm-hmm

**Mike Conover** [54:01]
... and these people made the right call and were able to cross-correlate all of these different sources, and this is the bet we're gonna make. I think that process of idea generation is absolutely automatable.

**Alessio** [54:14]
Mm-hmm.

**Mike Conover** [54:14]
And the question of like, do you ever get somebody who just sets the system running and it's making all of its own decisions like that? And it is truly like, um, doing thematic investing-

**Alessio** [54:26]
Mm-hmm

**Mike Conover** [54:26]
... or more of the like what a human analyst would be, um, kind of on the hook for as opposed to like HFT.

**Alessio** [54:33]
Yeah.

**Mike Conover** [54:33]
Uh, but the ability of models to say, "Here is a fact pattern that is noteworthy, and s- we should, we should pay more attention here." Because if you think about the- the matrix of like all possible relationships in the economy, um, it grows with the square of the number of-

**Alessio** [54:55]
Mm-hmm

**Mike Conover** [54:56]
... facts you're evalu- or like polynomial with the-

**Alessio** [54:58]
Yeah

**Mike Conover** [54:58]
... polynomially with the number of facts you're evaluating. And so if I, um, if I want to make bets on AI, I think it's like what are, what are, you know, ways to profit from the, the rise of AI?

It is very straightforward to take a model and say parse through all of these documents and find second order derivative bets and say, "Oh, it turns out that energy is like very, very adjacent to i- investments in AI and may, may not be priced in the same way that, uh, GPUs are."

**Alessio** [55:30]
Mm-hmm.

**Mike Conover** [55:31]
And a derivative vet-- of energy for example, is, um, long duration energy storage. And so you need a bridge between renewables, which have fluctuating demands and the compute requirements-

**Alessio** [55:42]
Mm-hmm

**Mike Conover** [55:43]
... of these data centers.

**Alessio** [55:44]
Yeah.

**Mike Conover** [55:44]
And I think that... And I'm, I'm telling this story as like having witnessed Brightwave-

**Alessio** [55:50]
Right. Yeah, yeah, yeah

**Mike Conover** [55:52]
... do this work. You can take a premise and say like, "What are second and third order bets that we can make on this topic?" And it's gonna come back with, "Here's a set of reasonable theses."

**Alessio** [56:02]
Mm-hmm.

**Mike Conover** [56:02]
And then I think a human's role in that world is to assess, like, does this make sense given our fund strategy? Does this-

**Alessio** [56:08]
Yeah

**Mike Conover** [56:09]
... is this coherent with the calls that I've had with management teams? There's this broad body of knowledge that I think humans, um, sort of are the ultimate like synthesizers and deciders.

**Alessio** [56:22]
Mm-hmm.

**Mike Conover** [56:22]
And like maybe I'm wrong. Maybe the world of the future looks like... And the, the AI that truly does everything, that is a wor- I think it is kind of a singularity-

**Alessio** [56:34]
Mm-hmm. Yeah, yeah, yeah

**Mike Conover** [56:34]
... like factor where it's like really hard to reason about like what that world looks like.

**Alessio** [56:38]
Uh-huh.

**Mike Conover** [56:38]
And I... You asked me to speculate, but I'm actually kinda hesitant to do so-

**Alessio** [56:41]
Yeah

**Mike Conover** [56:42]
... because it's just the forecast, the, the hurricane path-

**Alessio** [56:47]
Right

**Mike Conover** [56:47]
... just diverges far too much to-

**Alessio** [56:48]
Mm-hmm

**Mike Conover** [56:48]
... have a, a real conviction about what that looks like.

### Open Source

**Alessio** [56:51]
Um, awesome. I know we've, we've already taken up a lot of your time. Um, but maybe one thing to, to touch on before wrapping is-

**Mike Conover** [56:58]
Sure

**Alessio** [56:59]
... um, open source LLMs.

**Mike Conover** [57:00]
Yeah. Yeah.

**Alessio** [57:01]
Obviously, you were at the forefront of it. Uh, we recorded our episode the day that Red Pajama was open source-

**Mike Conover** [57:06]
Mm-hmm. Yeah

**Alessio** [57:06]
... and we were like, "Oh man, this is, this is mind-blowing."

**Mike Conover** [57:09]
Yeah.

**Alessio** [57:09]
"This is gonna be crazy." Um, and now we're gonna have a open source, uh, dense transformer model that is 400 billion parameters.

**Mike Conover** [57:17]
Yeah.

**Alessio** [57:17]
That I, I don't know if one year ago you could've told me that that was gonna happen.

**Mike Conover** [57:21]
Goodness. Yeah.

**Alessio** [57:21]
Uh, so what, what do you think matters in open source? What do you think, uh, people should work on? Uh, what are like things that people should keep in mind to evaluate, okay-

**Mike Conover** [57:31]
Sure

**Alessio** [57:31]
... is this model actually gonna be good, or is it just like cheating some benchmarks to look good? It's like is there anything there? Like, yeah. The, the... Oh, but th- this is the part-

**Mike Conover** [57:39]
Yeah

**Alessio** [57:40]
... of the podcast where people already dropped off if they wanted to, so they wanna hear-

**Mike Conover** [57:43]
Sure. Sure. Yeah

**Alessio** [57:43]
... the hot takes right now.

**Mike Conover** [57:45]
Yeah. I mean, I do think that that's another reason to have your own private evaluation corpuses-

**Alessio** [57:49]
Mm-hmm

**Mike Conover** [57:50]
... is so that you can objectively and out of sample-

**Alessio** [57:53]
Right

**Mike Conover** [57:53]
... uh, measure the performance of these models. Um, and again, like sometimes that just looks like giving everybody on the team 250 annotations and saying-

**Alessio** [58:03]
Mm-hmm

**Mike Conover** [58:03]
... "We're, we're just gonna-

**Alessio** [58:03]
Right

**Mike Conover** [58:03]
... grind through this," and like you, you know, you have to tell does this meet... The other thing about like doing the work yourself is that you get to articulate your loss function precisely.

**Alessio** [58:12]
Mm-hmm.

**Mike Conover** [58:13]
What is the thing that I... What do I actually want the system to behave like? Do I prefer this system or the, you know, this model or this other model? Um, yeah, and I think the, the work around over- you know, overfitting on the, the test I think is like that 100% is happening.

Um, one notable

in contrast to a year ago, say, and, you know, the,

the incentives, the economic incentives for companies to train their own foundation models I think are diminishing.

**Alessio** [58:43]
Mm-hmm.

**Mike Conover** [58:43]
So the like w- window in which you are the dominant pre-train, and let's say that you spend five to $40 million-

**Alessio** [58:55]
Mm-hmm

**Mike Conover** [58:55]
... you know, for like a call it kind of a commodity-ish pre-train.

**Alessio** [59:00]
Yeah.

**Mike Conover** [59:00]
Not... And a 400 billion would be another-

**Alessio** [59:02]
Mm-hmm

**Mike Conover** [59:02]
... sort of-

**Alessio** [59:03]
Well, it costs more than 40 million.

**Mike Conover** [59:05]
Yeah. Another leap. But the kind of thing that, you know, like a, you know, small, uh, you know, multi-billion dollar-

**Alessio** [59:10]
Mm-hmm

**Mike Conover** [59:10]
... mom-and-pop shop might, might be able to pull off. Um,

the benefit that you get from that is like, I think, diminishing over time.

**Alessio** [59:20]
Yeah.

**Mike Conover** [59:20]
And so I think fewer companies are going to make that capital outlay. Um, and I think that, you know, there's, there's probably s- some, some material negatives to that.

**Alessio** [59:31]
Mm-hmm.

**Mike Conover** [59:31]
But the other piece is that we're seeing that at least in the past two and a half, three months, there's a convergence towards like, well, these models all behave fairly similarly.

**Alessio** [59:43]
Yeah.

**Mike Conover** [59:44]
And, you know, it's probably that the like training data on which they are pre-trained is substantially overlapping.

**Alessio** [59:52]
Mm-hmm.

**Mike Conover** [59:52]
And so it's generalizing a model that generalizes-

**Alessio** [59:55]
Yeah

**Mike Conover** [59:55]
... to that training data. And so it's, um, unclear to me that you have this sort of balkanization where there are many different models, each of which is good in its own-

**Alessio** [1:00:05]
Mm-hmm

**Mike Conover** [1:00:05]
... unique way. Um, versus something like Llama becomes like, listen, this is a fine standard to build off of. Um, we'll see. It's just like the, the upfront cost is so high.

**Alessio** [1:00:19]
Right.

**Mike Conover** [1:00:19]
And I think for the people that have the money, the, the benefit of doing the pre-train is now, um, less. The... Where I think it gets really interesting is how do you- differentiate these and all of these different behavioral regimes.

**Alessio** [1:00:33]
Mm-hmm.

**Mike Conover** [1:00:33]
And I think the cost of producing, um, instruction tuning and fine-tuning data-

**Alessio** [1:00:39]
Mm-hmm

**Mike Conover** [1:00:39]
... that creates specific kinds of behaviors, I think that's probably where the next generation of, of really interesting work starts to happen, and that's... If you see that m- the same model architecture trained on much more training data can exhibit, um, substantially improved performance, it's the, the value of modeling innovations.

It, it's, you know, for fundamental machine learning and AI research, like, there is still so much to be done.

**Alessio** [1:01:12]
Mm-hmm.

**Mike Conover** [1:01:12]
But I think that the much more-- like, much lower hanging fruit, I guess, is developing new kinds of, uh, training data corpuses that elicit new behaviors from these models in a sp- a specific way. And so that's where when I, when I think about the m- the availability to mu- Like a year ago, you had to have access to, uh, fairly high performance-

**Alessio** [1:01:37]
Right

**Mike Conover** [1:01:37]
... GPUs that were hard to get in order to get the experience of multiple reps fine-tuning these models.

**Alessio** [1:01:43]
Mm-hmm.

**Mike Conover** [1:01:43]
And what you're doing when you, you know, take a corpus and then fine-tune the model and then see across many inference passes what is the qualitative character of the output, you're developing your own internal-

**Alessio** [1:01:56]
Mm-hmm

**Mike Conover** [1:01:56]
... mental model for, like, how does the composition of the training corpus shape the behavior of the model in a qualitative way. A year ago, it was very expensive to get that experience.

**Alessio** [1:02:05]
Mm.

**Mike Conover** [1:02:06]
And now you can just recompose multiple different training corpuses and see, like, "Well, what do I, what do I do if I insert this set of demonstrations-"

**Alessio** [1:02:14]
Yeah

**Mike Conover** [1:02:14]
"... or I ablate that set of demonstrations?" And you can... That, that I think is a very, very valuable skill and one of the ways that you can have models and products that, um, other people don't have access to.

And so I, I think as more people-- as that s- those sensibilities proliferate because more people have that experience, you're gonna see teams that, um, release data corpuses that just imbue the models with new behaviors that are especially interesting-

**Alessio** [1:02:42]
Mm-hmm

**Mike Conover** [1:02:42]
... and useful. And I think that, that may be where, um, some of the, the next sets of kind of innovation differentiation come from.

**Alessio** [1:02:49]
Yeah. Yeah, when people ask me, I always tell them the, the half-life of a model is much shorter than a half-life of a dataset.

**Mike Conover** [1:02:55]
Yes.

**Alessio** [1:02:55]
You know?

**Mike Conover** [1:02:55]
Absolutely.

**Alessio** [1:02:55]
Like, I mean, the pile-

**Mike Conover** [1:02:57]
Yeah

**Alessio** [1:02:57]
... the pile is still around-

**Mike Conover** [1:02:58]
Totally

**Alessio** [1:02:58]
... and, like, core to most of these-

**Mike Conover** [1:03:00]
Yeah

**Alessio** [1:03:00]
... training runs versus-

**Mike Conover** [1:03:01]
Yeah

**Alessio** [1:03:01]
... all the models people trained a year ago. It's like they're at the bottom of the LMSYS leaderboard.

**Mike Conover** [1:03:06]
Yeah, and it's kind of-

**Alessio** [1:03:07]
It's just like-

**Mike Conover** [1:03:07]
It's kinda crazy. Like, I don't... Just the parallels to other kinds of computing technology where, like, the work involved in producing the artifact-

**Alessio** [1:03:16]
Mm-hmm

**Mike Conover** [1:03:16]
... is so significant, and the, like, shelf life is, like-

**Alessio** [1:03:21]
Right

**Mike Conover** [1:03:21]
... a week. It's like I-- You know, I'm sure there's a precedent, but it is, it is remarkable.

**Alessio** [1:03:27]
Mm-hmm.

**Mike Conover** [1:03:27]
Yeah.

**Alessio** [1:03:29]
Yeah. I remember when DALL-E was the best open source model. Um-

**Mike Conover** [1:03:32]
I, we-- We-- Yeah, but DALL-E was never the best open source model, but it w- it-

**Alessio** [1:03:35]
But-

**Mike Conover** [1:03:36]
... demonstrated a new s- something that was not obvious to many people at the time. Yeah, but we always were clear that it was never state-of-the-art.

### Hiring

**Alessio** [1:03:44]
State-of-the-art, whatever that means, right?

**Mike Conover** [1:03:46]
Sure.

**Alessio** [1:03:46]
Um, this is great, Mike. Anything that we forgot to cover that you wanna add? Any, um-

**Mike Conover** [1:03:52]
No, that's-

**Alessio** [1:03:52]
... call-- I know you're, you know, thinking about-

**Mike Conover** [1:03:54]
We are, we are hiring across-

**Alessio** [1:03:55]
... growing the team and yeah.

**Mike Conover** [1:03:56]
We are hiring across the board. Um, AI, uh, engineering, classical machine learning, systems engineering, distributed systems, front-end engineering, uh, design. Uh, we, we have many open roles on the team. Uh, we hire exceptional people. Um, we fit the job to the person as a philosophy, um, and would love to work with, uh, more incredible humans.

**Alessio** [1:04:20]
Awesome. Thank you so much for coming on, Mike.

**Mike Conover** [1:04:22]
Thanks, Alessio.

---

This library is powered by PodHood (https://podhood.com), the podcast website platform.
