# ⚡️Using RFT to Build Clinical Superintelligence

Latent Space · 2025-07-29

<https://addtry.com/fed71fa3-2d16-41ae-ba87-0987d9408442>

Brendan Fortuner, Head of Eng at Ambience AI, explains how the healthcare startup uses OpenAI's reinforcement fine-tuning (RFT) to build clinical AI assistants that help doctors automate note-taking and ICD-10 coding, saving up to two hours per day. Ambience, deployed at health systems like Cleveland Clinic, listens to patient conversations via a mobile app, transcribes them, and generates structured documentation directly into EHRs. Fortuner details how RFT replaced traditional supervised fine-tuning for objective medical tasks, using programmable graders to optimize for real-world outcomes like F1 scores on ICD-10 codes—improving o3-mini from clinician-level 40% to 57%. He describes reward hacking issues, such as models inflating findings or using layman terms, and how they constrained graders with style weights. The episode also covers domain expert vs. ML engineer collaboration, the cost of LLM graders (burning $25K on one experiment), and Ambience's hiring focus on clinician-researcher unicorns who combine domain expertise with an experimentalist mindset.

## Questions this episode answers

### How did Ambience use OpenAI's reinforcement fine-tuning to improve ICD-10 coding accuracy?

Brendan Fortuner explains that clinicians scored only 40% F1 on ICD-10 coding. Using RFT on O3-mini, Ambience reached 57% F1. The task maps doctor visit notes to one of 70,000 standardized codes, a tedious administrative burden. This early result shows RFT's promise for automating objective medical tasks with direct optimizer feedback.

[10:48](https://addtry.com/fed71fa3-2d16-41ae-ba87-0987d9408442?t=648000)

### What reward hacking issues did Ambience encounter when applying RFT to physical exam notes?

Brendan Fortuner describes two issues: the model inflated the number of findings to game precision, and it used layman terms like 'Grandpa has an unhappy heart,' degrading medical tone. They fixed this by updating the grader to penalize hallucinations and adding a style component weighted 75-25, showing how iterative grader design mitigates reward hacking.

[9:00](https://addtry.com/fed71fa3-2d16-41ae-ba87-0987d9408442?t=540000)

### What kind of talent is Ambience looking for to advance clinical AI?

Brendan Fortuner says they need top ML researchers for frontier work, but the real 'unicorn' is the clinician researcher archetype—people with medical domain knowledge, an experimentalist mindset, and a startup operating style. This blend is rare but essential for translating clinical feedback into model improvements and driving healthcare AI innovation.

[25:42](https://addtry.com/fed71fa3-2d16-41ae-ba87-0987d9408442?t=1542000)

## Key moments

- **[0:00] Intro**
- **[0:33] Ambience Overview**
  - [1:35] Ambience AI saves doctors up to two hours a day by automating note-taking and administrative tasks with its AI assistant.
  - [2:16] Epic holds over 50 percent market share in electronic health records, with Oracle's Cerner at 27 percent, says Brendan Fortuner.
  - [3:36] Fine-tuning LLMs is very sample efficient compared to supervised learning with millions of labeled images, says Brendan Fortuner from his Cruise experience.
- **[5:14] RFT Basics**
  - [5:21] Q: How did Ambience decide to use reinforcement fine-tuning (RFT) and build the infrastructure for it?
  - [6:31] Reinforcement fine-tuning (RFT) uses programmable graders, teaching models to maximize a score through trial and error, explains Brendan Fortuner.
  - [8:03] Brendan Fortuner: 'In basketball, it doesn't matter your form. What matters is you get points.' — why RFT optimizes end objectives.
- **[8:37] Reward Hacking**
  - [10:10] Ambience's RFT model initially reward-hacked by inflating findings and using terms like 'Grandpa has an unhappy heart,' needing grader fixes.
- **[10:43] ICD-10 Coding**
  - [12:03] Ambience used RFT to boost ICD-10 coding F1 score from 40% (human clinicians) to 57% with O3-mini, recruiting 18 physicians for evaluation.
- **[12:26] Beyond Scribing**
  - [13:05] Ambience is developing a patient-facing voice agent to handle medication reminders and administrative tasks, plus pre-charting summaries for doctors.
- **[14:18] RFT Economics**
  - [14:29] Brendan Fortuner warns that using an expensive LLM grader for RFT can burn $20K quickly, as Ambience did on a 100-example experiment.
- **[16:50] Clinicians & MLEs**
- **[18:33] Model Robustness**
  - [19:13] Brendan Fortuner says that while benchmarks report best metric out of 64 attempts, healthcare requires evaluating the worst event due to safety.
  - [20:32] Base models have 'medical school' from textbooks but lack 'residency' experience encoded in EHRs, making clinical data out of distribution, says Brendan Fortuner.
- **[22:00] Data Quality**
  - [23:20] Brendan Fortuner defines 'clinical taste' as the judgment to know which EHR data sources to trust, analogous to avoiding SEO spam on the web.
- **[24:33] Research Agent**
  - [24:47] Brendan Fortuner predicts a 'Devin for clinical AI' could autonomously gather user feedback, mine data, and run evals for medical AI development.
- **[25:19] Hiring**
  - [25:42] Ambience is hiring machine learning researchers and clinician researchers with both domain expertise and startup operating mindset, says Brendan Fortuner.

## Speakers

- **Alessio** (host)
- **Wix** (host)
- **Brendan Fortuner** (guest)

## Topics

Healthcare, Reinforcement Learning, Evals

## Mentioned

Ambience (company), Athena (company), Braintrust (company), Cerner (company), Epic (company), MedTech (company), OpenAI (company), Ambience Scribing (product), Health Bench (product), RFT (product), o3-mini (product)

## Transcript

### Intro

**Alessio** [0:04]
Hi, everyone. Welcome back to another Latent Space Lightning Pod. This is Alessio, partner and CTO at Decibel, and I'm joined by my co-host, Wix, founder of Smol AI.

**Wix** [0:12]
Hello, and today we are joined in a remote studio with Brendan Fortuner of Ambience. Welcome.

**Brendan Fortuner** [0:18]
Hey, thanks so much for having me. Huge, huge fan of the pod.

**Wix** [0:21]
So let's talk about Ambience. Um, uh, you know, I, I think, I think you are actually our first guest covering or, you know, working in the healthcare sector. There's a lot of doctor, medicine, AI scribe things out there.

What's the origin and what's your specific pitch for Ambience?

### Ambience Overview

**Brendan Fortuner** [0:36]
Yeah, sure, sure. So I'll tell you maybe the just the quick TLDR and maybe get into the workflow and, and, and share some of the problem, uh, for folks who are less familiar with healthcare. So the TLDR is basically, you know, Ambience is a healthcare company.

We're building an AI assistant for doctors, and we help doctors take notes. We help them complete, you know, administrative tasks, and the end result is we save them a bunch of time, and they can spend that extra time, like, caring for patients, right?

Our customers are actually, like, big health systems, so like Cleveland Clinic and UCSF, Ardent, Sean Muir, these are the kind of customers we sell to. But the end users are actually the clinicians themselves, doctors, nurse practitioners. The basic workflow with Ambience is we have a mobile app.

The doctors will take it into the room with the patient. You may have seen this, you know, if you're a patient. We'll then listen to that audio, we'll transcribe it into text, and then we'll, we'll call language models that are fine-tuned for medicine to generate all different types of unstructured and structured documentation, and then we'll automatically write it back into the electronic health record.

Uh, that's the system of record where doctors do a lot of their typing and note-taking. And the end result is, like, we'll save doctors, you know, up to two hours a day. We make them tremendously happier, um, and we help the health, health systems as well see more patients.

**Wix** [1:47]
Every time someone mentions EHR, I have to mention, like, it's a, it's a complicated system in, in the US, right? Like, that's basically the default system that is owned by one company. Is that, is that how it works?

Like-

**Brendan Fortuner** [2:00]
Right

**Wix** [2:00]
... you have to submit data to it, and, like, everyone works off of that system?

**Brendan Fortuner** [2:04]
Yeah. So there's, there's actually many electronic health records, at least 20 or so. Um, but there's a few big ones that our customers, you know, the big health systems run. Epic is obviously like the Goliath in the room, over fifty percent market share.

Cerner, that's another really big one o-owned by Oracle, maybe about twenty-seven percent. Uh, then there's Athena and, and MedTech. You know, Ambience, we integrate with these EHRs, and they have open APIs. They have some private APIs, but we try to make it seamless for clinicians.

**Wix** [2:30]
And you previously were working in self-driving with Cruise.

**Brendan Fortuner** [2:33]
That's right. Yeah. So I joined, I joined Ambience like three and a half years ago, um, to join the founders, Mike and Nikhil, my close friends. Uh, prior to this, I was at Cruise. I was working on self-driving cars and computer vision and lidar.

Um, I co-founded our machine learning platform team there, um, and then took some of that knowledge here to Ambience.

**Wix** [2:50]
Yeah. Uh, um, yeah. Part of me w-- thinks that this, this sort of LLM-centric problem is very different from the machine learning problem that you had at Cruise, but maybe it isn't because- So what's different and what's the same?

**Brendan Fortuner** [3:04]
Machine learning is still very frustrating. It's always been frustrating. Um, that's a great question. I think, like, over the last, you know, four years, there's been huge transformation in the industry and how machine learning has been practiced. Um, you know, before you might focus at Cruise on labeling massive datasets, millions of images, right?

Or lidar and, and fine-tuning models that are, like, a lot smaller. These days, the models are, are massive. Um, they generalize extremely well. You can use techniques like prompting to kind of few-shot them. Even fine-tuning itself is very sample efficient.

So it's definitely a lot faster to iterate. It's a lot faster to get into production, and that's actually really, really exciting. The thing that I'll say that they have in common, um, is it's still a data engine. It's still very data-driven, you know, um, machine learning, where what data do you have?

Like, what's the distribution? How do you collect really good annotations? How do you evaluate success? Some of these fundamentals haven't changed.

**Alessio** [3:58]
That, that makes sense, and talking about data being frustrating at the healthcare, it's like one of the, the number one places where, where that happens.

**Brendan Fortuner** [4:05]
That's right.

**Alessio** [4:06]
Can you-- can we just walk through very quickly the different pieces of, like, the Ambience experience, so to speak? Like, uh, where people start using it, you mentioned that doctors kind of bring it into the room. Uh, we can talk about some of the work you've done on reinforcement fine-tuning as well.

**Brendan Fortuner** [4:21]
Yeah, sure. Sure, absolutely. So the Ambience, the, the initial use case as I described, it's called Ambience Scribing. That's where the name Ambience came from. You know, that's... We listen to audio, and we take notes. But that's actually just, like, I would say five percent of, like, the total value that we can kind of offer, you know, once you have the audio, um, of a conversation, right?

And once you have access to the EHR and you can read and write to it. So Ambience is rapidly proliferating an actual AI platform. So why stop at just notes? Why not place-- help place orders or why not place, you know, ICD-10 codes or why not help, you know, maybe call patients on the phone and ask them if they've taken their medication?

So Ambience is slowly building out all these generative AI use cases in healthcare, you know, given the transcripts and the access to the EHR. In terms of the stack, I think it really depends on the use case. Ambience internally, we use prompting, we use chaining, we'll use RAG, um, we use fine-tuning.

It will use SFT and RFT.

### RFT Basics

**Alessio** [5:14]
Yeah. So you mentioned RFT specifically. Uh, I think it just went GA on OpenAI maybe a couple weeks ago.

**Brendan Fortuner** [5:20]
Mm-hmm.

**Alessio** [5:21]
Obviously, everybody's talking about reinforcement learning, but maybe most companies don't really know how to do it. Can you maybe talk about the RFT journey at Ambience? And like, when did you think-- when did you decide it was, like, a good fit, and then how did you spin up the whole infrastructure to get the right data for it?

**Brendan Fortuner** [5:36]
Yeah. This is a great question. Maybe just, you know, stepping back, just a quick, quick summary of RFT. So RFT is-- it's a RL-based method for fine-tuning language models and, and teaching them how to think and how to reason, you know, before answering questions.

Um, it's phenomenal for very objective tasks. You know, STEM, math, science, coding are kind of the flagship use cases. And it's the same technique that's used to train these state-of-the-art reasoning models like o3, you know, R1, uh, Claude 4.

But it's now, you know, very recently available to all developers and startups. Our foray into RFT has primarily been through the OpenAI kind of platform. So we're using those self-service, you know, APIs. The basic core concept, right, of, of the RFT, which I think makes it so special, is unlike supervised learning, where you create these big labeled datasets of thousands of examples, you know, drafted by humans, and you kind of train the model to imitate them and try to get the correct answer.

With RFT, you kind of replace those labels with these graders, uh, these programmable graders, these scoring functions, right? They'll output a score of zero to one, and the model learns to kind of hill climb and try to maximize these scores using RL, right?

So with the API, it, it starts with a grader, right? And there's different types. You could try, like, string match. You could try regex. You could do a fu-fuzzy match. Often you could do a unit test, run some Python code, or even use an LLM grader or even combine a bunch of graders together into a massive grader, you know, kitchen sink.

I've seen I've seen some crazy stuff. But I think the actual, like, you know, technique behind the scene, and we don't know exactly how OpenAI does it, but I think, like, from our perspective at a high level, the models, each epoch, they're generating four or sixty-four different candidate trajectories for every example in your dataset.

The grader scores them zero to one, and then the model kind of learns, like, what tokens did I generate that helped improve that score? And then reinforce those. Or, like, what tokens actually hurt me? And then, you know, don't, don't send those tokens as much.

And then through trial and error over thousands and millions of, of, you know, kind of iterations or trajectories, kind of figures out how to maximize the score, right? So Ambience, like, we found this-- This is attractive to us for many reasons.

In medicine, there's a lot of tasks that are fairly objective in nature, and we'll talk about maybe ICD. Also, there's a lot of reasoning. There's a lot of medical reasoning that's kind of difficult, uh, and complex for just, you know, some of the off-the-shelf pre-trained or SFT models.

And I also think, like, in machine learning, we're often optimizing these proxy metrics, right? Like loss. In the real world, nobody cares about loss. So as a machine as a machine learning engineer, you're, you're always looking for ways and techniques to kind of optimize for the end objective.

Like, what do you really want? Like in basketball, like, it doesn't matter your form. What matters is you get points, right? Um, and I think that's what RFT brings to the table. It's like it gives you a way to optimize for the end objective, um, and Ambience cares about that in healthcare.

And the second, I think, is, like, it's, it's tremendously sample efficient, right? Each example sort of blooms into dozens of labels, right? And trajectories. You can squeeze, like, X more signal, right, out of a, out of a dataset.

And I think those two things are, like, tremendously profound. I'll pause there, but yeah, there's some cool stuff.

**Wix** [8:37]
Something worries me about data efficiency. We all want it until it works against us. So sometimes it's possible to learn too much from one sample. You know what I mean?

### Reward Hacking

**Brendan Fortuner** [8:46]
I saw a paper on that recently. It's like

**Wix** [8:49]
RLVR-

**Brendan Fortuner** [8:49]
It's this RLVR thing?

**Wix** [8:50]
Yeah.

**Brendan Fortuner** [8:50]
Yeah.

**Wix** [8:51]
That, that one is more e-elicitation and, like, when, uh, being quand. Uh, but, like, did you have examples where it went wrong and, like, what would be your advice there, you know? Like, that, that can help others try to do the same thing.

**Brendan Fortuner** [9:03]
Yeah, I couldn't-- I think the interesting discussion is maybe-- I mean, there's many things to talk about. The reward hacking, I think, is a thing. I think in the physical exam use case, we can kind of talk about that, but we used an LLM grader for the first time.

And whenever you're using, like, an LLM grader, the task is, like, a little bit more prose or a little longer form generation. You could be very vulnerable to this. The models are super clever. They're incentivized to win, but they'll cheat, and they'll do weird things.

**Wix** [9:27]
Mm.

**Brendan Fortuner** [9:27]
So yeah, maybe I'll, maybe I'll define, like, define what is a physical exam and, like, how, kind of the reward hacking that we saw. So in medicine, in the room with the patient, the doctor will do this thing called a physical exam.

They'll look at your elbows, they'll test your reflexes, they'll listen to your lungs. At the end of the visit, they go back to their desk, and they fill out this, like, structured form. But for each body system, what do they observe?

You know, is the heart, you know, good or not? And so we frame this as an RFT task. We had the model generate this structured data of a bunch of findings, right? And we used an LLM grader to kind of measure the accuracy.

We were optimizing for, like, precision and recall. You know, how many of the findings did it get? Did it hallucinate any findings? And we kind of saw, like, two things. The first was the model started to inflate the number of findings to kind of gain precision.

So it's like, "I'm gonna generate a bunch of the same thing redundantly, and I'm gonna get great precision," right? Uh, that was one thing. Um, so we had to kind of, like, update the grader to kind of constrain the model.

Um, the second thing, there was, like, this tone degradation. So the model started using these layman terms like, "Grandpa has a, has an unhappy heart." We would get crucified if we put that into a medical note, right? So what we did is in the grader, you know, in addition to just the content and, like, the semantic accuracy of what it's saying, we also started to add style.

And we kind of weight them, like, seventy-five, twenty-five, and over time you can kind of harness and, and, and, and get the reward hacking to go away. When-

### ICD-10 Coding

**Wix** [10:43]
What were the, some of the evals, especially, yeah, you mentioned the ICD coding, which is, from my understanding, how the insurance coding works. How did you eval whether or not the model was good?

**Brendan Fortuner** [10:56]
Yeah, it's a great question. So this is, this is the one, um, the ICD-10, it started as a blog post, right, on, on OpenAI's website. Um, and then it got picked up by the news media like CNBC and Inc, and, and my grandpa started texting me about it.

He's like: "Oh, like, you know, I saw you in the news. Like, when is AI gonna replace doctors?" Right? And I was like, "Gramps," like, "AI's not gonna replace doctors." So the task was ICD-10 coding. Um, this essentially, after the visit with the patient, the doctor has to go back to the EMR, and they have to select these things called ICD-10 codes.

They're essentially like an international system of, like, diseases and conditions that have these normalized codes, right? So you'll say in the visit, the doctor will say, like, "Oh, you have left ear pain," right? But the doctor later needs to go map it to this code, which is normalized.

There's about seventy thousand of them, and they look like H six five dot one nine five dot, you know, acute otitis media, comma, recurrent, comma, left, right? Like that's-- Doctors didn't go to medical school to do that, and they're not that great at it.

So what we did is we, we recruited about eighteen physicians out in the wild. We created this, like, you know, eval set annotated by four expert annotators each, you know, to kind of make sure that the quality was extremely high.

And, and the clinicians we-- using F1 score, were scoring, like, let's say around forty percent, right, on the F1 score, uh, which is surprisingly low, lower than we thought. We were able to use RFT to kind of hill climb and get that, you know, get a small model, O3-mini, um, up to around, like, fifty-seven percent, um- There's still plenty of hill to climb, but we just thought that was, like, a really interesting result and just early signs that this could be framed, you know, as an RFT task.

### Beyond Scribing

**Wix** [12:26]
Okay. So, like, what all else do you do, right? Like, so I, I think there's, there's a broad surface area of products, and I think there's a, there's an unlimited amount of demand for help. Like, what, what are you hearing from doctors?

Like, what do they need?

**Brendan Fortuner** [12:40]
Yeah. I think when you first get into a health system and into healthcare, like, you just learn how many problems are kind of unsolved because there hasn't been as much, like, startup innovation in that space. It's, it's very difficult to get in.

There's a lot of, like, HIPAA requirements and regulations, and health systems are traditionally very conservative buyers. And then you have the EMRs, which aren't making life easier. So now that we're getting in, like, our clinicians and our health system, you know, partners are telling us about all these different use cases, right?

So, like, for example, you know, we have, have, like, early prototypes of, like, a patient-facing agent. One of the problems after the doctor visit is patient, they, they don't take their medication. They don't even go pick it up.

They don't go to the labs. And this creates, like, hundreds of thousands of hours of phone calls and inbox messaging, right? And, and it's very time-consuming, very expensive. So Ambience, you know, is like, "Well, why don't we, like, use one of these very powerful voice agents to do this kind of more administrative task and just report back to the nurse," right?

That's an example. Um, there's another really interesting one actually is, you know, in certain specialties like oncology or cardiology, before they go into the visit with the patient, they could often spend, like, ten minutes or up to thirty minutes or sixty minutes looking at the chart.

They're gonna look at labs and images and, you know, other data about the past visit. So Ambience actually, you know, we have product where we, we call it pre-charting. We'll create this, like, summary of, like, here's what happened in the past notes for those visits to save them a bunch of time.

You know, we even create a visit agenda for them so they know, like, what questions to ask the patient. Um, and then, you know, make the visit more productive. Those are just two, but there's, like, dozens and dozens of use cases.

**Wix** [14:08]
Yeah. This patient side reminds me of the times that I've gone to the doctor about medicine and never t- never followed the plan, let's, let's say.

**Brendan Fortuner** [14:14]
That's right.

**Wix** [14:15]
Exactly. And it's, yeah, it's my fault, but, like, I, I could get some help.

**Alessio** [14:18]
I saw in the prep notes that we had, there was one worst story around burning 25K on a one grader when it comes to-

### RFT Economics

**Brendan Fortuner** [14:25]
Yeah.

**Wix** [14:25]
Code exam. Yeah

**Alessio** [14:25]
... exams.

**Brendan Fortuner** [14:26]
Yeah.

**Alessio** [14:26]
So would love to, to chat, chat through that.

**Brendan Fortuner** [14:29]
It's definitely relevant to, like, startups and, and, you know, if you're just toying around. Um, this gets into the cost thing. So it really depends on, you know, the dataset size. It depends on, you know, what type of grader that you're using, right?

String match would obviously be, like, less efficient. But with, like, SFT, let's say you're using the OpenAI, you know, to do some supervised fine-tuning, you'll probably have, like, you know, maybe a few thousand examples. The job takes a few hours.

It costs you, like, 100 bucks, right? With RFT, maybe you have, like, 100 examples. The job takes a few days, um, and it costs thousands of dollars. One way to burn money really fast is to use a very, very expensive grader.

So I think one of our first experiments we ran, it was, like, only 100 examples, right? We quickly burned, like, immediately 20K just on the grader alone. So just, like, a watch-out for, for, you know, uh, builders out there.

**Alessio** [15:12]
What, what's your observability stack for LLMs? Like, are you tracking the spend both on, you know... When, when you're, like, setting up an experiment, you know, what does that, the harness look like, so to speak?

**Brendan Fortuner** [15:25]
Yeah. Great question. So I think we use, like, a variety of techniques. One of the cool tools that we use internally is Braintrust. Um, I think they're doing some incredible work over there, building, like, a tool for domain experts.

They do give some built-in observability. They let you deploy on-premise. It's a fantastic technology. Um, highly recommend folks checking that out. Um, but there's a lot of, like, areas where that stops, where you kind of have to build some custom tools, right?

It could be something like data mining tools or automated ring release tools, right, to release safely and some, you know, maybe automated monitoring tools. And we've had to build a lot of stuff in, in-house as well.

**Wix** [15:56]
Uh, we had, uh, Ankur from Braintrust on the podcast, and then he's speaking at the conference next week. What is... What are you looking for in eval tooling? Is there something that is really, really good for you that you want to shout out?

And maybe what's missing?

**Brendan Fortuner** [16:08]
Really good question.

**Wix** [16:09]
Wow. Chain of thought there.

**Brendan Fortuner** [16:11]
I'm just reasoning. Yeah, so I think it really depends on, it depends on the task. You know, if you have a very objective task and it's, it runs as a unit test or a string match, like, a script is actually quite good, and engineers are very good at automating scripts.

I haven't seen, you know, traditional issues on that side. But if you're gonna have h- like, domain experts who are non-programmers, if they're gonna review, that's actually kind of challenging. It's not obvious how to do that. Um, there's ergonomics is a big issue.

Like, they need to see the outputs side by side or multiple side by side. I think Braintrust is making tremendous progress there, right, and even allowing you to kind of incorporate some automation. Um, but I still think it's very early days, you know, in the ergonomics, you know, the full IDE for a domain expert, like, early days.

**Wix** [16:50]
Yeah. I'm curious what-- whether the domain experts really use these kinds of tools or do they need something else? That's, it's something that everyone wants, but I haven't talked to enough domain experts to know.

### Clinicians & MLEs

**Brendan Fortuner** [17:02]
Yeah. Yeah, I think this, this is, like, the, the billion-dollar question, right? I'd love to hear what you, you all think. Do we need machine learning engineers anymore, or, like, can you just train with domain experts alone? Um, I, I think the answer is both.

I think, I think the answer is both. I think the domain experts, like in our case, clinicians, they're really good at, like, debugging model outputs, meeting with users, distilling that feedback into something actionable, maybe annotating or, or doing evals.

But they don't necessarily have, like, you know, the right intuitions about what techniques to even try in the first place, and I think that's where machine learning engineers and researchers come in, and they kind of guide, guide the domain expert teams.

They see something manual, they do it 10X faster. That's kind of how I see the interplay of these teams working.

**Wix** [17:42]
Yeah, I think, like, there's the sort of test environments, uh, where you sort of... Let's say you hire domain experts, like clinicians, to be enablers for you, like a Scale AI or whatever, but, like, also that's pretty artificial, and I think you want, like, demonstrated behavior from just day-to-day work.

Like, that's kind of where... Like, there's, like, kind of like a copilot, like, experience, right? Like, you don't take the, just the raw thumbs up and down. Like, you, you actually see, like, how much of that was useful and stick around.

Yeah. I, I don't, I don't know how to, uh, design it beyond that, but I am looking for, like, okay, like, no, you don't need any of that. You know, like, here's the shortcut.

**Brendan Fortuner** [18:19]
Yeah. I don't think, I don't think there's any silver bullet. Um, I think it's a combination of like automation, but still like humans reviewing output's fundamental. You know, fundamental. The, the sniff test, the vibe check, like huge part.

### Model Robustness

**Wix** [18:33]
Okay. Um, s-segments on hallucination, I don't think we've like dealt with it specifically. Obviously, I think a, a number one question that people have on medical use cases of LLMs is what do you do with hallucinations?

**Brendan Fortuner** [18:46]
Yeah, I think this is interesting. I have, I have some takes here. So I think like base models out of the box, they actually, they really still struggle in healthcare in many ways. One of the issues is this like robustness problem.

So the models are like geniuses, right? But they can still like erratically fail a non-trivial percentage of the time and, and like fumbling agents, it doesn't work in healthcare. This is a patient safety critical environment, right? I think one of the problems is like the benchmarks, they always report like the best event.

So they report like, you know, what's the best metric they got out of sixty-four attempts? But in healthcare, we're more interested in like, what's the worst event, right? Um, and I think like, you know, Karan and team, uh, you know, at Health Bench, at OpenAI are doing like wonderful work here.

They're doing some great work at Stanford. But we really need to start to think about that. If you plot out the worst event for a lot of benchmarks, you're gonna be like really surprised. Like as an example, another thing we see a lot in the base models, and this is why they require like a lot of fine-tuning, patient-- you know, they start to make these medical inferences.

They're so smart, but they start to like infer things that the doctor didn't actually explicitly say. You know, for instance, a patient will say like, "I'm feeling sad and, and stressed out, difficulty sleeping," right? Uh, you know, the base model is gonna be like, "Patient reports increased anxiety and depression."

But wait a second, that-- those are diagnoses, right? The doctor didn't actually confirm those things. The models are gonna jump to these conclusions in unsafe ways. I think the, the last thing I'll say about this is I, I actually think clinical, real world clinical data is out of distribution.

I think as much as the models generalize, if you have no access to that data, it's really hard to learn. I think that the reasons are maybe twofold. The first is a lot of realistic clinical data is inside these walled gardens, right?

It's inside Epic and Cerner, and they're locked down for patient privacy reasons. But this means there's not a lot of realistic clinical data on the internet, right, to feed into the models. The second I'd say is actually like, you know, there's a lot of tribal knowledge in medicine.

So the, the, the base models, they've went to medical school, they've memorized the textbooks, but there's also this thing called residency. It's like two, four, eight years where you get all this hands-on experience, right? Working with, you know, attendings and there's a lot of learning that happens there that's encoded in the EHR, but not actually encoded on the internet.

**Wix** [20:48]
Um, yeah. That actually is the same analogy that people use for the coding models. We did an episode with the OpenAI Codex team and they were like, "Yeah, you, you went to college and you have a-- you, you have basic knowledge of software, but here we need...

are giving them like the first few months, maybe the first year of being on-the-job training."

**Brendan Fortuner** [21:07]
Yeah, exactly.

**Wix** [21:07]
Yeah.

**Brendan Fortuner** [21:08]
Exactly. It's so huge.

**Alessio** [21:11]
Do you feel like things like, uh, Health Bench, uh, that OpenAI announced are steps in the right direction to maybe fix some of these like, uh, you know, model evaluation lacking? Or do you think like the same issue is-

**Wix** [21:21]
Also what is-- Uh, so I haven't read Health Bench. What, what-- like maybe you could also squeeze in an explanation of what it does.

**Brendan Fortuner** [21:26]
Yeah. So, so Health-- It's a new, um, open source dataset released by OpenAI. You know, Karan and, and a lot of amazing researchers and participants kind of built this dataset. I think it has a few thousand examples. And these are like very realistic kind of healthcare tasks that they've put together and, and given out to the community, um, which I think is phenomenal.

And I think like they're thinking about it the right way. They're thinking about it like these academic datasets that we've been kind of saturating for a long time are no longer useful. Models are gonna ace medical exams, and they're just trying to figure out, how can we create these real realistic, like, clinical cases?

And I think that's like the first step, and there's still a big hill to climb, and that's great to see.

**Wix** [22:00]
Yeah, I mean, broadly on data sharing, I don't know if you've come across Tanishq Abraham. He used to work on MedArc, um, at, uh, with Stability AI, and now he's working on Sofont, which is kind of the, the spin-out version of it.

### Data Quality

**Brendan Fortuner** [22:11]
Mm-hmm.

**Wix** [22:12]
And he ac- we actually had a really interesting conversation around like sometimes we, we don't mind like anonymi- anonymized data sharing if we can train this more natively. We just need more data. Like just, like this is obviously the single most privacy sensitive issue for data on planet Earth.

At the same time, we also need a lot of data, man. Like what are you gonna do?

**Brendan Fortuner** [22:32]
Yeah. This is, this is the billion-dollar question. I can also share some more anecdotes. Like once you get that data, I didn't tell you it was high quality data. In fact, like it's actually-

**Wix** [22:41]
Yeah. Patients lie

**Brendan Fortuner** [22:42]
... tremendously messy.

**Wix** [22:44]
Patients lie. Like watch House.

**Brendan Fortuner** [22:46]
Yeah. It's tremendously messy. Um, you know, and, and I think that creates like new problems when you even start to get that data. And I think the, the reasoning models are tremendous when you give them really accurate context.

But imagine giving them like messy kind of context and they have to have taste, they have to have clinical taste, and that's something we're still learning about as an industry.

**Wix** [23:03]
What is clinical taste? Just judgment of what's right, what's wrong?

**Brendan Fortuner** [23:07]
Yeah, exactly. Like what source should you trust?

**Wix** [23:10]
Mm.

**Brendan Fortuner** [23:10]
It's what source should you trust? Similar to like, you know, if you ask, you know, o3 to go on the internet and find the right travel kind of destinations, it's just gonna kinda get like tricked by whatever has the best SEO.

And then in healthcare it's the same. Like, you know, what is an example of a EHR kind of data element that actually is messy? It's like you shouldn't trust this one. You should actually trust these other, you know, paragraphs over here.

So that's what I mean by like clinical taste and intuition.

**Wix** [23:33]
Have you had-- So sometimes reasoning on wrong premises can lead you to wrong results, obviously. And I wonder if you've had conversations with OpenAI on are there different kinds of reasoning or is reasoning just a generally-- Is it just IQ or is there like seven different types of IQ and maybe OpenAI has spikes on math and coding IQ, but not so much on the medical IQ?

Like how, how distinct are they or how correlated are they?

**Brendan Fortuner** [24:00]
That's a really good question. I think-- I don't know how it works underneath the hood. I do know that when you kind of plot out base model performance on some medical tasks, like ICD-10 coding between like, you know, preverse generations and, and new reasoning generations, there's actually not like a big leap.

I think like what we're recognizing is, look, like, if you don't actually train on the distribution of data in these verticals, you're actually not gonna see gains. So I don't know what patterns are shared. I'm sure there's reasoning patterns and mental models, but I also think there are probably mental models, right, in healthcare specifically that these models need to learn before they can be effective.

**Alessio** [24:33]
Yeah. Thinking about patterns, also one ideas that we kinda talked about in the prep is like- So do you have Devin for clinical AI research, kind of like moving more towards in job assistance to like more long-running experimental things?

### Research Agent

**Alessio** [24:45]
Yeah. Can you share more about that?

**Brendan Fortuner** [24:47]
Yeah, I think the thought here is our engineering team is, like, seeing tremendous productivity gains, right, with coding assistant and coding agents. Why can't we do the same thing for clinical research? I think there's definitely a world where you have a voice agent talk with an end user, distill that feedback into, let's say, like a prompt or maybe a PRD for what we should be doing, and then you can have the agent, like, go mine the data, go look at the edits, distill that analysis, you know, run evals, even do annotation.

I feel like this is fully possible, um, and it's something I'm excited about. I'm not saying we built it yet.

**Alessio** [25:19]
Great. Yeah. Any parting thoughts, call to actions? Um, you know, you got the audience listening to you.

### Hiring

**Brendan Fortuner** [25:26]
Super happy to be here. I, I appreciate you both. I love what you're doing. Huge fan of the pod. Yeah. Thank you for the time.

**Alessio** [25:31]
Yeah. Thank you, Brendan.

**Wix** [25:34]
Y'all hiring?

**Brendan Fortuner** [25:34]
Absolutely hiring. Hiring engineers, clinician researchers, machine learning engineers, machine learning researchers, you betcha.

**Wix** [25:42]
Maybe let's focus on a specific role and why it's what you're really looking for that is hard. What's a skill that pe- that you want that you cannot hire for?

**Brendan Fortuner** [25:49]
This is a good question. I do think there's machine learning research talent, the best in the world, that's definitely hard to get but super valuable as we start to push some of these frontier use cases. That's number one.

I also think there's, like, a really interesting skill set in, like, a clinician researcher archetype. So these are the folks with the domain knowledge, but how can you find the ones that also have, like, an experimentalist mindset, that have, like, a startup work operating system?

It is-- That's definitely surprisingly hard, and I think that's, you know, maybe one of the gold kind of unicorn kind of archetypes that's, that's really important to us.

**Wix** [26:20]
So like a, it's like a Venn diagram of startup person and also somewhat domain expert, somewhat ML person, right?

**Brendan Fortuner** [26:26]
Exactly.

**Wix** [26:26]
Like those three.

**Brendan Fortuner** [26:28]
Yep.

**Wix** [26:28]
All right. Well, I think, like, being specific about that intersection helps you hire more because you're more cle- you're more clear about, like, you know, this is, this is what we're doing. So I, I wanted to elicit that out of you.

But thank you for all the important work you're doing. I'm sure you're saving lives. Not all of us can say that, so...

**Alessio** [26:41]
Thank you, Brendan.

**Brendan Fortuner** [26:42]
Cool. Thank you both. Keep up the great work.

---

This library is powered by PodHood (https://podhood.com), the podcast website platform.
