# Why Compound AI + Open Source will beat Closed AI — with Lin Qiao, CEO of Fireworks AI

Latent Space · 2024-11-25

<https://addtry.com/775b1a85-ebc4-44c2-a265-19ce0591b11a>

Lin Qiao, CEO of Fireworks AI, argues that compound AI systems combining multiple open-source models will outperform closed monolithic models like OpenAI's. She explains Fireworks' evolution from a PyTorch platform to a full-stack inference and customization engine serving 40+ customers including Cursor. Qiao details their distributed inference engine, Fire Optimizer for quality-latency-cost tradeoffs, and upcoming o1-like model built on open-source foundations. She defends their quantization approach after public criticism from rivals, emphasizes specialization over general intelligence, and invites developers to test their free LoRA adapter hosting.

## Questions this episode answers

### What did Fireworks CEO Lin Qiao announce about their new reasoning model inspired by OpenAI o1?

Fireworks CEO Lin Qiao announced a new model inspired by OpenAI's o1, designed as a declarative system to close the quality gap. It is already on LMSYS, with users speculating it might be from Meta or Gemini. Lin acknowledges OpenAI's team caliber but argues inference scaling laws are the future, and their approach leverages expertise in inference optimization to approach OpenAI's quality without training from scratch.

[29:38](https://addtry.com/775b1a85-ebc4-44c2-a265-19ce0591b11a?t=1778000)

### Why does Fireworks AI believe compound open-source AI systems will beat closed AI models?

Fireworks CEO Lin Qiao defines compound AI as integrating multiple specialized models across modalities with APIs and databases to overcome hallucinations and knowledge gaps in single models. He argues open-source expert models will beat closed-source generalists, like human specialization, and Fireworks' platform simplifies composing such systems while optimizing quality, latency, and cost.

[14:04](https://addtry.com/775b1a85-ebc4-44c2-a265-19ce0591b11a?t=844000)

### How did Fireworks AI respond to being called out by competitors over quantization quality?

Lin was caught off guard when a competitor publicly called out Fireworks' quality, calling it unfair. She explains Fireworks offers various quantization schemes tailored to different workloads, balancing quality, latency, and cost. After third-party benchmarks like AA's evaluation, Fireworks' quality appeared strong. The company wrote a blog post clarifying their approach, emphasizing one-size-fits-all comparisons are biased.

[46:45](https://addtry.com/775b1a85-ebc4-44c2-a265-19ce0591b11a?t=2805000)

## Key moments

- **[0:00] Intro**
  - [0:52] Fireworks CEO Lin Qiao recounts facing the Silicon Valley Bank run and accidentally deleting critical data in the company's first two years.
- **[2:09] PyTorch Origins**
  - [3:39] Lin Qiao explains that PyTorch began as a researcher-focused framework, ignoring production, which drove massive open-source adoption at Meta.
  - [4:59] Lin Qiao says it took five years to evolve PyTorch into a production-grade framework across Meta's vast AI use cases.
  - [6:36] Lin Qiao founded Fireworks after seeing the industry's pain in the AI-first transition, aiming to support companies beyond hyperscalers like Meta.
- **[6:59] Fireworks Genesis**
  - [7:11] Lin Qiao pivoted Fireworks from a PyTorch cloud to a GenAI platform after the launch of ChatGPT in late 2022.
  - [8:36] Lin Qiao bets on inference over training for GenAI, citing that inference scales with world population while training scales with researcher count.
  - [9:23] Fireworks launched its public platform in August 2023 as a distributed inference engine with OpenAI-compatible API.
  - [11:56] Swyx notes that OpenAI's API format has become the industry standard, with even Google's Gemini adopting it.
  - [12:08] Lin Qiao sees two competing standardization forces: OpenAI's API format and Meta's Llama Stack for open-source models.
  - [13:02] Lin Qiao acknowledges that Llama Stack's success depends heavily on community adoption, and Fireworks is providing feedback to Meta.
- **[13:52] Compound AI**
  - [14:16] Lin Qiao explains Fireworks' shift to compound AI, from a single inference engine to a multi-product platform with optimization and customization.
  - [17:00] Fireworks expands its model catalog to include audio transcription, vision models, embeddings, image generation, and soon text-to-video with Mochi.
  - [18:29] Lin Qiao argues that compound AI systems are necessary to solve real-world problems, combining multiple models with APIs to reduce hallucination.
- **[20:27] Inference Engine**
  - [20:44] Lin Qiao details Fireworks' distributed inference engine: it chops models into pieces, scales across GPUs and regions, and adapts to workload bottlenecks.
  - [22:59] Swyx clarifies Fireworks' value: developer-friendly endpoints with custom inference kernels like Fire Attention for most large language and vision models.
  - [24:08] Lin Qiao argues Fireworks' competitive edge lies in delivering low latency and low cost for interactive GenAI applications, optimizing for developers.
  - [26:32] Fireworks' Fire Optimizer continuously improves quality, latency, and cost by customizing inference setups for each customer's workload.
  - [29:14] Lin Qiao announces a new Fireworks model on LMSYS that replicates OpenAI o1's reasoning quality, sparking speculation it's from Google or Meta.
- **[29:38] New Model**
  - [31:59] Lin Qiao agrees that the scaling law is shifting from training to inference, and Fireworks' early investment in inference optimization pays off.
  - [32:40] Lin Qiao predicts open-source models will close the quality gap to closed models through specialization, already excelling in coding and function calling.
  - [33:55] Bitter lesson vs specialization: Lin Qiao argues expert open-source models will outdo generalist closed models, drawing an analogy to human specialization.
  - [35:35] Fireworks is still debating whether to expose reasoning traces for their new o1-like model, balancing transparency and intellectual property.
  - [36:51] Lin Qiao shares early tests of their o1-like model: handling deep medical queries and generating complex reasoning DAGs when asked to define AGI.
- **[38:39] Team & Culture**
- **[40:38] Cursor Partnership**
  - [40:38] Lin Qiao praises Cursor for choosing to partner with Fireworks for inference, pushing them to develop a more intense global inference stack.
- **[46:22] Quantization Drama**
  - [46:22] "It's not a good way to compete... we want to compete fairly."
  - [49:20] Lin Qiao responds to price wars in open-source model hosting, asserting Fireworks aims for margin by delivering differentiated value beyond raw model access.
  - [49:55] Lin Qiao highlights Fireworks' cost advantage: focusing on inference optimization without the enormous training cost amortization that closed model APIs face.
- **[51:38] Underrated Features**
  - [51:38] Lin Qiao highlights Fireworks' Multi-LoRA feature, allowing users to upload LoRA adapters and serve them at the same cost as the base model.
  - [54:02] Lin Qiao invites developers to join Fireworks' Discord for early access to new features, direct feedback, and upcoming office hours.
- **[55:36] Closing**

## Speakers

- **Alessio** (host)
- **Swyx** (host)
- **Lin Qiao** (guest)

## Topics

Inference, Language Models

## Mentioned

Fireworks (company), Hugging Face (company), Lambda (company), Lepton (company), Meta (company), OpenAI (company), Replicate (company), RunPod (company), Together (company), Cursor (product), Fire Attention (product), Fire Optimizer (product), Firefunction (product), Llama (product), Llama Stack (product), Mochi (product), Multi-LoRA (product), PyTorch (product), TensorRT (product), o1 (product)

## Transcript

### Intro

**Alessio** [0:04]
Hey, everyone. Welcome to the Late in Space podcast. This is Alessio, partner and CTO at Decibel Partners, and I'm joined by my co-host Swyx, founder of Small AI.

**Swyx** [0:12]
Hey, and today we're in a very special, uh, studio inside the Fireworks office with Lin Qiao, CEO of Fireworks. Welcome.

**Lin Qiao** [0:19]
Yeah.

**Swyx** [0:20]
Oh, you should welcome us.

**Lin Qiao** [0:20]
Yeah, welcome. You both.

**Swyx** [0:24]
Uh, yeah, thanks for having us. Uh, it's, uh, unusual to, to be in the home of a startup, but, like, it's also-- I think our relationship is a bit unusual with, compared to all, all our normal guests.

**Lin Qiao** [0:34]
Definitely. Definitely. Yeah, it's-- I'm super excited to, to talk about very interesting topics in that space with both of you.

**Swyx** [0:41]
You just celebrated your two-year anniversary yesterday.

**Lin Qiao** [0:43]
Yeah. It's quite a crazy journey. We circle around and share all the crazy stories across these two years, and, uh, it has been super fun.

**Swyx** [0:52]
Yeah.

**Lin Qiao** [0:52]
Uh, all the way from we experienced Silicon Valley bank run-

**Swyx** [0:57]
Right.

**Lin Qiao** [0:58]
To we delete some data that shouldn't be deleted

operationally. We went through massive scale where we actually are busy getting capacity to... Yeah, we, we learned to kinda work with it as a team, uh, with a lot of brilliant people across different places joining the company. Uh, it has really been a fun journey.

**Alessio** [1:25]
When you started, did you think the technical stuff would be harder or the bank run and then the people side? I think there's a lot of, like, amazing researchers that wanna do companies, and it's like the hardest thing is gonna be building the product, and then you have all these different founder things.

So, uh, were you surprised by it? What has been the, your experience the last years?

**Lin Qiao** [1:43]
Yeah. To be honest with you, like, my focus has, has always been on the product side, and then after product, go to market, and I didn't realize the rest has been so complicated, operating a company and so on.

But because I don't think about it, I just kinda manage it, so it's done. So I, I think I, I just somehow, like, don't think about it too much and, you know, solve whatever problem coming our way, and it worked.

**Swyx** [2:09]
So let's-- I guess let's start at the pre-history, like the pre, the, the sort of, uh, initial history of Fireworks. You ran the PyTorch team at Meta for, uh, a number of years. And, uh, we previously had Sumith Chintala on, and, uh, I think we're just all very interested in, like, the, the, the history of GenAI.

### PyTorch Origins

**Swyx** [2:26]
Maybe not that many people know, like, how deeply involved FAIR and Meta were in-- like, prior to the current GenAI revolution.

**Lin Qiao** [2:35]
Mm. Uh, my background is, uh, deep in distributed system, database management system, and I joined Meta from the data side, and I saw this tremendous amount of data growth, which cost a lot of money, and we're analyzing what's going on.

A-and it's clear that AI is driving all this data generation. So it, it's a very interesting time because when I joined Meta, Meta is going through ramping down mobile first, finishing the mobile first transition, and then starting AI first.

And there's a fundamental reason ab-about that sequence because mobile first gave a full range of user engagement that was, has never existed before. And all this user engagement generate a lot of data, and this data power AI. So then the whole entire industry is also going through, falling through, uh, the same transition.

When I see, oh, okay, this AI powering all this data generation and, and look at, uh, where's our AI, AI stack, there's no software, there's no hardware, there's no people, there's no team. I'm like, "I, I want to dive it up there and, and help this, uh, this movement."

So when I started, it's very interesting industry landscape. There are a lot of AI frameworks. It's a kind of proliferation of AI frameworks, uh, happening in the industry. But all the AI framework focus on, like, production, um, and they use a very certain way of defining the graph o-of neural network a-and they use that to drive, uh, the model activation and, uh, productionization.

And PyTorch is completely different. So they could also assume it, and he was the user of his product, right? And he, as a researcher, faced so much pain using existing AI framework. This, uh, this is really hard to use, and I'm gonna build something different, uh, for myself.

And that's the origin story of PyTorch. PyTorch actually started as the framework for researchers. Don't care about production at all. And as it grow in terms of adoption-- Uh, so the interesting part of AI is research is the top of funnel for production.

There are so many researchers across academic, across industry, and they innovate, and they put their results out there in open source, and that power the downstream productionization. So it's brilliant for Meta to establish PyTorch as a strategy to drive massive adoption in open source because Meta internally is a PyTorch shop, so it's, it create a flying wheel effects.

So that's kind of the strategy behind, uh, PyTorch. But when I, uh, took on PyTorch, it's kind of at the cusp of Meta established PyTorch as the, the framework for both research and production. So no one has done that before, and we have to kind of rethink how to architect PyTorch so it can really sustain production workload, the stability, reliability, low latency, all this production concern was never a concern before.

Now it's a concern, and we actually have to adjust its design and make it work for both sides. And that took us five years, uh, because Meta has so many AI use cases, all the way from ranking recommendation that's powering the business top line, or ads ranking newsfeed, video ranking to site integrity, detect bad content automatically using AI, to all kind of effects, translation, image classification, object detection, all this.

And also across, um, AI running on the server side, on mobile phones, on AI VR devices, wide spectrum. So by the time, um, we actually basically managed to support AI across ubiquitous everywhere across Meta. But interestingly, through open source engagement, we work with a lot of companies.

It is clear to us, like, this industry is start to Take on AI first transition. And of course, Meta as a hyperscaler always go be ahead of the industry, and we feel like-- it feels like when we start this AI journey at Meta, there's no, no software, no hardware, no team.

For many companies we, we engage with, uh, through PyTorch, we feel the pain. That's the genesis why we feel like, hey, if we, uh, create Fireworks and support the industry going through this transition, it will be a huge amount of impact.

Of course, the problem that the industry is facing will not be the same as Meta. Meta is so big, uh, right? So it's kind of skewed towards extreme scale and extreme optimization. The industry will look different. But we feel like we have the technical, uh, chop and we have seen a lot.

Would li- love to kind of, uh, drive that. Uh, so yeah. So that's the... That's how we started.

**Swyx** [6:59]
When you and I chatted about, like, the origins of Fireworks, it was originally envisioned more as a PyTorch platform, and then later became much more focused on generative AI. Is, is that, is that fair to say? Like a-

### Fireworks Genesis

**Lin Qiao** [7:11]
Right. So-

**Swyx** [7:12]
What was the customer discovery here?

**Lin Qiao** [7:13]
Right. So I would say our initial blueprint is, say, we should build a PyTorch cloud because a PyTorch is library and there's no SaaS platform to enable AI workloads. Like-

**Swyx** [7:26]
Even in 2022, it was existing?

**Lin Qiao** [7:29]
I would not say absolutely no, but, uh, like cloud providers have some of those, but it's not first class citizen, right? Because at 2022, there are still, like, TensorFlow, is massively in production and, uh, this is all pre-GenAI, and the PyTorch is kind of getting more and more adoption, but there's no PyTorch first, um, SaaS platform existing.

At the same time, we are also a very pragmatic set of people. We really want to make sure from the get go, we get really, really close to customers. We understand their use case, we understand their pain points, we understand the value we deliver to them.

So we want to take a different approach. Instead of building horizontal PyTorch cloud, we want to, uh, build a verticalized platform first. And then we talk with many customer. And interesting, uh, we start a company September 2022, and, uh, October, November, then OpenAI announced ChatGPT.

And then boom, then when we talk with many customer, they are like, "Can you help us, uh, working on the GenAI as- aspect?" Right. So of course it-- there are some open source models. It's not as good at that time, but people are already, like, putting a lot of attention there.

Then we decide, uh, that if we're gonna pick a vertical, we're gonna pick GenAI. The other reason is all general model are PyTorch models. Um, so that's another reason. We believe that because of nature of GenAI, it's gonna generate a lot of human consumable content.

It will drive a lot of consumer, prosumer, developer-facing application and product innovation, guaranteed, right? We're just at the beginning of this. Our prediction is for those kind of applications, the inference is much more important than training because inference scale is proportional to the up limit our world population.

And training, training scale is proportional the amount of researchers, of course, each training round could be, uh, very expensive. Although PyTorch support both inference and training, we decide to laser focus on training-- uh, inference. So yeah, so that, that's how we got started.

And we launched our public, uh, platform August last, last year. And when we launch, it's a single product. It's a distributed inference engine with simple API, OpenAI-compatible API with the many models. We started with LLM, and later on we add a lot of model.

Fast-forward to now, we are a full platform with multiple product lines. So we'd love to kind of dive deep into what we offer. Uh, so but that's a very fun journey in the, in the past two years.

**Alessio** [9:50]
What was the transition from you start focus on PyTorch and, like, people want to understand the framework, get it live, and now I would say maybe most people that use you don't even really know much about PyTorch at all.

You know, they're just trying to consume a model. From a product perspective, like what were some of the decisions early on? Like right in October, November, you were just like, "Hey, most people just care about the model, not about the framework.

We're gonna make it super easy." Or was it more a gradual transition to the model library you have today?

**Lin Qiao** [10:17]
Yeah. So our product decision all based on who is our ICP. And, uh, one thing we want to acknowledge here is the GenAI technology is disruptive. It's very different from AI before GenAI, so it's a clear leap forward.

Because before GenAI, the companies that want to invest in AI, they have trained from scratch. There's no other way. There's no foundation model. Doesn't exist. So that means they need to start a team, first hire a team who is capable of crunch data.

There's a lot of data to crunch, right? Because training from scratch, you, you have to prepare a lot of data. And then they need to have, uh, they need to have GPUs to train, uh, and then you need to start to manage GPU.

So then it becomes a very comp-complex project. Uh, it takes a long time and not, not many company can afford it, actually. And the GenAI is a very different game right now because it is a foundation model, so you don't have to train anymore.

That make AI much more accessible as a technology. As an app developer or product manager even, not a developer, they can interact with GenAI models directly. So... And our goal is make AI accessible to all app developers and product engineers.

That's our goal. So then getting them into the building model doesn't make any sense anymore with this new technology, and then building easy accessible API is the most important. Our, um... Early on when we got started, we decided we're gonna be OpenAI compatible.

It's just kind of very easy for developers to adopt this new technology, and we will manage the un- underlying complexity of serving all these models.

**Swyx** [11:56]
Yeah, OpenAI has become-

**Lin Qiao** [11:58]
The standard.

**Swyx** [11:59]
... the standard. Uh, even to-- as we're recording today, Gemini announced that they have OpenAI compatible APIs.

**Lin Qiao** [12:05]
Interesting.

**Swyx** [12:06]
So then we just need Anthropic to fall in line, and, and we have everyone fall in line.

**Lin Qiao** [12:08]
Yeah. That's interesting because, um, we are working very closely with Meta as one of the partners, and, um, Meta announced... Meta, of course, is kind of very generous to donate many very, very strong open source models, expecting more to come.

But also they have announced Llama Stack-

**Swyx** [12:25]
Yeah

**Lin Qiao** [12:25]
... uh, which is, uh, basically Standardize the upper level stack built on top of Llama models, so they don't just want to give out models and you figure out what the upper stack, they instead want to build a community around the stack and then build, build a kind of new standard.

I think there's, there's interesting dynamics in play in, in the industry right now. One is more like standardized across OpenAI because they are kind of creating the top of funnel or standardized across Llama because this is the, like, mostly used open source model.

So I, I think really a lot of fun working at this time.

**Swyx** [13:02]
I've been a little bit more doubtful on Llama stack. I think you've been more positive. Basically, it's just like the Meta version of whatever Hugging Face offers, you know, or TensorRT or BLM or whatever the open source opportunity is.

But like, I... It... To me, it's not clear that just because Meta open sources Llama, that the rest of Llama stack will be adopted, and it's not clear why I should adopt it.

**Lin Qiao** [13:26]
Mm-hmm.

**Swyx** [13:27]
So I, I don't know if you have-

**Lin Qiao** [13:27]
Yeah. It's very early right now. That's why kind of we're, we'll work very closely with them and give them feedback. The feedback to the Meta team is very important, so then they can use that to continue to improve the model and also improve the higher level stack.

I think the success of Llama stack heavily depend on the community adoption, and there's no way around it. And, uh, I, I know Meta team would like to kind of work with a broader set of community, but it's very early.

### Compound AI

**Swyx** [13:52]
One thing that, uh, after your Series B, so you, you raised from Benchmark and then Sequoia, I remember being close to you for at least your, your Series B announcement. You started betting heavily on this term of compound AI.

**Lin Qiao** [14:04]
Mm-hmm.

**Swyx** [14:04]
It's not a term that we've covered very much in the podcast, but, uh, I think it's definitely gaining a lot of adoption from, like, Databricks and the Berkeley people and all that. What's your take on compound AI? Why is it resonating with people?

**Lin Qiao** [14:16]
Right. So, uh, let me give a little bit context why, why we even consider that space.

**Swyx** [14:22]
Yeah.

**Lin Qiao** [14:22]
Uh, so we-

**Swyx** [14:23]
Because, like, pre-Series B you were not... There was no mention.

**Lin Qiao** [14:25]
Yeah.

**Swyx** [14:25]
And now it's, like, on your landing page.

**Lin Qiao** [14:28]
So it's kind of very organic evolution from when we first launched our public platform, we are single product and we are distributed inference engine, where we do a lot of, uh, innovation on customized CUDA kernels, ROCm kernels, uh, running on different kind of hardware, and, uh, build distributed disaggregated execution, inference execution, build all kind of caching.

And so that is one. So that's kind of one product line is the fast, most cost-efficient inference platform. Because we wrote PyTorch code, we know-- we basically have a special PyTorch build for that, uh, together with a custom kernel we wrote.

And then we work with many more customer, we realized, oh, the distribution inference engine is... Our design is one size fits all, right? We want to have this inference endpoint, then everyone come in and they... No matter what kind of form and shape or workload they have, it would just work for them, right?

So that's great. But the reality is they-- we realized all customer have different kind of use cases. The use case come in all different form and shape, and, uh, the end result is the data distribution in their inference workload doesn't align with the data d- distribution in the training data for the model, right?

It's a given, actually, if you think about it, because researcher has to guesstimate what is important, what's not important during, like, in preparing data for training. So because of that misalignment, then we leave a lot of quality, latency, cost improvement on the table.

So then we are saying, okay, we want to heavily invest in a customization engine, and we actually announced it called Fire Optimizer. So Fire Optimizer basically help user navigate a three-dimensional optimization space across quality, latency, and cost. So it's a three-d- dimensional curve.

And even for one company, for different use case, they want to land in different spot. So we automate that process for our customer. It's very simple. You have your inference workload, and you inject, um, into the optimizer along with a objective function, and then we spit out a inference deployment, uh, config and the model set up.

So, like, it's your customized setup. So that is a completely different product. So that product thinking is one size fits one, different from one size fits all. And now on top of that, we provide a huge variety of state-of-art models, uh, hundreds of them varying from text to large, state-of-art large language models.

That's where we started. And as we talk with many customer, we realized, oh, audio and, uh, text are very, very close. Many of our customers start to build assistant, all kind of assistant using text, and they immediately want to add audio, audio in and audio out.

So we support transcription, translation, speech synthesis, text-audio alignment, all different kind of audio features. It's a big announcement. We're gonna... You should have heard-

**Swyx** [17:25]
By the time this is out.

**Lin Qiao** [17:25]
By the time this is out.

**Swyx** [17:26]
Yeah.

**Lin Qiao** [17:27]
And the other area is vision, and the text are very close with each other because a lot of information doesn't live in plain text. A lot of information live in multimedia format, live in images, PDFs, screenshots, and many other different formats.

So oftentimes to solve a problem, uh, we need to put the vision model first to extract information and then use language model to process and then send out results. So vision is important. We also support vision model. Various different kind of vision models specialize in processing different kind of source a- a- and extraction.

And we're also gonna have another announcement of a new API endpoint with support for people to upload various different kind of multi- multi-media content a- and get the, uh, extract very accurate information out and feed that in, into LLM.

And, and then of course we support embedding because embedding is very important for semantic search, for RAG, and all this. And in addition to that, we also support text to image, image generation models, text to image, image to image, and we're adding text to video as well in our portfolio.

So it's very comprehensive set of model catalog that build on, run on top of our optimizer and distribute inference engine. But then we talk with more customer, they solve business use case, and then we realize one model is not sufficient to solve their problem.

And it's, it's very clear because one is the model hallucinate. And many customer, they... when they onboard this GenAI journey, they thought, "This is magical. GenAI is gonna solve all my problems magically." But then they realize, "Oh, this model hallucinates."

It hallucinates because it's not deterministic, it's probabilistic. Um, so it's designed to always give a answer, but, uh, based on probability, so it hallucinates. And that's actually sometimes the feature for creative writing, for example. Sometimes it's a bug because, hey, you don't want to give misinformation.

And different model also have different specialties. To solve a problem, you want to ask different special model to kind of decompose your, your task into multiple small task, narrow task, and have a expert model solve that task really well.

And of course, the model doesn't have all the information. It has limited knowledge because the training data is finite, not infinite. So model oftentimes doesn't have real-time information, it doesn't know any proprietary information within enterprise. It's clear, you know, how in order to really build a compiling application on top of GenAI, we need a Compound AI system.

Compound AI system basically is gonna have multiple models across modalities along with APIs, whether it's public APIs, internal proprietary APIs, storage systems, database system, knowledge systems to work together to deliver the best answer.

**Swyx** [20:08]
Are you gonna offer vector database?

**Lin Qiao** [20:10]
We actually heavily partner with several big vector database providers.

**Swyx** [20:15]
Which is your favorite?

**Lin Qiao** [20:16]
Um, they are all great in different ways, but it's public information, like, MongoDB is our investor, and we have been working closely with them for a while.

**Alessio** [20:27]
When you say distributed inference engine, what do you mean exactly? Because when I hear your explanation, it's almost like you're centralizing a lot of the decisions through the Fireworks platform on, like, the quality and whatnot. When you mean distributors, like, you have GPUs and, like, a lot of different clusters or, like, you're sharding the inference across single-

### Inference Engine

**Lin Qiao** [20:44]
Right. Right, right, right. So, um, first of all, we run across multiple GPUs. But the way we distribute across multiple GPUs is, is unique. We, we don't distribute the whole model monolithically across multiple GPUs. We, we chop them into pieces and scale them completely differently b- based on what's the bottleneck.

We also are distributed across, uh, regions. Uh, we have been running in North America, EMEA, and Asia. We have regional affinity to applications because latency is extremely important. We are also, uh, like, doing global load balancing because a lot of application there, they quickly scale to global population.

And then at that scale, like, different continent wake up, wakes up at different time, uh, and you want to kind of load balancing across. So all the way... And we, we also have-- we manage various different kind of hardware SKU from different hardware vendors, and different hardware design is best for different type of workload, whether it's long context, short context, long generation.

So all these different kind of type of workload is best fitted for different kind of hardware SKU, and we can even distribute that across different hardware for a workload. So yeah, so the distribution actually is, is all around in the full stack.

**Swyx** [22:03]
At some point, we'll show on the YouTube the image that Ray, I think, has been working on with, like, all the different modalities that you offer. Like, to me, it's basically you offer the open source version of everything that OpenAI typically offers, right?

Like, I don't think there is... Actually, if you do text to video, you will be a superset of what OpenAI offers because they don't have Sora. Is that Mochi, by the way? Is, uh-

**Lin Qiao** [22:25]
Mochi.

**Swyx** [22:26]
Mochi, right?

**Lin Qiao** [22:26]
Yes.

**Swyx** [22:27]
Yeah.

**Lin Qiao** [22:27]
Mochi, and there are a few others. Like, there... I would say, uh, the interesting thing is I think we're betting on the open source community is gonna grow, like proliferate. This is literally what I'm seeing.

**Swyx** [22:40]
Yeah.

**Lin Qiao** [22:41]
And it's, there's amazing video generation companies.

**Swyx** [22:45]
Yeah.

**Lin Qiao** [22:45]
There is amazing audio companies. There... Like cross border, the innovation is off the chart, and we are building on top of that. I think that's the advantage we have, like, uh, compared with a closed source company.

**Swyx** [22:59]
I think I want to restate the value proposition of Fireworks for people who are comparing you versus, like, a raw GPU provider, like a RunPod or Lambda or, you know, anything like those, which is, like, you create the developer experience layer and you also, like, make it easily sort of scalable or serverless or, you know, as, as an endpoint.

And then I think for some models you have, uh, custom kernels, but not all models.

**Lin Qiao** [23:24]
For almost for all model, for all large language models, all the models-

**Swyx** [23:28]
So you just write kernels all day long?

**Lin Qiao** [23:29]
And the VLMs.

**Swyx** [23:32]
Uh, yeah.

**Lin Qiao** [23:33]
Yeah, yeah. Yeah. And almost for all models we serve, we have-

**Swyx** [23:35]
And, and so that is called Fire Attention?

**Lin Qiao** [23:37]
It's called Fire.

**Swyx** [23:38]
I don't remember the, the speed numbers, but apparently much better than VLM, especially on a concurrency basis.

**Lin Qiao** [23:44]
Right. So Fire Attention is specific for, um, mostly for language model, but for other modalities we'll also have a customized kernel.

**Swyx** [23:52]
Yeah, I think the, the, the typical challenge for people is understanding, like, that has value, uh, and then, like, the, there are other people who are also o- offering open source models, right? Like, your moat is, is your ability to offer, like, a, a good experience for all these customers.

But if your existence is entirely reliant on people releasing nice open source models, other people can also do the same thing.

**Lin Qiao** [24:14]
Right. Yeah. So I would say we build on top of open source model foundation, so that, that's the kind of foundation we build on top of. But we look at our, the value prop from the lens of application developers and product engineers.

So they want to create new UX. So what's happening in the industry right now is people are thinking about completely new way of designing products. And I'm talking to so many founders, it's just mind-blowing. They help me understand- Existing way of doing PowerPoint, existing way of coding, existing way of managing customer service.

It's actually putting a box in our head. For example, PowerPoint, right? So PowerPoint generation is we always need to think about how to fit into my storytelling into this format of slide one after another. And I, I'm gonna juggle through, like, design together with, you know, what story to tell, but the most important thing is what's our story- storytelling lines, right?

And why don't we create a space that is not limited to any format? And those kind of new product UX design combined with automated content generation through GenAI is the new thing that many founders are doing. What are the challenges they are facing, right?

Let's go from there. One is, again, because a lot of product built on top of GenAI, they are consumer, consumer developer facing, and they r-request interactive experience. It's just a kind of product experience we, we, we all get used to, and our desire is to actually get faster and faster interaction .

Otherwise, nobody want to spend time, right? So get... And then that requires low latency. And the other thing is the nature of consumer, consumer developer facing is the, uh, your audience is very big. You, you want to scale after product market fit quickly.

But if you lose money at a small scale, you're gonna bankrupt quickly. So it's, uh, actually a big contrast is I actually have product market fit, but when I scale, I scale out of my business, right? So that's kind of very, very funny, uh , way to think about it.

So then have low latency and low cost is essential for those new application and product to survive and really become a generational company. So that's the design point for our, uh, distributed inference engine and the file optimizer. File optimizer, you can think about that as a feedback loop.

The more you, uh, feed your inference workload to our inference engine, the more we help you improve quality, lower latency further, lower your cost. It basically becomes better. And, and we automate that because we don't want you as app developer or product engineer to think about how to figure out, uh, all these low-level details.

It, it's impossible because you are not trained for-- to do that at all. You should kind of keep your focus on the product innovation. And then the Compound AI, we actually feel a lot of pain as the app developers, engineer, they-- There are so many models.

Every week there's at least a new model coming out. Okay.

**Swyx** [27:10]
Tencent had a giant model this week. Like you guys-

**Lin Qiao** [27:13]
Yeah, yeah, I saw that. I saw that. Um-

**Swyx** [27:16]
Like five hundred billion parameters.

**Lin Qiao** [27:17]
Yeah. Um, so they're like, "Should I keep chasing this or should I forget about it," right? So and which model should I pick to solve? What kind of sub-problem? How do I even decompose my problem into those smaller problems and fit the model into it?

I have no idea. And then there are two way to think about this design, right? I think I talk about that in the past. One is imperative, as in, uh, you tell-- you figure out how to do it, right?

You give developer tools to how to detect how to do it. Or you build a declarative system where developer tells what they want to do, not how. So these are completely two different designs, right? So, like, analogy I want to draw is in the data world, the database management system is a declarative system because people use database, use SQL.

SQL is a way you say what do you want to extract out of database, what kind of result you want. But you don't figure out which node is gonna-- how many nodes you're gonna run on top of, how you read file in your disk, which index you use, which produ- You don't need to worry about any of those, and database management system will figure out, generate the new best plan, uh, and execute on that, right?

So database is declarative, and it makes it super easy. You just learn SQL, which is learn a semantic meaning of SQL, and you can use it. Imperative side is there are a lot of ETL pipelines, and people design this stack system of with triggers, with actions and, and you detect exactly what to do, and it fails, and then how to recover.

So that's the dec- imperative system. And we have seen a range of system in the ecosystem, like, go different ways. I think there are value of both. There are value both. I, I don't think one is gonna subsume the other.

But we are leaning more into the philosophy of build a declarative system because from the lens of app development and product engineer, that would easiest for them to integrate.

**Swyx** [29:08]
I understand that's also why PyTorch won- ... as well, right? That's some-- This is the one of the reasons-

**Lin Qiao** [29:13]
Ease of use

**Swyx** [29:14]
... people came.

**Lin Qiao** [29:14]
So yeah, focus on ease of use, and then let the, let the system take on the hard challenges and complexities. So we follow-- We expand that thinking into current system design. So another announcement is we will also announce a, our next declarative system is gonna be appear as a model that has extremely high quality.

### New Model

**Lin Qiao** [29:38]
And this model is inspired by o1 announcement from OpenAI. You should see that by the time we announce this or, you know, soon.

**Alessio** [29:46]
Trained by you?

**Lin Qiao** [29:47]
Yes.

**Alessio** [29:48]
Is this the first model that you trained in, like, this scale or-

**Lin Qiao** [29:52]
It's not the first. Uh, we actually have, uh, trained a model called Firefunction. It's a function calling model. It's our first step into Compound AI system because function calling model can dispatch a request into multiple APIs. We have pre-baked set of APIs the model learned.

You can also add additional APIs to-- through the configuration to let model dispatch accordingly. So we have a very high quality function calling model that a-already released. We have actually three versions. The latest version is very high quality.

But now we take a further step, that you don't even need to use function calling model. You use our, uh, new model we're gonna release. Uh, it will solve a lot of problem, uh, approaching very high, like, open AI quality.

So I'm very excited, uh, about that.

**Swyx** [30:41]
Do you have any benchmarks yet or-

**Lin Qiao** [30:43]
We have benchmark. We're gonna release it Uh, hopefully next week. We just put our model to LMSYS, and people are guessing is this a next Gemini model or a Meta's model? People are guessing. That's very interesting. We're, like, watching the Reddit, uh, discussion right now.

**Swyx** [31:00]
I mean, I have to ask more questions about this. When OpenAI released o1, a lot of people asked about whether or not it's a single model or whether it's like a, a chain of models, and Noam and basically everyone on, on the Strawberry team was very insistent that what they did for reinforcement learning chain-of-thought cannot be replicated by a whole bunch of open source model calls.

Do you think that that is-- they are wrong? Have you done the same amount of work on RL as they have, or was it a different direction?

**Lin Qiao** [31:30]
I think they take a very specific approach where I do-- The caliber of team is very high, right? So I, I do think they are the domain expert in doing the things they are doing, but I, I don't think there's only one way to achieve the same goal.

We're on the same direction in the sense that the quality scaling law is shifting from training to inference. We are definitely on the s-- For that, I fully agree with them, but we are taking a complete different approach to the problem.

All of that is because, of course, we didn't train the model from scratch. All of that is because we build on the shoulder of giants, right? So the current model available we have access to is getting better and better.

The future trend is the gap between the open source model, closed source model, it's just gonna shrink to the point there's not much difference, and then we're on the same level field. That's why kind of, I think our early investment in, in inference and, uh, all the work we do around balancing across quality, latency, and cost pay off because we have accumulated a lot of experience there, and that empower us to, to build and release this new model-

**Swyx** [32:36]
Yeah

**Lin Qiao** [32:37]
... that is, um, approaching OpenAI's quality.

**Swyx** [32:40]
I guess, like, the question is, what do you think the gap to catch up will be? Because I think everybody agrees with open source models eventually will catch up. And I think with four-- then with Llama three to three one four five B, we closed the gap, and then o1 just reopened the gap so much and it's unclear.

Obviously, you're saying your model will have-

**Lin Qiao** [32:58]
We're closing that gap

**Swyx** [32:58]
... similar. Yeah, but you think, you think like in the future it's gonna always gonna be like months?

**Lin Qiao** [33:02]
So here's the thing that's happened, right? There's public benchmark, it is what it is. But in reality, open source model in certain direct-- di-dimension already on par or beat closed source model, right? So for example, for-- in, in the coding space, open source models are really, really good.

And, uh, in function calling, like Fire function is also really, really good. So it's all a matter of whether you build one model to solve all the problem and you want to be the best of solving all the problems or in the open source domain, it's gonna specialize, right?

All these-

**Swyx** [33:35]
Yeah

**Lin Qiao** [33:35]
... different model builders specialize in certain ar-- narrow area and it's logical that they can be really, really good in that very narrow area. And that's our prediction is with specialization there will be a lot of expert models really, really good and even better than like one size fits all open source-- uh, closed source models.

**Swyx** [33:55]
Yeah. I, I think this is the, the core debates that I am still not one hundred percent either way on in, in terms of compound AI versus normal AI, 'cause you're basically fighting the bitter lesson.

**Lin Qiao** [34:10]
Look at the human society, right? We specialize and you feel really good about someone specializing doing something really well, right? And that's how our-- like when evolved foundation time, we're all generalists, we do everything-

**Swyx** [34:22]
Yes

**Lin Qiao** [34:22]
... in the tribe to now we have been specialized in different domain. So I-- my prediction is in the AI model space it will happen also.

**Swyx** [34:29]
Except for the bitter lesson. You al-- you get short-term gains by having specialists, domain specialists, and then someone just needs to train like a ten x bigger model on ten x more inference, ten x more data, ten x more model perhaps, whatever the, the current scaling law is and then it supersedes the-- all the individual models because of some generalized intelligence/world knowledge.

You know, like I think that, that is the, the core insight of the GPTs, the GPT one, two, three that was-

**Lin Qiao** [34:57]
Right. But the scaling law again, right, the, the training scaling law is because you have increasing amount of data to train from and you can do a lot of compute, right? So I think on the data side we are approaching the limit and the only data to increase that is entirely generated data.

And then there's like what is the secret sauce there, right? Because you have a-- if you have a very good large model you can generate very good synthetic data a-and then continue to improve quality. So that's why I think even OpenAI they are shifting from the training scaling law into inference scaling law and it's the test time and all this.

So w-w-- I definitely believe that's the future direction and that's where we are really good at and in doing inference.

**Swyx** [35:35]
Couple questions on, on that. Are you planning to share your reasoning traces?

**Lin Qiao** [35:39]
That's a very good question. We are still debating.

**Swyx** [35:44]
Yeah. Yeah.

**Lin Qiao** [35:46]
We're still debating.

**Swyx** [35:47]
I would say if you, for example, it's interesting that like, for example, SWE-bench if you're want to be considered for ranking you have to submit your reasoning traces and that has actually disqualified some of our past guests like Co- Cosine was doing well on SWE- SWE-bench but they didn't want to leak those results.

So that's why you don't see o1, o1 preview on SWE- SWE-bench because they don't submit their reasoning traces.

**Lin Qiao** [36:07]
Mm-hmm.

**Swyx** [36:08]
And obviously it's IP but also if you're gonna be more open then that's one way to be more open. So your model is not gonna be open source, right? Like it's gonna be a endpoint that you provide-

**Lin Qiao** [36:18]
Yes

**Swyx** [36:19]
... that's. Okay, cool. And then pricing also the same as OpenAI just kind of based on-

**Lin Qiao** [36:25]
Yeah. This is-- Oh, I don't have actually information.

**Swyx** [36:28]
Yeah.

**Lin Qiao** [36:28]
Everything is going so fast we haven't even think about that yet. Yeah, I should be more prepared so

**Swyx** [36:33]
No, no, no. I mean this is live. I mean, you know, it's, it's, it's nice to just talk about it as it, as it goes live.

**Lin Qiao** [36:38]
Mm-hmm.

**Swyx** [36:38]
Any other things that you're like you want feedback on or you're thinking through? It's, it's kinda nice to just talk about something when it's not decided yet about this new model. Like, I mean it's, it's gonna be exciting, it's, it's gonna generate a lot of buzz and-

**Lin Qiao** [36:51]
Right I'm very excited about to see how people are gonna use this model. So there's already Reddit discussion about it, and people are asking very deep medical questions. And it seems the model get it right, surprisingly. And, and internally we're also asking models to generate what is AGI, and it generate a very complicated DAG thinking process.

So i- it's-- and we're having a lot of fun testing this internally. But I'm more curious, like, how will people use it, how-- what kind of application they are gonna try and test on it, and that's where we'll really like to hear feedback-

**Swyx** [37:30]
Yeah

**Lin Qiao** [37:30]
... from the community. And also feedback to us, like, what works out well, what doesn't work out well. What works out well but surprising them and, uh, what kind of thing they, they think we should improve on, and those kind of feedback will be tremendously helpful.

**Swyx** [37:45]
Yeah. I mean, so I've been a production user of, uh, Preview and Mini since launch. I would say they're very, very obvious jumps in quality, so much so that they, they made Claude Sonnet and 4o just... Like, they, they made the previous state-of-the-art look bad.

Like, it's really , it's really that, uh, stark, that difference.

**Lin Qiao** [38:04]
Mm-hmm.

**Swyx** [38:04]
The number one thing I, I actually, you know, just feedback or request, feature request, is people want control on the budget because right now-

**Lin Qiao** [38:13]
Yes

**Swyx** [38:13]
... in o1 it kind of decides its own thinking budget.

**Lin Qiao** [38:16]
Yeah.

**Swyx** [38:16]
But sometimes you know how hard the problem is, and you want to actually tell the model, like, "Spend two minutes on this"-

**Lin Qiao** [38:23]
Yes

**Swyx** [38:23]
... or spend some dollar amount. Maybe it's time and maybe it's dollars, I don't know what the budget is-

**Lin Qiao** [38:26]
Yeah

**Swyx** [38:26]
... but they do.

**Lin Qiao** [38:27]
That, that makes a lot of sense. So we, uh, we actually thought about that requirement, and it should be at some point we need to support that.

**Swyx** [38:35]
Yeah.

**Lin Qiao** [38:35]
Uh, not initially, but that makes a lot of sense. Yeah.

**Swyx** [38:39]
Okay. So that was a fascinating overview of just, like, the things that you're working on. First of all, I realized that, I don't know if I've ever given you this feedback, but I think you guys are one of-- like, one of the reasons I agreed to advise you because, like, you know, I think when, when you first met me, I was kind of dubious.

### Team & Culture

**Swyx** [38:53]
I was like-

**Lin Qiao** [38:54]
Who are you?

**Swyx** [38:55]
... there's Replicate, there's Together, there's, like, a Lepton. There's, like, a whole bunch of other players. You're in very, very competitive fields. Like, why will you win? And the reason I actually changed my mind was I saw you guys shipping.

**Lin Qiao** [39:08]
Mm.

**Swyx** [39:09]
You know, I think your surface area is very big. The team is not that big.

**Lin Qiao** [39:12]
No.

**Swyx** [39:13]
No.

**Lin Qiao** [39:13]
We are only 40 people.

**Swyx** [39:14]
Yeah. And now here you are trying to compete with OpenAI and, you know, everyone else. Like, how-- what is the secret?

**Lin Qiao** [39:21]
Um, I think the team. The team is the secret.

**Swyx** [39:23]
Oh, boy. So there's no, there's nothing I can just copy, just-

**Lin Qiao** [39:29]
No.

**Swyx** [39:30]
Yeah.

**Lin Qiao** [39:31]
I think we all come from very aligned on the culture, uh, 'cause most of our team came from Meta-

**Swyx** [39:38]
Yeah

**Lin Qiao** [39:38]
... and many startups, so we really believe in results. One is result, and second is customer. We're very customer obsessed. Um, and, uh, we don't want to drive adoption for the sake of adoption. We really want to, uh, make sure we understand we are delivering a lot of business values to the customer, and we are-- we really value their feedback.

So we would woke up middle of night and deploy some model for them or shuffle some capacity for them. Um, and, uh, yeah, over the weekend, no-brainer. So yeah. So that's just how we work as a team.

**Swyx** [40:19]
Yeah.

**Lin Qiao** [40:19]
And, uh, the caliber of the team is really, really high as well. So, like, as plugin, we're hiring. Uh, we're expanding very, very fast. So if we are passionate about, uh, working on the most cutting-edge technology-

**Swyx** [40:34]
Yeah

**Lin Qiao** [40:35]
... in the generative space, uh, come talk with us.

**Swyx** [40:38]
Yeah. Let's talk a little bit about those, that customer journey. I think one of your more famous customers, Cursor. We were the first podcast to have Cursor on, and then obviously since then they have blown up. Cause and effect are not related.

### Cursor Partnership

**Swyx** [40:48]
But but you guys, uh, especially worked on a fast apply model where you were one of the first people to, to work on speculative decoding in, in a production setting. Uh, maybe just talk about, like, what was the behind the scenes of working with Cursor.

**Lin Qiao** [41:03]
Right. I would say Cursor is a very, very unique team. I think the unique part is the team has very high technical caliber, right? There's no question about it. But they have decided, although, like, they-- uh, many companies building coding copilot, they will say, "Well, I'm gonna build the whole entire stack because I can."

And they are unique in the sense they seek partnership, not because they cannot, they're fully capable, but they know where to focus. That to me is amazing. And, uh, of course, they want to find a bi- best partner, so we spend some time working together.

They are pushing us very aggressively, uh, because for them to deliver high caliber product experience, they need the latency, they need the interactive, but also high quality at the same time. So actually we expanded our product feature quite a lot as we supporting Cursor, and they are growing so fast and we massively scaled quickly across multiple regions, and we develop pretty high in-intense inference stack, almost like similar to what we do for Meta.

That-- I think that's a very, very, uh, interesting engagement. And through that, there are a lot of trust being built as in they realize, hey, this is a team they can really partner with and they can go big with.

That comes back to, hey, we are really customer obsessed and all the engineers working with them, there's just enormous amount of time syncing together with them and discussing. And we're not big on meetings, but we are like Slack channel always on.

**Swyx** [42:33]
Yeah.

**Lin Qiao** [42:33]
Uh, yeah. So you almost feel like working as one team. So that, I, I think that's really highlight of.

**Swyx** [42:39]
Yeah. For those who don't know, like, so basically Cursor is a VS Code fork, but most of the time people will be using closed models. Like, I, I actually use a lot of Sonnet. So you're not involved there, right?

It's not like you host Sonnet or you have any partnership with it. No, like you're involved where Cursor is small or like their, their propri- their, you know- House brand models are concerned, right?

**Lin Qiao** [42:59]
I don't know what I can say-

**Swyx** [43:00]
Okay

**Lin Qiao** [43:00]
... but the things they haven't said.

So, as the founder, hard to-

**Swyx** [43:04]
I'm just saying, like, it's, it's very obviously the dropdown is 4Revo-

**Alessio** [43:06]
Yeah, yeah

**Swyx** [43:07]
... Cloud, and then Cursor, right? So, like, I assume that the Cursor side is the Fireworks side, and then the other side is they're calling out the other. Yeah. Just kinda curious. And, and then, like, do you see any more opportunity on like the...

You know, I think you made a big splash with, like, 1,000 tokens per second. That was because of speculative decoding. Is there more to push there?

**Lin Qiao** [43:26]
We push a lot. Actually, when I mentioned a file optimizer, right? So as in we have a unique automation stack that is one-size-fits-one. Uh, we actually deployed to Cursor early on. Basically optimized for their specific workload, and that's a lot of juice to extract out of there.

And we see the success in, in that product is actually can be widely adopted, so that's why kind of we, we started a separate product line, uh, called the File Optimizer. So speculative decoding is just one approach. And speculative decoding here is not static.

We actually wrote a blog post about it. There's so many different way to do speculative decoding. You can pair a small model with a large model in the same model family, or you can have Eagle heads and, and so on.

So there are different trade-offs of which approach to, uh, take. It really depends on your workload. And then w-with your workload, we can align the Eagle heads or Medusa heads or, you know, small, big model pair much better to extract the best latency reduction.

So all of that is part of the File Optimizer offering.

**Alessio** [44:24]
I know you mentioned some of the other inference providers. I, I think the other question that people always have is around benchmarks. So you get different performance on, like, different platforms. How should people think about... You know, people are like, "Hey, Llama 3.2 is X on MMLU."

But maybe, you know, using speculative decoding, you go down a different path. Maybe some providers run a quantized model.

**Lin Qiao** [44:48]
Yeah.

**Alessio** [44:48]
How should people think about how much they should care about how you're actually running the model, you know? Like, what's the delta between all the magic that you do and, like, what a raw model?

**Lin Qiao** [44:58]
Okay. So there are d- two big development cycle. One is experimentation, where they need fast iteration. They don't want to think about quality-

**Alessio** [45:06]
Yeah

**Lin Qiao** [45:06]
... and they just kind of want to experiment with product experience and, and so on, right? So that's one. And then it looks good, and they want to kind of post-product-market fit scaling, and the quality is really important, and latency and all the other things are becoming important.

During the experimentation phase, it's just pick a good model. Don't worry about anything else. Make sure even, like, GenAI is the right solution to your product, and that's the focus. And then post-product-market fit, then that's kind of the three-dimensional optimization curve start to kick in across quality, latency, cost, where you should land.

And to me, it's a purely a product decision. To many product, if you choose a lower quality but better speed and lower cost, but it doesn't make a difference to the product experience, then you should do it. So that's why I think inference is part of the validation.

**Alessio** [45:58]
Mm.

**Lin Qiao** [45:59]
The validation doesn't stop at offline eval. The validation is kind of will go through A/B testing, through inference, and that's why we kind of offer various different configurations for you to test which is the best setting. So this is the, like, traditional product evaluation.

So product evaluation should also include your new model versions, um, and different model setup into the consideration.

### Quantization Drama

**Swyx** [46:22]
I want to specifically talk about what happens a few months ago with, uh, with, uh, some of your major competitors. I mean, you know, all of this is public. What is your take on what happened? And, and maybe you want to set the record straight on how Fireworks does quantization, because I think a lot of people may have outdated perceptions, or they, they didn't read the, the clarification post on your approach to quantization.

**Lin Qiao** [46:45]
First of all, it's always a surprise to us that, uh, without any notice we got called out.

**Swyx** [46:51]
Specifically by name, which is-

**Lin Qiao** [46:53]
By name

**Swyx** [46:53]
... normally not what-

**Lin Qiao** [46:55]
Yeah. In, in a public post and have certain interpretation of our quality. So I, I was really surprised. And, uh, it, it's not a good way to compete, right? We want to compete fairly, and oftentimes when one vendor give out results on interpreting another vendor, it's always extremely biased.

So we actually refrain ourself to do any of those, and we happily, happily partner with third party to do the most fair evaluation. So we are very surprised, and we don't think that's a good way to figure out the competition landscape.

So then we react... I think when it comes to quantization, the interpretation, we wrote a actually a very thorough blog post because, again, uh, no one size fits all. We have various different quantization schemes. We can quantize very different parts of the model, uh, from ways to activation to cross CPU communication to, like, consi- they can use different quantization scheme or consistent across the board.

And again, it's a trade-off. It's trade-off ac-across these three-dimensional quality, latency, and cost. And for our customer, we actually let them, like, find the best optimizer point, and that's kinda how... And we have very thorough evaluation, uh, process to, to pick that point.

But for self-serve, then there's no only one point to pick. There's no, like, customization available. So of course we... Depends on, like, what we, we talk with many customer, we, we, we have to pick one point. And, and I, I think the end result, like, AA published, uh, later on AA pub-published a quality measure, and we're actually-- we look really good.

So I, I don't- I wouldn't... That's why what I mean is I will leave the evaluation of quality or performance to a third party and work with them to find the most fair benchmark approach and methodology. But I'm not applaud of approach of calling out specific names and critique other competitors, uh, in a very biased way.

**Swyx** [48:57]
Databases, uh, happens as well. I think you're the more politically correct one, and then Dima is the more, uh- ... something like that. Miss you on Twitter.

**Lin Qiao** [49:09]
We, uh, yeah-

**Swyx** [49:10]
He's like the Russian, uh-

**Lin Qiao** [49:11]
We, we partner. No, actually, all these reactions we build together.

We play different roles. At least.

**Swyx** [49:20]
Another one that I wanted to... Uh, uh, on just the last one on the competition side, there's a perception of price wars in ho-hosting open source models. You are, you're... A-and we talked about, uh, the competitiveness of the market.

Um, do you aim to make margin on open source models?

**Lin Qiao** [49:37]
Oh, absolutely yes. So but I, I think it, it really... When we think about pricing, it's really need to coordinate with the value we're delivering. If the value is limited or there are a lot of people delivering same value, there's no differentiation, it's, there's only one way to go, is going down, right?

So through competition. If I take a big step back, there is pricing from... We are more compared with, like, closed model providers, APIs, right? The closed model provider, their cost structure is even more interesting because we don't have any-- we don't bear any training costs.

**Swyx** [50:11]
Yes.

**Lin Qiao** [50:11]
And we focus on inference optimization, and that's kind of where we continue to add a lot of product value. So that's how we think about product. But for the closed source API provider, um, model provider, they bear a lot of training costs, and they need to amortize the training costs into the inference.

So that create a very interesting dynamics of, yeah, if we match pricing there, and I think how they are gonna make money i-is-

**Swyx** [50:36]
Right. Yeah, yeah

**Lin Qiao** [50:36]
... is very, very interesting.

**Swyx** [50:38]
So for listeners, OpenAI's 2024, four billion in revenue, three billion in compute training, uh, two billion in compute inference, um, one billion in research, compute amortization, and 700 million in salaries. So that is like

a lot, I mean, a, a lot of R&D.

**Lin Qiao** [51:02]
Yeah. So and I think Matter is basically like snake at zero.

**Swyx** [51:06]
Yeah.

**Lin Qiao** [51:07]
So that, that's a very, very interesting dynamics we're operating within. Uh, but coming back to inference, right? So we are, again, as I mentioned, our product is we are a platform. We're not just a single model as a service provider as many other inference providers, like, they are providing single model.

We have our optimizer to have it customized towards your inference workload. We have a Compound AI system where it significantly simplify your interaction to high quality and low latency, low cost. So those are all very different from other providers.

### Underrated Features

**Alessio** [51:38]
What do people not know about the work that you do? I guess, like, people are like, okay, Fireworks, you run model very quickly, you have the function model. Is there any kinda like underrated part of Fireworks that more people should try?

**Lin Qiao** [51:51]
Yeah. Actually, one user post on x.com, he mentioned, "Oh, actually, Fireworks can allow me to upload the LoRA adapter to the service model with, at the same cost and use it at same cost." Nobody has provide that. That, that's because we have a very special, like...

We, we, we rolled out Multi-LoRA last year actually. And we actually have this function for a long time, and many people has been using it, but it's not well-known that, oh, if you fine-tune model, you don't need to use on demand.

If you fine-tune model with LoRA, you can upload your LoRA adapter and we deploy it as if it's a, a new model, and then you use-- you get your, uh, endpoint and you can use that directly, but at the same cost as the base model.

So I'm happy that user is marketing it for us.

**Swyx** [52:43]
Yeah.

**Lin Qiao** [52:44]
They discover-- He discovered that feature- ... but we have that for last year. Uh, so I think we-- I think to, um, feedback to me is, uh, we have a lot of very, very good feature as, as, um, Shawn just mentioned.

We-

**Swyx** [52:57]
I've been an advisor to the company, and I didn't know that you had speculative decoding released, you know?

**Lin Qiao** [53:03]
We have prompt caching way back last year also. We have many... Yeah. So yeah, so I, I think that is one of the underrated, uh, feature. And if, if there are developers, um, you are using our self-serve platform, please try it out.

**Swyx** [53:16]
Yeah, yeah, yeah. The LoRA thing is interesting because I think you, you also... Like, the reason people add additional cost to it is not because they feel like charging people. Like, normally, in normal LoRA serving setups, there, there is a cost to dedicating, loading those weights-

**Lin Qiao** [53:32]
Yeah

**Swyx** [53:32]
... and dedicating, uh, a, a machine to that inference. Uh, how come you can avoid it?

**Lin Qiao** [53:36]
Yeah. So, so this is kind of our, our technique called Multi-LoRA. So we basically have many LoRA adapters share the same base model.

**Swyx** [53:45]
Yeah.

**Lin Qiao** [53:45]
And, uh, basically we significantly reduce the f-memory footprint of serving, and one base model can sustain 100 to 1,000 LoRA adapters. And then basically all these different LoRA adapter can share the same, like, do-drive the same traffic to the same base model, where base model is dominating the cost.

So that's how we optimize, uh, that way, and that's why how we can manage the tokens per, uh, dollar, um, million token, uh, pricing the same as base model.

**Swyx** [54:13]
Awesome. Is there anything that you think you want to request from the community or you're looking for model-wise or tooling-wise that you think, like, someone should be working on and it's-

**Lin Qiao** [54:24]
Yeah. So we really want to get a lot of feedback from the application developers who are starting to build on GenAI or, you know, on the-- already adopted or starting about thinking about new use cases and so on to try out, um, Fireworks first.

And let us know works out really well for you and what is your wish list and what is, what sucks, right? So, uh, what is not working out for you and we, we like to continuously improve. And for our new product launches, typically we want to launch to a small group of people.

Usually, we launch on our Discord first to have a set of people use that first. So please join our Discord channel. We have a lot of communication going on there. Again, you can also give us feedback. We'll, we'll have started office hour for you to directly talk with our DevRel and engineers to exchange more, more notes.

**Alessio** [55:17]
And you're hiring across the board.

**Lin Qiao** [55:19]
We're hiring across the board. We're hiring front engineers, infrastruc- cloud infrastructure engineers, backend system optimization engineers, applied researchers, and, like, researchers who has done post-training, uh, who has done a lot of fine-tuning and so on.

**Alessio** [55:34]
Cool.

**Swyx** [55:35]
Cool. That's it.

**Alessio** [55:35]
Cool.

**Swyx** [55:36]
Thank you.

**Lin Qiao** [55:36]
Awesome.

### Closing

**Swyx** [55:37]
Thanks for having us.

---

This library is powered by PodHood (https://podhood.com), the podcast website platform.
