# Building an open AI company - with Ce and Vipul of Together AI

Latent Space · 2024-02-08

<https://addtry.com/98d94665-d063-40c5-9ac6-8152517d3df4>

Together AI co-founders Vipul Ved Prakash and Ce Zhang explain why openness is core to their mission, detailing their journey from Apple (Vipul) and Stanford research (Ce) to building an open AI platform. They discuss RedPajama’s evolution into a modular dataset with 40 quality signals, the need for 5,000 tokens/second inference speed, and their investment in state space models like Mamba and Hyena as alternatives to transformers. The company, which runs 7,000-8,000 GPUs (mostly H100s), sees training as a larger workload than inference, with fine-tuning driving top models. They advocate for independent inference benchmarks, publish open research like FlashAttention, and keep some software proprietary. With 38 employees and 45% researchers, they are hiring across the stack, from CUDA to DevOps.

## Questions this episode answers

### How many GPUs does Together AI have?

Vipul Ved Prakash disclosed for the first time that Together AI has close to seven to eight thousand GPUs, mostly split between A100s and H100s. They build a disaggregated, decentralized AI cloud by combining data centers around the world rather than relying on hyperscalers, and the fleet is growing monthly with plans to add next-generation hardware.

[27:15](https://addtry.com/98d94665-d063-40c5-9ac6-8152517d3df4?t=1635000)

### Why is Together AI investing in state space models like Mamba and StripeHyena?

Ce Zhang explains that transformers may not reach their target of 5,000 tokens per second, and state space models offer a path to cheap long-context and high throughput by minimizing state size. They are exploring hybrid architectures—mixing attention layers with state space layers—to match or beat transformer quality. StripeHyena was built to rival Mistral, while Mamba probes pure state space scaling. Vipul adds this research is essential for more efficient inference.

[55:07](https://addtry.com/98d94665-d063-40c5-9ac6-8152517d3df4?t=3307000)

### What is the significance of Together AI aiming for 5,000 tokens per second inference speed?

Vipul Ved Prakash states that many applications consume model output without a human reading it, and current speeds of 30–70 tokens per second cannot meet that demand. Achieving 5,000 tokens per second would dramatically increase throughput per GPU, lower latency for agentic workflows, and enable new user experiences where answers appear instantly. He notes no infrastructure today satisfies these needs.

[1:04:25](https://addtry.com/98d94665-d063-40c5-9ac6-8152517d3df4?t=3865000)

## Key moments

- **[0:00] Origins**
  - [0:42] Together AI was founded in June 2022 with a mission to build open and independent AI systems
  - [2:28] Vipul Ved Prakash left Apple to found Together AI, blending Apple's UX polish with open-source ideals
  - [3:05] Vipul Ved Prakash created Vipul's Razor, an open-source collaborative spam filter, before founding Cloudmark
  - [5:02] The scaling laws paper signaled a profound new era where algorithms improve with scale, says Vipul
  - [5:43] Chris Ré introduced Vipul Ved Prakash to Ce Zhang, leading to the formation of Together AI
- **[10:17] RedPajama**
  - [10:17] RedPajama was created to reproduce Llama's data recipe with an open dataset so the community could learn from mistakes, says Ce Zhang
  - [12:15] Cerebras built Slim Pajama by deduplicating Together's RedPajama dataset, improving data quality
  - [15:25] DSIR selects training data subsets by measuring similarity to a validation set, making model training more efficient, explains Ce Zhang
  - [16:34] State space models are necessary to reach 5,000 tokens per second inference speed, predicts Vipul
  - [25:12] Top five inference models on Together are fine-tuned open-source variants, and training is the bigger compute portion, says Vipul
- **[25:19] Inference Business**
  - [27:23] "We have close to seven to eight thousand GPUs today, mostly A100s and H100s, and it's growing monthly," says Vipul
  - [29:57] The GPU market is very tight with two-to-three-month lead times, and 4-5 million GPUs will be sold in 2024, states Vipul
- **[32:10] Inference Moats**
  - [32:10] Ce Zhang calls for an independent industry benchmark for inference services, like the TPC for databases
  - [34:58] Together's inference moat comes from 50 proprietary techniques, with only universally useful ones open-sourced, says Vipul
  - [37:50] Together's 'disaggregated cloud' aims to abstract away the diversity and reliability of underlying hardware for users, says Ce Zhang
  - [41:27] Ce Zhang predicts the next 10x inference improvement will require co-optimizing algorithms, model architecture, and systems together
  - [43:00] Together vs Anyscale benchmark: Ce Zhang argues benchmarks must incentivize technical progress, not marketing
  - [45:48] Vipul wants an independent third party to run inference benchmarks, noting variables like time to first token depend on location
- **[49:25] Beyond Inference**
  - [49:25] Together offers a spectrum of fine-tuning and training services, from serverless fine-tuning to scaling custom models, says Vipul
  - [52:10] Ce Zhang believes embeddings are far from mature and will see huge improvements in size, speed, and quality for RAG and model training loops
- **[55:07] State Space Models**
  - [55:07] Together's Stripe Hyena and Mamba represent different stages of state space model exploration, aiming to surpass transformer quality and efficiency
  - [58:02] Stripe Hyena targets beating Mistral with a hybrid architecture, while Mamba pushes a pure state space model to 3B parameters, says Ce Zhang
  - [1:01:08] The optimal mix of transformer, Hyena, and Mamba layers in hybrid architectures is unknown and requires systematic scaling studies, says Ce Zhang
- **[1:04:25] Speed**
  - [1:04:25] Vipul says 5,000 tokens per second inference is needed for machine-to-machine apps and new UX possibilities, not yet achievable today
  - [1:06:31] Together is moving toward serverless inference and fine-tuning to let developers start without committing to thousands in compute, says Vipul
  - [1:07:51] Together provides $25 in free credits, enough to build and deploy an AI app completely for free, notes Vipul
  - [1:08:14] Ce Zhang reports that combining fine-tuning and RAG works better than either alone, as seen at Together's hackathon
  - [1:09:01] Together AI, with a team of 38, is hiring across CUDA, research, systems, and front-end, welcoming candidates without prior AI experience
- **[1:11:57] Wrap-Up**
  - [1:11:59] "I think we sort of need a framework for what the world looks like with advanced intelligence systems, rather than just doomerism," says Vipul

## Speakers

- **Alessio** (host)
- **Swyx** (host)
- **Ce Zhang** (guest)
- **Vipul Ved Prakash** (guest)

## Topics

State Space Models, Inference, Compute

## Mentioned

AnyScale (company), Mosaic (company), Together AI (company), A100 (product), Flash decoding (product), FlashAttention (product), H100 (product), L40s (product), Llama (product), Mamba (product), Mistral (product), RedPajama (product), State Space Models (product), Stripe Hyena (product), Together Embeddings (product), Transformers (product)

## Transcript

### Origins

**Vipul Ved Prakash** [0:00]
Hey, everyone. Welcome to the Latent Space Podcast. This is Alessio, partner and CTO in residence at Decibel Partners, and I'm joined by my co-host Swyx, founder of Smol.ai.

**Swyx** [0:10]
Hey, and today we have-- we're together with Together.

Welcome to the studio, guys.

**Vipul Ved Prakash** [0:16]
Thank you.

**Ce Zhang** [0:17]
Thank you.

**Vipul Ved Prakash** [0:17]
Thanks for having us.

**Swyx** [0:18]
Maybe you, you guys wanna-- I don't know how you typically give self-intros, but, um, uh, does anyone want to go first? Like, uh, h-how do we, how do we get our audience acquainted, uh, especially to, like, who's speaking?

'Cause there's a, you know, it's unusual for us to do a four-person pod.

**Ce Zhang** [0:32]
Yeah. Hi, everyone, I'm Ce. Yeah, so I'm one of the co-founder of Together and the CTO working with the team on the technical things. Yeah.

**Vipul Ved Prakash** [0:39]
I'm, uh, Vipul Ved Prakash, co-founder and CEO of Together.

**Swyx** [0:43]
I always consider you guys as one of the sort of all-in-one, um, companies. I-I always wanna say labs, but I, I feel like you're not a lab. Um, what is the sort of origin of Together, and then w-what is it today?

I feel like you used to be Together like XYZ.

**Vipul Ved Prakash** [0:59]
Right.

**Swyx** [1:00]
And then now you're Together AI.

**Vipul Ved Prakash** [1:02]
You know, I, I think fundamentally Together is about open and independent AI systems. We think this is, um, you know, one of the most consequential technologies of our time. And when we started the company in June 2022, uh, you know, our focus was to build a platform for, you know, open source, independent, user-owned AI systems.

Uh, y- one way to think about it is, uh, big labs, frontier model labs have built their own platforms for developer platforms for their models. We think of Together as a platform for everything else, um, you know, whether these are open models, whether these are models being built by companies that are owned by them.

Um, and, you know, our, our sort of XYZ roots, you know, we have a fairly deep decentralization and open ethos that, uh, kind of reflects in all our, um, you know, platform and strategy and business. Uh, and we also-- the way we structure our, our, our cloud is, um, by combining data centers around the world instead of, uh, you know-- Uh, we are today not located in hyperscalers.

We have built a, a footprint of, uh, you know, AI supercomputers, uh, in this sort of a disaggregated, decentralized manner.

**Swyx** [2:28]
Uh, I know before Together you were Apple, so you go from, like, the most walled garden, private-

**Vipul Ved Prakash** [2:34]
Right

**Swyx** [2:34]
... we don't say anything company to we want everything to be open and everybody to know somebody. What maybe did you learn from, like, the Apple way of being super closed and polished, and maybe what are you taking now to Together to make it open, but also a very nice developer experience?

**Vipul Ved Prakash** [2:50]
One sort of my, my, you know, background has been in open source for, for a long time. I, uh, one of the first things I created was a s- collaborative spam filter. Um, you know, this is back in the day.

Uh-

**Swyx** [3:04]
It's called Vipul's Razor.

**Vipul Ved Prakash** [3:05]
It's called Vipul's Razor, and, uh, it became, uh, quite popular, and, uh, the, the first company I founded called Cloudmark was built around, um, you know, taking open source and building, uh, you know, building both an open side of it and a commercial product around it.

I think Apple is, uh, you know, sort of very focused on providing this amazing experience to, uh, to its customers with, you know, most of the technology sort of hidden behind the product. Uh, and certainly the focus on fluidity and, um, you know, applying complex technology to make, uh, everyday things simple is something that Apple does really well.

And, you know, that's been a sort of big part of how we think about our developer platforms. I think it informs it. Uh, the, the other thing is that during my, um, years at Apple, we, you know, worked a lot on deep learning, and one of the things that was sort of very viscerally accessible to me was how well these systems worked.

We, you know, we built an open domain Q&A system. Um, this was, uh, based on, uh, Facebook's LSTM paper in, in 2016. And, uh, it was remarkable because we, we had a parallel system based on sort of information retrieval techniques, which is extremely complicated, uh, didn't work that well.

And, uh, you know, this thing we wrote in a week was just, uh, incredible performance. So I, I think some of those experiences, uh, at least for, for, for me personally, sort of, uh, were creating this roadmap of how, you know, how important and powerful this, uh, technology is.

And, you know, when, um, the scaling laws paper was published, that was very clear. Like it was in some ways something very profound. We've never had algorithms that improve in capabilities but scale out. So this is almost a, you know, new era of computing.

And, uh, so, so that's been, I think, the influence of Apple. My, my years at Apple, uh, really

f-for me, like crystallized the, the value of, uh, what we are doing at Together.

**Swyx** [5:37]
Mm-hmm. And how did you decide to join forces? Because you, uh, did a postdoc with Chris Ré at Stanford. You know, we already had Tri Dao from Together, and we talked about hazy. Um, how-- uh, what was like the meeting of the mind of, hey, I come from like the more technical postdoc, um, assistant professor background, and Vipul, you had a more product thing.

Uh, what got you excited to like build this now? You know-

**Ce Zhang** [6:02]
Yeah

**Swyx** [6:02]
... there's so many people.

**Ce Zhang** [6:03]
Yeah. So I think- Um, so we have been working on this together with Chris in the essentially last, like, 10 years, right? So it was, like, a machine learning system 10 years ago was like probably the graphic model, right, and then convolutional neural network, and then all the foundation model that we see today.

But if you look at this, I think that fundamentally, the system we are actually optimizing is actually not that different. It's always about data movement across essentially all the stacks, right? So when you do distributed, like, computing, it's about communication across different machines.

When you do, for example, flash attention, it's about data movement at a different essential memory hierarchy, right? So we have been doing this es- in the last 10 years and seen the field start grow, grow, grow. So we kind of feel the current kind of, kind of this, like, wave of technology is actually the perfect time to actually bring all the research essentially into something real.

And we are super lucky that we got introduced to Vipul, right? So, and, uh, and, uh, yeah, and then we tr- hope to join forces and, uh, bring this to real world.

**Swyx** [7:03]
Mm-hmm.

**Ce Zhang** [7:04]
Yeah.

**Swyx** [7:05]
Yeah. Uh, uh, it's, uh, very interesting that, um, you know, like, it's an unusual team of, like, sort of research and industry. Like, you've been a, like, a third or fourth time founder now.

**Ce Zhang** [7:14]
Third time founder, yeah.

**Swyx** [7:15]
Third time. Um, and, and so, like, what, what is your first order of business when you, like, set up Together? Like, how, how do you sort of put something like this together? Oh my God, I'm gonna use this word so much.

**Vipul Ved Prakash** [7:27]
I think, um, the... You know, I, I, I feel AI companies are really kind of driven by research. And, um, it was, uh, actually, like, uh, you know, Chris and I had been talking about, about how to reduce the cost of building models.

That was-- We felt that, you know, there aren't really big data moats around foundation models. They are built from a subset of the web. Um, what is difficult is the cost of capital to build these, and ha- what- one of the ways in which you can reduce this cost is by making more efficient systems.

Uh, so, you know, uh, with that, uh, it was really about finding the right, uh, set of co-founders and team. In fact, when I-- Chris introduced me to Si, and, uh, I think within the first five minutes of talking to Si-

**Swyx** [8:25]
Mm-hmm

**Vipul Ved Prakash** [8:25]
... I was like, "This, you know, we are, we are starting this company." Um, and, you know, our, our early focus was thinking about, uh, this more sort of disparate set of resources, uh, you know, GPUs around the internet.

Can we use those to build a model? And, uh, we really have to compress communication for, you know, um, when we do, uh, gradient averaging. There's just a lot of traffic, and if you can reduce that somehow, uh, you sort of open up the possibility of using cheaper compute, uh, y- you know, across the network.

And Si's research, uh, for a decade has been in that, in that subject. Um, and, uh, y- you know, and from there, finding, you know, other folks in the network, I think there is generally a lot of excitement and philosophical alignment around what we are doing, which, uh, uh, you know, we publish papers, we publish, uh, open source libraries and code.

We build open models. Uh, a- and I think the, uh, a lot of people in academia in, um, you know, machine learning and NLP, that's really what they want to do. So I, I, I think that's been, uh, really a k- kind of kernel for, you know, composition of the company.

And we've-- we are lucky to have, you know, at this point attracted some of the best researchers in the field. So I, I, I think that's the most important thing. And, you know, the rest of it is sort of driven by a s- a couple of these philosophies around independent systems and decentralization and, um, you know, and, and good developer interfaces.

You wanna make it accessible. That's, uh, you know, just as important. Uh, and the rest follows from there, I think.

### RedPajama

**Swyx** [10:18]
I wanna try and fill in some of the blanks in the history of Together. I think people come on your website today and they say, see you-

**Vipul Ved Prakash** [10:25]
Mm-hmm

**Swyx** [10:25]
... raise a hundred million dollars series A. They're like, "Wow, these guys are, like, super legit company." But it, it feels like, uh, RedPajama just came out a year ago. Uh, I remember we had Mike Conover in the studio who ha- who had built DALL-E at Databricks, um, and you-

**Ce Zhang** [10:40]
The same day. Yeah

**Swyx** [10:41]
... yeah, you announced it literally the morning we were recording, so we were, like, in the studio on our phones looking at it, and it's like, wow, this is, like, the first time now there's, like, a good curated dataset to do open, uh, pre-training.

So maybe let's start from there. Like, what was the motivation behind it? Why did you decide to do that? It's-- Datasets are one of the things that most people don't wanna work on. They just wanna do models- ...

not datasets.

**Ce Zhang** [11:03]
Yeah. So yeah, first one is not the first, right? So I think it's actually built on a whole bunch of amazing effort community already have. For example, Eleuther have the pile, right? There's a whole bunch of amazing datasets have, like C4, right, from Google, right?

So I think really got inspired by the impact those, like, datasets have on the community, right? So I think when we did RedPajama, it was a time that people are really fascinated by Llama, the model.

**Swyx** [11:27]
Mm-hmm.

**Ce Zhang** [11:27]
Like Llama 1, right? Which feel like decades ago.

**Swyx** [11:29]
Mm-hmm.

**Ce Zhang** [11:29]
Right? But it's kind of people are really excited about the quality, right? So that's really, like, a big shift in people how s- how to think about open model. People start to see hope, right? So but the one problem of Llama is the data recipe has been described in, in a pretty detailed way in the paper, but the data is actually not there.

**Swyx** [11:47]
Mm-hmm.

**Ce Zhang** [11:47]
Right? So and our original thinking is: how about we take the recipe and we try to do our best effort reproduction, and try to put it out such that we can learn from our mistake in the reproduction together, right?

So that's essentially the original thinking behind RedPajama. Uh, we have been pretty happy and excited about- What community are building-- can, kind of build on it, right? For example, there's a dataset called Slim Pajama, right, which do deduplication-

**Swyx** [12:15]
Deduplication

**Ce Zhang** [12:15]
... over our data, right?

**Swyx** [12:16]
From Cerebras. Did they talk to you before?

**Ce Zhang** [12:17]
Oh. Oh, yeah, yeah, yeah. Yeah.

**Swyx** [12:18]
Yeah.

**Ce Zhang** [12:18]
So yeah, so we are very good, very good friends, so we can discuss about technical perspective. We are pretty excited because I think it's kind of why we do Red Pajama in the first place, is that people can actually build not only models, but also datasets, essentially over that piece of artifact.

**Swyx** [12:34]
Yeah.

**Ce Zhang** [12:35]
So that's actually what inspired us to do the r- like, first version of Red Pajama dataset.

**Swyx** [12:40]
Yeah.

**Ce Zhang** [12:40]
Yeah.

**Swyx** [12:41]
And then you released V2 maybe two months ago.

**Ce Zhang** [12:43]
Yeah.

**Swyx** [12:44]
Uh, 30 trillion tokens.

**Ce Zhang** [12:46]
Yeah, 30 trillion tokens. So I think what's exciting about Red Pajama V2 is not only the number of tokens, but we start to kind of learn from Red Pajama V1. So one thing that we learned was that data quality is really the core, right?

So you want to take this couple trillion token dataset and try to bring them down maybe to one trillion or two trillion, right? The way that you are actually filter them, deduplicate them, is not something that kind of pre-decided before you see the application, right?

So you kind of want to have a modular framework to think about data quality, right? Given the application, let's automatically or may- maybe semi-automatically try to come up with a way to filter it down, right? So that's why in Red Pajama V2, we kind of overlay the dataset of, like, 40 different pre-computed quality signal.

**Swyx** [13:35]
Yes.

**Ce Zhang** [13:35]
Right? If you want to reproduce your best effort, like C4 filter, it's kind of like, like 20 lines of code.

**Swyx** [13:41]
Yeah.

**Ce Zhang** [13:41]
Right? And this open up this opportunity, you can actually put different filter together, learn the combination of filter. We are very excited to see what community actually come up with-

**Swyx** [13:50]
Yeah

**Ce Zhang** [13:50]
... using Red Pajama V2.

**Swyx** [13:52]
It was retrospectively so obvious that this is a good idea, that I wonder how come more datasets don't do this, which y- you just release, you release the dataset in, uh, with, with all these toggles that you can turn on and off.

**Ce Zhang** [14:05]
Yeah.

**Swyx** [14:05]
Right? That, and you can sort of tune up and down the quality in, in ways that you believe is important to you. Um, yeah, I, I just, uh, it makes so much sense now in retrospect. 'Cause everyone just publishes their, like, their pipeline and then the end result.

But what about all the intermediate stages?

**Ce Zhang** [14:19]
Yeah. Yeah, so I think-- So there are multiple things there. So it's, so, like, first one, like, I don't think we are the only one, like, doing that.

**Swyx** [14:27]
Oh, yeah?

**Ce Zhang** [14:28]
For example, like, uh, Doma from AI tool, right? They have this very flexible format to actually put in those quality signals, right?

**Swyx** [14:35]
Oh, okay.

**Ce Zhang** [14:36]
So, so, so think, like, uh, like we are actually calling it something, right? So you can actually load Red Pajama using their tool, right? Like, the whole thing should work, right? So, so I think one fundamental thing that changed in the last year, essentially, in the beginning when people think about data, is it's always like a byproduct of, of the model, right?

You release the model, you also release the data, right? The dataset is there for you to essentially to show people, ah, if you train on this data, you get a good model. But the thing will start to change is when people start building more and more those models, people start to realize, like, different subset of dataset is kind of available for different applications, right?

**Swyx** [15:16]
Mm-hmm.

**Ce Zhang** [15:16]
The data becomes something you want to play with.

**Swyx** [15:18]
Yeah.

**Ce Zhang** [15:18]
Right? So I think we are kind of lucky that we, we happen to release Red Pajama right at that point, that we get this opportunity to actually learn from that. Yeah.

**Swyx** [15:27]
Yeah. Um, and you guys have a custom model training platform on, on Together too. Uh, you have a bunch of stuff in there for data selection, like, uh, the SIR and things like that. How, how did you decide to work on that versus-- Because you first started with, like, some of the fine-tunes on, on Llama.

Uh, do you see a lot of interest there? And I know you've been doing a lot of research on, um, state space models and, uh, other transformer alternatives. Like, do you also see that as something you'll keep working on this year and push more people towards?

**Vipul Ved Prakash** [15:59]
Yeah, I mean, we, you know, and if we think of how to make, uh, training more efficient and building models more efficient, part of that is being able to select the right dataset. And that's what-- this is why you have signals, DSIR.

You can start with a small, you know, uh, uh, a small dataset and find similar documents, build models with that. So we think it's an important part of the kind of model build tooling that, uh, uh, y- you know, sort of widely useful for people building different kinds of models.

Um, similarly, you know, we are running into, uh, y- y- you know, the limits of how fast you can make transformers. Uh, and, uh, you know, we w- we want inference at five thousand tokens per second, right? And, uh, I don't think we will get there with transformers, and we need, um, uh, you know, we need, need to learn longer sequences.

Uh, data, again, becomes very, very expensive with transformers. So the, our work on space state models, uh, and all the research that we are doing there, and hopefully other labs will pick up on, on this and, and, and, you know, uh, make it a kind of important target, uh, for, for optimization.

But we think that, you know, open source is a great place for this. Um, we can provide these recipes for data and for training to our customers who are building, y- you know, custom models themselves. And, uh, you know, we, we are quite, quite excited about the sort of progress we are seeing there.

**Swyx** [17:44]
Mm-hmm. Do you have some of these models available for inference on Together? Can people play around with a-

**Ce Zhang** [17:50]
Sure do, yeah.

**Swyx** [17:51]
Yeah.

**Vipul Ved Prakash** [17:51]
Yeah. They are, they are available for inference on our serverless platform.

**Swyx** [17:56]
Cool. Yeah, uh, um, actually, so, uh, I always try to be the person who asks about acronyms in case, you know, people want to understand. Uh, DSIR, uh, should we explain importance resampling, you know, that kind of stuff?

**Ce Zhang** [18:08]
Oh, yeah. So DSIR essentially, uh, it's a fundamental idea. So it's one of the paper from, uh, from Percy, right? So essentially, uh, if you know what you are doing- You can actually use that as a very strong signal about what data to put in to insert training process, right?

So that's essentially the fundamental idea, right? So, and then more concretely, right, so there are actually different version of like DSIR, right? So one version is like you have validation side, right? You can actually somehow measure the similarity between the validation side and also your pre-trained corpus, and essentially subs- like the subset.

Uh, and often there's actually like less targeted version of DSIR where you'll say, "Yeah, maybe Wikipedia is, is actually a very good corpus. Let's try to find more Wikipedia," right? You can think about that as one way to...

You can think about it in two ways, either as a way to, to come up with different weights for different data slices or, and, uh, like a, yeah, so, or as a like filter type of step. Yeah, for a data side, or think about that as like data augmentation, right?

So yeah, so that's how, yeah, that's how we think about DSIR.

**Swyx** [19:12]
Got it. Um, that makes sense. Uh, I, uh, will have to read the paper to understand a little bit more, because when you, when you say things like, "We have to know in advance what we are trying to do with the model, then we do importance resampling," that is against the principle of general intelligence, right?

Like the point is to train AGI.

**Ce Zhang** [19:30]
Well, I mean, it depends on- Yeah, so it depends on what do you mean by being general or generic, right? So I think, I mean, you can always take a meta-learning perspective that we know the distribution of a task that we cares about, right?

So you can always go kind of up in the ladder of how general the whole thing is, right? Uh, but also for many of the customers that we are, we are actually talking to, right, they have kind of very targeted application, right?

The benefit you can get out of that is you could build a better open model, often smaller, often easier to do inference, if you know what you want, right? So I think the whole trade-off would be, and the x-axis would be how generic the whole thing will be.

The y-axis would be not only the top accuracy, but also a whole bunch of, uh, the deployment cost, right, the, the size of the model, right, the, the, the, the robustness of the model. So I think different people will navigate the space in different way, and we want to be the platform essentially whatever problem that you want, uh, we have a solution for you.

**Swyx** [20:30]
Uh, one more thing on data before we go deeper on state space models. Um, are we running out of data? Is, is thirty trillion, can we go in order of magnitude? Can we go five orders of magnitude? What, um, how do you, how do you, do the both of you think about how much data we have and how much we need?

**Ce Zhang** [20:51]
Yeah. So I think that's a very, very good question. So, so, so I think... I don't think we are running out of data on Earth. Right?

**Swyx** [21:00]
Yeah.

**Ce Zhang** [21:00]
So think about it globally, right? Yeah.

**Swyx** [21:01]
Training data. Training cor-

**Ce Zhang** [21:02]
Yeah, yeah.

**Swyx** [21:02]
Training class data.

**Ce Zhang** [21:03]
Yeah, yeah, yeah. So, so I think, I mean, some of them are not accessible, right? But, but I do think, uh, there are many organizations in the world have enough data to actually train like very, very good models, right?

So I mean, they are not publicly available, right? But, uh, there are people who actually have access to those, right? So, so I think, uh, in general, right, so if you think about the, the data in the open space, right?

So, so, so, so I guess that was specifically that you actually mean whether we are running out of data.

**Swyx** [21:37]
Yeah.

**Ce Zhang** [21:38]
So I, I do think there need to be some way, right, that, uh, people who are training open models get connected, uh, with essentially data that's not internet data, right? So I think that channel need to be opened up for the open model to gather more data, right?

But I'm kind of on the optimistic side that, uh, the society will figure out a way that we can train open models that's beyond this internet data. Yeah.

**Swyx** [22:07]
Beyond internet meaning books?

**Ce Zhang** [22:09]
Uh, I mean, there are a lot of those, right? Books, right? Transcripts, right? Videos, audios, right? So there's a whole bunch of like data sources that we are not integrating into open like data set, right? So, and maybe they shouldn't be open, right?

So I think the community need to figure out a way, uh, yeah, like the best balance, yeah, such that we can have open models, um, and, but on the other hand, uh, also have a, a reasonable collection of data that-

**Swyx** [22:39]
Yeah

**Ce Zhang** [22:39]
... we can actually use.

**Swyx** [22:40]
Yeah. I've, um, I think a lot of people think that, um, there's, there's a theory that Whisper was released so that you could transcribe YouTube and then use that as a source of tokens. Then I talk to other researchers who are like, "You know, YouTube has very low quality tokens.

You know, do you want your model to talk like a live streamer from YouTube?" "Because that's what they're gonna do." Um, so it's not clear, like what, uh, what the, the quality of this, uh, this data could be.

I don't know.

**Ce Zhang** [23:07]
Yeah. I guess that's-

**Swyx** [23:08]
It's an interesting open question.

**Ce Zhang** [23:09]
Yeah. I guess that depends on your application, right?

**Swyx** [23:11]
Yeah.

**Ce Zhang** [23:11]
So I think as a platform, right, so our goal is, uh, whatever application that you have, uh, yeah, so we have a platform that, uh, you can actually achieve your goal, right? So there are definitely applications that kind of make sense to speak like YouTube, right?

**Swyx** [23:26]
Yeah.

**Ce Zhang** [23:26]
So, but there are probably also other application that kind of more on the formal side, right? So I think there are going to be a diverse collection of models-

**Swyx** [23:34]
Yeah

**Ce Zhang** [23:34]
... both open and close, right? So, and we kind of want to be the engine that power that.

**Swyx** [23:38]
Yeah, for sure. For sure. I think it's, uh, I think it's just like there's a lot of people who own data sources who are doing the locally optimal thing, and humanity as a whole is losing out. So like New York Times is suing OpenAI.

Um, you know, Stack Overflow shut down their API, Reddit shut down their API, uh, X, you know, made their own model, right, on, on Twitter data. Uh, is-- we're just gonna have all these like tiny little gardens of data that it would be useful in a general model, but everyone's just trying to make their own model, and it seems like globally suboptimal.

**Vipul Ved Prakash** [24:10]
Yeah. I, I, I think you need, you need to have some kind of marketplace for figuring out how to get this, uh- Um, you know, data into models and have-- I, I, I think we'll see-- increasingly see more of that.

Uh, and, uh, you know, I, I, I think there's a positive aspect to it, too. There is a incentive for creators to participate in a system which is sort of more fair relative to, um, uh, you know, the capture of value by, uh, any AI company that's, uh, taking their data.

Uh, but, but I, I, I agree. I think this is a big open problem that, uh, uh, that needs to be solved and, um, you know, I, I, I, I hope there will be, you know, serious efforts around it.

**Swyx** [24:57]
Yeah. Yeah.

**Alessio** [24:58]
Um, let's talk about the most precious resource on planet Earth, GPUs. Um, you, you have a lot of compute obviously, uh, but you also have a lot of product pieces. You have inference, you have fine-tuning, you have pre-training.

What's the split in terms of usage? Do you see most people are just running inference on off-the-shelf models? Do you see maybe some last mile fine-tuning? Uh...

### Inference Business

**Vipul Ved Prakash** [25:22]
I would say right now the top five models on our inference stack are probably all fine-tuned versions of open models. Um, and we're seeing-

**Swyx** [25:34]
Who, who fine-tune them? You fine-tune them?

**Vipul Ved Prakash** [25:35]
Uh, either they were, they were fine-tuned by our customers-

**Swyx** [25:39]
By your customers

**Vipul Ved Prakash** [25:39]
... uh, y- you know, either on our platform or off our platform. And, uh, we are generally seeing that. But, you know, that is the sort of trend where you can get better quality on your, uh, on your task by, um, may-- sort of now easily adapting these models to, uh, to your data.

We also have over twenty big model builds happening on the platform with our customer, um, so we, we see a lot of training. Uh, and, um, it's also somewhat surprisingly a more continuous kind of workload. We would say-- we had sort of imagined that this would be more episodic.

You train a model and then you do inference. But what we find is, you know, people train a model, and then they train the next version, and then the next version, which sort of grows in scale. Um, so it's starting to-- it's, uh, uh-- I would say training is still the bigger portion, um, but inference is, uh...

In some ways, inference is super linear to model quality-

**Swyx** [26:44]
Mm-hmm

**Vipul Ved Prakash** [26:44]
... and as the models are getting better, there's more and more inference.

**Swyx** [26:49]
Yeah, um, because they're more useful.

**Vipul Ved Prakash** [26:51]
Yeah, they're more useful.

**Swyx** [26:52]
Uh, so okay. So in-- uh, so training is bigger. This is actually consistent with what we've heard from Mosaic, that, um, you know, people think that training is sort of like a one-time deal. You do one big run, and then you're done.

**Vipul Ved Prakash** [27:02]
Right.

**Swyx** [27:02]
Um, it's never true.

**Vipul Ved Prakash** [27:04]
Right.

**Swyx** [27:05]
Um, and, and so, uh, it-- uh, I, I'm interested in, like, putting some numbers. Uh, I-- and I don't know what you have disclosed or what you, what you want to disclose, but, like, how, how many GPUs do you have?

Right. What is the equivalent amount of compute that you, that you have? Because I, I understand that your GPU setup is different than what pe- people typically think of, like, a giant data center somewhere, right?

**Vipul Ved Prakash** [27:24]
I don't think we have shared this number publicly. It's, uh, you know-- So this will be the first time, I guess. Like, uh, we are-- we have close to seven to eight thousand GPUs today. It's growing monthly. Uh-

**Swyx** [27:39]
Like what class of GPU are they?

**Vipul Ved Prakash** [27:40]
They're mostly A hundreds and H hundreds.

**Swyx** [27:42]
Okay, got it.

**Vipul Ved Prakash** [27:43]
Um, and, uh, probably more, I think, split towards H hundreds now. And we are-- you know, we'll be sort of, uh, uh, building this best of class, uh, uh, hardware. Um, so as there are other versions of these-

**Swyx** [28:03]
Mm-hmm

**Vipul Ved Prakash** [28:03]
... coming out later this year, we plan to have those in the fleet as well.

**Alessio** [28:08]
Um, I know when we talked last year, uh, you were also using some of the supercomputers, um, by the Department of Energy. There was kind of like a lot of, uh, random GPU compute in-

**Vipul Ved Prakash** [28:19]
Right

**Alessio** [28:19]
... in the world. Have you seen that kind of getting timed out? I, I think maybe a year ago, people were like, "Oh, yeah, you can use this GPU computer, uh, that is gonna be end of life." Um, has the bar changed to give access to those resources?

**Ce Zhang** [28:32]
Yeah. So I think from our perspective, it's actually getting, getting better. Yeah, so from the community perspective, because many of the institutions in the world, they are actually investing on hardware.

**Alessio** [28:45]
Hmm.

**Ce Zhang** [28:45]
Right? So for example, we are working with one of the, like, institute, like, in Germany called Hessian AI, right, which give us a lot of help on the compute side. So they start to have this, like, very big GPU cluster, and they are actually sharing that with the community, right?

They start to have-- And it's not super big, right?

**Alessio** [29:01]
Mm-hmm.

**Ce Zhang** [29:01]
But, but also not a small one, right? So you start to see this, like, different labs that sort of pop up, right? And, uh, because of the power of the community, they start to actually share that. So we actually find as a researcher today, it probably easier for them to actually get a GPU than, than last year.

Yeah.

**Alessio** [29:20]
Interesting. And then for you to buy them, what's the state of the market right now? Is it still extremely hard to get any? Do you have Jensen's phone number? Do you have, like, a GM phone number? Do you guys get, like, the DSIR because you are, like, under ten thousand?

**Vipul Ved Prakash** [29:36]
NVIDIA is obviously motivated to, to help us, uh, both, both as an investor, and we are their customers. Um, I would say the market is very tight still, and, um, it's likely going to be this way for a while, uh, is our, um-- is my sense.

That the demand for, you know, AI computing has just kind of ramped up very, very quickly, and it will take a while for supply to catch up.

**Swyx** [30:09]
Can, can you describe how tight it is in-- let's say compared to, like, a year ago, two years ago? Um, what, what do you mean when you say tight? Like, you, you-

**Vipul Ved Prakash** [30:16]
Like-

**Swyx** [30:17]
The things you want, you, you can't get?

**Vipul Ved Prakash** [30:18]
You can't get them immediately.

**Swyx** [30:20]
Yeah.

**Vipul Ved Prakash** [30:20]
They are sort of, you know, um, minimally like two to three months out. Um- Any inventory that shows up tends to clear very, very rapidly.

**Alessio** [30:32]
Yeah.

**Vipul Ved Prakash** [30:32]
Um, and you know, I've-- we've, we've-- we obviously sort of look at this in a very detailed and analytic way. Uh, y- y- there is four to five million GPUs that will be sold this year-

**Alessio** [30:48]
Per year

**Vipul Ved Prakash** [30:49]
...Nvidia and, and others buying. And if you think about, um, you know, five hundred and twelve to a thousand GPU cluster for a company, that's four thousand to eight thousand companies, right? So it's, uh, um, in some ways a very small number.

Uh, in other ways, this infrastructure-- the cost of this infrastruc- the cost of the GPUs will be, you know, eighty to hundred billion dollars. And then you layer servers and data center space and electricity on top of that, that's, you know, close to two hundred and fifty billion dollars worth of kind of, uh, compute, which, when you compare to the cloud computing of today, you know, AWS's last year was eighty-eight billion dollars in revenue.

So this is a, a really kind of a build-out happening of AI hyperscalers. Uh, it is much more disaggregated, um, and it's very, very global. So, uh, you know, we think that GPUs are going to be, uh, sort of a precious resource for a long time, and using them optimally-

**Alessio** [32:09]
Mm-hmm

**Vipul Ved Prakash** [32:09]
...is very valuable.

### Inference Moats

**Alessio** [32:10]
Yeah. Yeah. W- uh, our friend Dylan Patel from SemiAnalysis, he wrote a, a post about the inference market recently and obviously mentioned-

**Vipul Ved Prakash** [32:18]
Right

**Alessio** [32:18]
...you guys. Um, in his post he said, "Our model indicates that Together is better off using two A100 eighty gig system rather than a H-H100-based system. Uh, the temperature and performance testing also pointed Together utilizing speculative decoding." Any thoughts?

Is Dylan right? Or what's-

**Vipul Ved Prakash** [32:37]
What is his, what is his model, man? What, what does he know that we don't know?

**Alessio** [32:40]
Yeah, exactly. I wanna, uh, uh, I wanna know, I, I guess, like, from the outside, and sometimes we even do it, we try and speculate on what people are actually doing. So for the first time, now we have a former guest- ...writing about a current guest.

So we wanna know what, what you guys thought and maybe what are some of the misconceptions that people from the outside have on what it takes to run like a, a GPU cloud today.

**Vipul Ved Prakash** [33:01]
Big fan of Dylan's, by the way. I am-- I'm, uh, uh-- I religiously read, um, SemiAnalysis. I think there were some errors in that analysis. In particular, uh, um, we were trying to decode it, and one of the things we noticed is that it assumed that input tokens weren't being priced.

So I think that may have been an error in the model. Um, I also don't think that, uh, there's this assumption that people are running this at a loss. Um, I, I think it's very expensive. You can't do that for very long.

Uh, and th- there are trade-offs in terms of, you know, batch sizes you use and, um, uh, the kind of tokens per second performance that is, you know, kind of system trade-offs. We've done a lot of work. Uh, you know, this is one of the key areas of research for us.

So our inference stack is a combination of, you know, fifty different sort of tricks and techniques. Um, and we think there's a lot of room for optimization here. So, um, you know, whichever hardware provides better performance, whether it's H100 or A hundreds or L40s, uh, we can sort of measure price performance on, uh, uh, you know, particular hardware-

**Alessio** [34:27]
Mm-hmm

**Vipul Ved Prakash** [34:27]
...and we tend to use that for that, that model. Or, um, you know, in some cases, certain, uh, uh, certain customers have data streams which, uh, can be then optimized for a particular configuration regime. So we do, we do fairly detailed work on, y- you know, how to make this more efficient.

And so i- it's hard to, from the outside just-

**Alessio** [34:54]
Mm-hmm

**Vipul Ved Prakash** [34:54]
...uh, you know, looking at memory bandwidth and estimating, uh, what's, what's actually happening.

**Alessio** [35:02]
How much of these fifty tricks are you keeping to yourself, and how many are you gonna open? Because, uh, we had Treedao, obviously-

**Vipul Ved Prakash** [35:08]
Yeah

**Alessio** [35:08]
...FlashAttention too, is open source. He mentioned he'd love to come work at Together because of how much you care about open source. Um, yeah, how do you weigh that as a CEO and CTO?

**Vipul Ved Prakash** [35:20]
I think a lot of it is open, right? Uh, yeah. Uh, FlashAttention, flash decoding, et cetera. And we publish, um, you know, something that's very really universally useful. It's going to produce better open source AI. We tend to, you know, publish as open source.

I think on the inference stack, there are open source inference stacks, uh, which are pretty good. And it gives us, you know, definitely today it gives us a competitive advantage to have the best one. And so we're not sort of rushing out to release everything-

**Alessio** [35:58]
Mm-hmm

**Vipul Ved Prakash** [35:58]
...about it. It's, uh, not overall that additive to open source out there. Uh, and it is particularly useful as a business for us to, you know, uh, provide best price performance. So we, you know, we make these decisions, we have discussions.

Uh, we-- uh, anything that we keep closed, we generally talk about it quite a bit and decide, like, this is the piece that is closed for today-

**Alessio** [36:26]
Mm-hmm

**Vipul Ved Prakash** [36:26]
...and it may not be the case, you know, six months from now. It may not matter as much. Um, yeah.

**Swyx** [36:34]
Wow.

**Ce Zhang** [36:34]
Yeah. Yeah, so I think, um, being open is kind of very important, right? So I think the whole company actually built on this idea that, uh, open model going to be a kind of... There's going to be ecosystem built on our open models, right?

So... And that's also how we are really lucky to attract this top group of talents to actually join us because of the dream and the w- like, mission that we have on our side to really facilitate the open ecosystem, right?

So I think in general, it's like, I think all the ideas should be, should be open, right? So that's why we publish papers, right? We, uh, talk about ideas, right? So I don't think it makes any sense to keep idea, like, close, right?

So, so there are some software artifact that are kind of really deeply embedded into our kind of own kind of, like, stack. It kind of only useful when you are trying to build a disaggregated cloud, right? So that part, right, so we are, uh, kind of...

Yeah, so that's like maybe at some point that we're going to be open, as Vipul said, right? But at this moment, right, so we are kind of busy actually building it, right? So that's probably kind of getting to the picture about when that piece going to be open, right?

But I think on the research side, the ideas and, uh, for our people to publish things, I think that's really, really important, right? So I think that's how we get talent. That's how I think we as a company going to move the field forward.

**Swyx** [37:58]
Hmm. I, I noticed that you never used the word federated, uh, learning or inference. Is there, is there a distinction that you draw?

**Ce Zhang** [38:05]
Uh, so I mean, it's definitely not intentional, but, but I think federated learning is, have been used in so many different ways by so many different people, it start to lose a very precise meaning about what that really mean, right?

If you go back to the original Google paper of federated learning, I think that's very different from what people are talking about today when they say federated. Yeah, we kind of want to be really precise about it.

**Swyx** [38:28]
And so your term is disaggregated.

**Ce Zhang** [38:30]
Uh, yeah. So as an infrastructure, right? So that's disaggregated, right.

**Swyx** [38:34]
W- I- Aren't most clouds disaggregated? Like, what's different about...

**Ce Zhang** [38:39]
So, um, so, so I think there are different way. So, so one way is that, uh, most of the cloud are disaggregated, but some of that is actually being exposed to the, to the user, right? If you go to AWS, you do know which region you are in, right?

So I think one thing that we are trying to do is you have this, like, disaggregated cloud, not only about, like, location or, like, geographically where they are, but about this, like, reliability and also this diversity of this, like, infrastructure, right?

So, and we want to build a reliable high quality layer over that, that user actually don't know, right?

**Swyx** [39:17]
Hmm.

**Ce Zhang** [39:17]
What's actually happen under the cover.

**Swyx** [39:18]
Hmm.

**Ce Zhang** [39:19]
Right? So I think that one of the difference, uh, that we are, like, uh, of the way that we are thinking about infrastructure. Yeah.

**Swyx** [39:26]
Yeah. A bit closer to Cloudflare than AWS. Yeah.

**Ce Zhang** [39:29]
Yeah, one way to look at it.

**Swyx** [39:30]
Yeah. We have one question here, which I, we'll just throw out. It's, it's kind of fun. Um, so going back to this sort of inference stack, uh, piece, um, maybe if you had to pull out, like, a call for researcher or call...

Or just, like, point out interesting areas of work that you're, that you're interested in, what pieces of the stack have the most opportunity for im-improvement?

**Ce Zhang** [39:49]
Yeah. So, so, so I think the way we are thinking about the inference stack is... So there are multiple things can happen, right? So you can do better algorithms like speculative decoding. Uh, you can change the model architecture.

Uh, you can go really crazy on the, on the, on the system side, right? And you can also code it on the hardware, right? So it's not really clear innovation on single dimension will get you there. Yeah. So the key thesis on our side is if you only push on one direction, you are going to reach diminishing return really, really quickly.

Yeah, there only that much you can do on system side, only that much you can do on algorithm side. I think the only big thing that are going to happen is when you ask all those dimension to actually compound, right?

So to have algorithm, model, and system all come together. So I think that's how we reach the next, like, 10 times improvement on inference, right? So I don't think there's a single dimension that is particularly important, but looking at this space in a joint way, right?

Try to co- kind of, kind of co-optimize jointly multiple dimensions, I think that's going to be really important for the company to look at. Yeah.

**Swyx** [40:59]
Yeah.

**Vipul Ved Prakash** [40:59]
Yeah, we often see, uh, I see numbers from the, from the team, and you have these multiple methods. Not all of them compound.

**Swyx** [41:06]
Hmm.

**Vipul Ved Prakash** [41:06]
So you mix these together, it's, you know, still similar results, and some combination of them will have this incredible effect that, uh, uh, is really, really super interesting. So it's, uh, very systems, uh, you know, very kind of broad systems approach to it that's the, the most effective.

**Swyx** [41:27]
Um, I think I finally get the name of the company, like-

**Vipul Ved Prakash** [41:30]
Bring it together, yeah.

**Swyx** [41:31]
Everything needs, everything needs to be optimized together.

**Alessio** [41:33]
Um, all right. Uh, just quickly, how does all this work change, just like some of the architectures change? I know a mixture of experts, like speculative decoding, is a little less efficient because of memory bandwidth. Um, how much of it do you invest when it's a maybe model specific improvement versus more horizontal thing?

Um, also you're researching different architectures. So how much do you want to spend time optimizing what's state-of-the-art today versus what's coming next?

**Vipul Ved Prakash** [42:01]
We, we do spend time on what's state-of-the-art today, uh, as well as what's next. It's, uh, um, you know, the value we get from doing specific optimization, even, even for, you know, what works well for a particular model on A hundreds with a particular bus, uh, versus H hundreds, it's a worthwhile investment for us.

Uh, so we will go down fairly deep into a specific architecture and specific hardware. Um- You know, w- the, the, the, it, it, it does also inform what works better where, and you don't have to take the same approach for, you know, every model, um, and every sort of hardware setup.

We can take these different approaches, and we do have these multiple systems now. We know that this, you know, system B is better for, uh, mixed raw, and system C is going to be better for stripe tying or Mamba.

**Swyx** [43:04]
Before we move on from inference, uh, we need to talk about Anyscale drama. Uh . So we're, we're actually having Sumeet on the podcast tomorrow, who also talked about, uh, kinda came to your guys', uh, support-

**Vipul Ved Prakash** [43:20]
Defense

**Swyx** [43:20]
... about how, yeah, how important. It's not just like, oh, Together is saying this benchmark's not good because they look bad in it. Uh, how... I guess, like, it's a hard question to ask, but, like, h- why did you decide to just come out and say it, and how maybe does that also reflect the values that you guys have about open source and openness and kinda like being transparent about what's, what's real, and maybe hopes for standardizing some of these benchmarks to, to make it more clear?

**Ce Zhang** [43:51]
Yeah. So, so, so, so I think first one is like, so it's a great service Anyscale is doing for the community, right? So, so, so I mean, it's very hard to do benchmark. The moment do benchmark comparing end players, right?

N minus one will be unhappy.

**Swyx** [44:05]
Yeah.

**Ce Zhang** [44:05]
You have two tables and maybe n log of them are happy, right? So, so, so it's a very great thing they are doing. And in some of the work that we are doing, we actually use, uh, uh, RMO perf, right?

So, so it's, it's a great thing that, that they're actually doing. So, so I think one thing that about benchmark is, and probably the professor part of me are talking, is a good benchmark should think about how it's going to incentivize the field to actually move forward, right?

So if the benchmark really become kind of standard, how are people going to over-optimize to the benchmark? Because people are going to do that. And when people are doing that, what are we actually trying to incentivize, right?

**Swyx** [44:46]
Mm.

**Ce Zhang** [44:46]
Will that move the world to a better place, or will that essentially have every single player focus on marketing or spending time or money on something that actually do, do not matter on technical side, right? It's very hard to actually strike a balance, right?

So I think, uh, the reason we kind of try to give feedback on the benchmark is kind of want to, yeah, so want to open up the discussion about how does the industry should come together and define maybe a common way that we compare with each other, right?

So like how, like database people doing TPC, right?

**Swyx** [45:19]
Mm.

**Ce Zhang** [45:19]
Maybe we should have something actually similar, right? So we are trying to start some of the conversation. So just... it's not really like we jump out to say it's not good, uh, because, I mean, there's no way we can have a perfect benchmark, and that's not really exist, right?

So just try to kickstart a conversation that maybe we should come together and do something that, um, the community agree and align with the benefit the end user are going to get, right?

**Swyx** [45:45]
Mm.

**Ce Zhang** [45:45]
So just get the conversation started. Yeah.

**Vipul Ved Prakash** [45:48]
Yeah, no, I, I, I've, I've spoken to the Anyscale team after that, and I think they had really great intentions. Um, and partly, I think it felt like the, you know, it felt like very objective. Um, but, uh, and everyone sort of had a reaction to it because it just didn't match their-

**Swyx** [46:09]
Mm-hmm

**Vipul Ved Prakash** [46:10]
... uh, benchmarks that we've all run internally against different services. Um, but I, I, I think, uh, you know, a common industry benchmark run by s- an independent party versus one of the vendors, uh, you know, uh, uh-

**Swyx** [46:27]
Is there one that you would point to?

**Vipul Ved Prakash** [46:29]
I, I don't think one exists today. Uh, I, I think there should be. We're having some conversations about, uh, someone setting one up.

**Swyx** [46:37]
Yeah.

**Vipul Ved Prakash** [46:37]
Um, and you know, there, there's lots of interesting aspects of this. You know, time to first token is a function of where the test was run from. Um, there is different load on these services at different times-

**Swyx** [46:49]
Mm

**Vipul Ved Prakash** [46:49]
... of the day and, and, and, you know, weekday or weekend, so you have to measure that well. And I think if all of that were done very well by an independent source, that would be a very useful service to customers and, and, and, and the services themselves.

**Swyx** [47:09]
Yeah. I'll point people to artificialanalysis.ai-

**Vipul Ved Prakash** [47:12]
Yeah

**Swyx** [47:12]
... uh, which is a new one that recently emerged. I don't know if they've done it right.

**Vipul Ved Prakash** [47:16]
I-

**Swyx** [47:16]
It looks like a side project of-

**Vipul Ved Prakash** [47:18]
Right

**Swyx** [47:18]
... a couple people. Uh, but I think it's in all the providers' interest to work with them and ensure that there's an independent third party that's measuring these things, right?

**Vipul Ved Prakash** [47:27]
Yeah.

**Swyx** [47:27]
At, at least on the baseline. For me, what's, what's worrying is more of what, what Sumeet's saying was, which is, um, do these benchmarks skew things in ways that customers, um, might not be mindful of? Like, what are, what are these things overemphasizing that we might be, uh, that, that we might be missing?

And, and, and I don't really know. Uh, it seems like a lot of these services, a lot, a lot of different services bundle together their c- their version of quantization as well, so that means there's performance trade-offs, right?

You're not comparing apples to apples, the same model itself, even though it's like a-

**Vipul Ved Prakash** [47:59]
Right

**Swyx** [48:00]
... Llama variant or whatever. So what, what do people trade off? They trade off latency, they trade off price. Obviously, the, the, those are the first two, but what else, right? What, what are th- what factors matter in a inference business?

Right? It's an open, open question.

**Ce Zhang** [48:13]
Yeah. So I think there's also the throughput, right? So there is the time to first token, right?

**Swyx** [48:18]
Yeah.

**Ce Zhang** [48:18]
So, and then there are things that users do not often see. For example, the reliability, right, the capacity, right? So that also have impact on, on user experience at global scale, maybe not a single query, right? But in aggregation, you can also see a whole bunch of, like whether you are emphasizing p fifty, p ninety-five, right?

So there's a whole bunch of things that you can actually play with. And of course, there is, there's also quality, right? So there are different ways to actually make the whole thing faster, uh, sparsification, quantization, or combination of those, right?

So yeah. So there are so many things to actually play with, so they probably need a benchmark that- The protocol is transparent- ... to make sure, like, it's very clear what we are doing, and a whole bunch of check on the quality to make sure we are putting the right group of service in the same table.

Right. So, and s- and so I think then essentially the user can actually navigate the space.

**Swyx** [49:10]
Yeah.

**Ce Zhang** [49:10]
Right. So I think that's going to be good for everyone.

**Swyx** [49:12]
It's very important field, and, you know, I think, I think hopefully there, there's a good third party that emerges from this. Um, so I just want to touch on one more piece, which is I, I think, um, I think I am appreciating from this discussion that fine-tuning is a bigger part of your business than I thought.

### Beyond Inference

**Swyx** [49:26]
Um, the other big player in fine-tuning is, is Mosaic, or, uh, well, Mosaic is more training, but, like, there, there's a bunch of other players in the fine-tuning space. Like, what-- If I was a prospective fine-tuning customer, what do I come to you with?

Uh, do I come to you with my custom data and that's it? Uh, do I also have to, uh, write the fine-tuning code? Um, what level of engagement do you do with your customers?

**Vipul Ved Prakash** [49:49]
I think across the spectrum. So there are, um, y- you know, our customers are training models, pre-training models from scratch, and they, many of them will bring their data sets and, uh, you know, use our infrastructure and training stack to train their models.

There are others who, uh, you know, have trained smaller models and want to scale up, uh, scale up across infrastructure or scale up across data, so we'll sort of help them do that. Um, we will have customers who are sort of initially started a little bit more consultative.

They have a particular task and idea in mind, and we will help them get from there to the data set and, and, you know, the right model to achieve that data. So it's a, it's, it's a spectrum, and, um, you know, our goal is to-- We're trying to productize as much of this as possible so that, uh, the whole process can be fast and, and scalable.

Uh, I would say there is a lot more understanding around fine-tuning now. Like even the last six months, there are, you know, source tools, recipes, um, literature, podcasts- ... uh, Discord channels where, uh, people are figuring out. And it really is in, in many ways one of the successes of open source is you have small collectives of, you know, um, uh, engineers who have created-- who are now creating the top models on open source leaderboards and have tried out all sorts of different sort of, you know, data recipes, creating synthetic data, um-

**Swyx** [51:41]
Merging models

**Vipul Ved Prakash** [51:42]
... merging models. Uh, so it's, uh, that's really fun to see, and, um, I think that, that sort of agency that exists now is, um, uh, I- exciting and that is w- we, we see a lot of that sort of being applied into, uh, products and-

**Swyx** [52:04]
For sure

**Vipul Ved Prakash** [52:04]
... you know, more sort of commercial, um, more commercial models that people are deploying in their applications.

**Swyx** [52:10]
Um, and then just to, I guess, wrap up the Together, it's almost becoming like a platform-

**Vipul Ved Prakash** [52:16]
Products. Yeah

**Swyx** [52:16]
... as a service.

**Vipul Ved Prakash** [52:16]
Product suite.

**Swyx** [52:16]
You know? Uh-

**Vipul Ved Prakash** [52:17]
Right

**Swyx** [52:17]
... because then you release Together Embeddings. Um, how did you get, uh, ninety-two point five accuracy on thirty-two K retrieval? Um, and do you think we're kind of like getting to embeddings or just like we did everything that we could, you know?

We're getting to, like, the most optimized it's gonna get, and then we should just focus on models and inference? Or do you think there's still room there, um, to improve?

**Ce Zhang** [52:39]
Oh, I don't think we haven't, like, even got started on embedding.

**Swyx** [52:42]
Mm-hmm.

**Ce Zhang** [52:42]
Yeah. So I think, so I think there are so many things. So, like, embedding is really fundamental for many things. For example, RAG, right? So deep in application, so that's how people bring knowledge in. That's also the fundamental piece when you want to build a better model, right?

So that give you this understanding about what actually get into the model. You can actually use that to actually build a better data set, get a better model, then get better embedding, you start this loop, right? Without the good embedding, the loop is now closed, right?

So I think both on the quality side, how to embed more, like, dedicated semantics, like, into those vectors, right? How to deal with negation, for example, right? So and, uh, how can you make the whole thing really, really fast, right?

So I don't think we have, like, scratched the surface, like, even, like, even a little bit. So, so, so I think for the next couple years, yeah, we will see a whole bunch of new embeddings, uh, maybe of different size and, uh, much, much faster than today.

So I think, yeah, so I think it's a very active research area. I think people should invest more. Yeah.

**Swyx** [53:46]
Yeah. I was surprised to see, um, I think Jina or... Yeah, there's JinaAI.

**Ce Zhang** [53:51]
Yeah.

**Swyx** [53:51]
And then there's, um, another guy, Teng, um, Teng Yu's Voyage.

**Ce Zhang** [53:55]
Yep.

**Swyx** [53:56]
Um, just they're, they're only-- they, they are coming out as startups purely focused on embeddings.

**Ce Zhang** [54:00]
Yeah. Yeah. So yeah, so, so I think it's a very, very important piece of the, the system, right?

**Swyx** [54:06]
Yeah.

**Ce Zhang** [54:07]
So you people haven't focused on, a lot on them before, and they should definitely start to do that.

**Swyx** [54:12]
Yeah. Why are the Chinese universities so good at embeddings? You, you know, you know what I mean, right? Like the BGE and-

**Ce Zhang** [54:19]
Yeah, yeah, yeah. So actually, I don't know. Yeah. So I think embedding is, is something that... I don't know. We just, like, released our first embedding model, so we're still trying to learn how to build a, build embedding model.

Yeah. So ask me again in six months.

**Swyx** [54:35]
Okay.

**Ce Zhang** [54:36]
I'll probably have more insight about how to build a better one. Yeah.

**Swyx** [54:38]
I just noticed that you-- So, um, Ada 02 was used to be at the top of the MTB chart, and then it's just, like, sliding down and down and down and, and all the new models are coming out of China for some reason.

**Ce Zhang** [54:47]
Yeah.

**Swyx** [54:47]
And I'm like, I don't know what's going on there.

Um, okay, cool. Um, so we, we cannot leave this discussion without talking about State space models. Uh, well, first of all, uh, how much of the company is dedicated to research? Like it's, it's obviously, like, not production quality yet, but like-

**Vipul Ved Prakash** [55:03]
I think it's, it's like forty, forty-five percent I was counting this morning

### State Space Models

**Swyx** [55:08]
That's huge.

**Ce Zhang** [55:08]
Yeah.

**Swyx** [55:09]
Yeah.

**Ce Zhang** [55:09]
So that's, uh, pretty big.

**Swyx** [55:10]
That's a big investment. Yeah.

**Ce Zhang** [55:11]
Yeah.

**Swyx** [55:11]
Okay. Well, I mean, it looks like it's paying off, so, you know. Uh, but, uh, so and then, and then high level, um, I will, I will confess or admit or mention, uh, for the listeners who are also s- similarly skeptical, I did not use to care about long context because I was like, you know, 30K is enough, 100K is enough, right?

I'm not, you know, modeling DNA sequences or anything like that. Why do I need long context? Um, and, uh, I mean, I'll, I'll first of all, I'll throw that open to you, but second of all, I think what Mamba did for me was change that perception of that it's only about a long context.

Like, you, you, like, the only reason you want some quadratic architectures is for long context. Actually, that's not true, it is also just more efficient to train, period, right? I'll just leave that open to you. Like, what, what's the motivation, uh, that people should keep in their heads?

**Ce Zhang** [56:00]
Yeah. Yeah, so I think there are, there, there are multiple things, right? So one thing is that, I mean, the moment a model can do for long context well, so it often means that it's kind of cheaper. Yeah, so I mean, that's why it can do long time.

I mean, in principle, transformer can do long context, it just very expensive, right? So I think what those, like, uh, state space model is trying to do is try to push the size of the state, right, uh, like as small as possible.

That's why it can do long context, right? Uh, and try to kind of like decouple this like ch- quadratic dependency, right, to make sure you can have a much better execution pattern, right? So all of those, like one direct consequence of those is you can do long context really cheaply, but on the other hand, also introduce a whole bunch of benefit even you are not doing long context, right?

So I think that's actually probably, like equally important, right? Because state gets smaller, you can do really large batch size, right? You can actually be very faster, right? So yeah. So and uh, another thing is like one of the hypothesis that we have is for them in, like in Stripe Hyena, it start to have a hybrid architecture, right?

It has, part of it, it has like state space model, and part of it is, is still a transformer, right? So different component probably deal with different things kind of better, right? So maybe by putting them together, by thinking about how information propagate, right, over this whole horizon of this context, you can probably get a even better quality model than transformer, right?

So I think that's why we are kind of invest a lot of things, right, on those models, not only for the context, which is very important, but also for a whole bunch of benefit it could get. Yeah.

**Swyx** [57:45]
How, you know, how, how should people treat the distinction between Mamba and Stripe Hy- Hyena? Like, what's the strat- what's the point of releasing these two as separate models? Um, is, is, is one like sort of the together proprietary one, and then the other one is like the more open research one?

**Ce Zhang** [57:58]
Yeah. So I think it's, it's pretty much a different stage of exploration.

**Swyx** [58:02]
Okay.

**Ce Zhang** [58:02]
So they kind of have different hypothesis, uh, when we try to, uh, when we try to build those. Yeah, uh, like first one, they are different view about state space model.

**Swyx** [58:10]
Okay.

**Ce Zhang** [58:10]
One is Hyena, another is like Mamba, right? They're actually different architecture.

**Swyx** [58:13]
Different families, yeah.

**Ce Zhang** [58:14]
Yeah. So, so, so, so when we build Stripe Hyena, right, so the curiosity that we had is how good can we-- So c- so what is the highest quality non-transformer model we can ever build? Yeah, so the goal of Stripe Hyena is try to see whether we can match Mistral.

Yeah. And by finding well whether we can, uh, outperform that in, in, in some way, right? So it has a very, very strong baseline that, that we are trying to beat. So that's why this hybrid thing like, like get in the picture, right?

And for Mamba it's kind of more the curiosity was, yeah, so how far can we push for pure architecture, right? So like then we start from this very system like from small to large, right? Like, like, uh, all the way to three billion, right?

So baseline was essentially the best three billion model. So I guess at different stage of exploration, at some point I think they are going to converge. We actually learn different things like when building different models. Um, I think they are just like this intermediate stage in the exploration at different point.

Yeah.

**Alessio** [59:21]
You mentioned the hybrid architecture. Is that the model grafting that you mentioned in the Stripe Hyena post, um, where you mentioned you can have transformers and, um, and not together? Like, it, this is a concept that I hadn't heard before reading about this, so I think most people's mental model is like transformers or something else, it's not transformers and-

**Ce Zhang** [59:45]
Yeah.

**Alessio** [59:45]
... something else. Um, how do you train a model that is hybrid? Is there any difference in like how you construct your data sets? Um, is there any difference in then how you'd run inference on it? Um, how should people think about, um, starting research in this field?

**Ce Zhang** [1:00:00]
Yeah. Yeah, so, so, so we were also very surprised. Yeah, so when we, when we come up with this hybrid architecture, right? So, so the way to think about it is like you have different layers in the neural network, right?

**Alessio** [1:00:10]
Mm-hmm.

**Ce Zhang** [1:00:10]
So like the state space model for some layer will already give you the benefit. For the other layer, um, they could be transformers, right? They could give you this more global view of the sequence, but for maybe for other layer don't have to have that, right?

Then can have all the other things that kick in, right?

**Alessio** [1:00:28]
Mm-hmm.

**Ce Zhang** [1:00:28]
So we don't know what is the optimal mixture between different architectures. I mean, in principle, we can have Mamba, Hyena, and Transformer, all those things, right, come together, right? And then you can see what makes sense. Uh, we have no idea what is optimal doing that.

So what we are excited about is now the community have a whole bunch of building blocks that they can actually like playing like a Lego, right? So just put together and, and, and see what happen, right? So we are kind of very excited about that.

So and, uh, yeah, we are in the process of try to learn more like, like, like, like, uh, about this architecture, and, uh, when we know what we are talking about, we will definitely share with the community about how to do that in systematic way.

Yeah.

**Swyx** [1:01:08]
What are we still unsure about? Like, why don't we just- ... you know, put all the money in the world in training these things now. Like, what, what is left to figure out before we scale this thing?

**Ce Zhang** [1:01:20]
Yeah. So, so like if you look at how Transformer, like, has been developed, right, in, in the, in the last, like, five to 10 years, right?

**Swyx** [1:01:26]
Yeah.

**Ce Zhang** [1:01:26]
So people don't start from, like, you have this attention to all you need the paper, and then let's put all the money in and then train that, right?

**Swyx** [1:01:33]
Yeah, yeah, yeah.

**Ce Zhang** [1:01:33]
Always start from this very systematic, uh, understanding about the scaling, about data quality, uh, about essentially the limits, right? So I think for state space model from, from the labs to, to the real world, you're going to go through the same process.

But of course, the second time doing that is kind of easier, right?

**Swyx** [1:01:53]
Mm-hmm.

**Ce Zhang** [1:01:53]
So, so, so by that saying, there's no way we can get rid of this systematic style of studying scaling law, study what data to put in, right? So what's the impact of different data sizes to the data, to, yeah, to the final model quality.

Yeah.

**Swyx** [1:02:06]
Do, do you expect that the data inputs will be different than-

**Ce Zhang** [1:02:11]
I don't know.

**Swyx** [1:02:12]
Okay.

**Ce Zhang** [1:02:12]
So I mean, that's... But I wouldn't take that for granted that they should be the same.

**Swyx** [1:02:16]
Huh.

**Ce Zhang** [1:02:16]
Right? So that's one of the hypothesis that... So, so we, so we have no opinion on that because I think that's the result of the study, not the assumption. Yeah, we, we do not need to assume that.

**Swyx** [1:02:28]
Okay.

**Ce Zhang** [1:02:28]
Yeah.

**Swyx** [1:02:28]
Scaling laws and data. Anything else like architectural that we are not sure about? 'Cause now you have this selection mechanism that you're-

**Ce Zhang** [1:02:35]
Yeah, so-

**Swyx** [1:02:36]
... pretty happy with.

**Ce Zhang** [1:02:36]
I mean, first one, how to mix them, right? So, so and, uh, and second is what is the architecture? So, so if, so if you look at Transformer, right? So one very interesting piece there is people optimize also the hardware to, yeah, to make sure that things run very fast, right?

They're very efficient kernel. They're very efficient hardware, and then that add another boost, right, for the Transformer architecture, right? So, so I think that sounds, yeah, that sounds that should happen for state space model, uh, which architecture is kind of easier kind of to run on the hardware, right?

So go, so go, things going kind of faster, you can put more data. It add another dimension in the scaling law, right? So I think we just need to plow the whole space and just... So be really systematic from small model to, to one billion, three billion, seven billion, just go all the way up, right?

So I wouldn't jump around in the space. I would just, like, be patient and just, like, be systematic. Uh, and, uh, yeah, I think we'll get there. Yeah.

**Swyx** [1:03:38]
Yeah. Well, looking forward- ... for more research from you guys to-

**Ce Zhang** [1:03:41]
Yeah

**Swyx** [1:03:41]
... figure that out. So one dimension which we didn't talk about, at least we talked about long context, we talked about efficiency, uh, but speed is very, speed is also very important. Um, a good inference provider provides, let's say, 70 tokens per second, and then, you know, maybe that's faster than less good, uh, inference providers that are more like 30 tokens per second.

But that's the, that's the rough range, right? State-of-the-art today. Um, y- that's around the human speaking speed. Um, human reading speed is what? 200 words per minute. Uh, words per minute. Yeah, it's word, words per minute. Anyway, so like why do we need 5,000 tokens per second is, is what is my question back to Vipul.

And may- maybe is, is this something that is a emphasis for research as well, or is this more just an inference-only thing?

**Vipul Ved Prakash** [1:04:25]
You know, there, there are applications that are, you know, consuming the tokens that are produced from end models.

### Speed

**Swyx** [1:04:31]
Yeah.

**Vipul Ved Prakash** [1:04:31]
So they're not necessarily being read or-

**Swyx** [1:04:34]
By humans

**Vipul Ved Prakash** [1:04:36]
... heard by humans.

**Swyx** [1:04:36]
Yeah.

**Vipul Ved Prakash** [1:04:36]
Uh, so that's, that's a place where we see that level of, uh, requirement today that, uh, really nobody can quite satisfy. Um, you know, there is kinda think about how do you,

as intelligence grows, how do you sort of increase the bandwidth of, uh, you know, how do you reduce the latency of it? Um, if we can do 5,000 tokens a second, the same card can produce y- y- the throughput of that card goes up significantly, um, and can support, you know, support more, more applications.

So I think it's important from that perspective. Um, and then there are, it opens up new UX possibilities. Once you can get sort of an immediate answer from a model, uh, it starts working in a different way and, um, you know, new types of applications will be created.

We are, uh... We rarely run into users, except for perhaps those feeding this into a text-to-speech, uh, model where, um, you know, they say that, "Okay, slower is better," or like, "We don't need more, more performance." So I, I think there is, uh, I, I think this may just be fundamentally very, very slow today in general, and we're just sort of used to that speed, and that will change once, uh, you know, these models can get faster.

**Swyx** [1:06:03]
Yeah, 5,000 tokens per second is, uh, I don't even imagine. Like, well, it's a lit... It makes me worried a bit that the machines will be communicating at a much higher bandwidth than us. Uh, but yeah.

**Vipul Ved Prakash** [1:06:14]
I mean, they do, they do that already.

**Swyx** [1:06:16]
They do that already.

**Vipul Ved Prakash** [1:06:16]
It's not in natural language.

**Swyx** [1:06:17]
They do that already.

**Alessio** [1:06:19]
Awesome. Anything we missed about Together as a product? Um, we're gonna talk about the hackathon you just did and whatnot, but, uh, any last product thoughts?

**Vipul Ved Prakash** [1:06:31]
I think one of the big sort of focus of, focuses of our product is to become more and more serverless, like have AI run in a, AI development run in the serverless manner. Uh, and we are there now on inference, uh, also on fine-tuning.

You know, we are pushing to do that on training. Um, and that is, you know, we think, um, if there is a sort of y- you know, developer experience message, that's probably the big one, is where you have enough flexibility.

You don't have to sort of commit to, um, you know- Thousands of dollars of compute before you can start using open models. We s-- we really wanna change that and, uh, make it really as easy as possible to get started.

**Swyx** [1:07:24]
Yeah. W-when I first signed up for Together, I had-- I, like, left an instance up running, and I just, like, ran out of my credits immediately.

**Vipul Ved Prakash** [1:07:31]
Yeah. So, you know, and, and we changed that whole model now. Uh, so you, you never run into that issue. And that was, you know-- And I, I think the response to that has been amazing, is, uh, um, we also provide, you know, twenty-five dollars free credits, which is a large number of tokens depending on the model you're using.

Uh, and you really can build an app, or you can do a f- you know, you can do a fine-tuning and run that model and build an app on Together for free, basically. Um, and, and really pushing further in that direction.

**Alessio** [1:08:06]
You just did a hackathon at AGI House about fine-tuning versus RAG for open source. Any learnings, recaps from it?

**Ce Zhang** [1:08:14]
Yeah. So I think one thing that we kind of learn is, like... So I think the hackathon was phrased as, like, uh, uh, something versus something, right? But I think the combination of those works really well, right? It's like, like, yeah, so, so I think, like, combining all those techniques all together, right?

So will give you essentially another boost, right? So that kind of one thing that we learn on the, on the technical side. Yeah. And also, we are very kind of excited about the excitement of the audience, right? So I think people are really kind of using the platform and building something.

They're really cool.

**Swyx** [1:08:49]
Yeah.

**Vipul Ved Prakash** [1:08:50]
It's always surprising to us what people build.

**Swyx** [1:08:52]
Yeah.

**Alessio** [1:08:53]
Um, is there something you're focused on this year? Hiring, building engineering team? What should people that wanna work at Together, um, know?

**Vipul Ved Prakash** [1:09:01]
Y-y-you know, all those things. I, I, I think, um, hiring is, uh, a pretty big topic. We are, um, thirty-eight people in the team, um, and we are hiring across all the areas. Um, you know, if you're a, a CUDA and kernel hacker, we're-- we have lots of exciting projects.

Um, if you're a, a researcher, you like to build models, we have exciting projects. If you work on systems and infrastructure and the cloud layer, um, you know, we, we do a lot of work there. And, uh, as, as well as sort of front-end and developer experience and applications.

Uh, so really kind of across the board. We have, I think, twenty-plus postings on our, uh, job openings on our site. Uh, and, and, and folks who are passionate about open and, you know, AI. Um, I also say if you, if you, y-y-you know, people looking at Together, they don't necessarily, for all the postings, have to have experience, um, you know, professional experience working in machine learning or AI.

Um, uh, many of the systems people are sort of doing this for the first time, and they can apply their, uh, you know, systems expertise to, um, to the kind of things that we are doing. And we can, we can teach people AI, uh, uh-

**Swyx** [1:10:31]
Mm-hmm

**Vipul Ved Prakash** [1:10:31]
... as long as they have expertise in other areas.

**Swyx** [1:10:33]
W-will, will you call out what kind of expertise you're looking for? Like, we, we definitely have systems people listening, so.

**Ce Zhang** [1:10:39]
Oh, I mean the whole, the whole stack, right? So, like, all the way from the-

**Swyx** [1:10:43]
Like Kubernetes? I, I don't know.

**Vipul Ved Prakash** [1:10:44]
Kubernetes, yes.

**Ce Zhang** [1:10:45]
Yeah, Kubernetes.

**Swyx** [1:10:46]
Yeah.

**Ce Zhang** [1:10:46]
CUDA.

**Swyx** [1:10:46]
What else?

**Vipul Ved Prakash** [1:10:46]
CUDA.

**Ce Zhang** [1:10:47]
Yeah. So, and, uh, DevOps, right? So that's a-

**Swyx** [1:10:50]
Yeah

**Ce Zhang** [1:10:50]
... that, yeah, that's a big thing.

**Swyx** [1:10:51]
Is that, like, what, Terraform? Like, Pulumi?

**Ce Zhang** [1:10:53]
Right. Right. Yeah.

**Swyx** [1:10:55]
Okay.

**Ce Zhang** [1:10:55]
Yeah. And the, and all the way to machine learning systems, right? If you want to, like, like, to hack over, like, VRM, TGI, right? That's great.

**Swyx** [1:11:02]
Okay.

**Ce Zhang** [1:11:02]
If you want to play with different fine-tunes or build the models, like, development algorithms, right? Essentially the whole, the whole stack.

**Swyx** [1:11:10]
Okay.

**Ce Zhang** [1:11:10]
All the way from application to-

**Swyx** [1:11:11]
That's very broad.

**Ce Zhang** [1:11:13]
To system, right.

**Swyx** [1:11:14]
Okay.

**Ce Zhang** [1:11:14]
So yeah, so I think that, like, like, uh, uh, so a fun thing about the company is, like, we have this very diverse collection of expertise and talents in the company.

**Swyx** [1:11:24]
Yeah.

**Ce Zhang** [1:11:24]
Uh, and the goal is to really try to innovate at every single layer.

**Swyx** [1:11:27]
Okay.

**Ce Zhang** [1:11:28]
And then have them all compound together and, yeah.

**Swyx** [1:11:32]
Yeah. Every- doing everything together, that's why the company is, uh, named this way. Like, uh, no, seriously, I, I, like, didn't really get the company naming until now. Like, yeah, makes sense.

**Alessio** [1:11:41]
Um, awesome, guys. I know we kind of bent on the lightning round in the, in the last few episodes, but I, I think for you two, the, one of the questions we used to ask is, like, what's the most interesting unsolved question in AI?

Um, so maybe another way to think about it is if you weren't building Together, what would you be working on?

### Wrap-Up

**Ce Zhang** [1:11:59]
Yeah. So if not building for-- uh, if I'm not building Together, I'll be a professor. I, I mean, let me do all the, like, whole, like, whole bunch of other things without justifying as being useful. We used to work on quantum machine learning for a while, right?

**Swyx** [1:12:13]
Wow.

**Ce Zhang** [1:12:13]
So, so, so I think that's cool. Right. So I think, um, I'm very excited about... So, so I think IoT is going to become very interesting. Yeah. So yeah, so I know people have been saying that for, yeah, for the last couple decades, right?

But I think very excited about how does technology like, like Starlink, right? So, so, like, change the communication between different edge devices and the, like, like, like, all those machines, and the new battery coming out, right? So I think that could be very cool.

Yeah. So if not building Together, probably, yeah, spend some time thinking about how to compress communication even more, given all the satellite communication stuff. Yeah.

**Vipul Ved Prakash** [1:12:56]
I, I think sort of to the first question of what is the most-- what, what's one of the more important open questions, um, the one thing I think about is that we sort of need a framework of thinking about, uh, you know, what the world looks like.

With, um, advanced intelligence systems in it. I think we have had this, uh, very, um, you know, sort of a doomerism, uh, view of it, uh, really kind of informed by science fiction, you know, dystopian science fiction and Terminator.

And I don't think we have a kind of a positive or, or a realistic really, um, framework coming from, you know, experts in the field. Uh, so I, I think, I think that's a pretty important question because that really gives us a roadmap of where this industry should go.

And, um, uh, you know, I'm, I'm hoping that, uh, some of the, you know, industry drama this last year maybe is sort of pointing us in that direction. Um, and solving that is sort of, I think, im-important in, uh, kind of a, i-i-in a, in a meta way.

Um, I'm actually not sure what I'd be doing if I was not doing Together. So I, I think I'm doing the perfect thing that's like, this is the-- this is, you know, really, uh, um, my dream job, and I, uh, have ev-every day, this is kind of what I want to do, and I expect that's going to be the case for a very long time.

**Alessio** [1:14:45]
Awesome. Uh, thank you guys for coming on. This was a lot of fun.

**Vipul Ved Prakash** [1:14:49]
Thank you so much.

**Swyx** [1:14:50]
Thank you.

**Vipul Ved Prakash** [1:14:50]
Yeah.

---

This library is powered by PodHood (https://podhood.com), the podcast website platform.
