LALatent SpaceOct 1, 2025· 29:14

⚡️Raising $1.1b to build the fastest LLM Chips on Earth — Andrew Feldman, Cerebras

Andrew Feldman, CEO of Cerebras, joins Latent Space to discuss their $1.1B fundraise at an $8.1B valuation and their wafer-scale chip that delivers 20x faster inference than NVIDIA's B200 GPUs. He explains how their architecture uses SRAM instead of HBM, providing 2,625x more memory bandwidth by eliminating the narrow straw between compute and memory. Feldman details the decision to accelerate sparse linear algebra rather than specialized convolutions, enabling support for transformers and diffusion models unseen during design. He discusses the explosive growth in AI inference demand, the shift from closed-source to fast open-source models, and the importance of speed—citing Paul Graham's observation that ChatGPT's slowness drives users away. The conversation covers enterprise trends in the 10-30B parameter space, the complexity of building data centers that pull gigawatts of power, and the often-overlooked routing and caching systems that make AI work seamlessly.

  1. 0:00Raise
  2. 1:27Context
  3. 3:50Investors
  4. 5:08Speed
  5. 8:41Large Chips
  6. 12:07Workloads
  7. 13:41Trends
  8. 18:52Cloud
  9. 21:39Deep Dive
  10. 25:37Vision
  11. 28:40Outro

Powered by PodHood

Transcript

Raise0:00

Swyx0:03

Okie dokie. We're in the remote studio celebrating a big, big raise with Andrew Feldman of Cerebras. Welcome.

Andrew0:11

Well, thank you for having me.

Swyx0:13

It's such an honor to finally have you here. Uh, we've been sort of poking around, um, this sort of Cerebras ecosystem for a while. I run the AI Engineer Conference. I run Latent Space. I'm friends with, uh, your entire marketing team, of course.

And also more recently, I joined Kognition, which is a, a big customer and, and friend of, of Cerebras. Uh, but here, I mean, let's, let's get the big news right out of the top. What are you announcing today for your fundraise?

Andrew0:35

So we announced a, a $1.1 billion fundraise that we had completed. It was done at an $8.1 billion post-money valuation, and it was led by Fidelity and Atreides Management. So we, we, we're announcing that we raised the most money on the highest valuation with the best investors in, in the category.

And so, uh, it's a good day. I, I think when you raise money, a lot of excitement, and then you sort of put that down and move ahead to keep building cool things and executing. So that, that's today.

We're, we're doing our high fives, and tomorrow we go back to work and keep building cool things.

Swyx1:16

Yeah, I like-- I love the s- the spirit of the thing. So help me contextualize. You know, like, billions of dollars is such a impossibly huge number for people to understand. How does a round like this come together?

How do you contextualize this in the, in the, in the scale of, uh, what Cerebras has done?

Context1:27

Andrew1:32

Sure. I think we were founded n-n-nine and a half years ago, and I, I think, you know, in, in early 2016, I don't, I don't think we could've i-imagined, uh, the, the AI market behaving like it is today.

And I, I think we were, uh... You know, in 2016, we, we met with Sam Altman and Ilya Sutskever at OpenAI, and they were an idea, and we were PowerPoint, right? That's amazing. And what AI was doing was identifying cats in pictures.

And, um, I think they ended up investing in us, both of them and, and, and many of their colleagues in our early rounds. But I, I think what happens is that you, uh, you begin building things. You, your, your, your vision is how can you, can you make your customer's life better?

How can you delight them with their experience or delight their customers, right? How can your infrastructure play a role? And we, we saw early on that AI would be a monstrous market. We probably guessed wrong that it-- we probably didn't think it would be as big as it is, but we, we, we thought that if you did extraordinary engineering, if you tried to do something that no one had ever done before, uh, you could do something really special.

And from there, you know, we shipped our first product in, in 2020. We delivered the next generation a couple years later. We delivered the, the third generation a little less than a year ago. Customers a-around the world have been deploying and, and benefiting from, from our performance.

Really focused on performance, both for training and for inference. You think twenty times faster than NVIDIA B200 GPUs, and it's, it's been an amazing run. And so the way a round comes together is you decide you're gonna go out, you engage at this stage with investors who do both public and private investing.

Like any other sales process, you go and you visit them, and you tell them your story. We were massively oversubscribed, and we were able to, to choose sort of the best of the best to, to, to be in our consortium.

Swyx3:37

Yeah. Uh, so I, I'm gonna flash up the, uh, the announcement on screen. Mostly we talk to researchers and engineers, more s- uh, more a technical podcast on our side. I mean, while we're at this, while we're talking about a fundraise, let's contextualize a bit.

Fidelity is a big name. I-- Atreides, I don't, I know less. What is a good investor to you? What, what, what, what makes a good investor versus a not so good investor?

Investors3:50

Andrew3:58

I think it, it depends. There, there's no single rule. It depends on the stage. I, I think in your early stages, you need people who, who understand the risks and challenges of being an early-stage company. And in our Series A, we, we had Benchmark, who is perhaps the, the best in the business, along with Foundation Capital and Eclipse, extraordinary companies.

I think in later stages, as you get close to IPO, you're looking for, uh, a very different type of investor. You're looking for an investor who primarily does public markets. You're looking for an investor who is extraordinarily thoughtful, is a leader in their field.

And I, I, I think that's what we achieved with, with Fidelity, Atreides, Tiger Global, Valor. Each of these are, are leaders. They're looked to as extremely knowledgeable in the category. A-and, you know, p-people, people take notice when, when the smartest of the smart money in-invest behind your idea and your company.

Swyx5:08

Yeah. That's, that's amazing, and congrats on all those. Let's dive in more sort of the technical side, obviously, 'cause we, we speak to, more to sort of the engineering crowd. Uh, we talked about twenty times faster. That's obviously not, like, a strict number, but you, you see, uh, this is-- I've seen this chart evolve over time.

Speed5:08

Swyx5:23

Like, you, you started out with the Llama number-

Andrew5:25

We did

Swyx5:25

... and then you sort of branched out into all the others.

Andrew5:28

That's right.

Swyx5:28

So obviously the growth of open source and the growth of these, like, really large open source models, mostly MOEs as well now, uh, has really benefited you. You talked with, um, Ar-Artificial Analysis. I'm familiar with them too. I also have them for, for the AI Engineer Conference.

Was it a surprise that you, uh, had, like, such much, like, higher performance numbers than the competition? Yeah.

Andrew5:54

No.

Swyx5:54

Tell, tell me about that.

Andrew5:56

So inference performance comes from memory bandwidth And the memory bandwidth is the limiting factor in inference performance. Remember, in order to, to generate a token, to generate a word, all the weights have to move from memory to compute.

If you're constrained right there, your inference is slower. And we have two thousand six hundred and twenty-five times more memory bandwidth than the GPU does. We use a different type of memory. We use SRAM, whereas they use a, a flavor of DRAM called HBM.

And so we were built to solve this problem. We were built... Every ounce of, of our design and our architecture was designed to provide huge amounts of memory bandwidth. Maybe the, the, the way to think about memory bandwidth, imagine you, you've got your cup of tea there, but imagine a, a cup that holds Coke.

The cup is your memory capacity, and the Coke you pour into, or the Coca-Cola you pour into the, the cup is your data. Now, the rate at which you can get Coke into your mouth is a function of the straw.

Right? It's not a function of the size of the cup. It's a function of the size of the straw. And G- GPUs have a very narrow straw between compute and memory. We chose a fundamental different architecture. We got rid of the straw altogether.

We built a chip the size of a dinner plate, so we could put all the memory onto the chip. We could use SRAM. And this said, "All right, we've got our Coke in our cup. Get rid of the straw and just put it right on our mouth."

And, and that's the way you get huge amounts of data. That's exactly right. That's exactly right. Moving. Were you just pretending to drink tea? Did you pour Diet Coke into a teacup? Is that what you're doing?

Swyx7:49

No, I didn't. Uh, the, the... I happen to have a Coke can next to me, but I, I'm actually, I'm actually sick, so-

Andrew7:54

Okay

Swyx7:54

... this is why I'm drinking tea to, to stay, uh, to have a voice.

Andrew7:56

To stay hydrated.

Swyx7:57

But-

Andrew7:57

Good

Swyx7:57

... to stay hydrated.

Andrew7:59

So we weren't surprised every ounce of our design, of our architecture was, was built to solve this particular problem and, and enable the type of speed that, that your guys at, at Cognition have, have benefited from, the guys at Mistral in the closed source world benefit from.

We serve Le Chat from Mistral. Meta has benefited from, IBM benefits from, Alpha Wi- Alpha Sense, and hundreds of other customers around the world benefited from this extraordinary speed.

Swyx8:30

Yeah. Fun fact, I used to work at Sentieo, which was an Alpha Sense competitor, and they merged into Alpha Sense. So I'm also a shareholder of Alpha Sense.

Andrew8:36

Well, you, you've been at, you've been at our customers your whole career. This is awesome. Th- th- this is awesome.

Large Chips8:41

Swyx8:41

So, uh, I mean, I, I think, like, you obviously made the right bet on, on the large chip and, uh, you know, there's a, there's a number of, uh, sort of people in the large chip sort of industry.

You know, you're not, you're not the only ones. Like, what, what has made you stand out so far? You know, I think, like, what has been hard for a non-hardware person like myself to understand is, you know, like, a, a lot of people who are, who are sort of competing with NVIDIA and bet- betting on, like, more on, uh, on, uh, on-chip memory and all that, like, it's hard to distinguish between what the strategies are, and I think over time as well, the, your different generations, you are developing more and more of an un- understanding of what the market really wants.

So I just wanna get a sense of, like, how you see this market, the, the large chip market.

Andrew9:21

Sure. I, I think the, the, the w- the way to think about this is that there are really two types of memory. There is DRAM and, and HBM, whi- which they can hold a lot-

Swyx9:30

Yeah

Andrew9:30

... but they're slow. And th- that's what NVIDIA chose, and they chose that primarily because originally that was the right choice for graphics. Lots of capacity, but slow. On the other hand, there's SRAM. SRAM is the exact opposite.

It can't hold very much, but it's blazing fast. And so what we, what we saw was that if we built the largest chip in the seventy-five-year history of the compute industry, if we built a chip that was fifty-six times larger than the largest GPU, we could stuff it to the gills with this fast memory.

All right? We could fill it with the fastest memory and thereby overcome its limitation of not having a lot of capacity and just benefit from its blazing speed. Now, if you build SRAM-based solutions based on little tiny chips, you, you need to use hundreds or thousands to do work.

So if you build a normal-sized chip, say eight hundred square millimeters, right, the size of a postage stamp, and you wanna support a trillion-parameter model or six hundred billion-parameter model, you might need four thousand chips, all right? Right, that, that's a lot of complexity, that's a lot of wires, that's a mess.

And so-

Swyx10:46

Yeah, you need interconnects, uh, as well

Andrew10:47

You need interconnects. It limits the things you can do. It makes all sorts of cool AI techniques like speculative decode extremely difficult. Whereas if you have a giant chip, you might only need a handful. A- and so I, I think that, that's one of the, the, the ways to, to differentiate.

I, I think also just go up and look at artificial analysis. Wherever we are, we're the fastest, uh, not by a little bit, but by a lot. And the reason we like them as a, as a benchmarker is they go to your cloud anonymously every day, do a little work, publish the results, right?

There's no gaming, there's no months of preparation for the test that advantages big companies like MLPerf. Every day they go, and this is exactly what your customers get. Right? There, there's no gap between, uh, the benchmark and what your customers actually see.

Swyx11:41

Yeah. It's all... It's like, you know, MLPerf is kinda like the Olympics. Like, you really prepare for it.

Andrew11:46

That's exactly right.

Swyx11:47

And it's, like, a once in a lifetime, you know? And then, and, and, like, the, the, uh, artificial analysis is more like Apple Maps, like-

Andrew11:54

That's right

Swyx11:54

... estimating traffic for you, you know?

Andrew11:56

Right. You know, like, it, it doesn't really matter how, how fast you are at the drag strip if every day you go to the grocery store, right? You can... So that, that's the way we think.

Swyx12:07

So, um, and then, uh, fol- look, look follow up on, on sort of the inference, uh, versus training split. I, I'm getting the sense that, like, most of these are just trained on, on non-Cerebras silicon and then inference on silicon...

Workloads12:07

Swyx12:20

on, on Cerebras silicon. Do you see that persisting going forward? Are you specializing in inference? Do you want to d- would it be a drea- uh, a goal for you to have more training workloads, or do you not care?

Andrew12:31

We support both training and inference, and we deliver solutions on premise and in the cloud. So we are perfectly happy to, for our customers to use us to train or do inference, and we do a, a great deal of work with large enterprises in training, with Mayo Clinic, with GlaxoSmithKline, with the U.S.

military, with the, the Department of Energy. With our, our customers in the Middle East, we've trained leading models, language models in, in Arabic, in Catalan, in Hindi. What, what's happening in addition to the, the suck- giant sucking sound for fast training is that, that these models are being used, right?

I mean, we, we, we make AI with training, and we use it with inference. And the, the amount of AI being used right now is extraordinary. It's just off the charts. And I, I think there is certainly a, a, a, an, an unquenchable demand right now for fast inference.

Swyx13:25

Yeah, I, I think that the numbers I had was something i- internally for Google with the TPU workload is something like a two to three ratio of training to inference. And I wonder, I, I feel like for Cerebras it might be like, I don't know, one to five in favor of inference and-

Andrew13:39

I think that's not an unreasonable guess.

Trends13:41

Swyx13:41

Yeah. Okay, cool. Amazing. Uh, and then what do researchers tell you they, they want? Like, uh, you know, I, I... You s- I'm looking at these, these charts of Qwen 3, Llama, Llama 4, o- OSS GPT, and obviously there's more.

The GLM just came out today and I, I know there's a lot of work, uh, that's going on between our companies on, on, on that, on the open model side. What trends are you seeing and what do you need to support on the silicon side for, for those trends to happen?

Andrew14:04

I think the, while we design silicon, we also build the system and the software layer. And many of our customers in- including Cognition benefit from us via our cloud offering. And so it's a sort of a, a soup to nuts solution.

The trends are pretty clear. I think for sort of AI companies like, like Cognition, like all your competitors, like AlphaSense, like dozens of others, they are trying to replace closed source models with very, very fast open source models, and they're trying to drive the open source accuracy dramatically up, and they're doing that with fine-tuning.

They're doing that with all sorts of, of interesting ML techniques. I think the closed source community has come to realize that, that being slow is not ideal and that there's a, a very sort of interesting tweet by Paul Graham, and he said something like, "I'd use Google half as much if ChatGPT weren't so slow."

Right? And that's very interesting that, that, that if you make your customers wait, they leave you, right? They, they go and use something else. And so I-

Swyx15:09

Yeah, and they-

Andrew15:09

Go ahead

Swyx15:09

... they're not, they don't necessarily tell you about it too. They, it's just that-

Andrew15:12

No, that's right

Swyx15:12

... they're voting with their feet. Yeah. Yeah. They, they're voting with their feet and like it's, it's a slow, like maybe like the 20% of my, my work goes over. It's not the full thing. I'm not rage quitting.

I'm just choosing more one over the other.

Andrew15:25

That's exactly right.

Swyx15:26

Yeah.

Andrew15:26

That's exactly right. We see at the large enterprise level, particularly those who have large data assets, a desire to, uh, train their own models and to go a little smaller, say in the 10 to 30 billion parameter category.

I, I think, uh, especially the very large companies, enterprises have, uh, concerns about some of the, the, the data that was used in training some of the open source models. And so they, they wanna take the architecture of the open source model, randomize the weights and begin again, and do that with data sets that their own legal teams have approved.

So we, we see that in the large enterprise category. I think areas that are, are of interest are, uh, sort of securing agentic work, right? I think the, uh, when you, when you spin off dozens of agents, the sort of attack surface of the, uh, of the solution expands geometrically.

And so there's a lot of work being done to think about how one might secure, uh, an agentic flow like that. I think still the, the bulk of the models are in the transformer category, uh, for language. We continue to see innovation in, in diffusion for, for images.

I think many of the, the, the new models are, uh, tweaks to some of those core foundation architectures.

Swyx16:44

Yeah. Uh, uh, it's, it's weird how like, you know, the transformer paradigm has stood strong for quite a long time. I do think like there are some discussions, you know, I just recorded a podcast with, um, uh, with AI21 yesterday where they, you know, they've been really exploring hybrid models.

But I feel like at least, at least the paradigm of like deeper models or like, you know, like more MoEs, I think that, that's not really gonna go away. They, they might have some innovation in some part of the attention for-

Andrew17:12

That's right

Swyx17:12

... a variant that you might be using, but like nothing, uh, nothing that you have to particularly take care of at the hardware level.

Andrew17:18

Yeah. I, I think, you know, we designed this architecture before, uh, transformers were invented. And one of the, one of the hardest things in, in computer architecture and one of the first things you do is you decide what you're not gonna be good at.

I'm gonna build a chip, what am I not gonna be good at? We, we're not gonna be good at general purpose compute. We're not gonna be good at graphics. We are going to be good at AI work. Second really hard question in computer architecture is what should I be general at and what should I be specialized at?

And in 2016 and '17, as we were setting this architecture, we could have said what we wanna do is embed three by three convolutions into the circuits. And that would have made, right, three by three-

Swyx18:01

You're very good at that. Yeah.

Andrew18:03

That's right, very good. But you would've been, you would've been carrying a tax if you did a one by one convolution. You'd have been carrying a tax if you did anything other than a three by three. We said- That was not where we wanted to be specialized.

What we wanted to do was work underneath that. We wanted to work at the linear algebra level. We wanted to accelerate sparse linear algebra. And what that did is that enabled us to support models that we had never seen when they came out, right?

We'd never seen transformers, and yet we're the fastest at transformers by twenty X. We'd never seen diffusion models, yet we're the fastest at diffusion models. Getting that right is one of the very, very important decisions in computer architecture, and one I can take no credit for whatsoever.

Sean and Michael, JP and Gary, who are my co-founders, they got that right really early on, along with the, the, the, the ASIC team.

Cloud18:52

Swyx18:52

Yeah. Awesome. So a-and then finally, you know, I think you are, you, you were mostly selling chips directly to, to, to your customers until about a year ago when you basically started sort of your self-serve cloud offering. Uh, now, now there's a dedicated effort on Cerebras Code, which is doing very, very well.

Andrew19:09

It is.

Swyx19:09

How do you feel about that? Like you're kind of competing with your own customers a little bit.

Andrew19:13

You know, I, I think the following. The way we look at the market is as follows. Un-until about a year and a half ago, AI was sort of a novelty, right? You, you, you would... It was curious, and it was interesting, but it wasn't useful every day.

And about a year, year and a half ago, we began to see changes where it became useful in a very different way, and we, we dedicated resources to the inference side. We thought we could improve and stand up a cloud to deliver inference.

And in sort of early September of, of, of last year, we, we stood up our, our cloud, and we had an extraordinary reception. It was just overwhelming. And sometimes too much good i-is a challenge too. You know, we had downtime, and we had to fix that, and there was just so much demand, um, we had to get better.

And so step by step, we, we improved and, and our quality jumped up and, uh, we, we did some of the right things, we did some of the wrong things and then fixed them, um, and step by step.

And, and what we found is that there's just this aching demand right now for, uh, inference that doesn't make your customer wait. Now, w-when we did that, we saw an opportunity to continue to learn from our customers, so we, we stood up a Cerebras Coder.

We have no interest in being an IDE. That is not gonna play to our strength. That, that's your strength. That's the strength of you guys and your competitors. That is not who we are. On the other hand, one of the things we, we were really looking to do is get enough adoption so we could watch utilization, we could hear directly from customers what features we, we were lacking, what we needed to do better and deliver to those who wanted sort of raw performance, maybe without the IDE, right, an offering.

I, I don't think we have any interest in competing with, with you guys or, or your competitors. We-- We're, we're not gonna be very good at that. What we're good at is the processing of tokens unbelievably quickly. And the, the user experience and, and the engagement engine around, that's for-

Swyx21:20

That's okay

Andrew21:20

... for other app makers.

Swyx21:22

Yeah.

Andrew21:22

Whether it's you guys or AlphaSense or, or Mistral or, or, or any of the others, that's what is really important that, that, that you guys do.

Swyx21:31

Yeah. Incredible. What do you th- wish people asked you more about? You know, like I th- I'm sure you do a ton of these media interviews. I probably asked, you know, some of the same questions, but I just have to, you know?

Deep Dive21:39

Swyx21:39

Uh-

Andrew21:40

No, I-

Swyx21:40

What do you, what do you wish people understood better?

Andrew21:43

I think I wish they asked more about the team. I, I think they're always asking about me, and I, I think-

Swyx21:48

Well, you're the star, you know.

Andrew21:50

No. I, I think, think I'm the thin edge of the wedge. I, I don't think that's, uh, the star. I, I think first there were, there were five founders, Sean Lee, Gary Lauterbach, Michael James, JP Fricker. They were the technical visionaries.

I have market insight. I have product management insight. I have... They were the technical visionaries, and, and what they put together is extraordinary. I think people don't ask me, "Well, where'd you get this collection of extraordinary engineers?" And, you know, dozens of them have worked for us and with us before.

And so we, we began with one of the world's leading chip teams. We built around them extraordinary talent. We added hundreds and hundreds on the, on the compiler side and the software stack. Pe-people don't ask about that, and that's a pity because there's just an extraordinary amount of hard work that, that goes into building a product.

I, I think tho-those are some examples that people don't ask about. I think people don't ask about the complexity of standing up a, a data center that pulls as much power as a small city, right? I think that-that's really interesting things are going on in the data center business.

We just opened a new facility in Oklahoma City. It's our fifth in 12 months, um, and it's amazing the amount of power these data centers are pulling. And, you know, it's very easy to hear people talk about tens of megawatts or gigawatts.

You know, if you look at a city like Pittsburgh, it's a top ten city in the US, it draws three gigawatts. So you're, you're pulling sort of a third of Pittsburgh to two buildings. That's amazing, right? And the complexity-

Swyx23:22

And even like heating, heating gets, uh, to be a question, right?

Andrew23:25

Heating, cooling, and the complexity there is extraordinary. And so I, I think people often sort of miss those things. And finally, I think the easiest thing to miss is sort of th-these collection of things that, that are on top of AI that are ease of use, that are unbelievably annoying if you don't get them right.

Right? I, I don't think prompt caching is an AI feature. I, I think managing your router, where you distribute tokens to your partner, and the technology involved in integrating your routing with your token processing partner, these are all things that are extremely interesting and have, can have a huge impact on performance and that almost nobody talks about.

Swyx24:01

Sorry, uh, you have a router that, that is part of your offering? Shouldn't that be-

Andrew24:05

No

Swyx24:05

... our, our job?

Andrew24:06

Yeah. No, no, no.

Swyx24:07

Okay.

Andrew24:07

That is your job.

Swyx24:08

Okay. Okay.

Andrew24:09

In, uh, a close it... No, no, no. It's not talked about enough. I think you're exactly rou- right. I, I think the distribution of tokens and the techniques used so you don't crush your, your partner, the various techniques that can be used upstream of us is an area that isn't talked about enough at all.

And there's a great deal of, of really interesting insight there and where a lot of data engineers work. And, you know, it's, it's sort of in the background rather than it should be in the front ground. We were doing work with, we were doing work with one of the hyperscalers and for a product they have, and they said it, "Look, we used your, your system, and it didn't make us that much faster."

And so we said, "Well, that's a surprise because we're twenty, twenty-five times faster on this workload." And what, what ended up happening i- is that other steps in their data processing slowed the whole system down.

Swyx24:59

Yeah.

Andrew25:00

A- a- and-

Swyx25:01

So they were not compute constrained. Yeah. It's always-

Andrew25:05

That's right

Swyx25:05

... it's always IO bound, compute bound, IO bound, compute bound, you know?

Andrew25:06

That's right. There were other challenges that they had and that they had to work through, and by everybody assumed it was, was the token processing, and it, it wasn't. And so I think-

Swyx25:17

Well, I mean, honestly, that's a bad customer because they just didn't know what they were testing for. They, they, they blamed you.

Andrew25:22

I think that happens a fair bit in this world. Blame the new guy. I mean, that is, uh, that is not uncommon.

Swyx25:31

Yeah, totally. Um, okay, so I- I'm already past time, but I'll, I'll, uh, I'll ask one, one more question if, if you-

Andrew25:37

Of course

Swyx25:37

... are in the mood to, to humor it. We're in an age of, like, very big infra build-outs. You know, I, I was at Meta Connect. I, I met with, uh, with Zuck and Alexander Wang, and, like, we talked a little bit about, like, the, the infra build-out that, that, that Meta's doing.

Vision25:37

Swyx25:49

OpenAI has Stargate, which is, like, basically terraforming Texas to be a data center.

Andrew25:55

West Texas. They've decided to annex West, West Texas, yeah.

Swyx25:58

What is the Cerebras grandmaster plan for, like, you know, like, the, the next ten years? Like, what is the most ambitious version of where you're going?

Andrew26:06

Look, I, uh, I, I think that, you know, ten years ago, uh, NVIDIA was a twenty billion dollar company.

Swyx26:14

Oh my God.

Andrew26:15

I mean, in two thousand and fourteen-

Swyx26:17

Twenty?

Andrew26:17

Yeah, two thousand and fourteen they were ten billion. In two thousand and fifteen they were twenty billion. It's really hard to, to have a ten-year vision, uh, in this. Uh, and it's a very interesting question of when a market's moving very, very quickly, what is the right l- length of time to have a vision over, right?

Because if the market landscape is changing extremely quickly, the longer your vision, the more likely it is to be wrong.

Swyx26:41

You have, you have to have a plan to plan towards, you know? Like-

Andrew26:43

That's right. So, um-

Swyx26:45

So-

Andrew26:45

You... That's right

Swyx26:45

... it's, it's more the act of planning.

Andrew26:46

No, no, what you said is extremely smart. Rule-changing rules.

Swyx26:49

Yes.

Andrew26:49

Plan creating milestones. I think you're exactly right. Look, we, we see the following. Number one, we see AI inference continuing to grow exponentially, and we know why. The size of the AI inference compute market is the number of people using AI times the frequency with which they use AI times the amount of compute they use each time they use AI.

And what we know is that every week more people are using AI, they're using it more often, and they're trying to do more interesting things with it that take more compute. We see no end to this. Absolutely no end.

And so over the next three to five years, we see extraordinary continued growth. That's number one. Number two is that for AI to deliver on its promise to be embedded in our lives, it must be fast. There aren't things that are embedded in your life that make you wait ten or fifteen minutes to get a good answer.

Th- those are proof of concepts. Those aren't products. And whether it be on the coding side, whether it be on the agentic side, whether it be on drug design, whether it's in healthcare, right? You want the results of that MRI quickly, right?

You don't want to wait weeks, right? You want, you want the results of your pathology test. You wanna know whether you have bad news, whether it's cancer. You don't want to wait weeks to get the results back, right?

If you could get it back while you're at the office, that is a much better medical experience. All right. And it's true for the entire range of things that AI can, can do for us. And so we, we believe there will be continued phenomenal emphasis on speed.

And finally, I think the transformer, the current architecture, has a way still to run. I, I think we, we will see that for several years more.

Swyx28:40

Yeah, totally. Okay. Well, tho- those are great macro trends to hang our hat on. Uh, definitely co-sign that. Well, congrats on your, your huge fundraise. Uh, this was a great first meeting, uh, for, for Latent Space. Uh, I hope we get to do this in person sometime, some other time.

Outro28:40

Swyx28:52

But, uh, thanks for taking some time out of your, your big day.

Andrew28:55

Well, thank you and, and the Latent Space team for, for inviting Cerebras o- onto your show. We really appreciate it.