LALatent SpaceAug 18, 2025· 47:56

⚡️Accelerators @ 3x NVIDIA H200 perf, Made in the USA - Thomas Sohmers + Mitesh Agrawal, Positron AI

Thomas Sohmers and Mitesh Agrawal of Positron AI argue memory bandwidth, not compute, is the true bottleneck in AI inference, and their accelerator achieves 93% memory bandwidth utilization—triple NVIDIA's efficiency—enabling 70% faster token generation at 150W. The founders, both Lambda Labs veterans, shipped an FPGA product 15 months after founding, then raised a $51M Series A for an ASIC in late 2026. Their hardware requires no recompilation: it ingests raw binary weights from NVIDIA training and outputs an OpenAI-compatible API. Positron focuses on the decode phase of transformers, where memory-bound matrix-vector multiplication dominates, and already counts Cloudflare and Parasail as customers. The company sells systems directly, prioritizing capital efficiency and ROIC over operating its own cloud.

  1. 0:00Intro
  2. 4:34Memory Bound
  3. 7:52Roofline Plots
  4. 14:13Utilization Edge
  5. 16:51Fast to Market
  6. 22:51Power & FPGA
  7. 29:02FPGA to ASIC
  8. 40:55Decode Focus
  9. 45:34Hiring

Powered by PodHood

Transcript

Intro0:00

Alessio0:03

Hey everyone. Welcome back to another Latent Space Lightning Pod. This is Alessio, partner and CTO at Decibel, and I'm joined by Sweeka, founder of Smol AI.

Swyx0:11

Hello, hello. Uh, today we are joined by Mitesh and Thomas from Positron AI. Welcome.

Mitesh Agrawal0:17

Thanks, Sean.

Thomas0:18

Thanks for having us.

Mitesh Agrawal0:19

It's good to be here.

Swyx0:20

And also special shout out to Maxime from DFJ for putting us in touch. So, uh, I, I don't know who-- which of you want, you want to go first, but both of you have a background in some of the, the top like, you know, chip companies in the world.

I, I think, I think mostly Lambda Labs, but also a little bit of Groq. How do you find your way to Positron? What's, what's sort of like the founding thesis and the founding story?

Thomas0:41

Yeah. Um, I guess, uh, you know, just starting my journey a little bit. You know, I got my start in the semiconductor industry back in 2013. I started my first company called Rex Computing, uh, that was focused on the DSP for like, uh, mobile base station workloads.

And I, I started that after receiving the t- the Teal Fellowship. So I was, uh, you know, 17 with the crazy idea of starting a semiconductor startup. Uh, and, uh, thankfully was able to, you know, uh, convince, uh, some folks to join me.

And, uh, we raised, uh, you know, two million dollars and, uh, ended up taping out, um, you know, first chip when I was 19. Um, and, uh, uh, had, you know, leading for, for that time, um, the silicon for, um, for that, uh, signal processing workload.

Sadly didn't really get commercial success with it, but, uh, it was a amazing experience to be able to actually produce silicon from scratch. You know, after that, I worked on a couple cryptocurrency ASICs before joining a friend's startup, uh, called Lambda that, uh, I was principal hardware architect at, and, uh, built out the first couple data centers for Lambda Cloud.

And that's where I had the opportunity to work with Mitesh. I actually was the one that introduced him to Lambda originally that, that had him join before I was an employee there. And Mitesh and I have been housemates and just very good friends for over a decade.

After my stint at, at Lambda, I wanted to get back into actual semiconductors, and so I, uh, joined Groq, the, the company with a Q, and I was there as a director of, uh, technology strategy for, for two years.

And, uh, then in early 2023, decided to leave and, and start Positron. And I originally tried to recruit Mitesh to be CEO of Positron back then. You know, he was growing Lambda crazily at that time. Uh, so, uh, I was actually, you know, uh, CEO of Positron for the first, uh, 18 months or so, uh, and finally was able to, to bring Mitesh over and, uh, he joined at the beginning of this year as CEO.

I'll CTO.

Mitesh Agrawal2:37

Yeah. And kind of for me, a lot simpler. I think, uh, Thomas mentioned he introduced me into Lambda Labs. Uh, you know, Steve and Michael really loved working with the Balaban brothers, right? Really grew the company from zero dollars in revenue all the way to where it is now, well over half a billion in ARR, right?

And, you know, worked there for almost a decade. Did all types of roles. Uh, you know, just a very early employee. I was employee number five there, so did engineering, sales, support, uh, you name it, and, uh, was CEO and head of cloud.

The main thing for me was just, uh, really focused on building out the cloud at Lambda, you know, all the way, and then was primarily focused like in one word, like on growth, so revenue, operations, data center, you know, megawatts, uh, vendor relationship and vendor procurements, uh, engineering and engineering headcount.

And primarily for me, I think... I mean, L-Lambda is still growing massively, right? So it's-- I, I don't want to imply that I left because, you know, anything. I still have great connections there and, and really, really am proud and, and very, you know, have huge admiration for what they're continuing to build.

For me, I think, you know, the two reasons really why I left was, one, is really could see a little bit around the corner at Lambda. You know, we worked with frontier model labs, enterprise companies, you know, AI startups.

Was seeing how the model architectures were growing, uh, the requirements around memory architecture, memory requirements capacity and, and bandwidth, and really wanted to get to contribute in an underlying technology stack. You know, Lambda is, is a great company built on cloud services layer.

Really wanted to work on the underlying compute that will power kind of the generations of the future. And, and then second thing was, you know, my background is chemical engineering, so study fabrication, really wanted to get into, back into the silicon space and, uh, build that out.

So joining, you know, this year itself, January 2025, I was, uh, uh, I was talking with Thomas in December. I had my first kid in, in November, and then Thomas was like, "Well, you already had one major change.

Why not make it two major changes in your life?" And, uh, decided to join in. So that's, that's a quick summary of the background.

Swyx4:23

Yeah, pretty amazing. Also very unusual to have a, you know, it's like kind of like a handover of CEO that you've been trying to recruit for a while. That's, that's a fun story.

Mitesh Agrawal4:30

Yeah.

Swyx4:30

But we could, we could, we could go technical. Uh, Alessio, I, I think I snapped you out of-

Alessio4:34

Yeah. I, I was curious if you guys thought about sticking to the software side at all, or if you were busy looking at all the software improvements and you thought that was kind of like a local maxima that was kind of capped, and you thought the hardware had to be rebuilt, or was it just a matter of-

Memory Bound4:34

Thomas4:51

Yeah

Alessio4:51

... you have experience in the space. We think we have a unique angle. We should just go do this.

Thomas4:55

Yeah. I would say for me, I, I've always been a hardware guy. Um, you know, I started tinkering with electronics when I was like five, six years old. While I programmed a bit as, you know, as a kid in, in, in school, I was always so much more interested in, in physical things that, uh, could actually burn or blow up in my face.

Um, so I think the natural sort of gravity, but also just looking at the actual application space itself You know, being a, a huge believer in, you know, the, the capabilities and, and future growth of, of this industry, the biggest limiting factor for that being, you know, distributed at scale.

You know, the, the, the quote of, you know, "The future's here. It's just not evenly distributed yet." I think that the limiting factor for that is, you know, the, the cost and from both, you know, dollars and, and power perspective for deploying the, these applications.

And so starting the company was really focused, how can we get the best performance per dollar and performance per watt for these workloads? And with my, my experience on, you know, going all the way back to my, my DSP work to, uh, uh, you know, the time at Lambda, et cetera, seeing that the kind of way that the whole industry had been going for, you know, the decade before ChatGPT was not really going and, uh, you know, optimizing the, for the right things to what transformer inference really needed because that the, the actual real underlying hardware, um, or the, the underlying algorithms have very different hardware requirements than, than what was needed for CNNs and, and, you know, the first wave of, you know, our modern machine learning, you know, explosion.

So it was seeing that everyone else was focusing on the wrong things. They were just trying to have more and more flops when memory bandwidth, memory capacity were the real, uh, real bottlenecks. And then that, that realization that, hey, everyone that has been focusing on that is, like, completely missing these sort of neglected parts of computer architecture and, and processor design that, uh, thankfully I, I got exposed to through, through my career.

Mitesh Agrawal6:51

Yeah. And for me, I-I'll keep it kinda one-liner, you know, the, the bitter lesson, right? You know, compute really has been the driving force for a lot of just amount of compute and, and the way you can, uh, make it efficient has been the driving force for a lot of model improvements.

This is not to say, I mean, I personally don't believe the software at, at local maxima. I think it's more around our backgrounds, you know, what we are good at. That's, that's obviously the driving factor there. But and many people that are doing amazing work on model improvement sides, both on small and, and, and frontier-edge model, we're seeing it today.

But really, like the, what the bitter lesson has taught us is like the more amount of compute you can provide, the better it is for all the people that are working on software to drive in more utilization, to drive more intelligence from it.

So yeah, that, that was for me, it's still the case. It's like, how can you just drive more compute in the world, both from flops, memory, and deployment perspective.

Alessio7:38

Cool. Um, yeah, no, that's great. And I know you touched on some of this, but maybe this is a good time to talk about what you guys are doing that is unique and maybe contrapost that both to, you know, NVIDIA and the kind of traditional ones as well as Groq.

Roofline Plots7:52

Alessio7:52

Obviously, you worked there. Cerebras, and then there's kinda like the long tail of all the other GPU alternatives things. But just give people a lay of the land of like what's working, what isn't, and what you guys are doing differently.

Thomas8:03

Yeah. Um, I guess I'll actually share my screen just to cover a couple things at a high level that, um, can set the stage a bit. Uh-

Alessio8:17

Surprising how many clicks it takes to share your screen.

Thomas8:19

Yeah. Uh, it's amazing. We're on the verge of AGI and, you know- ... we still can't get presentations or calls to work properly. Um, yeah. So, uh, just to, to show here, uh, you know, some, some of the viewers here may be familiar with, uh, roofline plots.

Um, but this is, you know, pretty abstracted way to be able to look at, um, you know, one of the key differences between, um, you know, uh, really convolutional neural networks and transformers. More specifically trying to focus in on, on them at, at inference time.

Alessio8:50

Yeah.

Thomas8:51

So this showing, you know, uh, operational or arithmetic intensity, um, um, as a fraction of, you know, the, the, uh, over the performance. The, the real key thing here that this is showing is that, um, in the top right corner here is example of a, one of the most compute bound problems you could have of doing your inner product matrix-matrix multiplication for CNN, where depending on the exact parameters can be anywhere from five hundred to a thousand flops for every byte that actually needs to be moved.

And on the other side of this chart, you have the case of a transformer where when you're actually, you know, doing attention or really just any case where you're doing, you're fundamentally doing matrix vector multiplication rather than matrix-matrix multiplication.

And in that case, you are, are now actually completely memory bound. For every step in, in the sequence that you're doing, you're now having a one-to-one ratio of flops to, to bytes that have to be moved. And so in that case, you know, the, the really big realization at the beginning of Positron was, hey, this looks a lot like when you're doing QAM-256 in a, um, you know, mobile base station for, for LTE, uh, just in terms of what actually needs to be moved.

And the systems built for doing that, you know, namely utilizing FPGAs and specialized DSP chips for that have extreme focus on the memory system and the ability to actually saturate whatever DSP compute elements you have, uh, and be able to f-feed through at that one-to-one memory-to-compute ratio.

And if you look at sort of traditional big iron compute systems, the last like major compute, compute platform that had that balance of memory-to-compute ratio was the Cray 2 supercomputer. So, you know, we're, we're pushing, you know, forty plus years of computing development, uh, you know, that has been focused on just providing more and more flops, just, just, you know, raw math compute capability while not balancing that by actually having memory scale in the same way.

And-

Mitesh Agrawal11:11

Well, capacity and memory bandwidth.

Thomas11:13

Yeah. Uh, yeah, I'm really mainly talking on the memory bandwidth side here. Uh, but yeah, that this, this was like the key thing that we saw, like, okay, everyone had been optimizing for the CNN portion of, of the workload.

And the other thing I should point out here is when you're doing training for, for transformers, you're also in this very compute-bound case. In the case when you're doing training, that what is matrix vector in inference because you're, you're generating one token at a time and you're just appending that to your KV cache.

In training, you already have the full sequence of, you know, whatever you're, you're feeding in from a dataset perspective. So that gets to be a matrix, and it is compute-bound. But it is this particular case of transformer inference that is just massively memory bound.

Mitesh Agrawal12:04

Can I, can I also just spell out that this also applies to Groq and Cerebras and other sort of big chip companies, right?

Thomas12:11

Yeah. I mean, this is just fundamental to the algorithms.

Mitesh Agrawal12:14

Right. I, I mean, just 'cause most people are familiar with NVIDIA, uh, H200, so I just wanted to-

Thomas12:19

Yeah, yeah

Mitesh Agrawal12:19

... uh, clarify.

Thomas12:21

Yeah. I mean, uh, this really is kind of hardware agnostic from, from perspective of the operational intensity-

Mitesh Agrawal12:27

Okay

Thomas12:27

... of the algorithm. Um-

Mitesh Agrawal12:28

Okay

Thomas12:29

... uh, you know, the, the way in which different people approach that, and I would say Groq, Cerebras and others have tried to take-- have tried to say that, "Hey, we're providing more memory bandwidth," but typically that comes at a massive, uh, trade-off of them not having the memory capacity.

I'll, I'll get to the kind of the memory technology in a second, but the, you know, very quick thing I'll just throw up here is, you know, just trying to visualize the difference between matrix-matrix and matrix-vector. And so when it comes to like the actual memory accesses that happen, you know, when you're doing a matrix-matrix multiply, the, the simple way to say this, and it doesn't take into full effect the tiling and, and more complexities there.

But, you know, the, the simple high school math version of this would be, you know, for every element in matrix B, you get to reuse the same element in matrix A when you're then making matrix C. While when you're doing matrix-vector, you actually need to change the elements in both matrix A and vector B for every single operation that you do.

So the, the second level of this is that matrix-vector multiplication is basically uncacheable. When you're doing transformer inference, matrix A is the weights of your, uh, model. And so if you're talking about model weights that are tens of gigabytes, hundreds of gigabytes now going into ter-terabytes of memory, you're not going to be fitting that in any on-chip cache.

And so even with-- if you had perfect cache prediction, it just won't work. And so a lot of the, the memory architecture decisions of the past few decades just go completely against what you need for, for this workload.

Utilization Edge14:13

Thomas14:13

So fundamentally kind of getting to, to Positron here, you know, the, what we're sh-show-showing here with, you know, A100 and H100 is that sort of in NVIDIA's absolute best case scenario for an H100, you know, for Llama seventy B here, uh, you know, they're able to achieve about twenty-nine percent of the theoretical memory bandwidth in a device.

So you pay for three point three five terabytes per second memory bandwidth for an H100, and in practice, you're only actually getting about a terabyte per second. And the A100 comparison here is interesting because in most of these cases, they're actually the percentage of theoretical memory bandwidth actually has gotten worse generation over generation.

And we'll see exactly where Blackwell ends up. But all indications are even though they've, you know, more than doubled the theoretical memory bandwidth going from Hopper to Blackwell, the actual percentage of theoretical that you can achieve is again, going to be, uh, less than, than the previous generation.

And really what Positron's whole mission is about is getting absolute maximum memory bandwidth utilization possible. And so our fundamental architecture is enabling us, you know, today with hardware that we're shipping right now to be achieving, you know, ninety-three percent of the theoretical memory bandwidth of our device consistently across all use cases.

So you can see that this degrades in the NVIDIA case as you scale up the number of concurrent users, but you would see the same degradation as you scale the context length or scale the model size. You know, our architecture, the key things is that we have, you know, a, um, our, our core compute elements are designed to be optimized for one-to-one flops to memory bandwidth, so we can actually feed and completely saturate our device.

And, you know, even when you then scale out and, you know, to external memory, you know, we're able to, to fully utilize that. And basically that missing seven percent that we've got is effectively just due to the, the refresh on the DRAM, which, uh, can't really get around right now.

So, you know, what that actually results in is like today we're, uh, you know, able to achieve about, uh, you know, seventy percent higher performance than NVIDIA with the, the cards that we're shipping today, significantly lower power and price point.

Yeah, that's, uh, and just to re-emphasize, you know, this we're actually shipping.

Fast to Market16:51

Mitesh Agrawal16:51

Yeah.

Thomas16:51

So-

Mitesh Agrawal16:51

I think, I think the question I was asked around, like the difference between not the NVIDIA, but the alternatives, like that's a big thing. It's like we got to market shipping within fifteen months of founding of the company with only the seed round raised and, and again, are now planning our next generation of silicon within eighteen months of the first generation.

So like we are not only matching NVIDIA on the speed of performance and performance per dollar and performance per watt, where we are really trying to outcompete there, but also coming up with the generations of the product, uh, 'cause they're, they're setting the pace.

They're the trendsetter, right? So, so we're keeping up with that, which is kind of classic going away from the case of a lot of silicon and ASIC companies of the past decade, where they've taken three to five years to really get out a product in the market-

Thomas17:29

Yeah

Mitesh Agrawal17:29

... and then taken another three to five years to get the second product in the market.

Swyx17:34

But I want to talk about the technical achievement, but actually this is an operational achievement, right? The fact that you've been able to, to ship. What's your secret?

Mitesh Agrawal17:42

So I think one is small team, I would say. I think it's, it's, it's, it's, it's, it's-

Swyx17:45

Tiny team

Mitesh Agrawal17:46

... like, I think the part about I wanted to say with small team is like, with small team, you're forced to focus on applications that you wanna ship out. And for us, from day zero, it was about, you know, accelerating the linear algebra of matrix vector map.

And, and second of all, look, from a marketing term, we call it like falling within the NVIDIA ecosystem. But like from both of our past experiences, I can tell you that most of the, the silicon providers, you know, whatever technology they might come up with or, or whether it's good or bad, I think the biggest hurdle has been actually really convincing people to use it because there is always change of workflow of how they integrate into their existing inference work systems or engineering ecosystems, whether that's, you know, figuring out some compiler, compiling the same model again, or some kind of backend system which is complicated, doesn't scale out.

And for us, like we were from, from early days, we really focused on getting a product out in customer's hand so that we can actually show proof of concept and show really our architectural, like memory architecture and the silicon architecture working and, and showing the performance improvements.

And then second was on the same point, making it easy for people to really work those systems without really having to change any of their engineering workflows. So, you know, basically what that means in actual world is when we say part of the NVIDIA ecosystem, what it means is we basically take the raw binary weights output from when you train your model on an NVIDIA GPU, you get a .pd, .safetensor file that you upload on Hugging Face, and, and that's your raw binary weights for, for your model.

So we take those, that file into that, our hardware, and we can actually output an OpenAI compatible API that people can then send their tokens in and tokens out. So from a workflow perspective, you can imagine you have an NVIDIA server running and you have a Positron system running, and you can have them side by side, and all you have to do is just direct the tokens to whichever system you want to, and it'll have the same kind of workloads running without having to really change or without having to recompile the model for a Positron system or so on.

Thomas19:35

Yeah. So like the really key thing there was, you know, from the get-go, we knew that the, the, you know, having a single step that a user had to take was going to be a, a, the, the ultimate, uh, barrier to, to getting any sort of adoption.

And, you know, be it a lot of the, the companies that you mentioned and, and, and others, I mean, even big guys like AMD have more or less, you know, flopped many times in their, their pathway to try to, to get people to use their, their, um, their hardware.

The approaches that people have taken before have been just saying, "Well, we're going to build a PyTorch backend." And AMD had their own separate, uh, you know, non-mainline PyTorch, uh, distribution for years that was always six to nine months behind any new PyTorch releases.

But I, I would say AMD's done an amazing job catching up in the past, you know, one, two years. But still, you have so many other libraries and functions in the PyTorch ecosystem which just aren't built around being able to work with AMD, and it makes a lot of people's workflows basically fall flat.

And that's before you get into the mess with drivers and everything else in, in the ROCm side of things. Um, so you have that on, on one hand, and then you have others, like Intel was trying to push oneAPI for Gaudi and, you know, all of their other stuff.

Um, a lot of the chip startups, so like a big part of their story was that, "Hey, we made a super compiler, and our compiler will just automatically be able to take anything and get it to run." And there was different levels of that actually working or not.

But I would say the, the core premise there was flawed from the beginning. If you are requiring a user or having yourself as the, the company needing to actually recompile-

Mitesh Agrawal21:11

Yeah

Thomas21:11

... uh, a workload, that's already one step too far, even if you assume it works perfectly. And so from the very get-go for Positron, we said, "Hey, we're only really going to care about, you know, transformer models." And we were betting on the Hugging Face transformer library and sort of the framework that, that was created there.

And we designed our hardware to directly ingest-

Mitesh Agrawal21:30

Yeah

Thomas21:31

... those raw binary model weights. So rather than having like we don't have a compiler whatsoever. There's no compiler, there's no translator, no tooling that's involved in actually taking those, those and getting that to, to, you know, for your, your, uh, your, your common, you know, Hugging Face transformer models to be able to then have that execute on our device.

And so we truly made it a zero-step process to take weights that you're already running on an NVIDIA-based platform and have that run on us. So like the bits, I, I, you know, as the CTO, I can, you know, I like to be technically precise in that, you know, the way that we say that we're in the CUDA ecosystem, it's by nature of the fact that we are ingesting the, the output of everything that exists in, in the CUDA infrastructure.

So the fact that we are betting that people are going to continue to train on NVIDIA for at least the foreseeable future, we're-- since we're able to, uh, you know, and I'll say, I really hope others are able to be successful in, in providing competition against NVIDIA.

But given the realities of things, we're going to assume people are going to still be buying a whole lot of NVIDIA hardware. We want to be able to have it be a zero-step process for people to be able to take models trained and that they're using on NVIDIA, and then be able to deploy that on Positron and get the performance, get the, the performance per dollar and performance per watt benefits, uh, from that.

Power & FPGA22:51

Swyx22:51

Yeah. Uh, I, I think that's a very reasonable, uh, business bet and, and, you know, thereby technical bet. One of the questions I was fed was just actually a little bit, you, you always mention, I think, the per, per watt, uh, efficiency.

You're like a third, uh, uh, of the, of the normal power consumption, and, uh, maybe that's underappreciated, but like how much actually power consumption is a part of inference cost.

Mitesh Agrawal23:15

It's-

Swyx23:15

Uh, apparently FPGAs are not... They're, they're supposed to be more power-hungry. Like how, how do you, how do you get there?

Mitesh Agrawal23:22

Yeah, I think so just one thing on the power consumption to your point about underappreciation. I think for, for us, like we've run a cloud business, right? So when we look at the cost structure, it's always you have CapEx, so that's performance per dollar, like what you're paying for the chip.

And then you have OpEx, which is performance per watt or energy that, that you're doing. And that's why, by the way, like just wanted to very explicitly say why we focused on those two metrics is we wanted to drive both the CapEx and OpEx either- Efficiency is higher or, or, like, basically their cost structure's lower, right?

Like that, that's kind of how we really focused on. And with the FPGA, I mean, in general that's true, but I think that's kind of where our own silicon design and memory architecture, and actually finding the right FPGA was the critical point to, to really drive that.

Thomas24:04

Like, well, I, I would just say that it really depends on what application use case.

Mitesh Agrawal24:08

Yeah.

Thomas24:08

If you're comparing an FPGA power to a cell phone, then yeah, the FPGA uses a lot more power. But like our cards are only using a hundred and fifty watts.

Mitesh Agrawal24:15

Fifty, yeah.

Thomas24:15

So it's really, you know, how you're using it, where it's being used, et cetera. Like, I'll say, like FPGAs traditionally have had like three or four main areas. One, you're using them, you know, internally, typically in a semiconductor company as a prototyping vehicle for, you know, whatever custom silicon you want.

Two, they got very popular in the HFT space to be able to design algorithms in... And maybe I should explain FPGAs for people that don't know. But, uh, you know, FPGA stands for field-programmable gate array. It's, uh, basically the, the abstract way to think about it is it's a special type of chip that you can kind of think of as being a, a sea of transistors, and you can actually program and, you know, and, and reflash the device to have different wiring connections of it.

That's not an actually technically factual statement, but it's the, the easy way to think about it. And so FPGAs can enable you to take any chip design that you have, and you can implement that on the FPGA, and rather than having to tape out with great expense and time and everything, you can try out those designs on the FPGA at very low development cycle cost and, and time.

That being said, that FPGA, in order to simulate like one gate, uh, that you would have in, in a normal piece of silicon, is using an order of magnitude more gates to, to simulate that. Um, they are inefficient from that standpoint.

But, um, I, I would really be pointing back to the-- what I was making before of that, you know, everyone else has been focused on that five hundred to a thousand FLOPs per byte world, and they're, they are that imbalanced that we can use an FPGA that was not designed for doing this workload.

Like, that hardware is, is very general from that perspective. But implementing our unique architecture on that, we're able to get these huge efficiencies because we're focused on that specific transformer memory bound use case.

Mitesh Agrawal26:07

Yeah. And then, and it kind of reverting back to your question around like, how did we launch so quickly? Like, I mean, relying on FPGA was, was the way-

Thomas26:14

Yeah

Mitesh Agrawal26:14

... that we got a product out so quickly. With our next generation, it's gonna be a dedicated silicon where, you know, the inefficiencies of FPGA, such as the just the brute force specs on both FLOPs and memory capacity, memory bandwidth, where those are lacking compared to the latest GPUs or things like those, will all be addressed, will be improved upon in so much so that our next generation silicon will have the highest memory capacity by almost, you know, four to five times than, than the-

Thomas26:39

More than that

Mitesh Agrawal26:40

... than, than any of the existing silicon at that point of time. So we are gonna be coming out with our ASIC in, in like later twenty twenty-six. Should-- It will have more memory capacity than any other silicon in late twenty twenty-six or in twenty twenty-seven, actually.

And, and so that's kind of where we are showing that, hey, you know, we've-- with our architecture, we can run it even on an FPGA this fast and, and provide those perf per dollar and perf per watt. So imagine when, when dedicated for a system designed for that architecture, how much more, uh, we can, we can, uh, perform there and then show performance.

Thomas27:12

Yeah. I, I think, like, the really key thing is, like in terms of FLOPs, we have about twenty teraFLOPs of, of, uh, you know, technically it's FP nineteen. It calls it TF thirty-two, and I, I don't wanna give them credit for having a very misleading...

So the-- in NVIDIA's TF thirty-two number format is a nineteen-bit number format. They just call it thirty-two bit because they-- it's closer to FP thirty-two in precision than BF sixteen. So they just said, "Why not just double the number?"

I'm, I'm exaggerating a little bit, but, but it's a nineteen-bit format to TF thirty-two. You know, we have that as our, our number format in, in our FPGA, which gives us better precision in our, our, uh, actual multiplication, um, which gives us, you know, a lot of, uh, benefits from a quality of result perspective.

But, you know, our FPGA is literally, um, you know, twenty teraFLOPS versus about a petaFLOPS on the H100. And we're still providing about seventy percent faster token generation rate. So, uh, it kind of goes to show that the FLOPs that NVIDIA spends so much silicon area on and so much power and everything else is, is not well utilized.

Mitesh Agrawal28:25

For, again, for transformer inference.

Thomas28:26

For inference.

Mitesh Agrawal28:27

Again, those FLOPs are needed for the training side.

Thomas28:29

Yeah.

Mitesh Agrawal28:29

Which is why when you're using the same chip across the board, like that's kind of where we see our opportunity into, into the inference efficiencies that we're driving.

Thomas28:35

But yeah, I will say for our second generation product, we're, you know, in the petaFLOPS of, of compute range and everything else. It's, you know... To go back to your original question of how did we do it so fast?

You know, small team, not much sleep, et cetera, et cetera. But, you know, a core, core element, like just philosophically different, is we decided, "Hey, we can actually use FPGAs for doing this because our architecture is so unique and targeted to this problem in a very different way than, than other approaches."

FPGA to ASIC29:02

Alessio29:02

What's the penalty, so to speak, to go from FPGA to ASIC? And then there's kind of the question of, well, maybe I should just get a ASIC with like the actual model in it, which I think obviously you shouldn't do because the models evolve and all of that.

But what are you exactly burning on the ASIC, and then what are you giving up by moving away from-

Thomas29:22

Yeah

Alessio29:22

... from FPGA model?

Thomas29:24

Yeah. So I would say another big difference between us and, you know, there's other startups out there that have, you know, said that they're burning or, you know, etching the, the transformer architecture into silicon, and that's their approach.

We're very opposed to that kind of philosophy. Like, fundamentally, what Positron has built is a linear algebra accelerator, which is optimized for matrix vector math in particular. And I would say, you know, more importantly, the fact that, you know, we achieve this, this, uh, you know, massive memory bandwidth, um, to that- Pretty general compute ar- uh, architecture.

So, like, I think it's pretty foolish for anyone in the just- industry right now to be saying that, like, transformers are going to be absolutely 100% the thing that gets us to AGI, or even if they do, that there isn't a better architecture.

And, you know, doing any of that, like, hardening for specific model things, I, I don't think lasts more than, you know, two or three months at the rate that the industry moves at. And so... But the thing that I would, you know, be willing to bet, you know, uh, you know, good money, you know, the company on-

Mitesh Agrawal30:27

Yeah, exactly

Thomas30:28

... um, is, is the fact that, uh, um, uh, m- you know, fundamentally linear algebra, uh, like b- building a good linear algebra processor, um, matters just as much in, you know, 1955 as it does in 2025 as it will in 2055.

Uh, having, you know, solving that as a core problem is what we've done and what we're focused on. And the, the software layers that enable you to get these models to run efficiently on that, um, you know, is, is a lot of our secret sauce to, you know, not, you know, harden the hardware in a way that makes it, um, uh, too optimized for just the thing that is hot right now, but really just build something that's good at the actual underlying math.

Mitesh Agrawal31:14

Is there something that will make you change your mind? Something in maybe the model architectures or anything like that?

Thomas31:21

I would say that the, like... I, I totally believe that there are application use cases that at our high enough volume, enough demand, et cetera, that, you know, a company can say, "Okay, we wanna build this super optimized chip that uses single digit watts or something.

Like, something that is so optimized because we know we're going to build a billion of these things, and we want it at the lowest cost and the lowest power for just this one thing." I believe that those applications can exist.

I've not seen something that, like, would meet those criteria for me that would be worth the, the NRE and, and everything involved to, to do that. And if you're trying to build something that's general purpose to saying, "Okay, building a transformer specific chip that is hardening things in a way that are very, very specific to current transformer architectural decisions," I think pretty much everyone in this space should know that this industry is moving astronomically fast and it's accelerating.

Mitesh Agrawal32:13

There, there's... I mean, obviously I don't have the example of what you asked, but there's a counter example actually to what Thomas is saying. I mean, even the attention mechanism changing when DeepSeek... Well, not changing, but, like, if DeepSeek opened it to the world, which was the MHLA stuff, I think if you had burned or etched in the, the architecture of doing transformer the way before that, you know, you might not be able to actually gain all of those efficiencies that the attention mechanism that DeepSeek announced to the world.

Thomas32:37

Yeah.

Mitesh Agrawal32:37

Right? So, so that's, that's in my mind a counter example why that wouldn't-- that is not the best approach, or at least we think that's not the best approach, right? So that's not what we are doing, is, is burning into a particular architecture.

Cool. Yeah, I mean, I, I think that this is relevant to the, always the discussion of the other ASIC companies like Etched. I don't know if you've seen the beef that George Hotz has had with Etched. Any comment there?

Uh, what, what... Is he, is he right? What, what is, what is he, what are they talking about?

Thomas33:05

George loves having beef with everyone. George and I have known each other for probably ten, 12 years now. I, I, I don't actually remember specifically all of his beef or anything. My, my, my opinion in general is it's a remarkably difficult thing to build any new silicon.

Mitesh Agrawal33:21

Yeah.

Thomas33:21

And I know I've met, you know, the Etched guys as well. You know, they're cool guys and, you know, I respect them tremendously for, for building a company and, and, uh, uh, tackling a hard problem. I think that the...

I would disagree with the idea that you should be hardening in that way, and I would say I'm skeptical of how much optimization you can even get with those things with that wouldn't be putting you in a crazy restricted corner.

But I'm, I'm looking forward to seeing what they develop. Like, I'm a technologist. I, I wanna see cool things in the world. You know, even if I think, you know, that we're building something that's going to be better in a more pure sort of linear algebra direction, I can totally believe that we'll have success in what we're doing and others will have success in their areas.

You know, I would say the same thing for any other company. Like, I would say just in my talking with, um, you know, the, the founders and, and people in, in all of the different semiconductor, you know, startups trying to go after Nvidia.

Success for any one of us is a success for all of us. Like, anyone that can, like, take a little tiny chink in Nvidia's armor is, uh, you know, a great win for eventually Goliath not collapsing. I mean, the reality is Nvidia is going to have a very, very good decade ahead for them, and the market is growing so fast that all of us in the space trying to, to take them on can be very happy with, you know, very, very small wins, uh, in the space.

Mitesh Agrawal34:41

Not a, not a huge fan of that particular argument. I am not. But, uh, obviously that, that, that's just, uh... I think the other way-- the way I like to present that one is actually, and, and it's not even just trying to be nice or anything, it's just really, like, look, the, the markets are growing so massively, and when certain inference applications will come out, they'll be at such a big scale that people will wanna minimize costs, right?

Because inference is what drives revenue, so that's actually driving true margins for companies, so they'll want to minimize cost. And if our architecture is suitable for certain industries much more than, let's say, Nvidia's or AMD's or any other silicon architectures, I think they would use us.

So, like, in a way, it's kind of like expanding the market for those companies to be much more efficient than th- what they would be on an Nvidia GPU or an AMD GPU. So I think, you know, from, from my perspective, it's not just like, "Hey, you're trying to capture, you know, small point 1% or 1%."

I think the idea, the goal is from a technical point that Thomas and I always talk about is obviously going and supporting 100% of the market and then really proving that out and proof is in the pudding, and then capturing as much of it.

But you have to start somewhere, and, and we're starting somewhere, which is, okay, we are focused on this really, you know, where, where we think today are the primary inefficiencies in AI inference workloads and really solving them and showing that, uh, for, for it.

Thomas said that, you know, we are accelerating matrix vector. It doesn't mean that we can't do matrix matrix. We can, and we can do it fairly well. It's just that the point becomes is like, are you really beating Nvidia on it from a perf per dollar?

And if you're not, then Nvidia is, is, is the existing thing, so people won't move on it. So today we won't be able to, but in, in the future, I think that's kind of as the company matures and, and if we get the opportunity to grow revenue to that scale that we can actually focus on these things then, then we will try to.

Thomas36:20

And, and I would also just say from, like, that competition scale, like, again, I'm very ho- like, hopeful that many people come out to market with great pieces of silicon. I would say from a competition angle, the key thing that everyone should really be looking at to measure success fundamentally comes down to revenue.

Mitesh Agrawal36:36

Yeah.

Thomas36:36

You can, you can announce the greatest piece of silicon. It could be a great... Have all, all of the right specs, et cetera. But- I would s- I, I'm a big believer in capitalism and, you know, the invisible hand.

And that's like the companies that everyone is trying to sell to, the companies that are building the future grand applications that we're all going to be using, they have all either raised money or they've generated, you know, their own revenue, and they have dollars, and they're going to commit those dollars to the, the infrastructure providers that actually generate value to them.

And so the, the biggest thing that, uh, everyone, uh, in the market, I think forgets when they hear so and so company has raised a huge amount of money. Raising money is one thing. Being able to deliver real value to the customers, to the end users, such that those customers are then-

Mitesh Agrawal37:24

Paying you

Thomas37:24

... you know, pay, paying you for, for that value, um, is, is the only true measurement. I mean, I'm really happy that, you know, we had our first revenue, like we, we sold our first system, you know, fifteen months in.

You know, we've deployed now multiple rack scale deployments. We're announcing... So I think with-- when this is being timed, we'll-

Mitesh Agrawal37:44

Can embargo this, yeah.

Thomas37:45

Yeah. Um, you know, we're, we're announcing two, two customers with our fundraising, but we have multiple others that we're hoping to announce, uh, in the coming, coming weeks and months.

Mitesh Agrawal37:54

So we have customers in Cloudflare and Parasail both using us. Cloudflare because of performance per watt kind of a story. Like, you know, they, they really love the fact... Look, they have data centers at the edge in this, in, in metropolitan cities, right?

So they can't really scale out the power that much, so they, they want to maximize performance per, for the given amount of power that they have. So that's kind of where they're really working with us. And then, uh, Parasail, which is, you know, much more focused on performance, raw performance because they're an inference as a service company, right?

So, so that's kind of where they're targeting us. And then by the time this comes out, I think, uh, we'd have our fundraise announcement, uh, done. Uh, you know, and that's, you know, we, we, we've raised in the history of the company so far that we publicly announced roughly twenty-four million, and then we just, uh, did a Series A close with, uh, Valor, Atreides and DFJ Growth, uh, for a roughly fifty-one million dollar Series A.

So, um, and, and, and that money for us is really like focused on coming up with the next generation silicon. We get to tape out, we get to sampling, we get to production, uh, with that money. And again, we, we really want to prove it in a capital efficiency way.

Look, I think that's what we've done so far. Ideally, we get the customers to pay for our, our production. Right, so like, like NVIDIA- ... we wanna sell systems. Like we wanna sell systems to customers. Don't ideally like, you know, wanna go into this cloud services where you're owning the, the infrastructure on your balance sheet.

That's kind of where a lot of the current ASIC companies have gone. You know, if you look at Cerebras and Groq and others, they've, they've really tried to do this, their cloud kind of portal. And, and for us, you know, it, it shows two things.

One is the difficulty of software for a third party to, to implement that. That-that's why they're kind of abstracting away that layer and, and providing it as a cloud API. And then second is to convince a third party customer, you really have to show the economic viability of your system, right?

Like that, "Hey, this, the return on investment on this system will be two years, three years, one year." And really have shown that. Like, you know, if you deploy NVIDIA GPUs, hyperscalers, NeoClouds, customers have all seen that they can make that revenue back over a certain timeframe.

And, and that's what we're really here. We wanna sell systems and show people can make that ROI, uh, back on that system. So that's the way we are approaching it-

Thomas39:58

I-

Mitesh Agrawal39:58

... and going to market.

Thomas39:59

I think the, the four most important letters, uh, are ROIC. Uh, um, you know, from-- It's, it's not just from, as Mitesh was saying, that from a customer's perspective, if they're deploying hardware, they're going to be choosing between be it NVIDIA, AMD, Positron or anyone else.

They're looking at it from an ROIC perspective to also in our, from, from our side, we don't wanna go and raise hundreds of millions of dollars. Like we, we had the opportunity to raise, you know, many, many times more than what we, we raised earlier.

I was adamant, like we raised originally a twelve million dollar chunk and then about ten million afterwards to, to get us to the point that we are today. And it was because I want to raise as little as possible to minimize dilution, and because I feel that companies that just raise huge amounts of money typically don't spend it very well.

And I have as a like really key metric for our success is how efficiently we can actually deploy that capital to creating real value that we can then provide to customers.

Mitesh Agrawal40:55

Yeah. Uh, amazing. Uh, love that messaging. One thing I wanted to double-check and then we can sort of, uh, wrap it is, is another way to phrase it is that a lot of the other folks are focusing on prefill optimization, and you're-- because of the linear algebra and, and sort of general matrix math capability, you will just decode a lot faster.

Decode Focus40:55

Thomas41:13

Yeah.

Mitesh Agrawal41:13

So, okay. You're-

Thomas41:14

So like-

Mitesh Agrawal41:14

I see you guys nodding.

Thomas41:16

Yeah. So like I would say definitely our, our, um, performance focus and where we're getting the, the best relative advantage over NVIDIA is on that, uh, generation or decode side, um, of, of it. So, you know, for, for any transformer LLM, you know, you have a prompt that you're sending in, um, and that prompt, uh, you know, of, of some amount of size.

Computationally, that forward pass is exactly like the forward pass when you're doing training, because you've got the whole sequence of tokens that you, you want to process. All of that is there for you upfront. Then when you switch into generation mode, then you're generating one token at a time, and, and that is, is the memory bound portion.

So yeah, I would, I would say, you know, we're still fully capable and, and perform well with, with, uh, prefill. But I would say the thing that we believe and, and I was so happy, you know, last year when, when all of the new reasoning applications happening.

'Cause if you, if you go back a year from, from today, in, in July of, of last year, the ratios of like input to output for, for LLMs were very, very heavily on, on input, where you could be doing, you know-

Mitesh Agrawal42:23

Ten to-

Thomas42:24

... ten, ten, fifteen to one ratio of input to output tokens.

Mitesh Agrawal42:26

Yeah.

Thomas42:26

But that has completely flipped. And it's-

Mitesh Agrawal42:29

Oh, my God

Thomas42:29

... obviously use, use case dependent. But if you look at all of like when people are running benchmarks on these new models, they're generating, you know, it's not just flipped. You could be generating like 100 reasoning tokens for every token that you had as a input prefill.

The other big thing, and like our big Belief in having massive memory capacity is the fact that, you know, if you're able to actually just cache the, the KV caches-

Mitesh Agrawal42:52

Yeah

Thomas42:52

... so if you're able to hold massive system prompts, massive data sets, et cetera, that are typically recomputed every time you, you send in a new request. If you are able to store that because you've got a massive amount of memory that you, you know, lying around, then that e-even makes it so that the true prefill compute that you're doing is significantly less.

Mitesh Agrawal43:11

Yeah. And then I think Thomas touched upon, like, token generation. If you also look at, like, other modalities like video generation, for example, right? And I think if you convert that into effective tokens or whatever, like, then... And, and if you look at the ratios there, that's, that's also a completely flipped script.

So, like, I think that's other part of it where we will see a lot more integration into our models, reasoning models, multimodal models. And I think, but as, as Thomas mentioned, I think we are definitely, like, where our true advantage shines through is, is, is generation.

So, like, time to last token, generation speeds, we are, we are way off the charts compared to any other existing solution out there.

Thomas43:45

Yeah. I'll, I'll just add there on, like, the video example. The other exciting thing to see over the past, uh, few months, year has been that, um, a lot of things have actually been moving away from diffusion to being pure auto-regressive-

Mitesh Agrawal43:58

Yeah

Thomas43:58

... uh, transformers for image and, and video generation. So, like, the latest, yeah, so V03 and, and since Image Gen-3 on, on Google's side have been pure auto-regressive moving away from diffusion.

Mitesh Agrawal44:08

Oh.

Thomas44:08

The latest-

Mitesh Agrawal44:09

I'm not sure we knew that.

Thomas44:11

Um, the, uh, Athena model from xAI on image generation side is pure auto-regressive transformer. The ChatGPT image gen, the, the, you know, that's pure auto-regressive transformer. And so all of those cases are now diffusion could be extremely compute bound because you're generating all of your pixels, you're doing compute on all of them simultaneously.

But when you've... are in auto-regressive mode, that is back to this one-to-one bytes to FLOPs ratio, which is where we-

Mitesh Agrawal44:38

Yeah.

Thomas44:39

The second part-

Mitesh Agrawal44:40

Amazing

Thomas44:41

... is, like, when it comes to video, like, literally right now the, the rumors, estimates I've, I've heard is that, like, for the eight seconds of, you know, V03 image generation is, you know, around, you know, 800,000 to a million tokens.

So if you think if, if you're just ballpark of saying it's 100,000 tokens per second generated, the actual input prompt to that video model is hundreds of tokens maybe, if that. So that, that's, uh, that ratio is just absurd when you get into modalities beyond text.

And that being said, I do think bringing in video and then saying, "Okay, I want a new version of it," like, the ratios will always adapt and it'll be different for different-

Mitesh Agrawal45:20

Yeah

Thomas45:20

... applications. But I think the key thing everyone needs to remember, you know, everyone calls this generative AI. I'm, I'm going to say it's, it may be reductionistic, but that's, uh, generative AI. Hey, generation is a lot more important than that prefill.

Mitesh Agrawal45:34

Um, awesome guys. This was great. Any call to action? I'm assuming you're hiring, you just raised a new round. What type of people are you looking for?

Hiring45:34

Thomas45:41

Yeah, I think, uh, that's, that's, that's a big call. I mean, we, we are a small team, 27 people right now at this moment. We are definitely heads down working on our next silicon, uh, right? And as with any silicon, you know, looking into hiring great, uh, you know, verification engineers, design engineers, silicon folks that have experience, you know, N5, N7, N3 obviously is, is very relevant to us, uh, from Frost Stones perspectives.

Really looking to get that team ready to be testing verification emulation and then when we actually get the tape out in hand, really running quickly to, to system integration. So that's kind of the big part of it. Other f- part that I really do wanna call out and not forget is the software engineering side of things.

I think, you know, we've, we've done well with the small team and the approach we have taken is really, really good. We bet on the Hugging Face transformers. We do wanna make sure that we are keeping up with the multi-modalities, uh, of, of the different models that are coming out.

Obviously, as I said, like, we are accelerating matrix vectors, so we're really good for transformer inference, but it doesn't stop us from supporting other model types as well. And so we really wanna build a very rounded software engineering team as well.

So looking forward to hiring that and yeah, we-

Mitesh Agrawal46:51

And-

Thomas46:51

... folks around the United States. Yeah. So we, we are a remote-

Mitesh Agrawal46:55

Yeah

Thomas46:55

... distributed company. Uh, we do have, you know, a few, uh, few offices right now.

Mitesh Agrawal47:00

Offices, yeah.

Thomas47:00

Um, but I would say that, like, really want to be able to bring people in that have no hardware experience, and especially those that have, you know, been absorbed in the machine learning space and really understand models and want to take it to the next level of seeing how do you actually get the, the most efficiency and how better models can be created given, you know, new hardware that comes out.

We're, we're definitely-

Mitesh Agrawal47:22

Yeah

Thomas47:22

... you know.

Mitesh Agrawal47:23

He's talking about the next step after like, you know, when you think about someone looking at our architecture and like, "Actually, you know what? I can make a model architecture to be a lot better because I have all this memory capacity or all this memory bandwidth available to me-"

Thomas47:34

Yeah

Mitesh Agrawal47:34

... uh, as well." Cool. Awesome, guys. Thanks for joining.

Thomas47:39

Thank you so much- Thank you

Mitesh Agrawal47:40

... uh, for having us and, and for everyone.