Transcript

Neil Movva - Making AI 10x Cheaper - [Invest Like the Best, EP.488]

Free .txt

0:00 Ramp is the only platform built to make your finance team leaner, faster, and better, saving businesses 5% annually on average so you can stay focused on growth. Ram customers grow revenu three point two times faster than the average American business. Visa, Versel, Cursure, Stripe, Notion, Eleven Lab, Shopify, and 70,000 other businesses all now run on ramp. Mine does too, and so should yours. Learn more at ramp.com slash invest. OpenAI, Cursor, Anthropic, Perplexity, and Vercell all have something in common. They all use Work OS. To achieve enterprise adoption at scale, you have to deliver on core capabilities like SSO, Skim, Rback, and Audit Logs.

0:38 Instead of spending months building these mission critical capabilities yourself, You can just use work oas APIs to gain all of them on day zero. That's why so many of the top AI teams you hear about already run on Work OS. WorkOS is the fastest way to become enterprise ready and stay focused on what matters most, your product. Visit workos.com to get started.

0:57 Felix by Rogo is a personal finance agent that turns a single prompt into finished client ready work using your firm's own templates, context, and standards. Send Felix an email like take these comments and turn them for me. Or update my tracker with the context of these emails, and Feel extends back finished PowerPoint decks, Excel bottles, and sourced research. Felix works the way your team already does, delivering work quickly and accurately around the clock. Learn more at rogo.ai slash feelix.

1:28 Welcome everyone. I'm Patrick O'Shaughnessy and this is Invest Like the Best. This show is an open ended exploration of markets, ideas, stories, and strategies that will help you better invest both your time and your money. If you enjoy these conversations and want to go deeper, check out Colossus, our quarterly publication with in-depth profiles of the people shaping business and investing. You can find Colossus along with all of our podcasts at Colossus.com. Mm-hmm. Chick O'Shaughnessy is the CEO of Positive Sum. All opinions expressed by Patrick and podcast guests are solely their own opinions and do not reflect the opinion of Positive Some.

2:02 This podcast is for informational purposes only and should not be relied upon as a basis for investment decisions. Clients of Positive Sum may maintain positions in the securities discussed in this podcast. To learn more, visit psum.vc. My guest today is Neil Nova, the founder of Sale Research. Sale is building what Neil calls a token factory, an inference company designed for a specific kind of future. One where AI agents run in the background for hours or days at a time rather than answering a human in real time. І на вол, latency matters last, and cost matters much more. And Neil has built the entire company. Around driving the cost of a token as low as it can possibly go.

2:41 What makes this conversation special is it is one of the most detailed tours I've ever done through the full stack of intelligence, the software, the chips, the power, and how the three connect. Along the way, we cover the trade-off between speed and cost that lives inside of every GPU, his scavenger strategy for buying the chips and power no one else wants. His contrarian view on NVIDIA. And why the premium the Frontier Labs charge for being three to six months ahead may not last. Please enjoy my conversation with Neil Mova. I think it's important early in these conversations to just say the thing, literally what you're building and what it does today. So maybe just orient us.

3:16 there with a brief description like literally what the system is that you're building and why it should exist. Cell research is a token factory. We have an API where anyone can send us. Requests where they can use large language models, open source large language models. for any task they want. We will serve those tokens to them at a price that is unbeatable in the market. We also support their ability to build agents on top of this. We host what we call saleboxes, which are long running agent virtual machines hosted in the cloud.

3:42 that are designed for agents that run for hours, days, or weeks. So I should think about you as a peer company to others that serve different kinds of inference. You're serving one specific kind of inference. And your goal is to be the absolute cheapest provider. an enabler of a certain kind of use of intelligence.

3:58 Exactly. The theme of our company is abundance. We want to deliver this new commodity of intelligence. To as many people as possible. at a cost that is sustainable for almost every industry. We think that whenever you make something ten times cheaper it's a new product category. We aspire to do that for tokens. We think it's so profound that the machine can think. And now our job is to make as many machines as possible in the world.

4:19 work towards thinking. So if you think about the theme of the day being token costs, is token cost the right way to think about this? Is there some other way you'd put it? To start with, absolutely, token costs. Today my North Star is.

4:30 I wanna have the lowest cost per token in the industry and do that by a mile. I don't think tokens are the final unit of work or intelligence. But they are what we use today. After tokens, you start to move more towards More outcomes. Which is like a vague direction. You can imagine, for example

4:45 Today when you consume tokens through an agent, you don't actually control how many tokens the agent reasons for. It can reason for a certain amount of time or it can call a certain number of tools. And increasingly I think we will have agents do some unit of work, take as many shots on goal as they can. And however many tokens they use to get there is going to be

5:02 A dependent variable depending on the task. So you think about like agents that self-administer token budget as opposed to a company sending a budget for how many tokens engineers can spend per month. Why is there an opportunity that you can tackle it seems like the entire world is oriented around.

5:18 More, better, faster, cheaper tokens right now. It seems like the world is trying to solve this problem very aggressively. What was the unique opening that you saw that's maybe the market's not being efficient and it's attempt to tackle this. So I think there's two things that are tailbones for our company. One is gotta be the rise of open source. I had to talk about that first. I I think

5:38 We're starting to see increasing number of our customers and the broader market care about owning intelligence. They want to have control, sovereignty over the thing that they depend on. That created a much more robust market for our customized models. Or even just like these vanilla open source models that no one can ever take away from you. You always have the weights, you always have the right to deploy them however you like. In that world, there's been a reasonably robust market for the past couple of years serving these models at large scale. The challenges.

6:04 All those companies, you could take your pick. Base ten fireworks together. They all focus on low latency inference. And they were pulled in that direction by one very important customer. Cursor

6:14 I think that that was the right choice about a year ago. And as of six months ago, it started to look like maybe low latency wasn't the only thing you wanted from an agent. You wanted more persistence, more long horizon tasks. And now it's to me very obvious that Future of genetic inference. Is

6:29 Log Rise and Tasks. You're gonna run. the machine for hours or days at a time. It doesn't matter if it's been set tokens at a hundred tokens per second. Maybe ten is just fine. If that comes with

6:39 corresponding advantages and efficiency. Why are you so s confident in that? To me it seems like I want everything as fast as possible. When you're waiting on it, you absolutely deserve the fastest answer possible. Yeah. My trick is I don't want you to be waiting on it. I want it to be proactive. I want it to be in the background. One way to say it is like the best latency is no latency at all. When you wake up in the morning, the work's already been done overnight. You didn't even have to ask for it. That's the dream. We're not quite there yet.

7:02 More importantly, I think the more you're in the loop as you prompt agents and wait for a response. In fact, you're the bottleneck. in having agent do more or less work. What we'd like is the agent to operate on more human time scales.

7:14 You don't manage your colleagues every five minutes. You ask them to do a high level task and you come back and check in. Maybe every day, but more likely once a week. And that to me is the future of human agent collaboration. More like human time skills. Say more about the early indications that this is happening and therefore you should be building this company.

7:30 So the first and most important thing is the idea of test time compute scaling. The idea that you can give an agent more time and it will give you a better answer. So that was theorized about two years ago now. I mean it wasn't really something that we could actually bet on until I would say late last year with Opus Four Five. Hope it's four five was the first agent that was at all suitable for longer horizon tasks.

7:50 was pretty mediocre and when it first came out, but you look at the more recent models and what we've done on open source as well. And you see that agents are capable of running for An hour at a time. I wouldn't say it's days, but definitely an hour is quite suitable today. Seeing that like average task length get longer and longer.

8:06 It doesn't take many points to have you draw out the exponential and see that agents are worth running for longer periods. What do you think will be the market share of long running agents in three years or something like this? I love this market because it's unbounded. There's no human in the loop, so you can consume as many tokens as you like in the background. versus human attention span. If you tell me to consume tenx as many tokens at codex or at quad code. I'm actually not sure if I can anymore. I'm already in the loop and locked in coding for most of the day that I'm at the laptop. What is unbounded is how many

8:35 Tokens can be consumed in the background or proactively. Long term, I think, you know, we're gonna end this year at maybe fifty fifty. background and real time workloads. But I see this going to ninety ten in favor background. What are your favorite examples of something that gets accomplished much better as a background task than as a human in the web task.

8:53 Most deep research. Most questions where you want to have a definitive answer over not a hundred sources, not a thousand sources, but ten thousand sources or more. If you want to build an authoritative index of information, like for example, one of our customers, Parallel Web Systems, seeks to do, they want to build an index over the whole internet. And they want to monitor the internet in real time for changes.

9:14 crazy exabyte scale task that you need a very different kind of intelligence or scale of intelligence to achieve. Deep research is a top category for us. And then increasingly we see cybersecurity following in this direction. If you think about Yes, there's so much code you can generate, but there's exponentially more ways to break that same code than it is to generate that code.

9:32 There are some great customers out there who are working very hard to find agents that can break any piece of software and Proactively patch them. When Fable first came out, for example, or Mythos first came out, basically there was this push in the cybersecurity community to run Fable against Every line of code we've ever written.

9:48 And look for bugs in twenty different ways, meaning you're looking for both memory errors, you're looking for business logic errors. And looking for like network vulnerabilities. All these things. And these are all actually things that you would write specialized agents for. You wouldn't just have Fable look at the source code once, you'd have it actually set up environments where you can pen test these applications. At some point people started to make this joke that security has become proof of work.

10:10 When you want secure software, it's really a question of how many dollars did you spend on anthropics APIs. trying to break into your software. That is the best indication for how secure it is, because that's the best tool in the world. And increasingly we found that open source models Well the frontier of intelligence here is quite jagged. It's not the case that Fable finds a superset of all bugs in software, you would find some bugs with a very small model that you don't find with a large model.

10:33 You'd find some bugs with haiku. That you wouldn't find with Hable and vice versa. So it encouraged this very diverse. approach to sampling and and trying to build cybersecurity agents that break software autonomously. Such that you can touch them.

10:44 If you were to get speculative and imaginative about The sorts of things that very cheap, very long running agents can enable. We talked about some very practical examples, deep research. Cybersecurity, et cetera. If you get a little bit dreamier about

10:59 the use cases new product category that this sort of inference will unlock. I guess the question is just like so what? If you're maximally successful. Dream a little bit about what that might enable. I think for individual users, what I'm excited about most is this idea of proactive intelligent agents.

11:17 You can imagine a Siri that is running in the background all the time to understand all the ema you received in a day, all the text messages you receive in a day, and has a much more encyclopedic view of your life and how to be helpful in that life. Right now there's still point solutions. You're gonna do a lot of prompting. Series non reproactive. Something we can fix with abundant inference.

11:34 If you trust the machine enough that it's reliable and also Trustworthy as in private. You might even imagine the machine can understand how you interact with it and proactively surface. your next action whenever you open your phone. Can we build a good model and what you're going to do next?

11:47 My estimation is yes, we totally can. And the key to that is incredibly cheap intelligence. You have to be willing to spend tokens. Without any promise of return. That is the unlock. the long lens view to take on this is that We have a form of intelligence that can tackle any verifiable problem.

12:04 Any verifiable problem means most software. It means a lot of formal Math proofs and similar. And it could also mean scientific. These are all

12:14 Relatively verifiable problems. And all those things currently have a dollar cost attached to them, essentially, that's a hidden one. It's like how many tokens could you possibly harness to make this work? We have started the Bring it within view. a dollar cost for these long horizon tasks that is

12:29 reasonable. It's not millions, it's thousands. And maybe it could be hundreds or even tens of dollars in the near future to have a definitive answer. to any scientific question. to any research problem. So if we dream about that future, we then become limited just by the questions that people can ask, basically.

12:46 Pretty much the questions we can ask. The models are on the cusp of basically taking even a high level question and chasing it down every possible follow up. You can have the model essentially take that on its own. And the question is, what is your token budget? And we will solve the token budget problem. What about nonverifiable? Tasks. Those are Basically the entire category of human taste into that category.

13:06 We have not solved human taste yet, and I don't know that it Fundamentally can be. I'm excited to be surprised here, but We are focused on very quantitative problems. We leave the quality of writing, we leave the beauty of art to people.

13:18 I think for individual users, what I'm excited about most is this idea of proactive intelligent agents. You can imagine a Siri that is running in the background all the time to understand all the ema you received in a day, all the text messages you receive in a day, and has a much more encyclopedic view of your life and how to be helpful in that life. Right now there's still point solutions. You're end of doing a lot of prompting.

13:38 Series non reproactive. Something we can fix with abundant inference. Vanta automates security and compliance for over sixteen thousand fast moving companies like Ramp, Cursor, and Harvey. Keeping them audit ready around the clock. It's the number one agentec trust platform.

13:54 And it now helps companies like yours watch for the risks that show up between audits. Across your vendors, your AI tools, and your whole environment. Every new tool your team signs up for, every vendor that turns on AI features is an opportunity for something to go wrong. And most security programs weren't built for AI's pace of growth. The Vanta agent works like a twenty-four-seven GRC engineer in the background, finding issues, drafting fixes for you, and cutting vendor assessment time.

14:19 By up to fifty percent. Whether you're a fast growing startup or a global enterprise, Vanta helps you earn and prove trust. Invest like the best listeners get a special offer for one thousand dollars off. advanta.com slash invest. RidgeLine is the first end to end system of record with embedded AI for investment management firms.

14:39 Running portfolio accounting, reconciliation, reporting, trading, and compliance. On one unified platform. Firms are moving off legacy technology and onto Ridgeline because of how far ahead's AI features are compared to anything else in investment management software. Which is why I believe that firms that come out ahead in the AI era will be the ones running on Ridgeline's unified platform. If you're serious about your firm's AI strategy, Rine should be part of that conversation.

15:04 You can request a demo at ridgeline.ai. All right, now let's talk about the very clever stack of solutions that you hope to build. Ultimately they have this giant Token factory supplier of extremely low cost intelligence. I think you think about this in terms of software, hardware.

15:23 And power. Talk through what your master plan is to approach this challenge that's so different from what others are thinking about doing. We always have to start with software. Where is the opportunity on today's data centers to improve efficiency? And the first thing we did was we tried to build the entire LM software stack.

15:40 around peak GPU efficiency, meaning we're using a bit of GPUs. We wanted to squeeze out more tokens from the same chip. than anyone else in the world. And that starts with the lowest level of programming, kernels. It's actually my background. I spent my whole professional life working on GPUs and kernels. NVIDIA was my first job while I was in college. I got to see how the tensor cores got to earn their right. To be on the chip.

16:00 What does that mean? Like what is a tensor core? Tensor core is a specialized unit on the GPU that accelerates matrix multiplication. Simple as that. There's been a long history of how we evolved that tensor core over time that we'll get into. And why is matrix multiplication so important? I cannot say that there is a divine truth of the inverse that explains why matrix multiplies seem to be the atomic unit of computation, but Well, it's a really succinct way to mix two blocks of numbers together and have them interact in some interesting way. That's as much as I can say about it. It is really convenient that linear algebra turns out to be a very compact representation of

16:32 Arbitrary relationships in data. Nvidia Great graphic company. had market share dominance in GPUs and gaming graphics for quite some time. And then starting in like the mid twenty tens, they started to actually start these like Skunkworks projects to make the graphics processor more suitable for machine learning tasks that they were tracking.

16:50 I remember actually reading some of the lab notebooks of some of my managers when I was at NVIDIA. They would visit these small ML conferences like ICML or Nurepse at the time. they would just take note of these papers, like, oh, this deep learning thing seems to be catch on. And what's really interesting is that these grad students are using gaming and video GPUs. in order to train their large models. We should double click on this and figure out what's going on here.

17:12 By twenty fifteen, twenty sixteen at least Jensen had the conviction to kinda double down on Hey, this usage of our chips is only gonna grow. Let's start allocating more and more precious silicon diaria. to this capability that seems to be emerging.

17:25 Let's put the first version of Tensor Cores on the chip. So we're talking about taking this gaming chip, which is designed for painting pixels on a screen. And adapting it to do Make cult applies.

17:36 It was early. you would be competing against the graphics teams, essentially, when you ask for more silicon area at any chip company. There's always competition for that. It it is something that the the the designers guard so carefully. you don't ever want to invest in the wrong technology because that's opportunity cost that you could have allocated to some other functionality. We fought tooth and nail and got just a tiny bit of diet, maybe like five, ten percent, something like that.

17:58 For the first generation of these chips. to get some amount of acceleration for basic convolutions, which were the fundamental operation for computer vision models in the day. And then we had a software team. I was trying to squeeze.

18:11 all the performance we could out of the chip. And I think on that software team, which is where I work, that's what actually taught me the most about the ethos that NVIDIA has around this term called speed of light. They always chase the speed of light for any piece of hardware that they make. It is so ingrained in every engineer's mind that If the machine can do it, we're going to push the machine to the frontier until it does what we think. And the speed of light is the edge of what's possible. The speed of light is the edge of what's possible, exactly. If we think the chip can run at this frequency and produce this many multipliers per cycle, we're gonna get there. We're gonna break every bottleneck and get to that peak level of performance. To this day, I tell all my engineers, we're chasing a hundred percent Speed of light.

18:46 I don't care about relative numbers for the competition. I only care about absolute numbers. What are we able to do on the ship? How do we achieve that? Before we leave that chapter of your time at NVIDIA, anything else beyond that? cultural touch point that change the way you think about things or That stood out the most about how the business ran back then or its culture. I have a ton of stories about it in video. I could tell you a few of them. One of my favorites is that On the tenure side.

19:06 a lot of people I worked with in NVIDIA in twenty fifteen, twenty sixteen are still there today. That company has incredible retention and these are the best engineers. On the silicon side, at least, I've worked with them my whole career. They're extremely, extremely motivated and passionate. They believed in parallel computing as a concept. through its various incarnations and have Loved seeing the chip evolve.

19:24 This is their life's work and they're extremely competent in that direction. They're also a very frugal company. Nvidia and all of Silicon Valley companies after two thousand and eight, they had some cutbacks and like perks. So no free lunch, for example. Nidia took it one step further. There was no free milk in the fridge. So if you wanted to drink coffee in N video. And you wanted some milk, you'd actually had to chip in a dollar every month to the milk club.

19:45 And then all Clubwood stock. Costco milk in the fridge. And I remember that. Distinctly. We don't do that at Cell. It's a frugality that permeates.

19:52 And so coming out of this time there, you get this experience of what it's like to develop more efficient usage of the underlying hardware through software. So link that to today's environment. The GPU is fundamentally a throughput machine. The GPU is happiest when you give it a lot of work to do and let it chew through that work at peak utilization of its compute units. But that's actually not the way that we've taken AI in the last couple of years. We've

20:15 pushed AI to be an interactive chatbot tool. is the most common form of AI usage today. In that world, you care a lot about actually spitting answers out to the person at the keyboard as quickly as possible, to your point about don't make the user wait. I want things as fast as possible. That's actually quite interesting for the GPU. It's very difficult to put the GPU in its happy path of being fully. when you're trying to spin out tokens quickly.

20:36 There's a fundamental trade off in the GPU between being throughput oriented or latency optimized. And everyone has chosen latency optimization because the shape of Usage was chatbot oriented. I believe that's the most profound change we're going to see in the next year. We're going to move away from chatbots to more proactive or background agents.

20:53 in that world, it makes a lot more sense to build a stack around throughput. Can you explain? Technically why the trade off between throughput and latency is unbreakable. Why can't we have both from the same Hardware. It's quite foundational in almost every system that you could ever possibly look at. There's always a trade off between getting a small amount of data through the system as quickly as possible and leaving a lot of buffer room for that. Or trying to run wide and slow. Narrow and fast or wide and slow is like a classic trade off in all computer science.

21:21 But for GPU specifically, I think there's one thing to focus on, which is there's this concept of like batching on the GPU. We want to group many users' work together. into a batch that we can run all at once on the GPU. That's the parallel processing of the GPU. We'd like to have a lot of parallel work to do. The thing is though, you're doing net more work when you run a large batch of compute together. So you might be filling all the units, but

21:44 Every step along the way as you carry a batch of work through the GPU. There's more work. To be done. any individual token or any individual user's request in that batch.

21:53 it's gonna spend a longer time on the GPU being carried with other people's traffic. Maybe the way to say it is If you want to get downtown in SF, you can take the bus or you can take A private transit.

22:04 have its own direct path as the crow flies or using exactly the roads that you want from point A to point B. A bus it's gonna have to serve many more people and it has to fundamentally do something that works for everyone. So it takes a slower path and it stops and and waits for other people. To get on and off. I think the bus versus car analogy is pretty accurate.

22:21 It's a great analogy. And step one for what you're trying to do is create the best possible bus on top of NVIDIA GPUs. Like that's step one of your Optimization. That's exactly right. It means we explore things like different parallelism schemes. And maybe that's another example I can give you is With NVIDIA GPUs, one of the things that they've really innovated on and done a great job with is the NV link, interconnect between GPUs. And in fact that N V Link system is so good that If you have a large

22:46 matrix multiply that you want to Perform faster. You can actually Cut that matrix multiply in half and shard it across two or more up to eight. Let's say N mini GPUs.

22:55 and have them all work on pieces of that larger musical deploy. And have them connect to the results together at the end. Produce their results back together at the end. This is a great, great way to cut the minimum latency of an operation. Each GPU is now doing one eighth as much work, let's say. And therefore it can finish faster, but not eight times faster. It's sublinear scaling. you'll use eight times more hardware.

23:16 But you won't get eight times the speed. you might get like four to five X the speed. You're not gonna get strong scaling. This is because of communication overhead. because every GPU is gonna be a little less efficient working on a smaller tile of work than a larger tile of work. It's the only way to speed up.

23:30 If you want the minimum latency possible, you can do that. But it is not the choice I would make, for example. I would prefer to use a different parallelism scheme, like expert parallelism or pipeline parallelism. We may do interesting things to overlap and hide the communication latency.

23:44 In a way that you would have less. ability to do that for a low latency service. So is there what you think about M V Link as a technology which improves latency performance. Yes. And only latency performance. Which will segue into the next segment of what we do differently as a company. But yes, M V Link is mandatory for low latency inference. So NVIDIA is excellent at low latency inference. And I'm telling you that we don't really care that much about low latency inference. So where does that leave us? I'm not holding my breath for other companies broadly to figure out N V Link. quickly. It's challenging technology to figure out. It's hard to scale. It's hard to productionized.

24:16 If I do have some other vendors chip. And it is good at the foundational compute components. It can still do matrix multiplies really well. It just can't communicate those results across its peers quickly. Well maybe there's room for that other chip. In my stack. as a

24:31 really, really good compute per dollar option. That's what I actually optimize for in most cases is how many flops does this chip have and how much is it gonna cost me per hour to own and operate. There are other chips that definitely rank higher than NVIDIA on flops per dollar. but they may not have as much interconnect. So it's my job to figure out. What parallelism scheme am I going to use that's going to make this chip suitable for inference?

24:52 It's not gonna be tensor parallelism in video. It's basically mandatory for that. But other techniques may work well for me. So before we leave the latency part of the story, can you comment on companies like Cerebrus or others that can perform incredibly fast.

25:06 operations. I'm curious like what you think about those approach of those companies, what might happen in the future. What is your prediction for the future of very low latency focused hardware? Cebrus Grock. And a couple others that are coming out of stealth now. I think have made it

25:20 Very interesting bet. On Not just building another GPU, but actually building a different kind of accelerator that focuses on a different memory hierarchy. They want to maximize the amount of S RAM. On the chip.

25:32 And use that as very, very fast. memory for for weights and K V cache. So S RAM versus DRAM, there's two ways to make memory for a chip. One is to integrate the memory on the logic die itself. Meaning you tell T SMC I want this many megabytes of storage on my chip.

25:46 And there's a way to build that. TSMC has a standard cell library you can use, and you can just print out A bunch of cells, best Ram. The problem with S RAM is it Takes a lot of area.

25:56 On the silicon die. So if you want to build a large die like let's say the Nvidia block wall at eight hundred millimeter square. If you made that whole dias Ram. It would be in the maybe like single digit gigabytes. It feels like it's not a crazy amount of of data storage.

26:10 Compare that to if you're willing to take a different process technology entirely. So not T SMC anymore, but now Micron, SK hinex, Samsung. They build D RAM, which is a whole different way to build memory that's more focused on capacitors than transistor cells. The standard way to build S RAM is what's called the six T transistor cell.

26:27 It's a stable transistor arrangement that allows you to write a bit to it and then it holds that state in that bit, regardless of whether you keep it playing. Well, you had to play some power, but It's holding that bit without any sort of like active Management. It's static. Now dynamic RAM D RAM.

26:43 It's dynamic because what you do to write some data is you write a charge onto a capacitor. And as soon as you write that charge into that capacitor. The charge is dissipating. The dynamic part of DRAM is that you must every 50 milliseconds or so refresh. every bit you've written.

26:57 So you're constantly juggling billions of balls in the air, essentially, billions of bits. have to be managed by a memory controller, which is reading and refreshing every bit on the DRAM. Now the benefit of that is you can get much, much higher density. And it's a whole different process technology. There's a ton of different trade offs, hence where we split. the D RAM manufacturing into an entirely different company like Micron, SK Hex, and Samsung. These are the best companies in the world to do this. They build DRAM.

27:19 And if you take D RAM from those companies and you stack it into many layers and you Kind of print them or Solder them around the main logic die. That you get from NVIDIA.

27:28 You can now get hundreds of gigabytes. Like Blackwell has two hundred eighty eight gigabytes of HPM capacity. Around the logic dime. And the logic die itself maybe only has like five hundred megabytes. Mass RAM. So it's

27:40 possibly multiple orders of magnitude, three orders of magnitude difference in density. for DRAM versus S RAM. So let's go back to Cerebrus. What are they doing? They see this problem. There's not really an obvious way to increase S train density on the chip, but Thing with SRAM is because it's so physically close.

27:56 to the logic gates that actually do the computation. the arithmetic logic units are right next to the S that they're gonna pull from. the Compute units that are doing the matrix multiplies can pull data from SRM at Mind boggling speeds.

28:10 streamers quits petabytes per second, twenty one petabytes per second for their waiver scale engine three. Compare that to HBM on an NBD Blackwell is uh Ten terabyt per second or so. in that range. So once again, many orders of magnitude difference. More capacity.

28:23 But Proportionally less bandwidth, essentially. What Cerebers does is they say that they were gonna take as many of these dies as we can. We're not gonna limit ourselves to the eight hundred millimeter retical limit the TSM C eight hundred square millimeter limit that TSM C imposes on us. We're gonna take the entire wafer and have

28:39 every die connect to every other die over scribe lines and we're just going to try to get as much S RAM as we can on the whole wafer and we can get to like let's say fifty gigabytes. of S RAM per waiver. And then we're gonna stack many wafers together in a pipeline or similar. And now we can have

28:54 Up to a terabyte of memory. Very, very fast memory. You do all that work. just to get to the ability to read data from S RAM at twenty one petabytes per second per wafer.

29:04 Therefore, you can now serve these language models at extremely high tokens per second. Because you can move the entire parameter count of a large model like Kimmy. You can move all that data. in and off the logic cores in about a millisecond or something like that. There you go.

29:18 you have a path to a thousand tokens per second. So what is your prediction for like that segment of the market? What happens to them is Some hybrid. We had to pair this Rubis chip where it's very strong. It's very, very good at fast access to memory.

29:33 with something that has more capacity for memory. It's true that you can take a one trillion parameter model like Kimmy. And fit it on. a large number of cerebrus diet wafers. But you can't

29:43 do something about the KV cache very easily. The KV cache is something that grows as people use the model more. And that is always dynamic. You don't even know how much KV cache you're gonna need. It depends on how many users you have and how many users you want to serve. Can you explain KV cache? K B cash. When you ever use a language model

30:00 Every token you send through the language model. stays in the context window of the language model for as long as you're having a conversation. We talk for a hundred thousand tokens. The hundred thousand than one token is still in the conversation. behind us and the model is referencing all the past conversation history in order to make better predictions about what the next thing we're gonna say is.

30:19 That K V Cash. It's a bunch of memory. You have to store a representation for every token that we send through the language model. And it frequently gets to be larger than the weights of the model themselves.

30:30 You have this. crystallized knowledge in the model weights. Yeah, the dynamic knowledge of the exact conversation we're having in the K B cache is the way I like to think about. And this is why sometimes people would observe deep in a conversation things start to degrade because there's some sort of technical problem. So the KB cache is quite interesting in that regard. The KB cache is an exact representation of everything that came before. We store all the information.

30:51 That We've seen in the conversation. However During training, the model Did not

30:56 get trained primarily on very long context conversations. It got trained primarily on let's say 8,000 token conversations or 16,000 token conversations. So if you take the model to 200,000 tokens, there was some training that happened at that context length, but it's not the model's like core strength. And so there's always been a challenge for the Frontier Labs to figure out. How do we make the model exactly as intelligent at ten thousand tokens as we expect it to be at two hundred thousand tokens? And it's gonna be a perennial battle for us. We've had one million context windows as a concept for years now. Anthropic was, I think, the first to hit that one million context window lent lent. I still, you know, use Slash compact in my quad code. well before one million context length. I don't think it's actually great to hit the full length.

31:34 These extremely fast, extremely low latency approaches ultimately are limited by this factor. Yes. You can do whatever you want for the weights. It's very possible to have much unbeatable performance on weight storage. However the K V cache is gonna be a big thorn on your side. So three years from now, five years from now, what role do you think these kinds of chips play? Like what sort of market share do they have in the heterogeneous chip market? Like Cerebrus and Grok and maybe a couple of others. You should think of them as accelerators. What they are really good at is being used in conjunction with a

32:04 more traditional GPU like device. that critically has this off chip memory built in. You want off chip memory for capacity. And on ship memory for speed. we want to hybridise these two things. So if you take transformers in the limit. You take a transformer to a million context length, what ends up happening is

32:20 You have this compute bound stage, which is the actual matrix multiplies for the what we call the MLP. which is where most of the model's knowledge, world knowledge is encoded. And then you have the attention layer, which is where we're dynamically adapting to the current conversation.

32:33 Attention. in the limit is usually bound. The MLP in the limit is compute bound at large enough batch size. I would say the original sin of Transformers is that you've taken this extremely fundamentally memory bound layer. And

32:46 juxtaposed it right next to a compute bound layer. It is very difficult to have a single chip that is gonna both compute operations and memory operations. The GP is quite balanced in this regard. But you had to choose one or the other.

32:58 Cerebrus has a very fast memory access for Something like a matrix will apply. And it's really good to host the MLP, the weights, essentially, on this Rubis chip. But the GPU has the capacity to scale to really long context lengths. You would like to put the

33:13 attention possibly on the GPU and the MLP on the Cerebrus chip. And I believe this is what's happening with NVIDIA and Grok. Can you refer a minute just on transformers? For people that again aren't deeply familiar with what this innovation was in twenty seventeen, what its strengths and weaknesses are and whether or not you think it will remain the dominant architecture Or a dominant architecture.

33:33 For the future of AI. What it did. It allowed us to learn on unsupervised data really effectively. Because transformers, what they're all about at the end of the day is taking any sequence, any arbitrary sequence of data. and trying to find patterns in that data. And they critically the attention operation, which is the headline component of transformers.

33:52 It allows the model to dynamically adapt. to what it thinks the most relevant. component of the sequence. Every step you take through a transformer. You are essentially like rewaiting.

34:04 the input that you looked at before. And figuring out which is most relevant for your next prediction. It's extremely amenable to learning arbitrary sequence data. And the most Interesting sequences of data that we produce on a regular basis is language.

34:18 And that's how we got to dominance in the language regime. To zoom out even further, I think what Mate Transformers really did well. Is that they scaled. Transformers make no such human prior. Transformers just say, Well There's gonna be a pattern in the sequence of data, and if there is a pattern, I'm gonna find it.

34:32 I'm going to throw more and more parameters at this problem until it works. Transformers benefit from a lot of the computer vision work too. For example, one of the challenges in computer vision was We had a hard time going from hundreds of thousands of parameters, which you get for linear models like support vector machines or other legacy machine learning models. Those had thousands of parameters.

34:50 Then we got to deep learning and got to tens of millions of parameters with computer vision. The biggest models were 150 million parameters was a huge model for computer vision. And now we routinely talk about trillions of parameters. And transformers are the link to go from millions to trillions of parameters. So if I think about the important units of scaling being data and compute. Does it stand a reason then that you think transformers will just stick around because that's the thing that we're good at?

35:12 Getting more of those two things. Transformers are such great sponges. You increase the compute available to Transformer by ten X and you'll get Some log improvement somewhere. And so far the scaling laws really work. They're really quite beautiful.

35:24 And to the point about What do transformers do really well? They extend to almost any data set you can throw at them. They're extremely powerful general learners. And I think what's especially useful about transformers over other techniques that we've tried to replace attention. Is Transformers. Represent

35:40 any pairwise relationship that you want. Any token in the sequence? Can't attend to any other token in the sequence. So if there's any relationship that's in the sequence at all, you're gonna find it with the transformer. Now it may be the case that you don't need

35:51 all to all modeling. You don't need every token to look at every other token. But If you need to, transformers give you that option. And until we know a better way to prune that space down. a better way to kind of have information modelling.

36:06 Be more selective. Attention is a very, very good operation. This is uh another trick that we learned in the computer vision days. One of the old Carpathi sayings is that If you have a new data set that you want to train a model for. Your first goal should be to over parameterize the model.

36:20 and try to overfit the data that you have to prove that there is a relationship that you can model or memorize, that your learning algorithm works. Then you can instill knowledge into the model. Once you can overfit, then you can compress. And the compression is how you get generalization. You don't want to actually memorize the data that you have in front of you. You want to generalize and therefore Once you

36:37 overfit the data set, then you can kind of work backwards and try to find the general patterns that fit into the smallest parameter count possible. What's your prediction for the future of data and riff on the importance of data in this whole story? I like the phrase that internet was a one time subsidy on data. We got it for free. It's extremely high quality. About thirty trillion tokens of high quality tax. three hundred trillion tokens if you

36:58 Take a wider view on what qualifies as good text. And we've basically looked at it all already. Models have seen the entire internet many times over at this point. And there is not a whole lot more to be done on human data from the internet.

37:10 The next phase of data in my mind is model self improvement through R L environ gems, basically. Now in fact we don't even benefit from getting more random user interactions with AI. It used to be that the new type of data that we cared about a lot was the interaction data from people using ChatGPT and giving ChatGPT signals on what they liked and didn't like.

37:30 I like the argument now that the median model that we serve is so much more advanced. Than like a random human giving feedback that The signal you get from random human preference. unconditioned human preference is not actually worth anything anymore. You want expert human preference at this point. The model has outgrown every day. Yeah, every day. Exactly. So the future of data to me is giving the model a hard, verifiable task.

37:52 And letting it. run in this gym where it's kind of isolated and it just has a problem and it can make progress on it, get measurement of whether it made progress on that problem or not. You can imagine coding problems are in this category. Math problems are also in this category. more and more we can just give the agent a computer, essentially, and have it act like it's a human worker. And

38:12 give it feedback on whether it's making progress towards the target outcome. that environment becomes the data. I think this is not a super different tree to take, but it's been really, really productive from what I've seen so far. And you think that just goes on for a really long period of time, or is that another like if I think about the internet as this one big block, like this is Another big block that will have its stay in the sun and We'll kinda get it all. And then we'll have to move on to something else. I think it's actually more profound than that. Basically the idea is that if you want artificial general intelligence, the best way to get there is to

38:41 Just keep stacking specialized intelligences until you have no more gaps to fill. The only thing you need to make sure you do to make this work is you must make sure that your task is verifiable. You need to give the model a self.

38:54 If you have that, you have the recipe for Self improvement on any task you like. I think you've seen this held up by the way Frontier Labs spend. They used to spend that much on data, now they spend a lot more on R environments. These environments absolutely capture that relationship of

39:08 Recursive self improvement on a verifiable task. Coming back to Your initial task of making existing hardware. More efficient.

39:16 By being more in control of what's going on at the hardware level through software. Keep going on. what you've done so far and what you wanna do. And then we're gonna jump to hardware and then jump to energy finally. I mentioned kernels. It's surprising. People think kernels are done. There are great people like Tridow who write excellent kernels and they form the bedrock of all of our modern deep learning. is built on flash tension.

39:38 Modern transformers are built on flash attention. But if you deviate from the happy path at all, if there's a new model that comes out that has a slightly different way to embed positional information. Like the change of the rope. system. Suddenly the kernel that we had is not suitable for this new model. And we may have to make a patch to this kernel.

39:55 I wouldn't say we're in the phase where we have to invent new kernels from scratch, but having the ability to quickly modify existing GPU kernel, sorry, a kernel by the way is uh general term for any program you run on the GPU. Historically kernels tend to be Put into a library where

40:09 Every kernel has a very, very scoped purpose. Typically you have a kernel for A matrix multiply. You have another kernel for even something as simple as addition. You want to add two tensors together, that's another kernel. And then increasingly we've started to fuse those kernels together.

40:23 So if I do imageable apply. And then I wanna add it to another matrix that I've also multiplied. maybe those two become one kernel and I just fuse the operations where sort of writing the data out to DRAM and then reading it back in just to do the addition. Maybe I can just do this.

40:37 Easily. Why are humans still doing this? Like it seems like the sort of thing that AIs would be exceptionally good engineering, more efficient kernels. Maybe that's where we're going and we're just not quite there yet. But if we aren't there yet, is that where we're going? If we're not there yet, why humans still doing this why why is treat out so well known, you know, it's a name I know. I don't want to speak for a tree, but what he taught me was You shouldn't write colonels. By hand anymore, necessarily.

41:00 I like to say we write kernels on the whiteboard. We go to the whiteboard, we describe what we think the machine should be doing. Then we succinctly describe that in in natural language. Two. A model.

41:09 And then the model is able to do the execution of Here is my input and output. Here is the strategy of how we want to dispatch this work onto the GPU. I'm gonna go implement. So we're doing the conceptual design. Exactly. I'm not sure exactly why models are not superb at doing this. I don't think this is like our moat or anything like that. I'm sure in six months' time we'll have much better models on kernel engineering. And I'm sure the labs would tell you that they already do a lot of their kernel engineering.

41:32 In a fully automated way. So software as an edge. I think about software as maximally near speed of light efficient usage of an underlying piece of hardware is going to trend towards not being an advantage for a company like yours over time. That's right. The rising tide of something like Mythos or GBT five point six Sol. That lifts all boats. It really does. I actually don't think there's a point in specializing.

41:53 To say we work on making the model better for just kernel engineering. I think that's actually not the most meaningful subset of coding in general. kernel engineering in particular, maybe there's some privilege information you inject into the prompt. That's like a useful way to steer the model to be better at writing kernels, but broadly speaking, yes, we're all downstream of the frontier in terms of this capability.

42:14 I always love this from the history of energy, there's always this pendulum between the raw source, let's say coal. And then if there's a certain amount of energy available inside of a hunk of coal. What percent of it we can harness and use. A big part of the history of energy was getting that number from ten percent to ninety five percent or whatever.

42:32 If I just think about it as Blackwell or something, and Blackwell is the piece of coal. What percent do you think we're at? How efficiently can we use an existing piece today? There's a lot of different ways to analyze that. in some level we are really efficient at optimizing the Performance.

42:48 When the GPU's doing the thing that it's most happy doing, which is a large Dimension. Magic will apply. That operation runs at seventy, eighty percent of peak utilization and it's limited not by software, but by power. The way NVIDIA quotes peak flops is

43:01 Well optimistic, you never hit that because of power throttling. 'Cause it'd have thermals. In practice, you don't spend the majority of your time in a transformer in that happy path where you're doing a large batch matrix multiply. Our job is to basically

43:14 build the engine around the chip such that we are feeding the GPU these large batches of work at all times. One of the most profound transitions we've had in the GPU world in the last year has been This moving of you don't program one GP at a time anymore. you should think about the whole rack. And maybe you should think about the whole cluster, the entire data center at a time. With NVIDIA, they've started shipping not just a single GPU or a single motherboard, but actually the whole rack system is something that they prescribe. They call it NVL seventy two.

43:39 their latest chip, the Grace Blockwell three hundred. that ships as a rack of 72 units, and it's an open race to figure out who can program the whole rack scale computer as efficiently as possible. And my belief is that that shape of compute is the future of efficiency and speed. In fact, the video does a great job of If you want the lowest possible latency, you should be using that chip. And if you want the highest possible throughput, you should probably also be using that chip. As of right now.

44:02 And it's all comes down to like this is a very new paradigm of programming. One of the things you hear is that the market for the best chips, Blackwells let's say, is like a drug market or something right now. There's all sorts of fascinating things happening to get as many of them as possible'cause everyone's so short.

44:17 It'll be to react to that analogy, is that what it feels like? Then also to talk about what the market is like for like not the bleeding edge chips. If I I'm willing to accept a slightly or moderately inferior chip. What's that market like? Let us into that world. Maybe it has a long term

44:33 view on all their chips. They see this immense demand for the black hole chips. They can do what other suppliers have done in the past, which is crank prices. meet the market supplying demanders will correct. They'll intersect at some point. And everyone will be technically happier. And the VDSCs if they just let the most deep pockets buy all the chips that Maybe

44:50 hurts them the long term if that customer ends up accruing a lot of more power. They understand the compute is power today. They're quite strategic about how they Allocate compute. That's the first thought. The second thought is that relationships matter a lot. Nobody wants to

45:04 have a huge order of a chip rental come in from this new startup that says oh yeah, I'm gonna rent ten thousand block wells for three years or five years. The startup has only been operating for months, typically. Who knows whether they're good for the money. The way you convince someone to give you access to compute is quite challenging these days and requires some pretty

45:22 Great relationships or Just incredible financial backing. To make this happen. on the NVIDIA side. And it's all because the scarcity is so high and demand is just off the charts.

45:32 No, for other chips. I wouldn't even call them inferior. I like to say there's no bad chips, there's only bad pricing. I will make any chip work at the right price. That's one of the ethoses of the company. Let's talk about AMD. A M D

45:43 I think great chips overall. the challenge is that people don't understand how to program them. Very well. I've been talking to you about how we have such a great kernel team. We're so serious about squeezing the performance out of the hardware. NVIDIA is pretty good at doing that for their own ships, frankly.

45:56 There's some alpha that we can squeeze out, but There's a lot more to be done other chips because the vendor does a little bit less work than Nvidia does to make the best kernels. Out of the box. Or even better. there is alpha and just other people have this perception that AMD is not as good as NVIDIA. That's music to my ears. I'm very happy for them to sleep on this chip and for me to buy as much as I can.

46:15 Now I think that that's not actually super true anymore. I think A and D is actually Somewhat popular amongst Some large buyers. I think publicly meta and opening I have. bought a ton of AMD chips, we're increasingly seeing that all the AMD supply is also being allocated.

46:28 But there's a long tail of other companies that are popping up. Net new companies are great. such as hatched or sombinova or A D matrix. All these companies are popping up. I think the main challenge for them is scale. Can they actually get enough waiver allocation for T S M C to

46:42 I'm about. Chips to making it to the market. If there's a new chip on the market. I'd like to know about it as quickly as possible and evaluate whether we can buy. A good fraction of that supply.

46:51 And so it's fundamentally an arbitrage for you. Like if you can be much better at eking out performance from chips that have received less attention, you can then resell that. at a margin and it can be a great business. Exactly. And I think that it's not the case that everyone else has a skill issue that they can't. Make these ships work as well. I think we're quite competent in this. I think we're

47:10 Probably one of the best teams in the world to use. multiple silicon architectures and be quite aggressive in chasing down performance in unlikely places. But yeah, I think it's the speed at which we're willing to kind of build our stack around a new chip. We don't have a huge amount of incumbency around Well, our data center providers are only stuck with this class of chip and it's gonna be a huge pain for us.

47:29 to deploy these net new chips. we have some very creative data center partners who are willing to move very quickly. And there's a new class of those that we can talk about. Most importantly, we don't shy away from the challenge. That's frankly. A big part of this is just saying

47:41 Yes, we love TPUs, we're gonna make TPUs work. Yes, we love training, we're gonna make training work. And if it doesn't work, easily we're gonna find a way to fit it in with a heterogeneous serving system. it will have a place. Every chip has a comparative advantage. We had to find that advantage and then squeeze it in that direction. Just as an interlude, before we get to Hardware data centers, energy, et cetera.

48:01 Which would be really fun part of the conversation. I'd love you to talk about your perception of The investor classes worry. Like you look at memory stocks. Or my current favorite is you look at the chart that plots the percent of the S P five hundred that's semiconductors.

48:16 Historically it was like two, three, four percent. Now it's nineteen, twenty, twenty one percent. And it just sort of looks like if you're a student of market history, you get all these things through time that are sort of reach some crazy near term peak and then collapse back to long term norms. That has all investors worried. A lot of people made a lot of money in Micron and SK Heinex and companies like this.

48:35 But everyone feels like long term like compute's a commodity, it will not represent a quarter or fifth of the entire market capitalization of the world. They're scared and that's the setup. Everyone acknowledges that there's a huge shortage. Everyone sort of feels like, ah, we'll figure it out and these things will revert back down to their normal place in capital markets. I'm curious what you think about That narrative.

48:56 I'm less of a student of history as more of a member of history. I was born in nineteen ninety seven and My mom worked at Intel in the run up to the year two thousand and the dot com boom and crash. I remember the time where Cisco was the most valuable company in the world and in Intel was close behind. And mostly draw parallels to that period of history from twenty five years ago to today. And I think the main difference is that

49:17 A lot of the investment in networking equipment historically was speculative. We anticipated this future demand for users that never came. I think what's interesting about Token consumption or AI consumption broadly is that It's no longer speculative. People buy tokens because they're immediately valuable to them. You don't

49:34 hoard tokens, you use them immediately. also even different from what we had two years ago, where there was a supply crunch for hopper generation chips in twenty twenty three, twenty twenty four. In that period, it was all training oriented spend, and training is inherently speculative. Now it's Everyone's instructed in caps on how much you can spend on cloud code.

49:52 It's a very, very different world. to be talking about inference spend and predicting inference spend to go up. I do think inference spent monotonically increases. There's no Speculation on inference.

50:02 Coming back now to your take on hardware. The unit level is interesting to me. Talk about chips, talk about racks. Talking about clusters.

50:11 I'd love to talk about data centers. You said you've had some interesting partners doing some cool things. Talk us about the present and future. of data centers as you see it. Because this seems like Obviously a critical thing for being able to serve all this inference is lots of innovation in this part of the world and obviously you're focused on it.

50:27 I think one of the themes in our conversation has come back to what is training versus inference. Like what is the difference between these two categories and you know what was different about two years ago being training oriented and today being inference oriented. And I think The most conservative players in the entire AS deck have got to be the infra players, whether that's data centers or even more conservatives, TSM C that chip infra people. Data centers They were built.

50:47 For training. Training is a superset workload over inference. You can make any training cluster work for inference, but maybe not vice versa. The difference there is networking. How much do you Invest in bandwidth between ships. And how large of a cluster do you need?

51:00 There's actually a diseconomy of scale to data centers in some way. It's way more expensive and difficult to build a hundred thousand GPUs in one data center than it is to build ten thousand, then it is to build one thousand. Now we just talk about how many megawatts or gigawatts do you have? And basically there's no way to build a gigawatt data center in the United States.

51:17 Easily anymore. Even hundred megawatts is increasingly hard. It's Basically impossible unless you're a very special set of customers. Ten megawatts is probably on the edge of what's possible today. And one megawatt, I would argue, is plentiful.

51:30 So there's this incredible lore on the market where You can find lots of aggregate power. But it will not be concentrated. And that was not interesting to anyone who's building training. data centers because you just assume all be in one spot for and no one wants to deal with cross data center training.

51:45 the market has some lag in it. I think that the market still assumes that we had to go shake down those hundred megawatt and ten megawatt data centers wherever we can find them. It's still the attitude I hear from a lot of data center developers. But increasingly we're seeing a few New thinkers. realize that inference is going to be suitable for these distributed one megawatt data centers.

52:04 We're quite in agreement with that. And we are very happy to buy small pools of compute across the United States. And use that as our inference fleet. Give us a sense of literal physical size of one megawatt versus ten. Yeah. Well so this guy

52:18 really wonky with the advent of liquid cooling. Now you can pack insane levels of power density into a single physical rack. A megawatt of compute. you'd imagine this like massive data hall, like a huge warehouse, basically. And now you can actually pack that into around like eight racks for the compute. Each rack is about the size of a refrigerator.

52:36 Can just imagine eight of them lined up. That's a mega. Your view would be that the future that you want to help build is a whole bunch of different ships that can be Used together. That you can buy your buyer to eke out the most per chip.

52:50 And that those chips can then be coupled in very small data centers. To just do inference. Those two steps of a whole bunch of random compute, some of which is cheaper than it should be. Your ability to eat more out of it and then small

53:04 units of expression in the data center. equals way cheaper intelligence. I certainly think so. Yes, there's a lot of ways to access cheaper flops if you're able to be creative with what you take and so one of the ways that I describe what we do. Is

53:17 We will buy Any chip. Anywhere in the world. for any duration of time. That is a level of flexibility and liquidity that

53:23 no one else has right now. We're very aggressive about putting our money where our mouth is and we will take any capacity. and find a way to make it work in our fleet. And that is a big part of our advantage today. Long term, we have to create more of that advantage by investing in.

53:37 these data centers that other people are gonna be skeptical of because What's gonna happen when you set this army of a thousand small data centers versus the one big gigabot data center. few things. You're not gonna have power redundancy quite often. You're not gonna have back up diesel generators on site. Those are all very expensive. We cut all that overhead. We're not even gonna have we're done a networking in a lot of cases. We're gonna put these in

53:55 facilities where we have good access to power, a single source of power. And we're gonna entrench one line of fiber to these data centers. But we're not gonna have three lines of fiber with redundancy and failover and SLAs. Just gonna go down sometimes. I won't be surprised if some of them get down to like ninety five percent uptime. Which is bad. Very bad. That's fatal, atrocious for anyone else that's couldn't survive in a a big lesson. You'd have basically zero buyers for a data center that has ninety five percent uptime.

54:19 I'm that first buyer. I will buy the representative time. And the reason for that is because of this background engine thing. There's things running in the background. You don't care. Partially, it's actually two things. One is that we have a really robust control plane that is gonna be fine handling any single failure in any single data center as long as it's not correlated with other data centers and I can just move the workload somewhere else. I'm cool with that.

54:39 The failures happen. At some rate and I am Basically linearly happy with a data center that's 95% of time versus 98% versus 99%. It's just linearly. Good or bad for me. You do need that async piece that I mentioned.

54:53 We serve these long horizon agents because what happens when a request fails is that I'm gonna have to go find a new GPU to put that request on. And it means that for that single turn of the agent's work. it's working for an hour, but then hits a roadblock because its GPU got pulled away. In that moment in time. that agent is gonna experience maybe like an extra minute or two or three.

55:12 Maybe even ten of latency. My argument is that it my customers don't care. Because their agent was running for hours. They're sleeping. Doesn't matter. Doesn't matter if like a single turn occasionally becomes a little bit longer. We tell our customers. look, our average throughput is gonna be very competitive.

55:27 but our P ninety nine, our ninety nine percentile latency. It's not gonna be controlled. It cannot be. And in return, I'll give you unbeatable economics. I think that's the right fit for background agents. Talk about power as a category.

55:38 What the innovation that you're seeing is, where do you think it goes from here? What are you seeing that's interesting, innovative, where do you think this goes? Okay, so I said I want ninety five percent uptime on my data centers. Could I even take eighty percent of time at the right price? Probably. And what does that mean? Well I'm a son of California. I love solar and wind.

55:55 I think solarmen power is way undertapped in the United States and the challenge has always been this intermittency. You would even consider solar and wind unsuitable for data centers because you have a persistent base load and an intermittent power source. What are you gonna do? Well I think we're actually Not that far from solving that problem. I am totally capable of tolerating a outage.

56:14 From a data center that's measured in even days or weeks, which is like the worst case nightmare scenario for a data center is that we're gonna have a long term outage because The wind doesn't blow and the clouds are in the sky. Fog is hanging over the valley for some time. That's the worst case scenario. It's in fact highly predictable.

56:28 And I can just call in capacity and some other place of the world whenever that happens. I'll just model the weather and figure out when my data center is gonna be offline. Movement work with somewhere else. And it's fine.

56:38 The trick is that it's gonna give me better access to power that no one else is gonna touch because it is so annoying. to deal with that kind of outage. And if my chips are cheap enough, they're probably not gonna be NVIDIA racks. But if my chips are cheap enough. I don't mind the capital cost of having idle ships.

56:53 I've heard you describe this entire system as like a scavenger strategy. That's right. Unpack that analogy a little bit. Well, first we scavenge Chips. And then we scavenge. Power for those chips. The idea is in both cases, I do not want to be bidding against Anthropic or Up and AI for compute capacity. I'm not gonna win against them and I don't want to. I want to be more creative and use the supply that they don't find legible today.

57:13 And over time I am mass enough. Aggregate. supply. I'm never gonna get concentrated supply. I will only get aggregate supply. And over time I build my aggregate factory that is unbeatable in economics.

57:24 We are building a factory. We're trying to build the best steel factory in the world. But it will come through mini mills, not through Large. Monolithic steel plants. And if I imagine the different versions of this, how vertically integrated you can be.

57:37 The extreme would be you own everything. So that it's a very capital intensive business. You own the power source, you build the data centers. You design your own chips. You control the software that eats the most out of those chips. And you sell the end finish token to your user. Your user is me.

57:55 And you just own the whole stack. But you can imagine many other permutations of the business. Where draw the line anywhere. You could be incredibly capital light, own nothing, and just be like the coordination plane. across all this stuff, the virtual scavenger. How do you think about

58:10 that question of which type of these businesses to be. You know, there's actually two parts of me to receive that question. One is the CEO of a company that needs to work every day and grow sustainable and as quickly as it possibly can. Yeah, there's a founder. And the founder is much more imaginative and loves this stuff. The founder in me wants to do everything. This is my entire life. I spent my entire life thinking about chips, power, energy. All I care about is this stuff. So of course I want to be maximally ambitious. I don't want to stop ever. I will never stop until I have

58:37 build the most efficient system from soup to nuts. You're doing real life factorial, basically. Very much so intelligence. Very much so. So that's like the emotional from the hard answer. On the CEO side, I think. We have to be more pragmatic. The capital we're looking at for owning everything is insane.

58:51 Software has high leverage, so we have to start with software, but ultimately Do we own power generation or can we get great power purchase agreements with Utilities. I'm more inclined to pursue Letting other people specialize in

59:03 the things that they're historically good at and then see if we can get to this scale. I think of it as like I want to get to the scale where I earn the right to take this under our wing. I absolutely think that there is Efficiency is to be gained everywhere in the stack. If you can break the assumption that people I would be buying from, they made assumptions about who their customers would be.

59:19 And I maybe break those assumptions. It's a pretty optimistic view. I think it's only possible because we're actually trying to underwrite the largest market for compute in the history of computing. We are actually gonna build So I mean millions, trillions of dollars of investment into inference.

59:36 Because of that focus. It makes sense to build a lot of things that are custom for inference. And it's my job to seek all the Places where that's possible and then as as it become obvious to me and my partners.

59:47 I look at my partners to build custom things for me and if they can't do it for me, I will do it myself. If you had to just zoom out on this entire system, yes software, hardware, energy, et cetera. And stack rank the places that you think that we are the most inefficient today. at producing useful Intelligent tokens. What does that list look like?

1:00:06 I think compute scaling is actually very efficient. You give me more flops and I will use more flops. I would say we're actually fairly judicious already with our use of flops. If you look at a modern M O E model. There are very few models that are more than ten percent dense, meaning ten percent of the possible number of experts you can activate are activated.

1:00:23 And I think the frontier models are close to like one percent. fairly sparse already. I don't think that we're wasting too much on the MOE side. People have been working with MOEs for quite some time. They're pretty good at squeezing MOEs. Where we are not good, his attention.

1:00:36 And it's used in memory. Specifically. The K V cash is quite uncompressed right now. I think if you look at the entropy in a K V cache, it's not earning its keep, like we're storing many kilobytes.

1:00:47 of data in the KB cache per token. And that's probably off by an order of magnitude or two. I don't know what the Frontier Labs do, but Deep Seek certainly publishes really interesting work. To compress that. Further and further. And they're making good progress. And I think the fact that they're able to make an order of magnitude in progress here.

1:01:03 Every year or so. signals that there's a lot more room. to go. If you zoom out further, I think that we actually don't marshal our compute effectively at all. We have all this compute in the world. The NVIDIA is pumping out five million black wall chips this year. Where are they all going? Are they all being used at all all the time? I certainly doubt it. I think that at some level we just need better orchestration of compute across the world. This is very difficult to do because a lot of the compute disappears into private pool as a compute that will never see the light of day.

1:01:26 And those GPUs sit. Very sadly idle. It pains me physically to see that those GPUs are just silicon and power. Went into that and it's just sitting idle. And I want to fix that.

1:01:36 how we organize and orchestrate the world's compute as a shared resource. And pack it more efficiently. I would estimate that. We all make fun of X AI for having, you know, some challenges with total flop utilization on its clusters. The reality for the rest of the world is it's far worse.

1:01:50 a ton of GPUs just sit in warehouses or sit in private pools allocated to a specific customer. You're attacking the efficiency of that very directly. That's way more effective, yeah. Fabs. What do you think is the future of fabs themselves?

1:02:06 I think everyone is wondering. Will the memory companies, will TSM C, will Intel And others. How will they expand capacity, basically? Well, we do it here in the US.

1:02:16 Rif on fabrication of chips themselves. Like if we could just snap our fingers and have a hundred times the chips In the stock today. We probably have way cheaper tokens. That seems like an important part of the universe to hear your view on.

1:02:30 Everything grows in balance with each other. If we snap our fingers double all those things. you might fix a T Sim T bottleneck, you're just gonna run into another bottleneck. you make twenty percent more chips than you have another bottleneck immediately. I will say though It is

1:02:42 Interesting, what they consider to be a must deliver, like an invariant that their customers me are always gonna want. Versus what I think as like a more fluid relationship. I think that if the Fab exposes more of their trade offs To me.

1:02:56 I'm able to make more intelligent decisions about what I can do. One of the most Interesting examples here is that Any fab has a lot of spread in their worst chip that comes out of the production line and the best chip that comes out of the production line. There's a lot of variance in how chips are made.

1:03:10 The question is. If you have a company like TSM C, they work very, very hard. What we call these process corners. We wanna keep

1:03:17 The worst chip. as close in characterization to the best chip. And then they go to great lengths to make that possible. But that means that they are adding a lot of controls in the process that maybe I don't need.

1:03:27 maybe I'm actually willing to find a place for that worst chip. you don't need to tighten the process control as much. Which takes more time and cost. Maybe I'm willing to take a lot more rejects. And I think for us it's like a more holistic optimization around

1:03:40 cost of the dyes, supply of the dyes, and then the cost of power and places we can put them. My whole goal is to so dramatically expand the supply of power. across the United States that I have a home. For a lot of ships that otherwise. would not have earned their place in a data center.

1:03:55 Can we talk about how you design the system of your own business? What lessons have you learned? You talked about some of the interesting NVIDIA lessons. Bring me into the culture. and how you structure a team and a business where this is the North Star. There's a lot of

1:04:08 In the limit thinking, we don't worry about the immediate Nature of When we start working on a model. The efficiency's not going to be very good. we think about like where we could end up in like a month.

1:04:19 Six months or a year's time. We don't accept the state of the Machines we work on is fixed. Even Something like the black wall chip. If we think that there's

1:04:27 some bottleneck that is holding us back from achieving this performance. It's very important to me that we understand and characterize that very well. And write it down so we can both A tell NVIDIA about it, or friends. And also to basically keep this in mind for future chips that we buy. We want to

1:04:42 Learn things that are invariant for us or the company. Long term. and kind of fold that into future decisions that we make. We're very collaborative. I think one of the most important traits that we look for are people who either

1:04:54 who are both good students and great teachers. A lot of our people on the team were TAs in college and loved the experience of sharing knowledge in this way. We do whiteboard sessions all the time. the collegial environment where everyone has something to teach and something to learn is extremely important for us. What are the attributes of people that you would want to hire that you think will be resilient to the work environment three years from now when more stuff is handled by machines. Curiosity. A hundred percent curiosity.

1:05:20 The one thing I cannot teach is love for performance. Love for digging into every microsecond the machine is working and understanding what's happening. on the machine at that time. That to me is the most important trait for a performance engineer.

1:05:33 It's what I look for. I don't look for lots of AI experience. I don't look for Cuda experience at all. That's actually a huge red herring. I mean CUDA as a concept or GPs as a concept have evolved so much in the last five years. There's no point asking for 10 years of experience. I want to teach that. But I cannot teach.

1:05:47 The love for performance engineering. That is what I seek. Can you give your assessment of the major labs One by one. But also then the relationship of closed source as a category to open source.

1:05:59 What you think is happening and will happen. In a line, I would say the labs pay an immense premium. to be three to six months ahead. Of everything else. I think that's probably still worth it.

1:06:08 I think it makes perfect sense for it. And through it to do what they do. There's a sensitive topic around distillation, which I think is a very core piece of the relationship between closed and open frontier. I'd like to offer an alternative view on that, which is There is the sense that distillation is theft.

1:06:22 that you are taking something from the frontier models when you distill on their outputs. And in fact. Even if that's not your intent, even if you don't ever try to scrape data from anthropic one thing I'll offer is that an increasingly large percentage of the artifacts we put out on the internet are AI generated.

1:06:38 You just look at GitHub alone. What percentage of repos created in the last year do we think were created by Cloud Code? Do we consider that to be distillation? Because that's probably all we need. I would not be surprised if you could train a fable class model only on the outputs. of code you consider good on GitHub that's open source. And certainly if we take the position that users own the outputs of their

1:06:57 Interaction with AI. And they choose to put that up on GitHub, which a lot of them do. We're gonna have latent distillation. For a long time. It seems fundamentally impossible for me. I don't think it's fundamentally possible to prevent the diffusion of information or model capabilities. It will happen. The question is just how fast.

1:07:12 And so then the question becomes Do scaling and improvement laws hold forever? For a really long period of time. And if they do then there's value to being three or six months ahead and that will just last as long as it lasts. And they can charge a huge premium for those tokens relative to a very cheap open source token.

1:07:29 Is that the right way to think about it? I think it's possible. I don't know that the premium for being three to six months ahead is gonna last that long. I mean, if you look at like enterprise deployments. They don't move at three to six month speed. A lot of enterprises are probably still on like four six, opus four six, or opus four seven. They don't adopt the bleeding edge rapidly. There's a lot of questions that people have around rolling out any change at all.

1:07:51 We're just so early in scratching the surface that I don't think there's any way to call a winner in this race. Certainly I don't even think this is a race that can be decided. Ever. It's a continual process.

1:08:00 Fundamentally, I don't think open source ever goes away. if there's a vacuum because you know, one liter steps out. A new leader will step in. There's too much incentive and too much tailwinds too. It gets easier every day to train a frontier class model. So your hope of what the future looks like is what balance between closed and open.

1:08:16 What balance between model companies doing everything because they have the advantage of owning the stack or whatever Anthropic can do that. It's like the new Google will just do that or something. What do you hope the future Looks like I want Abundant tokens and diverse harnesses. I want everyone to build their own harness.

1:08:32 Every company, every user. Even Make the engineers your own. Right. not that far away from that level of customization capability. I want people to own their intelligence.

1:08:41 And I want that intelligence to be customized. Probably not through weight fine tuning, but probably through more in context learning. That's a more technical detail, but No. underlying input to this abundance future.

1:08:52 is about cheap tokens. My job is to make the tokens as cheap as humanly possible. I will achieve that. And I will do it through every layer in the stack available to me. I love the supply side levers. I will use every chip, I'll use every source of power, and I will use every piece of land in the United States that's suitable for this. In return, people will have the incentive.

1:09:08 Two Explore what it's like to have abundant intelligence. We still treat the agent as a person that is expensive to consult and you should ask them when you have a hard question. That's not the way to think about intelligence. It's incredible that the machine can think. And we should try to get that into as many hands as as many people as possible.

1:09:25 You sit in such a unique seat and you have such a unique perspective on like what you're trying to do to make this feature a reality. What do you think are your most divergent views of the world versus your friends who are really well informed and interested in this stuff. What ideas of yours make your friends look at you like you have three heads? Most of the ideas on chips, I would say, you know, when I talk about Building custom chips and they asked me, Oh, so what's different? It's about sidestepping the HBM shortage and focusing on

1:09:50 more extreme offload to other forms of memory, such as flash. I'm quite passionate about that idea. everyone on my team knows that I keep banging the drum around like what would we have to change about the model architecture to make offloading KB cache to flash work. at a much greater level. That's in the community of like inference people. We have some divergent views on

1:10:08 what you can do if you design a system around serving at one to ten tokens per second, which is our whole North Star. More broadly, I think. There is this larger sense around How do people consume a trillion tokens per day? That's the world we want to create, the capability for them to do that. What's a trillion tokens like ground us in how much that is? A trillion tokens. Well, okay, and open AI pricing, that's at least five million dollars.

1:10:30 At the very least for five point five or five point six. Yeah. I think the dollars is probably the most good metric. Yeah. Yeah. millions of dollars. Yeah, so what's the world in which we consume what currently costs five million dollars per person per day? We were asking for at least three to six orders magnitude improvement in cost per token. Get that into five thousand, you probably have some customers.

1:10:49 And in fact I would argue that For some size of model we are approaching Five trillion tokens being measured in tens of thousands of dollars. And that's something that you could imagine running for a single job. Are you all worried that just like the average person just can't and won't do that?

1:11:04 Doesn't do that now with their own brain. There actually isn't that much demand for intelligence in the world. I never will believe in that. There's always demand for intelligence in the world. the on ramps to that intelligence are our challenge as a product community. I'm not a product person, so I cannot say uh had the best vision. You want to enable those people. I want to enable those people. I want them to never be held back by the sense that my free tier users.

1:11:25 I can't afford to give them this many tokens. And I hear that from my customers all the time. We want to fix that. What about the inverse question, not what you think is craziest, but like what consensus thing You think is wrong. One of the things I keep

1:11:36 Coming back to is this question of NVIDIA. I am bullish on NVIDIA in the short term and NVIDIA, you should never bet against them. They're always going to reinvent themselves. Fundamentally, I think one thing that surprises people is when I tell them that. Hey, if you look at Hopper to Blackwell to Ruben.

1:11:52 And you compare like for like What is the performance per watt of B float sixteen multiply? It hasn't improved all that much. Or you take that one step further and go to TSM C. If you look at TSM C five nanometer versus four versus three versus two. The performance per watt on these chips doesn't change like a dramatic amount.

1:12:07 Yeah. Consequence of this is people lose their minds over geopolitics, like what would happen if we lost access to D SMC for any reason. My contrarian take is that it wouldn't be that bad. Supply would take a shock for sure, but

1:12:19 The best processes that we have in the West. Like Intel, not that far behind. At worst, like maybe two X. Worse performance per watt. The gap is just far smaller than than you would make it out to be if you follow like the chipboard dialogue.

1:12:31 What else is happening in the AI world that is not in your path, meaning it's not like a component of this whole system that you would end up doing something in? that interests you most? Well we're fully downstream of models. So the model people get to decide. How to design their architectures. I don't have any input to open AI or Enthropic, but I can only pray that they go in the direction that is amenable to me. Or I have to like do my best to predict where they I think they're gonna go and build my serving architecture accordingly, both software and hardware choices.

1:12:58 They have I think the most interesting game. In some ways to play. Once again, this is going back to like the profundity of the machine thinking and how consequential it is to decide to use something like sparse attention versus dense attention or how consequential it is to like

1:13:12 use a different data type. We were training in v float 16, but now we can train in FP eight or FP4 lower precision data types. That is just an arbitrary choice, it feels like, but it has profound implications for what chips I can use and how I should build my hardware, think about the future. Compute. If you had a hundred entrepreneurs in a room, all of whom wanted to create some new compute startup.

1:13:32 And let's say they were specifically wanting to make hardware chips or systems or racks or whatever. What advice would you give them on How to orient their company, sort of like the type of company, not the specific choice they're making on a tech bet or something like this. Cause it seems like we're gonna try everything, and that will be great for the world. Some stuff will work. But if you had to give them advice on how to orient their business to be successful in this coming world,

1:13:52 What advice would you give them? It's all about the bottlenecks on supply chain. So you need to first convince me or convince an investor that you understand the like three to five bottlenecks that dictate modern chip supply. There's TSM C waiver capacity, there's HPM capacity, and there's Advanced packaging. Maybe a fourth one would be power.

1:14:09 Where will you get the power? How will you build these racks? you should have a great answer to each of those four bottlenecks and how you're gonna work around them because It's all arbitrary at the end of the day. You're building a chip because you think that NVIDIA has made some choices that are difficult for them to change, which is true. NBD makes a lot of choices that are difficult for them to change. They're not perfect, they're just really well balanced.

1:14:28 You wanna be spiky. You wanna pick something and say I think they've underpriced. the impact of how short we're gonna be on HPM. We're gonna push really hard in this other direction instead. Which

1:14:37 I do think is probably the thing to attack most. Why? There's no easy way to bring on a lot more fabs. of memory. So it's just gonna be a while until we have yeah. The the boys and boisey don't love

1:14:49 Huge cap X for cyclical. They've been burnt on that many times. Conceivably like because of that shortage, the world is just gonna route around it by making Everything else in the system more efficient? I think they're gonna make everything else more expensive.

1:15:00 Think that iPhones will cut their memory iPhones are gonna go up in price. We're just gonna deal with it. Why doesn't NVIDIA go all the way to the end and sell tokens, do you think? Nvidia is really smart about this. They don't compete with their customers. Nvidia takes a long view on everything. Why don't they even start with the Neo Cloud? Why don't they just sell computer at the backdoor? Well

1:15:18 NVIDIA is really good. Jensen is really good at making his friends billionaires. He's made Core Weave a million dollar company, a many billion dollar company. And There's no need for him to destroy that goodwill. He wants to create a diverse community of neo clouds and inference providers who are all jockeying. to create demand for NVIDIA such that

1:15:34 If any one of them decides to I don't know, vertically integrate or go with AMD or any other Option. He's got three more people. hungry to fill that position.

1:15:43 It's great to have competition. Amongst his buyers. My favorite closing question for everyone is what is the kindest thing that anyone's ever done for you? My immediate first thought is like all inventors that I've had over the years. It's a rare person who

1:15:55 takes a lot of time out of their schedule and makes it like their personal interest, essentially, to make sure that you understand something. Or teach you something or ingraining some value in you that Yeah. think that you're on the cusp of understanding, but just push you over the line for understanding.

1:16:08 lot of the people in NVIDIA that I mentioned earlier who instilled that like love of performance in me. But also my professors in college who I remember like my advisor in like sophomore year. I was very impatient student as I show up at his office hours and say I wanna build AI chips. I know what I wanna do. Why am I wasting time taking all these like other basic classes and networking and operating systems?

1:16:27 And he just laid out basically like the whole stack and showed me the beauty of understanding every piece in the puzzle. he took my entire path of like trying to focus on one piece of the system. And said that it's so rare that someone can actually understand the entire stack from the gate level silicon all the way to

1:16:44 Building a great internet skill service. You should aspire to be. Someone who over the course of your lifetime achieves that level of understanding. It is such a rare trait. That level expertise is so noble to chase. That stays with me quite a bit.

1:16:57 Amazing conversation. Thanks so much for your time. Thank you so much for having me. If you enjoyed this episode, visit Colossus.com. You'll find every episode of this podcast complete with hand edited transcripts. You can also subscribe to Colossus, our quarterly print, digital, and private audio publication featuring in-depth profiles of the founders, investors, and companies that we admire most. Learn more at Colossus.com slash subscribe. You know how small advantages compound over time that's true in investing and just as true in how you run your company. Your spending system is your capital allocation strategy.

1:17:50 Ramp makes it smarter by default. Better data, better decisions, better economics over time. See how at ramp.com slash invest. As your business grows, Vanta scales with you, automating compliance and giving you a single source of truth for security and risk. Learn more at vanta.com slash invest. The best AI and software companies, from open AI to cursor to perplexity, use WorkOS to become enterprise ready overnight, not in months. Visit WorkOS.com to skip the unglamorous infrastructure work and focus on your product. Bridgeline is redefining asset management technology as a true partner, not just a software vendor. They've helped firms 5x in scale, enabling faster growth, smarter operations, and a competitive edge. Visit Ridgelineaps.com to see what they can unlock for your firm. Every investment firm is unique and generic AI doesn't understand your process.

1:18:36 Rogo does, it's an AI platform built specifically for Wall Street, connected to your data, understanding your process, and producing real outputs. Check them out at rogo.ai slash invest.