Transcript

Gavin Uberti - Real-Time AI & The Future of AI Hardware - [Invest Like the Best, EP.356]

Free .txt

0:00 I know firsthand how complex the tech stack is for asset management firms. And seemingly every new tool and data source makes the problem even worse, adding more complexity, more headcount, and more risk. Ridge line offers a better way forward, one unified platform that automates away the complexity across portfolio accounting. Reconciliation, reporting, trading, compliance, and more, all at scale. Ridge line is revolutionizing investment management, helping ambitious firms scale faster.

0:25 Operate smarter and stay ahead of the curve. See what Ridgeline can unlock for your firm. Schedule a demo at ridgeline.ai. Hello and welcome, everyone. I'm Patrick O'Shaughnessy, and this is Invest Like the Best. This show is an open ended exploration of markets, ideas, stories, and strategies that will help you better invest both your time and your money. Invest Like the Best is part of the Colossus family of podcasts, and you can access all our podcasts, including edited transcripts, show notes, and other resources to keep learning at joincolossis.com.

1:00 Mm. Patrick O'Shaughnessy is the CEO of Positive Sum. All opinions expressed by Patrick and podcast guests are solely their own opinions and do not reflect the opinion of Positive Some. This podcast is for informational purposes only and should not be relied upon as a basis for investment decisions. Clients of Positive Sum may maintain positions in the securities discussed in this podcast.

1:23 To learn more, visit psum.vc. Mm. Last week we heard from ninety nine year old Charlie Munger on the show. Today my guest is twenty one year old Gavin Uberti, who dropped out of Harvard to build Etched, which is one of the most fascinating companies I've seen recently. The topic of our conversation is the ongoing revolution in artificial intelligence, and more specifically the chips and technology that powers these incredible models. To date, general purpose AI chips like NVIDIA GPUs have powered the revolution.

1:55 But Gavin's bet is that purpose built chips hard coded for the underlying model architecture Will dramatically reduce the latency and cost of running models like GPT four. Gavin thinks we're about to embark on what he calls the largest infrastructure build out since the industrial revolution, and I won't spoil what he thinks this will unlock for all of us. It is so uplifting to me that someone so young can be working on something so big. Please enjoy this great conversation with Gavin Uberti.

2:23 Devin, it's really hard to know how to tackle this enormous topic, but you've become one of my favorite people. to learn about the future of artificial intelligence, models, hardware, software, data centers. A million things that affect tons of public equities, tons of Soon to be applications. maybe an appropriate place to begin would just be one of your early investors told me you said this interesting line to him, which was something like

2:48 You felt like you were born too late to explore the world, but too early to explore the stars. And you were sort of depressed by this. But then maybe this whole world of superintelligence and AI allowed something in between or even more interesting than the stars. I would love to hear the origin of that. reflection of yours and what you think about it now. I strongly agree with this.

3:08 One can't help but think. that hundred years ago, two hundred years ago There were so many mysteries left to be solved. That low hanging fruit, so to speak. Somebody smarter could go into their uh whatever they had instead of a garage.

3:20 Think for a couple of years and come up with optics or physics. You can go and explore continents. Then as Technolop. That became No longer feasible.

3:29 That stuff had already been done. And yes, you could still do interesting and novel work. But it often required going to school for ten years. Before you could learn all the mathematics you needed to go do new mathematics. So yeah.

3:41 Chemistry. You had to go study. All that had come before. And Someday mankind will go to the stars.

3:49 We'll be able to explore other planets. And that fruit will be low hanging too. But that stuff is uh very expensive today. But I think AI is very unique. Because

3:59 Unlike mathematics and unlike chemistry. The field is so young. that we haven't had a hundred years to develop it. We haven't had a hundred years to solve these low hanging problems. A single person can make a major, major impact.

4:13 And This got even better. With the transformer. All this previous technology of convolutional networks and support vector machines was thrown away. in uh twenty twenty when GPT three came out.

4:24 Showing that transformers really are the Best way to do things. I am really emboldened. By this transformer revolution. I think it is the most interesting problem of our time.

4:34 And I think that's It is. Very accessible. Because there's no hundred years of history built up around it. Can you give me your

4:43 interpretation of the idea of superintelligence I don't think superintelligence Is some black and white thing. There are some people who say That one day the machine will be as smart as us and the next will be a hundred times smarter.

4:56 I don't think that's true. And the reason I don't is because to go from GPT two to GPT three to GPT four requires spending Exponentially more compute. So

5:07 I don't think you get God like superintelligence. Just overnight. I think you have to go build an enormous data center to support a thing like that.

5:16 Way, way bigger than anything on Earth today. I don't believe in this concept of fast takeoff. I think that's a good thing. We're gonna spend. To use building GPT five.

5:25 And then we're gonna build new centers for GPT six and that'll take three, five years. And then we'll keep iterating on that. Slowly building better and better and better ones. And eventually they'll begin to seem human level. Well begin to show human level.

5:37 Agenicness. And then they'll become better than humans in some avenues. GPT four already is a better lawyer than I am. So I think the superintelligence is Much more of a spectrum.

5:48 Than it is a binary thing. And eventually we'll look and we'll say, Well GPT twelve is obviously super intelligent. This thing's way smarter than any human. It'll be clear in retrospect.

5:59 But when you're in the middle of that curve. It looks flat. What role do you think humans will play At that stage. I mean, obviously like safety and alignment and all these things are

6:09 euphemisms for Managing fear. Of something that is a hundred times smarter than us, even if that takes much longer than we think there is no fast takeoff. I'm super excited.

6:20 Yeah, what's the balance of excitement and fear that you have? It is very hard on the excitement side. I cannot wait to see What kind of books? What kind of movies? What kinds of video games.

6:31 A machine that's Ten times smarter than me can build. You look at the films of somebody like Spielberg, a great director. Some jobalo like myself, you say, Wow.

6:40 That is incredible. Imagine what a director ten times better than Spielberg could put together. Wouldn't you love to see that? So I think there's so much to be excited about.

6:50 AI is making. Breakthroughs in science and technology. Providing some sort of universal basic income. And yes, that stuff's probably more impactful. But like

6:59 Viscerally, I am most excited. Art and games and movies and T V shows.

7:08 In that world, is there a role for people? Of course. Are we gonna do anything but be the ones that are entertained? Oh, I think chess is a great analogy here. The machines have been superhuman at chess for a

7:20 Twenty five years? And back when uh Deep Blue first beat Gary Kasparov. People said Tress is gonna die. Weather Through lack of interest or through rant and cheating.

7:30 And the opposite has happened. just has become more popular today than it was back then. Go as well. Machines. became super intelligent at go in twenty sixteen.

7:39 And yet there was no Death of all the human go players. People just said these things are better at it than us, but there's still a lot of value to be extracted by humans playing other humans. And even go a step back. Machines have been fatter than humans for a long time.

7:54 I can outrun a car. But yet. I still go running. People still run marathons. There is so much value, so much fulfillment.

8:02 In doing a passes other humans. Even if the machines are in some sense better. Can you explain a lot of the basic terminology to us. I wanna step back and think about this.

8:14 as an opportunity to really learn the basics of AI. It's the component technologies that make this possible. So that we have those building blocks to talk about what they might do for us in the world. And you've got lots of really interesting ideas about what might happen. But let's start with the transformer. So what is a transformer? What does architecture mean in this context if you're baking one of these things onto the silicon?

8:38 start with the basic terminology of why the transformer is such a key piece of technology or a key idea that's driven so much of what's exploded in the last year. Well at a high level transformers a sequence to sequence model. You put in a sequence, you get out of sequence. And the way that you're trying to

8:55 Termine's What conversion is done? But that's not a very satisfying answer. So let's go one level deeper. One of the other real breakthroughs

9:03 Or the transformer. was realizing that It takes a lot of data. Teach a machine.

9:11 how to be helpful, how to be honest, how to give these useful responses. But We don't have to train it. Unjust helpful. Honest and useful responses.

9:22 Trick they found is Imagine we take a machine. And we give it the first hundred words of a document. And we say, given these first hundred words. Guess what the next word is?

9:32 And then given those hundred and one words. Guess what the next word is. We play this game again and again and again. Trillions of times. Now, at the end of that, you're gonna get a machine that is bad at guessing the next word.

9:45 Maybe it gets it right. eighteen percent of the time. But To even get it right eighteen percent of the time. You have to have a huge number of concepts stored within that machine.

9:54 So they take these transformer models. They feed them a huge amount of text playing this guess the next word game. And at the end you get a model. That is able to predict. The next work in science papers.

10:04 But understanding some of science. Internet stuff. Having read a lot of web pages, and so on and so forth. And the second step on top of that.

10:14 What they call R L H F. Well you take the model that has been pre trained. with uh predicting the next word and has all these concepts stored somewhere inside of it. And then you ask it to hey. Mimic this text.

10:27 And it is helpful and honest. You give it some examples. Oh yeah. Debugging code and being helpful. And this second piece.

10:35 requires a lot more extensive data. Takes much less computational time than that first step. So after you stack these two things. You get your uh GPT model. That

10:44 Takes it over. I'll put the Next word. That it thinks this helpful honest system would say. And I think a good analogy here.

10:53 Is uh Children going to grade school. We want to go teach kids. How to be good investment bankers. How to be good computer programmers.

11:01 But you can't really just throw a toddler. into a computer programming course and expecting to get it right. That's too hard. You need to get them to have some Baseline stuff.

11:11 And this baseline stuff has to be Easy to grade. It doesn't necessarily have to be super relevant. So you put the kid in uh grade school. For

11:20 Thirteen years. Having them do tasks that are Not that related. But easy to get feedback on, much like the transformer. Goes to the uh pre-training step.

11:29 Where it takes in a bunch of words. Predicts the next one. It's not super relevant, but it is easy to grade. And after grade school, or after pre training. You have a model that understands some core concepts.

11:41 And then And only then. Jasoned it to college or uh RLHF. where you can put these other concepts on top of this and make it into a helpful honest assistant or a good computer programmer. It's pretty amazing how

11:55 Cogent. And incredible some of the answers you get. When you consider that what is going on is that it is predicting the next token or the next word over and over again, that it's not thinking ahead. of that token. It's sort of constantly guessing just the next one. This seems like an extraordinary limitation.

12:12 What other kinds of models. You've talked to me about tree search before, I'd love to hear your thoughts on that. Like I can't believe how powerful this thing is just by predicting a little bit ahead. Do you think we'll be able to get to a point where we predict a lot ahead and have yet more impressive answers? I do want to clarify a misconception a little bit.

12:29 Just because the models output words word by word. Doesn't mean that they're not planning ahead. They have all these internal layers and there's very strong evidence that there is a little bit of planning going on. Most like when you and I talk. I give one word at a time.

12:42 But I'm able to make internal claims that you can't see about what I'm gonna say after that. And As you would expect, you have some internal plan. About what words come next. It makes you better at guessing the next word.

12:55 So this pre training helps to develop this skill. Now that's said. I do think there is a lot of work to be done to uh Make these models. Not as auto regressive as they are right now.

13:06 Here's a good analogy. Imagine a chest engine. Like a stockfish. The way these work, the neural ones that are the most advanced to date Is not by just looking at a board.

13:17 And guessing the next position. Instead You're gonna try a huge number. Of possible sets of moves. You'll play.

13:25 Hundreds of thousands. Maybe millions of uh games against yourself and your simulated opponent. And at the end of all of those. You'll say, Okay.

13:33 This position I find myself in. How much do I like that? Then you can kind of backtrack. Saying, Okay, based on where I want to be. This was the best tree to get there.

13:42 I think that language modelling might go a similar direction. Well the model makes Couple hundred plans of hey You're the next Hundred things I could say.

13:52 I like this path the best. So I'm gonna go pick that. However, this hasn't really caught on quite yet. There's a concept called a beam search in the literature. But does some of us?

14:02 Game search, I think, is much more primitive. Usually has at most four beams, but they don't look very far ahead. But for small model, it can be kind of helpful. And beam search also requires four times more compute. Which is a very large overhead to pay.

14:15 So this beam search, tree search stuff. Has largely not caught on yet. Although With the rumors of uh Q star learning from Open AI. If you believe them.

14:25 That does suggest that. They are doing similar Tree search stuff. In there. That's where the star comes from. Maybe describe just for the uninitiated what Q Star is.

14:35 A rumor, I should say, around Q Star. Well, allegedly it is the combination of Q learning. In the A Star Tree Search. This Q learning being a way to do stuff on top of R L H F? And the A star being, hey, we're gonna pick which of these tasks.

14:49 Is the best. That said. I don't want to put too much credence in this. It really is just a rumor. And there are so many ways a thing like this could work. Talk to me a little bit about how you see the future unfolding from where we sit with

15:03 A handful of Very impressive. models we'll just use GPT four as the main one that I think people know about and have used before. I'm using this. fifty, a hundred times a day, like

15:13 pretty much constantly. Lots of really great use cases, but for the most part I'm in a chat window in their open AI's window, I'm putting something in, I'm taking something out. And that's kind of the extent of my use. It's super valuable. I'm happy to pay whatever it is, twenty bucks a month for it. But I think you see a world where these models get used in a lot more places. So I'd love you to just sort of

15:32 open our minds to how you think the future might or should unfold. And what the constellation of technologies will be that are required to get us there. Well I think the right way to think about Or the tech will go. is to think about how this tech got here.

15:47 And there are a number of Innovation that Brought us to this point. But the biggest By far in my mind is scale.

15:54 We have gone from a uh one and a half billion parameter GPT two model. to a hundred seventy five billion parameter GPT three model. to a multi trillion parameter GPT four model. And companies seem set to spend. Ten times more cash.

16:08 training their next generation of models as well. Making these models bigger Makes them better. And now I wanna call this out because this is not how AI used to work. Previously if you took a stats class in like twenty twelve, you'd hear about overfitting.

16:23 If you put too many parameters in your model. It will fit the data too closely. It will get worse results when you actually test it and use it in the real world use cases. Imagine that you're sort of teaching the test. You give it the answers to the test.

16:35 And that makes it do very well on the test. But it's not gonna do very well in real life. But for whatever reason, transformers seem to Not fall into the trap quite so much. They are, to quote one of my teachers in

16:47 College. The first capitalist AI. Adding more money. Makes them better.

16:54 So based on that. I think we are going to see GPT five, next gen models from Google and Athropic. Be smarter than GPT four is today. And that's gonna unlock some interesting use cases.

17:06 Eight. Or I think people have talked about for a long time. Imagine a GPT four. thinking to itself and taking actions autonomously. And to date.

17:15 People haven't had very much success here. In large part because GPT four is not quite smart enough to get out of loops. And to go do the kind of planning needed to Operating on its own. But I think we're almost there.

17:28 I think when we scale up with G P five. And models beyond that. We will see agent use cases becoming more mainstream. And I also believe that this Chapa interface But I love as well.

17:39 Is Definitely not the end. Use case for these models. And he Scenario.

17:47 Or you are reading chat the uh link of model writes token by token by token. Is not a use case. That is Expensive. The model's probably gonna be cheaper than

17:56 human eyeballs. Human time is very expensive. And the first thing you'd think to try. But There are real technological limitations that stop other use cases from a

18:06 Becoming the standard. Take for example speech. This podcast we're doing right now. I'd love to be able to talk to GPT four, much like I talk to a human. It makes such difference at talking versus descending text.

18:20 So much easier to be misinterpreted then. And while people have tried to build Transformer uh speech to speech models. It's challenging because of latency. You talk to it, and there's this pause.

18:33 Before you get your response back. And that really breaks the flow of the conversation. So I think we need tech that will Enable.

18:42 Stuff like tree search. by bringing the cost of running these models down by a factor of a hundred. I think we need Tech that will bring the latency down. By an order of magnitude.

18:52 So we can go have conversations with these models. I think we need tech. Allow them to generate a thousand tokens per second. And make that code writing. Take

19:01 Two seconds instead of twenty. So I think there are so many use cases beyond this chapa interface. But it's waiting to be built. The tech's not quite there yet. Can you talk a bit about robotics?

19:12 And By robotics I mean like physical machinery, technology I know you did some stuff with the robotics in high school. What's interesting to you about robotics vis a vis models and performance and latency specifically. I think robotics is a very interesting

19:28 Use case that's currently kinda stuck. Yeah. The before times. Robotics Is a field that

19:35 works very heavily on Feedback loops. on shard coding algorithms. having some wild loop that says, Okay, move the arm until it goes and welds and then do this again and Old tech kind of stuff.

19:48 People have tried to go in. But some AI models into these robots. But it's very hard because of this lack of training data. For images you can go feed it.

19:57 A million things from ImageNet, a million frames. For robotics, you can't have it break the robot a million times before it figures out how to walk. That's just not affordable. And well, some people have tried to uh Build simulations. Whether or not you go interact with the world.

20:12 These don't work super well. Because of the differences between your simulation and the real world. You need to understand the real world. Without actually having to break itself.

20:23 Hundred thousand times. But transformers, I think, are an interesting way to solve this problem. There's a great paper out at Google. Two years ago, I think. R P one.

20:33 proposed a robotics multimodal transformer. One of the other beauties of a transformer. input in almost any Modality.

20:42 And the modality is a thing like text. Or images. Or video. Or in this case, a robotic sensor input. You can train.

20:50 A model. on many different inputs. In this case. They pre trained the model. With

20:56 A huge amount of text. Well that built some understanding of the world from this guess the next word game. And then on top of that. Other than this robotics encoder stuff. So I think that

21:07 Robotics is a super interesting use case for these models. I think that when Transformers are put onto robots. They won't be trained just with the robotics inputs. I'll be pre trained like today's

21:17 And you'll add those other sensors later. It's tough though. You also need some very good response times for a Robotics model. Yeah. For text.

21:26 It's okay if it takes a second. Tua Go off your machine. Be processed and then come back. A second is too long if you're about to follow.

21:35 So latency is so, so critical for any kind of application here. It's not quite there yet. There's been some limited adoption. And the Tesla self driving cars today, they use vision transformers. So

21:47 Much like a language model, but it takes images instead. The internals are almost the same. But They haven't hit the mainstream for robotics. Largely because of this latency problem.

21:57 It seems like latency is like the thing to Focus on. I mean there's two roadblocks, I guess, to more huge applications, companies, use cases being built on top of this technology. One is just how smart the thing is, and that seems to be on a one way trajectory. I'd love to hear what you think the limits to that are. And the second seems to be just latency, like so many of the things you described, I would say we need

22:18 real time AI or something. That's conversation, translation, robotics. So many of the things you've described. Seem like real time is the thing that's really missing here. So maybe describe

22:30 How we can Get over that hurdle. I strongly agree that real time AI is the bottleneck here. And to go a step deeper. What do we mean when we say uh real time AI?

22:41 What do we mean when we say we need Better latency. We measure it in two ways. First. The number of milliseconds between when you submit your response.

22:50 And when you get your first token back. So you're hitting enter and you're seeing the first answer from chat G. And Then The time between subsequent tokens.

22:59 Now why does the first token take so much longer? In theory it's only running the same model. Well, that's because it's running. The model on the entire prompt. If you give it, say, two hundred token prompt.

23:12 And then has to run all those two hundred tokens before it can give you your answer back. And it's actually a little bit worse than this. Because imagine that you have some conversation six messages long. It's a very expensive.

23:23 To go store all that in RAM. I'll just waiting for you to write your response. It's usually cheaper to just recompute it. So if you hit enter. Your chat's already six messages deep. You and chat GPT talking to each other.

23:36 It's gonna have to feed that entire message history back into the model. Before it gives you your next tokens back out. So instead of being two hundred tokens. Even if that was your last message, it has to go compute this model on a thousand two hundred tokens. So

23:51 If we wanna bring this Initial delay down. We're gonna need ships that have way more Compute. And are able to use that really efficiently.

24:00 And I think the best approach here Is model specific chips. Imagine a chip. Where you take the transformer model. This family.

24:08 and burn into the silicon. Because there's no flexible Ways to read memory? You can fit Order of magnitude more compute.

24:17 And use it more than ninety percent utilization. This lets you solve this problem. And get the responses back in milliseconds instead of seconds. Talk a little bit more about The actual

24:28 machinery of that. Way of doing things. This makes me think of basics from the Bitcoin era when those got so popular.

24:36 maybe describe what an ASIC is, but I remember these chips that became ever more specialized at literally just Bitcoin mining just guessing hashes basically over and over again as fast as humanly possible. And anyone that had an ASIC. outperform someone that had a GPU, outperform somebody that had a CPU. So maybe just talk us through

24:53 Everyone's heard those three terms probably, but may not know like what the differences are and what the trade offs are. What are A six and how does this work? Well let me give it to you in the uh Context of Bitcoin. In the very old days.

25:05 Back when did Bitcoin Algorithm was first devised. People ran it on CPUs. And to go mine a bitcoin, like you're saying, you have to solve A very hard math problem.

25:15 The only way to solve this map problem is by guessing and checking answers very fast. A CPU. Can do many things. It can do this too. But it's not very fast about it.

25:25 CPU mining was originally the uh standard. People then said, Hey, let's run this thing on GBUs. A GPU. Has Several hundred many CPUs scattered across the uh silicon.

25:36 And because of this, it's able to go do many things in parallel. Each one of these cores in a GPU can do all kinds of different tasks. Once people wrote GPU miners, CPU mining became So uh Bad by comparison.

25:49 That you'd lose money. Running it. But then people said, Well, what if we took this a step further? On a GPU. Only a tiny fraction of that GPU is used for a mini bitcoin.

25:59 We don't need the caches, we don't need the IOs, we don't need be crazy instruction processing in L zero caches. What if we Specialized at ship. But

26:08 Every transistor on that chip. So it would just mind Bitcoin. And uh Bitmain and MicroBT did this in the uh twenty twelve, twenty thirteen era. We saw an explosion. And how fast people were able to solve these problems.

26:22 The Bitcoin algorithm made them much harder to compensate. And GPU mining became unprofitable overnight. Because of this. GPUs have not really been seen in Bitcoin mining for Almost ten years.

26:33 That's not to go say that GPUs were a Scrapped overnight. They found other use cases. They're flexible. They went to mine Ethereum.

26:41 Or when to go do AI, or when to play games. But Not Bitcoin. Or do you have an ASIC for something? It is just

26:50 not commercially viable to run on GPUs anymore. I think the same exact thing is going to happen. For Transformers. Previously, in like the twenty ten era of Bitcoin. There wasn't really enough volume.

27:02 Justify spending a hundred million dollars. To build a custom chip. But once Bitcoin became popular. There totally was. Then the tips came out.

27:11 Everything else becomes. No longer feasible. For Transformers. Even eighteen months ago.

27:17 There wasn't enough demand to justify spending. hundred million dollars to burn that algorithm into the silver. But now With chat GPT? Get up co pilot with character.

27:28 With all these models from Google and Anthropics. It does make sense. The first transformer ethics are coming. Can we talk about this progression? So I'm just gonna use I'll stop saying chat GPT just to abstract away, but let's say we're at model level four right now, and we're gonna go to five and six and seven.

27:44 You referred to the need to probably build things that don't yet exist. Not just chips, which is obviously what you're working on. Not just training the models, not just gathering data, but actually the infrastructure, the physical infrastructure to do this. What do you think happens to the data centers and power sources and

28:01 Things like bandwidth. It just seems so interesting that We're gonna have to build new stuff to enable each successive progression of new model. So describe how you think that iterative thing might Play out.

28:13 I think there's a great analogy here, actually. And that is a Semiconductor fabrication. The buildings that make microchips. These incredibly advanced facilities have incredibly complex supply chains.

28:26 Many have heard of it. T S M C. And the the machines they buy from ASML for three hundred million dollars. And the mirrors ASML bias from Carl Zeiss. Well, that's just the photographic machines.

28:37 A huge array of complexity goes into this stuff. And that means that building a new fab costs ten twenty billion dollars. And I think we're gonna see a similar thing. For a Models of level five and six.

28:50 This whole chain will have to be redone. If you want to go and Scale from four to five. have the same multiple of compute. Over a

28:58 Three to four or two to three. You have to have ten hundred times more chips. And that means ten hundred times more power. This is not gonna be a twenty megawatt data center. That's a two gigawatt data center.

29:11 It's a data center that consumes the entire energy output of a big power plant. And I think this is also going to be very centralized. If you compare, say The bandwidth of an entire

29:21 Undersea submarine cable. For them between a couple of uh D P R each other? The GPU racks have double the bandwidth. Of a literal.

29:31 under sea cable. So there are so much incentives. to centralized data centers. So I think we're going to see if you very, very large facilities.

29:41 That do all this. Now the same is true of semiconductors for reference. T SMC's fabulous. Are all very large, very self contained buildings. But so much else has to happen too.

29:52 Need way more power. He had ten GPUs. Then Sure.

29:59 None of them are going to fail. If you have ten thousand GPUs. Then Okay, maybe one fails sometimes. And then you reset to a previous checkpoint and you try again.

30:08 You have ten million GPUs, or ten million. Transformer A6. Then You're going to have one failing. Every couple of hours.

30:17 You're gonna have to figure out. How do I deal with one of these things? Craping out. without having to stall the whole thing. No one's solved that problem yet.

30:26 Or heat. We've been able to go cool data centers in the two hundred, three hundred megawatt range. What happens when Two gigabots to go into that building. Are all turned into heat.

30:36 How does that get pumped out? Is that just dumped into the atmosphere? What happens next? I really think wholesale industries Going to evolve.

30:44 Around Building these Massive, massive AI models. We've only seen the beginning of it. All that would sound so remarkable.

30:53 That heat that heat dissipation, the power, the bandwidth, number of chips. All of that I think if I understand correctly is especially needed in training. Each marginal inference is less of a centralized crazy power intensive thing. Although collectively obviously inference will get quite big too.

31:10 We haven't really talked about training. And I know you're working on chips for inference to allow you to do orders of magnitude faster and cheaper inference, which will power the apps. But what about training? What is actually going on? In training. Of one of these models for those that

31:26 don't understand that process. Well, so first I wanna go talk about inference for a second, then we'll build into training. Because I believe that inference is going to be run in similarly large power hungry data centers too. And here's the reason why. Only run a transformer model.

31:42 We have this huge number of parameters. And each one of these parameters is a number. And To use that number, we take in a number from our input. We multiply them together.

31:52 and we add them to a running total. So every one of those parameters. In the case of GPT three, hundred seventy five billion. Is loaded from memory once. Then used in a math operation works.

32:02 It turns out that loading a thing from memory is way more expensive. Then Doing the math. So how do we solve this problem?

32:10 Always say that. These weights. Are the same. of cross one user or two users or four users or eight users. So you batch together a huge number of queries.

32:20 Let me load him that way once. Yeah, we use it. sixteen, thirty two, sixty four times. And that's one of the really interesting things that a transformer ASIC can do. You can

32:29 Have a much, much larger batch. Not sixty four. But twenty five hundred. So we're able to go load that white in once.

32:38 Pay the expensive price. And then amortize expensive price over a huge number of users. Making inference much, much cheaper. And now while this sounds good in theory. This does mean that you have to run that model in a place.

32:51 Well you can have Huge number of users all kind of grouped together. So I think that means inference will be centralized in much of the same way as training. To go to training specifically. Why are these machines so power hungry?

33:03 Well In inference. You have to run the model. Fords. But we're framing you have to compute. how much each one of those little weights contributed to the

33:11 Final error. Do that you have to run the model backwards. Which takes it out double the computer. Additionally. Running uh forwards and backwards.

33:20 Requires a Different kinds of network primitives, different kinds of connectivity. Then just running forwards does. So training is a more challenging problem. But

33:29 The economics that make A transformer specific inference chip makes sense. I think will also apply to training. If you're spending say hundred million dollars on a AI model.

33:39 And it costs a hundred million dollars to build A new custom basic. Then No matter how good the custom ASIC is. You're not gonna make your money back.

33:48 But if it costs a billion. Train your model. A custom is it costs a hundred million? Make you twenty percent better. Then yes.

33:56 It does make economic sense to build that custom chip. If you're spending not one billion but ten billion dollars. It is a no brainer. You have to do it. Now.

34:06 Building that custom training chip. Is Not an easy thing to do. Right now. Going from

34:12 Perception. to uh mass production of a yeah microchip. Intake. Four or five years. And we can't wait that long for

34:20 Training our new models. I believe that too. Support this demand. support these custom basics for training that we will eventually need. There has to be a whole redesign.

34:30 I'd love to understand a little bit about how the cycle time for new chip ideation, design, and then delivery. has happened in the past because

34:44 the iteration speed of chips themselves. Will need to go a lot faster than it has in the past. in order for us to have all this great stuff that we've been talking about. And that that is a vector of innovation that's

34:59 It's as critical, but certainly like extremely critical alongside everything else. We've talked about so How's that worked in the past? If I'm NVIDIA or something, I'm whatever company designing some new chip. Talk us through like the major stages of chip design.

35:13 maybe even starting with the thing trying to be solved, how you arrive at the problem to be solved in the first place. So that's step one and step X is the chips are in a data center somewhere being used. And how long does each of those steps take? Like I'm sort of a novice on semis. And so I'd love to learn how it's done in the past, and then we could talk about how we might innovate to change that in the future. Well, there are a few steps involved in bringing a new chip to market. No matter what year you're building in.

35:40 First have to say okay. What's the problem we're trying to solve? And then from there you'll divide an architecture. And in microarchitecture.

35:50 of how your chip works at a high level. To solve this problem. And as part of this architecture work. You're going to get a number of blocks. You'll say, Okay, in order for me to do this giant matrix multiplication.

36:02 I'll need a major multiplication block. And that major multiplication block will be made up of many sub blocks that each do a Multiply add operation. It'll be made up of a control block. That sends a signal to the full matrix block.

36:14 And so on. Then you'll take these blocks. number of blocks of Chip has varies widely. But it can be as high as hundreds. And you'll give it to a RTL engineer.

36:24 Who will write code? In a hardware description language, kinda like a programming language. That tells the block how to behave. Then This block is verified.

36:33 Because it costs so much to manufacture semiconductors? You have to make sure there is not a mistake in the RTL block. before you go pull the trigger and tell a TSM C or whoever it is to begin cranking out wafers off their line. So a verification engineer. Or very carefully.

36:50 Write a test fence to check that that block behaves the way that it should. And they'll also check that their test bench actually Runs. Every single line. In that paralog file. They call that line coverage.

37:02 And I'll go a step further. That they're a log? That code. Will be compiled. To a net list.

37:08 This is how each Group of transistors talks to each other. And then you'll run toggle coverage on this. You make sure every one of those transistors will flip to uh either state. So that's what they call the front end of the check.

37:20 Now you have a net list that comes out of that. But now you had to put the on the actual silicon wafer. and lay them out. That's a hard problem.

37:30 So you'll uh give it to a back end engineer. We'll make sure that All those. Beyond the logic being correct. There's no issue where hey

37:37 This block takes too long. And it's not gonna be done by the time this other block has to uh Take his results. They call timing closure. And I'll make sure about it.

37:46 Oh ne yet chip. Especially a large chip. Process issues. manufacturing problems in the TSM C line.

37:54 One part of the chip behaves differently than another part of the chip? Even things moving a different speeds. It works no matter What

38:03 Kind of wafer is made. Fast or slow, what temperature it is. They call this getting all the corners closed. And a number of other checks too. There's not gonna be some weird radio frequency interference thing. There's not gonna be some pool.

38:16 some vacuum in the chip or other uh photosist will flow in. They check a huge number of props. And eventually after this back end process is all finished. Billa give it to A masked house?

38:27 Mask shop? We will take these designs. Of where your transistor goes. And I'll say alright. The transistors?

38:35 We need to go and Have silicon in these parts of the chip? But not these parts of the chip. If you ever heard of screen printing. The process kinda works like that.

38:44 From this File. The GDS two file. Which says where there should be and where there shouldn't be a transistor. They make a photo mask.

38:51 That will shine light on the parts with our R transistor. And not on the parts where there aren't. And then those photo masks are sent to uh Fab. Actually makes the chip.

39:01 When uses them. Two. Image? Onto those uh silicon wafers. After coating them with a chemical called photoresist.

39:09 That hardens when exposed to light. So That means that After all this is said and done. You have parts of the uh wafer?

39:17 That are hardened. In just the right pattern. for you to put your transistors down on them. So this is a very long process. real production samples coming off an assembly line.

39:27 For modern semiconductors. If we want to build a little bit of a little bit of a little The next generation of Transformer specific ASICs. For training and for inference.

39:35 This has to shrink. Now there's been a huge amount Of progress. To make this faster. Especially on older nodes.

39:43 Google in particular. Haza Driven a lot of this. They have a series of open source efforts. Two.

39:50 Automate Parts of this back end process. Because really. It's not that hard. It's just checking rules. It should be done by a computer.

39:58 And if you're building on a node like those from it. Building for Old technology? Also make the protection like two thousand five. You can have the backend done automatically in a day.

40:09 Est-il. Nine months by using one of these automated tools. And we haven't gotten to this point. With Enormous.

40:17 modern skill ships yet. Possible in theory. And I'm very hopeful that more smart folks will begin working on this problem. And same thing for uh Actually writing the chips.

40:27 doing this RTL coding. But today it's all done by hand. But For uh example C in Python. We have seen language models themselves.

40:36 begin to help out with it. GPT four is very good at both those languages. As is GitHub Copilot. Because there's so much code. And see and in Python available online.

40:46 So having access to these AI tools. Can make a developer much faster. Bruh. paralogue in these hardware languages.

40:54 That's not true yet. While software loves open source. Hardware really doesn't. So it's very, very hard to get. All this

41:03 Code. into your model as training data. Because it's all proprietary. And I'm very optimistic. that somebody will be able to go to

41:11 Synopsis and go to all the old Defenders. Just say hate. Your code for CPU is from like two thousand five. That's probably not very useful to you anymore.

41:21 Can I use that to train my AI model? And if enough companies are in this to say yes. Ideally to open source it, but probably not. Maybe just to go train that model. Then you can get a language model that was good at Verilog.

41:33 as good at Verilog as it is at SAC and Python. And use that to speed up. Writing some of this. RTL development.

41:41 What do you think the outside limits are on how fast this could happen. Where You said five, four, five, six years might be the Start to finish time now.

41:52 Where do you hope it is? Ten years from now. What do you think the upper limit is of how fast this iteration time could be? I really want to go Call out.

42:01 A lot of the uh Google led Silicon Ecosystem. I wanna call out. Global founders. The Skywater Foundry for open sourcing their PDKs to help make this process fast.

42:11 Now If you're a hobbyist. Who wants to build a chip on One of these older technologies.

42:17 You can do it. From conception. To putting it together a few transistors, taping it out. Weeks, months. Even if you're starting from absolute scratch.

42:27 That is so cool. They're able to do it for pretty cheap because they take all the hobbyist projects. Put'em together. On one mask that. You make a one photo mask that has all the chips on it.

42:37 And I'm really optimistic that This can be scaled up. Now it's definitely not easy. These are hobbyist chips on all nodes.

42:47 Have you no? hundreds of thousands of transistors on them. And the amount of effort it takes a computer to decide where to put these things. Is Super linear.

42:55 And the number of transistors. Going from one hundred thousand to eighty eight billion? Like you have on the NVIDIA GPUs. Is More than

43:04 Eighty billion over a hundred thousand times harder. So instead of being able to do it on one DC server. be suddenly a huge, huge data center. Working on this. Optimization problem.

43:15 To get it working. So I am Optimistic as a future. Where

43:20 Just like for a hobbyist. Building a chip on an old note today. We'll be able to get companies published tips on new nodes. By spinning up huge data centers. to do this back end process.

43:30 And training these language models to help them write the front end. I think the upper limit really is. four or five years of today.

43:39 A lot of work's gonna have to happen to bring that down. And it's also a problem where The first chips. That are made. With some faster process.

43:48 Probably aren't gonna work. People don't like rocking the boat. Because of how expensive it is. To build a ship. Only for it to come back dead.

43:57 A few people do it. It'll be the obvious solution. But nobody wants to be first. Why do you want to try? I mean why because it has to be done.

44:07 Someone's gotta do it. Someone's gotta do it. That's the only way that we get to our uh Ten billion dollar. Training runs.

44:14 That's the only way that we get to GPT six. That's the only way that we get. These AI generated books and movies and TV shows. That would make the artist of today.

44:24 Be impressed. Someone has to do it? And might as well be me. It's interesting to think about In this future world now.

44:32 The existing landscape of companies and stakeholders and how they'll start to interact with this technology as it becomes available. So when I went and looked up the top buyers of H one hundreds or something from NVIDIA. You see kinda exact ones you would expect. You see, I think Meta and Microsoft are the two top purchasers, Amazon's a big purchaser.

44:51 So these hyper scaling, hyper large technology companies that are either providing cloud services in the case of Azure or AWS, or their meta training their own models, open source models. Buying all these H one hundreds. Who do you think? Are the customers

45:07 in this different world where data centers are different, ASICs are the predominant chip. Fortune five hundred companies are all trying to find their version of applied generative AI. Do you think that the market map and market structure looks a lot different in terms of like Is it the same buyers buying these chips as buying the H one hundreds?

45:26 Again. That's a great analogy. In the world of chip making. Well the role of the chipmaking used to be pretty cheap. Anyone designing their chips probably have their own little chip fab.

45:35 Because it costs. Couple million dollars. And as these things became more and more expensive. All but the biggest players became priced out. Now

45:44 Only at two or three firms. T S M C Samsung and maybe Intel. The world state of the art ships.

45:52 I think the same thing will happen. for a generated AI and these negative scale data centers. Today. Many companies are trying to explore Probing the waters.

46:02 Because it is so cheap to do so. But as having a state of the art model becomes a hundred million billion, ten billion dollar endeavor. It'll be much more common. To outsource that.

46:14 to one of the few firms who will have one of these mega scale data centers. And don't get me wrong, you'll still be able to fine tune on top of these models these companies build. But they'll be much more like T S M C The

46:26 Here play AI generating house. And they will be some full stack integrator. I just think the world Cannot support.

46:34 a hundred different ten billion dollar models. But I can support. Three. So I see a world where You have these

46:41 few enormous leading edge state of the art development houses. Possibly. at the current hyperscalers, possibly at some new company. that specializes on just this kind of thing. And there'll be a tale.

46:52 Much like the R for semi conductors. Not every application needs. Leading edge chips. For example. The US government and US military.

47:01 care much more about chips being produced in a reliable American factory. And the tips being state of the art. So fabs like Global Foundry have accepted. They're not gonna be as good. As TSM C

47:12 But there's still a market. For them to dispens being American. I think we'll see a similar thing for these AIs. You'll have a few smaller data centers. Don't try to compete on being as

47:23 Purely smart. Ask the uh Big megaliths. But they'll find some niche. Obviously T S M C is one of the most interesting and intriguing

47:33 geopolitical Assets that We talk a lot about in the news. It's even a political issue. Do you think that these data centers and the control over the specific models becomes like the central political issue. Like there's a presidential campaign where

47:49 the candidates on the stage are mostly just talking about Who has the best models and where are they located and what's the security around them and nationalistic type issues. It's all seems almost like the start of a sci fi novel or something. If I said that we were going to have Threats of war over a chip fab. Ten years ago, that sounded like a sci fi novel too.

48:08 This has to be the conclusion here. Some countries are going to have these hyper scale data centers. And if you don't have one. It's gonna be really hard to catch up. China is famously trying to catch up on development.

48:22 They realize how critical that is for AI. But they did not get on the E U V train? When it was leaving the station. They're now on this older.

48:30 immersion technology. That's really holding them back. And while they have produced some uh great quote unquote seven nanometer chips. I do think they'll need getting better technology.

48:40 China is not state of the art for semiconductors today. And does not seem Like it's not like path to be either. And I think for data centers as well. The United States

48:50 It's very lucky to have. All the big hyperscalers. Very lucky to have. The big AI companies too. But we cannot.

48:59 Be uh passive about this. We're in a good state now. But that could change. So I think the export bans four state of the art AI chips. are really in the national interest here.

49:11 Can you explain the essay The Bitter Lesson. And then talk about the implications of the upper limit of how good these things could get. Based just on the available data in the world.

49:22 These things have been a little bit more than that. Trained on Insane amounts of data, but there is a finite amount of data in the world. There's more being created every day. But talk about the bitter lesson and the importance of data and the sort of upper limits of how good these things could get.

49:35 So the bitter lesson is an essay put out by Rich Sutton. great pioneer AI. In twenty nineteen. Before the whole uh transformer crates.

49:44 We're just looking at the field. And saw what would happen all the time. If somebody would take a model that worked. And they'd hard code some knowledge and do it. Uh, for example, somebody would say, Hey

49:55 Here's a great image recognition model. But it gets this one weird case wrong on occasion. I'm gonna hard code a hacky fix on it. And Inevitably.

50:05 doing this lack and more thick strategy. Works in the short term. you get a little bit better performance, you can go publish your paper all as well. But it never scales. The only techniques

50:16 But have routinely Worked. Are those that leverage the fact that compute is getting cheaper. For example. I mentioned chess engines before.

50:25 In the very old days, people would write chess engines by saying, Well A rook is worth five points. A bishop is worth three. I'm gonna add a rule to my machine. That says hey.

50:34 If you can go trade a bishop for a rook. That's a good trade to make. You should go make that move. And again, that works a little bit. But it doesn't really generalize.

50:45 So the strategy that eventually worked for chess. Enable the blue. To become the uh world chess champion. Was not Encoding human knowledge, not hard coding rules.

50:56 But by just leveraging A massive amount of competition. Deep looted search. Much like this. trees which we were talking about earlier.

51:04 It would test. hundreds of thousands of positions. Find the one that I'll like best, and then play moves to get there. And the same thing is happening for uh Artificial intelligence.

51:14 In the very old days, The Image recognition? People will say, Hey There's a block of white pixels.

51:21 In this kind of shape. It might be a bird. It might be a light bulb. They had these big trees of fixes for it. And that didn't really work. Didn't scale up.

51:30 But the thing that did scale up. is making neural networks very big. And feeding them huge amounts of data. So The lesson.

51:37 Is that the only Techniques that were really Advance the field of artificial intelligence. That are able to leverage.

51:45 The cost of compute getting cheaper. That are able to leverage These enormous data centers. And you mentioned data as well.

51:54 A key part. of the learning algorithms. Is that data. There's only so much data available in the world. Some people have raised the alarm that we're about to run out.

52:03 The models aren't gonna be able to go. And get the next trillion tokens. It doesn't exist. And yes. I do think that'll be a problem eventually. But we haven't even begun to tap into

52:14 The most valuable source of data. the video. I don't know about you, Patrick. But when I learned to walk. I learn to pick up objects.

52:22 I didn't do it from a book. I did it by Looking with my own eyes. And seeing hate. Does this behave the way that I think it does?

52:29 And we've begun to kinda dip our toes into it. People have dull. Vision transformer models. And usually what they do is they Train a model on text.

52:38 And at the very end. They add in the image part. Just as a little side feature. And I think that's wrong. I think Transformers tomorrow.

52:47 We'll be trained from day one. With enormous amounts of video data. So that solves the problem in the medium term. But eventually these models will get hundred times bigger.

52:57 That require a hundred times more data. They will exhaust. All that is available right now. But I think that's Softly.

53:04 And Generating your own training data. Is the final frontier here. One good analogy. is the machines that played go.

53:12 They try training at AIs. On human games? To make them better than humans that came ago. And it worked. We run out of those games very fast.

53:21 So Google did with the AlphaGo. is they made the model play against itself. And learn from those games. And the same trick. has begun to show some promise.

53:30 And natural language processing. But you have an AI. That writes out some output. Then you ask it to analyze that output and say, Was this good or was this bad? As a human.

53:40 You can do that. You can get better. That's why I Revise the things that I write. And I think that technique.

53:48 Generating their own training data. will uh allow us to scale even once we run out of all the images and video available on the web. When we get to that stage though, what's the objective function? In go and chest. We know

54:01 what winning is, we know what good and bad is. We can define it Suppose we could define the rules of physics or something in a Open ended. play yourself game that these crazy models might do it themselves. These are goal seeking things.

54:14 How do we define good objective functions in that case? Well, one of the beauties of language models, I think is that they learned these concepts themselves. You can test this. You can go ask GPT four. Here's a scenario.

54:25 Is the human behaving ethically in this scenario? The AI knows. What is good? What is ethical? And I got the understanding by reading.

54:34 All the text we've generated as a species. So Only Ask it to go say, Hey. Pick the response.

54:41 And is more ethical. Pick the response that's better. It is using that concept. That develops from Reading

54:48 All there is to read. Can you walk me through the most important big companies of today? And just riff on each a little bit. Like there's the obvious ones, there's the hyperscalers, there's open AI, there's Entropic, there's NVIDIA, there's all of these different key players. I'd love you to pick the ones that you think

55:06 are Most important Really just through the lens of your own strategic thinking. Like you're building a business. There's been enormous scale benefits in this last cycle of technology companies. It's staggering how big the big technology companies are. How do you think strategically about the role that they'll play, how as a young upstart business you might position yourself strategically or counter position yourself against these incumbents. Why won't those same five companies just win everything? Like this whole strategy game is really, really interesting to me. And I'm curious how you think about it or if you have certain lenses to approach it.

55:40 I know that it looks from afar. Like Microsoft and Google. And the Amazon Are dominating everything. And that you dominate.

55:47 But at one layer of this. enormously complex supply chain. Let's think about it from the top. At the uh very top here we have the state of the art AI models. being trained by folks like OpenAI and Anthropic.

56:01 These folks have Enormously good talent. I think that's the big moat. But They don't really run the compute themselves.

56:09 They get it from one of the hyper scalers. So we have the models at the top. Below this we have the Hyper scale. Tech giants.

56:17 Those are the ones people are scared of. Google Microsoft. And Amazon. Are the big three in this domain.

56:25 buying those chips, racking and stacking them, supplying data thing with power. And then selling that to these companies on top. Where do those companies get their uh chips from? Right now it is almost entirely NVIDIA.

56:37 But I think that's going to change. A and D Is coming out with a new M three hundred X next year. I think it's going to be a Viable competitor to the H one hundred.

56:48 And companies like us. Are trying to replace this idea of general purpose AI chips entirely. And built specialized ones just for the transformers. So I think that Us along with some of the uh general purpose guys.

56:59 become this layer underneath it. But us. NVIDIA and AMD. Cannot. Make the chips ourselves.

57:06 We have to go to a fat house. T S M C. Yes they want a choice. But the chip is just one piece here. on a ch like the M F hundred or on the age four hundred.

57:17 The memory actually costs more than the silicon itself. By more than a factor or two. Now These memory modules.

57:25 Only two or three companies as well. Samsung Ski Heinx. And maybe micron. Two of those are in South Korea.

57:33 And Macron's American. So these people are other Really, really critical, folks. And that supply chain. Yeah, folks talk all the time about

57:42 NVIDIA, TSM C, ASNL relationship. They don't talk enough about NVIDIA two Samsung, Nvidia to Eske Hines relationship. Because these memory modules are so critical for uh building the state of the art GPUs.

57:56 But then These companies themselves. Where do they get their machines from? And now we get to a uh

58:04 ASML is a much smaller company. Then T SMC. And there are so many others. That

58:10 Help make these chips as well. But I think that The supply chain does. Multiprovider on top. Then

58:16 Paperscaled data center below that? Then Chip designer below that? And ship fab below this. This is the fundamental stack.

58:24 Of the future. And now what I did not mention here Is the end application. What does the user see? I don't think this will be built by OpenAI and by Anthropic and by these model companies.

58:36 That'll be folks building on top here. This fifth level of a stack. And some of those folks will be uh Encumbents. big established firms that want to build some new AI product.

58:45 A lot of it will be startups. Who say hey. Chatbots are boring. I have this new interface. Speech to speech.

58:52 I have this dupe interface. Enormous amounts of document ingestion and querying. I have this new interface. AI powered search of the whole internet. So

59:01 I think there's the most opportunity. Well, I think it's an opportunity at all these levels. Building new a high products. Innovating on AI models. Building.

59:10 New kinds of data centers. Better. designed to operate these crazy high scales. Building. Specialised ships.

59:17 And don't make them fast for running these models. And fabricating those chips. In a way that is. Much faster than today's processes. Given that the application developers at the top of the stack.

59:28 Have kind of an open green field, which is really cool. I'm sure we'll see amazing things. The way they build competitive advantage and succeed Seems kinda clear, just like it has been in the past. You're operating at a much deeper part of the stack. How do you think about

59:43 Building a better product, but also doing so in a way that Three years from now, NVIDIA doesn't just get one of them and look at it and say, Okay, well We're the scale player here, we sort of have the right to win. In ASICs just like we did in GPUs.

59:56 How do you think about that strategically? And I know the world needs all these things and it's a big world and there's a big and growing pie. I don't mean to suggest anything different, but how do you think about competing at a deeper part of the stack where there's more supply chain, more interdependency, huge giants that move fast and so on? Let's be clear, NVIDIA is a fantastic company. And I believe they will eventually build. Specialized transformer chips.

1:00:18 Mostly we're doing. It just makes sense. But right now they're stuck in the innovators' dilemma. But the H one hundred Yeah, selling like hot cakes.

1:00:26 It has an enormous margin on it. And if they were to build. Not even a trans specific ASIC, but say two hundred for inference. Or eight four hundred with a little bit of the graphic stuff ripped out.

1:00:37 That would Reduce the margin they have? On the H one hundred. And would I'll be bad for them entirely.

1:00:44 They can't innovate? Or the vanish goes away. And this is especially true for specialized chips. Uh one of the biggest motes for NVIDIA. Is Cuda.

1:00:53 their software stack that is very, very flexible. graphics able to run AI. Crypto That is a great note.

1:01:02 In the world where General purpose chips. General purpose AI chips are dominant. But specialized ships like ours. are much easier to program. In a world

1:01:12 Or you burn the transformer architecture into the silicon. You don't need Cuda in the same way. If the world goes to specialized chips. NVIDIA kinda loses that moat. So while I do believe they will compete.

1:01:23 They're not gonna be first. The only people who can really innovate are startups like ourselves. So I see a world where we get to market first. And NVIDIA gets to market a year later or so. Now.

1:01:35 What happens in that year? I think again going to the crypto analogy makes a lot of sense. A lot of companies. Thirty, I think. tried to build Bitcoin mining ASICs.

1:01:45 They're much simpler than GPUs. You can tape them out much faster. But The companies that got to the space first Ended up winning. There it's now this duopoly.

1:01:54 Between a micro B T and Bitmain. But main they're testing the wires for a IPO. A couple years back. Value that more than

1:02:02 Fift billion dollars. So Because When a specialized ship move into a space. general purpose devices become no longer competitive.

1:02:11 D. Person who gets there first. is almost guaranteed to win. I think that to go back to us. This means we have to move as fast as possible.

1:02:20 Back when we began this company last year. Before transformers were big. Before this was obvious. Then yes, we could uh afford to be a little more relaxed. But today. People are saying why is a no transformer ASIC?

1:02:32 We must Move as fast as possible. Ship on time. As silicon that works. For us to win.

1:02:39 You mentioned the importance of CUDA for NVIDIA. We haven't talked about the software that you have to build at the same time as you're building hardware. What role does that software play for a company like yours? And the way that the chip itself can be programmed or tapped or accessed by other developers.

1:02:56 Say a bit about those two. Parts of this story. Why did other AI companies fail? People have tried to build general purpose AI chips in the past.

1:03:06 And have generally lost out to GPUs. Graphcore, Sanova, City Berson, Grok. Have had some modest success. have not taken over the world in the way that folks thought they would. I think this is two reasons.

1:03:19 First The software A second the chips generally weren't better. We'll do software first. It is

1:03:25 Extremely hard to do what NVIDIA does. Cuda has been developed. Over I think it came out in two thousand seven. More than fifteen years.

1:03:35 And because it is so old, it has become Stable. Reliable. The workers people can depend upon. I started trying to get into the space.

1:03:43 won't have that. We'll be putting out something a little bit more raw. That's much harder to use. And if you look at say the ease of compiling code for

1:03:53 Video. versus even say A and D. It's no contest. Buddha is great because it is Production great.

1:04:02 And The only way to solve this problem. Besides being old, I think. is to be uh less ambitious. If you're doing a simpler thing.

1:04:10 you can get to real production great quality much faster. And luckily for us and for specialized chips. Because There is so much less programmability on the thing. Because you're only taking a transformers to

1:04:21 Transformer compiled code. The software stack is Much, much easier. Now that's not to say I want to be complacent about it. We already have a prototype of this running internally.

1:04:32 Running with our chip in simulation. And because of this. We can begin iterating on that software. Getting it stable. reliable production grade.

1:04:40 So that when the chips come back. The software Works the way that you expect it to. But the second piece of why these Companies didn't.

1:04:48 Succeed. Is it they weren't better? If you look at say The amount of memory bandwidth that they have. Versus the GPUs they compete against?

1:04:56 It's the same, if not worse. If you look at the amount of flops. Compute they have. Again. It's the same, if not worse.

1:05:05 And that, I think, is a much bigger problem. than just the software itself. If you think about the performance improvement that you want to drive. You just said, Look, these things didn't work'cause they weren't better.

1:05:17 How much better do you want to be is it? Five times, ten times. There's gotta be more than ten times. I don't want to share exact numbers right now. We gotta have some secret sauce.

1:05:27 But Because we're able to Specialized You can get a huge amount. More compute.

1:05:34 On these. Chips. And let me just give a concrete example here. It takes about Ten thousand transistors to build

1:05:42 A fuse multiply add unit. That's the uh building block of any measures multiplier. And NVIDIA has Not that many on their chip. Only about

1:05:50 Four percent of the chip is the matri multipliers. Because to feed them requires so much more circuitry. That's what they cost being so flexible. Or as us. We will have more than

1:06:01 An order of magnitude. More of these multipliers. Order of magnitude more raw flops. And of course, turning that into a Useful throughput.

1:06:11 That is still not trivial. You need a very good memory bandwidth there as well. But The reason I can confidently say we will not fall into this trap and that we will be better. Is the I I already know the metrics on our chip. I already have PPA figures.

1:06:24 We will just have So much more raw compute. Substantially more than an order of magnitude. That's not just for compute, but also for latency. People talk a big game about hey, if the cost of compete goes down, you can do so much more stuff.

1:06:38 I do think that's true. But it's not what gets me excited. With Chips. that have twenty times the better latencies than the state of the art GPUs.

1:06:48 Because they're able to run a huge prompt. Total parallel. That lets you build products you could not otherwise build. Let you build this real speech to speech t machine.

1:06:59 That Injust. The whole internet. In uh seconds. Instead of ours.

1:07:05 And that That is what gets me excited. One of the things that's so clear to me is the role of Strategic leadership. in making things like this happen. And by that I mean

1:07:17 Usually in a person, a single person. Obviously like on the software side. Sam Alton seems to be A or the person that's really driven a lot of this forward. He clearly

1:07:27 early recognized an opportunity and galvanized talent and a team around it and has managed them to build something really spectacular. I sort of think of you like the hardware Yeah. How do you think about strategic leadership. How do you lead a technical staff? I don't know how technically

1:07:45 My sense is that he is more of a traditional Leader not in the deep, deep weeds. How would you characterize yourself. And what do you think are the key components of a good strategic leader?

1:07:57 In the hardware side of this equation. Given that we're already familiar with some of the leaders in the software side. Let me tell you a bit of the story of how we got here. And then I'll Bring the into

1:08:09 What I think a good leader has to do. When I was uh All my gappy year working for Octo ML, living as a digital nomad, versus off in his dorm room. We had this idea for hey. Let's build it. Transformer specialized ASIC.

1:08:21 that we get a huge amount of performance improvement. We went to it. Big industry effect. Who told us that nope. Not gonna work.

1:08:29 But also He said that Despite being a veteran industry that He wasn't the guy to know for sure. But he knew just the guy.

1:08:37 To tell us why our chip would not work. Mark Ross. His resident Downer. Mark Ross is a seriously legit guy. Kind of personally makes you a little bit scared.

1:08:49 He was a CTO of Cyper Sendy Conductor for a time. Which that sold for nine billion dollars. Twenty nineteen. The mark has shipped. think five chips that have done more than a billion in revenue.

1:08:59 So we go to Mark. We show them our tech. And he says, Well I would say this isn't gonna work. But it's too early to tell.

1:09:08 You guys should go build a functional simulation. You guys should go write a white paper. And then it come back to me. And we can talk about hey, here's the problem that you're gonna hit. Here's why this isn't gonna work. And you guys are gonna learn a lot by doing it.

1:09:19 And then maybe from that we can figure out how to pivot. And build a chip that is viable. It's over a Lot of long nights. And I do just that.

1:09:28 We write a technical white paper. We do the original architecture for the chip. We build our functional simulation. And we go back to Mark. And Mark says shit.

1:09:37 Just works. But you guys are college students, he says. This is sending conductors. Just to get started. You're gonna need two to three million dollars, and a hell of a lot more after that.

1:09:48 This isn't data caps, you guys have never built a chip before. Then we raised five and a half million. Went back to Mark. It has to He knew anybody to be the itchup architect.

1:09:57 After a series of conversations, He used to be a tip architect. And we talked about a retirement. So Mark Ross is now our chief hardware architect. And they're kinda the rest of the Domino still.

1:10:08 We got Azat, our VP of Engineering. Fantastic. Until uh V P before this. Who is known for building these?

1:10:17 Very large ships. On advanced nodes with teams. It's like a dozen people. That is not how that normally happens. One of the few good men there who can still build the ship on a budget.

1:10:27 Yeah, we got a Saturday. Co founder Orating. We got Reynold, the first chip guy at Cruz, and dozens of other bets from Google, Microsoft, Amazon, you name it.

1:10:37 What does a good leader do? Kind of tie this all back. It looks from afar like Silentman himself is pulling all the strings of open eye and masterminding the whole thing. But I think it's the folks under him.

1:10:48 who are doing the magic. The folk under them. Who are making that into a real product. The role of a good leader. To set the vision? Get the right people in those chairs.

1:10:58 And not running out of money. That's a I've heard stories though about you in meetings. Where you go to the one inch or one centimeter level on some very technical part of what's being designed.

1:11:12 And I'm curious how much you'll miss that if you give it up and you just do those other things. And or whether or not there might be a world where you being in the technical details, I think about like an Elon or something that Apparently can walk the SpaceX floor, the factory floor, and know every frickin' detail somehow. What do you think?

1:11:29 Is the prospect of being a more technical hands on Leader for you. I think there's a lot of value. And a leader being tech. Cool.

1:11:38 Understanding your team's problems, understanding why they're doing the things that they're doing. Is a requirement. For bringing in the right people. There aren't that many levers you can Push at the top. It's very coarse things.

1:11:50 Guy to bring in, you have to understand that problem. And that's what being technical gets you. This is why Great Brockman Open AI.

1:11:59 It's like being a CTO. You still see that guy coding. He can't understand. how far along the model is. Unless he himself has seen the internals.

1:12:09 Even though I love the folks who work here and I trust them completely. I can't understand. How far along is the signal verification? How far along is this block on our chip? I'm talking to that guy.

1:12:21 And I understand why he's doing what he's doing. And also, to be quite frank, it's fun. Any closing Thoughts about things In this world.

1:12:31 That would be most surprising. To investors to and customers out there to CEOs of other businesses. that aren't technology or aren't AI related.

1:12:43 You've gotten so deep in the weeds on what's going on and trying to drive a f certain future. Is there anything that would surprise people most that we haven't talked about yet? Just take it back to this idea of Transformer. It's becoming the way neural networks are.

1:12:56 Where it's not gonna be replaced, there's gonna be a building block. You're gonna build say retrieval augmented generation on top of that. To make the model smarter. She read Microsoft paper, well, they have agents living in a simulation.

1:13:09 You're gonna build memories on top of these transformers. To let them remember things. So a real key storage cache. I think that Folks have talked and talked and talked about

1:13:20 Ah, the transformers going away. much more productive line of reasoning is how do I build on top of this? And when I look out How's happening in this landscape? It is the folks who are it.

1:13:30 Building on top of it. We're building the most successful coolest products. AI games. Based on having this memory. Build conversational things.

1:13:39 Speech to speech. Although With the higher latency than I would like today. Or uh

1:13:45 Primitive Air Lawyers of the Future. So again, the transformer Is not just the next fad. This is a building block. What makes you most confident?

1:13:54 In that. When I asked around I heard lots of people that agreed, but why That architecture. Which obviously has led to all the stuff we're talking about, so obviously it's incredibly powerful. But why doesn't it Tap out.

1:14:06 And Two years, three years, four years, five years. Why might there not be a different architecture that just totally takes over? If the world was fair, there could be. But the world's not fair.

1:14:17 The transformer. By virtue of being the dominant architecture. It's suddenly the thing that has Wait. Way more.

1:14:25 humans thinking about trying to improve, trying to optimize. Not with people. Hardware too. NVIDIA is building a little bit of hardware support for the transformers to optimize certain kinds of memory loads.

1:14:37 Into the next GPUs. That is going to preferentially favor the transformer. over any replacement that comes out. Sure. A software.

1:14:45 NVIDIA has Three. very highly optimized transformer libraries. I mean if you want to run a transformer model. It'll be super easy.

1:14:54 Do you want to run it say some other kind of thing, state space diffusion. You have to make the model so much better. But it outweighs the factor of two three performance improvement you get. But optimizing so heavily. So

1:15:06 For a new architecture to go and beat the transformer. It has to be so incredibly good. But it will outweigh. the hardware disadvantage it will have. Because transformers are favored.

1:15:16 It will outweigh. software bandage that has because transformer libraries are so common. it will outweigh the end user application because end users are used to Working with these transformer models. That

1:15:27 It's just such a heavy lift. And again to go back to this neural network analogy. Random forests. Or uh S VNs. could have been the dominant AI paradigm, I think.

1:15:37 But now To go cash that up to the state of neural networks? Totally infeasible. One thing we haven't talked about is In the world where there's a GPT X

1:15:49 And maybe a llama Y and an anthropic Z. What about Building models that are specialists That rely on access to a very proprietary data store. So there's companies out there

1:16:04 Or people out there that have closed access to really valuable data. That Could make these models better. What happens there? Are do they train their own model? Do they just

1:16:14 somehow add their data to one of the big existing models. What do you think about A proliferation of models beyond the top Three, four of them. I think this data is very, very valuable.

1:16:26 But it can be learned later. And the reason I'm so confident. That's okay to learn this stuff later. Because humans do it. All the time.

1:16:34 If you go to work for, say, a Bloomberg company that has valuable financial data. You come in there having a good understanding of the world around you. But not having the Bloomberg secrets.

1:16:44 And you're able to get caught up pretty fast. I think the same thing will happen for these. AI models. Up an air will train. GPT X.

1:16:52 Will cost ten billion dollars. And then on top of that. Bloomberg will say, Hey. I'm gonna fine-tune that. With my uh

1:17:00 Bloom group oretary data. I'm gonna pay open AI. Ten million for this. And I'm gonna use that fine tuned version. Other companies.

1:17:08 say a semiconductor company that wants to build a semiconductor uh Based AI. For coding chips. Or Some of the investment bank, you name it.

1:17:16 They will fine tune on top. And the second piece of why I believe this is gonna be the future. Th there's not a whole lot of advantage. Training your own thing from scratch. Remember how we talked earlier?

1:17:27 About the pre training phase. When you do things that are tendentially related. But Easy to great. Then you do the RLHS on top of that to really make those capabilities come out.

1:17:37 Pre-training phase is way more expensive. And the RL H F stuff on top. The pre-training is gonna look the same. Regardless of whether you're AI for finance.

1:17:46 AI for law. AI for therapy. There's no sense. In a large company. Doing their own thing from scratch.

1:17:54 Doing the expensive thing again when it's gonna look the same. You just find him. Is there anything else to talk about in hardware? Edge computing or other things in the hardware world that

1:18:04 we haven't talked about that you think are gonna be important. Beyond just the chips and Training and inference. Is there a market for Smaller models.

1:18:12 I think absolutely. I don't think there's a market for medium sized models, though. I think there's a lot of demand for a thing that's super, super low latency. Going to the server and then going back. It's so

1:18:23 Slow, that it causes problems. Much like, for example, speech recognition today. It used to be the case that when you said Hello Google. And then to ask the question. It would send the audio to the data center.

1:18:34 It would transcribe it and it would come back. And the latency there was annoying. Should Google built AS speech recognition model on the device. For better latency.

1:18:43 Now that model is very small. And I do think we're going to see other Very small transformers. Four tasks, like shifting the next autocomplete word. Doing a little bit of drafting to help with your texts.

1:18:54 These things are not gonna be the crazy smart agents we're gonna have in data centers. However, I don't think there's a future for medium sized models. Hundred million am parameter range. That's just too big to run on a device.

1:19:07 If you're gonna run in a data center Might as well use a smarter model. And you might say, Well Hundred billion is going to be cheaper. But I think there are enormous economies of scale here.

1:19:17 Because The cost of Loading into those weights. So much higher than the cost of doing marginal extra computation. It is really easy to kinda get added to that batch.

1:19:26 So I think we're going to see Promising future for Edge. And a promising future for the biggest of the big and these hyperscale data centers. And nothing in between. Kevin, I look forward to, I don't know, hopefully decades of learning from you about all the stuff that's happening in this world. It does seem to just be the most, certainly for me, the most

1:19:44 interesting, entertaining, impactful. incredible explosion of technology that I've come across. And it's so cool that you're gonna be driving the hardware side forward as folks like Sam and others drive the software side forward. Always ask the same traditional closing question at the end of my interviews. What is the kindest thing that anyone's ever done for you? The kindest thing anyone's ever done for me.

1:20:05 This guy Gunnarmine. When I was a kid in uh middle and high school. I was very interested in engineering. But

1:20:12 Yeah. young. It's hard to go turn that into a Impactful results. Gunar. Dropped out of college in Germany.

1:20:19 Went to work for Microsoft for uh Couple of decades. And they decided that. Instead of working for Microsoft. He wanted to go teach high school.

1:20:27 Which is quite a pivot. He went and taught programming classes at my high school. Which I token they were good. But really. He stated.

1:20:35 Very, very late. To run the robotics program. To run. Designier stuff. Gunnar thought.

1:20:41 by a bunch of high school they're should build a nuclear fuser. Do turn it Deuterium into a Chilean. Real nuclear fusion.

1:20:49 It would generate power. These things haven't built before. But do you thought that'd be a good thing. much of high schoolers could do it was a little bit uh crazy. Gunnar mine.

1:20:57 Side of that. Uh my school needed a electron microscope. We didn't have the uh money to buy one. But he saw A

1:21:04 One thousand dollar broken. Joel eighty seven machine. That was a user for parks being sold. Yeah, Everett. He and some of my friends and myself drove over in a U Hall.

1:21:15 Picked up the machine. which weighed several thousand pounds, drove it back to the school. The next one Year and a half fixing it. We got it working actually.

1:21:24 It was a very old microscope. Had a Polaroid camera, if you can uh believe that. It would flash pictures one at a time. As it scanned across the image. Because each will be exposed.

1:21:34 They'd slowly be turned into the full image. And We ripped that out, replaced it with an Arduino. We Tried to go window.

1:21:40 Rip out the high voltage box, replace it with a newer one because it kinda broke sometimes. We got that working. We never figured out how to keep it from leaking. But that's what we had. The whole microscope would take a About thirty minutes.

1:21:52 Yeah, to go to the room. Turn the machine on, wait for the vacuum to be drawn, turn the diffusion pump on, wait for that to be pumped out. And it had to be done totally manually. Because the automated box had been broken.

1:22:04 Gunnar would go sit there and supervise us. And we messed around with that thing until like nine PM at night. Almost every night. And that Got me interested in engineering.

1:22:15 Got me to Harvard in some ways. And uh Got me to where I am right now. What an awesome Story. So incredibly unique.

1:22:23 What a cool guy. So many stories like this of someone taking an interest in teaching and sharing the world with someone so young. Amazing closing story. Gavin, thank you so much for your time. Yeah. My pleasure.

1:22:35 Um If you enjoyed this episode, check out joincolossis.com. There you'll find every episode of this podcast complete with transcripts, show notes, and resources to keep learning. You can also sign up for our newsletter, Colossus Weekly, where we condense episodes to the big ideas, quotations, and more, as well as share the best content we find on the internet every week.