Transcript
Etched - Building AI Hardware to Make Inference Faster and Cheaper - [Invest Like the Best, EP.480]
0:02 This show is an open ended exploration of markets, ideas, stories, and strategies that will help you better invest both your time and your money. If you enjoy these conversations and want to go deeper, check out Colossus, our quarterly publication with in-depth profiles of the people shaping business and investing. You can find Colossus along with all of our podcasts at Colossus.com. Trick O'Shaughnessy is the CEO of Postive Sum. All opinions expressed by Patrick and podcast guests are solely their own opinions and do not reflect the opinion of Positive Some. This podcast is for informational purposes only and should not be relied upon as a basis for investment decisions. Clients of Positive Sum may maintain positions in the securities discussed in this podcast. To learn more, visit psum.vc.
0:51 Mm. My guests today are Gavin Uberti and Rob Walken, the founders of Edged. A few years ago, when they set out to build a better AI chip than the largest companies in the world, almost everyone I called told me it could not be done. They've since done it, taping out a working chip on their first attempt and becoming the first hardware company founded after Chat GPT to do so. There you have more than a billion dollars of customer demand for their first product and have raised eight hundred million dollars to build it.
1:18 Edge to build chips and systems designed to run AI models faster and at lower cost. They started the company in twenty twenty three and the product is a complete rack for inference. The chip along with the boards, the power delivery, the interconnects, and all the manufacturing to produce it all. We talk about the technical bets behind their architecture, how they hired industry legends and paired them with elite twenty two year olds, and why they believe inference will become one of the largest markets in the world. I think you will find the story of what they have built hard to forget. Please enjoy my conversation with Gavin and Rob.
1:50 Alright, gentlemen, it's been Three years or so, Gavin, since you and I last did this, which is nuts. And at the time I was just wildly intrigued by your story and what you were gonna build. I didn't know a lot about chips at the time. I was considering investing in the company and so I was calling everyone I couldn't conceive of that could give me an opinion or something. And at the time Basically the consensus was These kinds of companies are not built by young people.
2:14 That The semis world the best companies are founded by forty, fifty year old people that have had a whole career's worth of experience, have learned all the problems, have shipped multiple chips. Two twenty one year olds are like not gonna do this. It's just not gonna work. It was indicative of a theme, which was nobody believes in us. That's obviously changed a lot now. You know, just walk the halls and talk to the people that have chosen to come work here, but
2:36 In the early days it felt like this was something that you had to face down. What was that like? facing that down where a a set of incumbents and an industry were the people and investors and everyone else sort of didn't believe in you. And what did that anneal in you to build the company the way that you are. Like what was the impact of that?
2:53 I think there's a certain level of naivety required. A chip better than every other AI chip ever built. And build a company to do it way faster than ever has been done. And we have the naivety.
3:06 But there's many times where we would say, like, why isn't this possible? And you know, really push on it. And it turns out that like everybody's answers are extremely siloed to a set of constraints that aren't true anymore. And the reality is the entire semiconductors and data center industry. Is built on Buffer.
3:22 And what I mean by that is every part of the stack, from the EDA tools to the power modules, to the circuit boards, to the chip design and standard cell, everything is built to be general purpose. for everything, not just in the data center, but IoT on the edge and so forth. And when you have a specific use case you're really trying to design for. You can change the constraints a lot. And I'll give you just a very simple example that we're not the only one who does.
3:46 Which is One of the things you care a lot about is the clock speed of your chip. It's proportional to the throughput of your system. When you are doing sign off for different timing, what clock speed you're actually gonna be able to run on when you tape out your chip. There's this concept called corners, which is, you know, what temperatures are you going to be able to run at this clock speed?
4:02 The default configurations for a lot of these EDA tools Assume that you're going to be running your chips in freezing temperatures. Now, I don't know about you, but I've never seen an AI data center with ice in it. So you know, we can feel pretty confident that our ships don't need to run at full speed at zero degrees Celsius. In fact, like they're never really gonna be running below eight degrees Celsius anyway.
4:20 And just by knowing that that's a constraint that doesn't matter, we can make a ton of changes throughout the entire system. That's like a very simple one, but there's many more that you get twenty percent here, fifty percent there, two X here, and these compound to a system that can be radically better for inference. I think you found two kinds of people. There are some folks who went purely on heuristics. Okay. Young founders.
4:40 They claim they can go beat. The biggest company in the world on performance. It cannot happen. And there is no thing you could go say to me. That would make me change my mind.
4:48 But there's also people out there who are of course skeptical. I'll spend the time, I'll do the work. And This is actually possible.
4:57 Like for example. One of her earliest, earliest supporters. with Mark Ross. And Mark. Was I
5:03 Very prestigious City Conductor expert. He used to be CTO of Cypress Semi, that sold for nine billion dollars. When we met him. We're just a couple of guys in a door. And we came to him and say, Hey, you want to go build
5:16 Hardware for inference. Maybe we can be much faster than Nydia. And Mark's like, No you can't, it will not work. But
5:24 Should you write a white paper, should you go ahead and build a functional simulation. And uh show me. So after a lot of very long nights. Went back to Mark and said, Hey Here's a simulation.
5:33 What do you think? And he was like, Huh? This works. But to go to a company like this, you'll need a large amount of capital. At least three million dollars even to get up.
5:42 You started. end up went ahead and raised five, raised a lot more after that. And then he was Again surprised. But got more involved.
5:51 And then he became an advisor. A half time advisor and eventually or full time CTO. They saw more and more of the development progress. And I think in general. This wasn't as filtered really heavily.
6:01 Fox you. Want to go ahead and be right regardless. Or she wanted to be a very truth seeking. And say, Sure, I'm skeptical. But I will go ahead and work through the numbers myself.
6:11 And if I can go figure out why this is possible. Well. Let's go build it. The specifics that you've made bets on. The way that you built this system are immensely interesting to me.
6:22 And because so many people are trying to do this now, build new chips that will do a better job of serving inference at massive scale. The world is interested in the research approaches the different architecture approaches that people are taking. to building a new AI chip. And I'd love you to just
6:39 Start by Describing what this thing is, what it does. But maybe more interestingly and more importantly the process that you went through to decide what That's to take what technologies to invent.
6:51 And compare and contrast those with What you've seen the rest of the marketplace try to do. Yeah, I think that Do you start with the product? We're not just building a chip.
7:01 We're building a full inference solution. And that means Iraq. That means the chip that means the power delivery into the chip. That means the board on which it sits that means the interconnect or the chip talk to each other. That means the production.
7:13 for this mass volume of wrecks. Really the production is the product. We think about how we get our advantage. There are two key parts of running inference. There's prefill?
7:22 And there's decode. Now we have two key tech match in both of these things. Prefill. And decode is then using that data. Generate output tokens.
7:31 When you go out and run a prefill, your key job is not to go predict tokens. You already know the text. Your job is to go and get the model's memory. What we call as KV cache into the right state. Then you can go ahead and run decode. with that same K F cache.
7:45 So we will often go do is we call PD disaggregation. Pre-filled decode disag. You'll have one cluster of surfers. Running these prefills. Well then transfer those model memories, those K D caches?
7:56 Over to the decode cluster. And then go ahead and uh use that cluster. The good turn the next tokens. So sort of like loading the gun and then firing it. Like if I think about it in super simple terms. Yeah. You got it. Yes.
8:06 Getting the model to remember the right things. And then using those things to go do tasks. Kann the people think about this market a bit lazily, or they say, are you prefill chip, are you deco chip, if you're decode chip, are you an HPM chip, are you an SRAM chip, are you three D D RAM chip? Are you using optics, using copper? When we started this, we just wanted to understand why extremely smart people were working on these different directions. We seriously looked at architectures like having a bunch of DDR m memory and like a shared memory pool and looking at advanced packaging to basically break out of the shoreline. We looked at things like is there ways to put memory dies on top of compute dies? In doing so, we realized that there's no free lunch. Everything has a trade off, right?
8:41 Three DRAM, you have a thermal issue, you have a supply chain issue, you have to figure out hybrid bonding, you have to figure out the flops, so now you're a decode chip. So we went through everything. both on the pre-fill and the decode side. In doing so, we realized there's a few design spaces that nobody had seriously tried to explore. Because they were never done in AI chips. And we asked ourselves, what are the actual metrics that are going to matter the most on The prefill side
9:03 The thing that matters is flops. And Flops Density. And people talk about flops often as a headline number, but in reality You should care about the flops you're getting when you're running real workloads. There's this concept called MFU or model flops utilization.
9:16 Which is, you know, for every peak flop advertised, how many cents on the dollar are you actually getting? And on GPUs You often get somewhere between twenty and fifty percent, depending on the workload. And actually You can provably not run at a hundred percent.
9:30 Because you have a thermal issue. Where as you increase the flop utilization You have more transistors going on and off. You draw more power. And the chip will self regulate and actually lower its clock speed to make sure it doesn't overheat.
9:41 So As we looked at inference, we said if we want way more flops. Because we want to run at way higher throughputs. We fundamentally need to solve the thermal problem before we even think about adding flops to the chip. If I just add more flops to a GPU today or another AI chip, I'm not actually gonna get more performance because it's just gonna thermal throttle. So Fundamentally, the essence of that is this concept of denard scaling.
10:02 which is voltage is quadratically proportional to power. So if I two X my voltage, my power goes up by four X. If I cut my voltage in half, my power goes down by a quarter. So we asked ourselves How could we run voltages lower than GPUs? And we talked to a lot of people about this. We flew out to s to Silicon Valley after dropping out and basically asked dozens of people and semiconductors and all these different chip companies how they did it.
10:23 And the answer we got was like you can't. You can't run at voltages lower than GPUs. And this was very dissatisfying because there were many different industries of chips that run at voltages lower than GPUs. Bitcoin miners run at under a quarter of the voltage of GPUs. So this is Obviously physically possible. The question is, are there issues with GPU architectures that make it unable to run at these voltages? And when we looked at the problem for a long time, we were able to create a new mechanism.
10:46 of running at much lower voltages a new type of power delivery that we call low voltage inference. And we think all AI chips in the future are gonna be low voltage chips. They're gonna have to cram way more flops in the same silicon area and without thermal throttling run at way lower voltages. Four decode. It is all a memory game. More memory bound.
11:06 You can load the model faster, load the K D cache faster. And serve more tokens per second per user. We think people ask the wrong question here. People often ask. How much memory bandwidth is on your chip.
11:17 Should be asking how much of a memory bandwidth is on your full scale up cluster. is add way, way more bandwidth. And a much lower latency from chip to chip. to our interconnects.
11:29 Allows us to be able to go serve. models at this much higher speed because you can go use the S RAM. An HBM from the full scale up cluster as a single pool. And that's our second key technical bet. what we call cluster scale memory.
11:42 And on TPUs today. The cluster memory bandwidth is often very badly utilized. Because the time to go hop from one GPU to another It's extremely long. For example, on black bell chips.
11:53 It can be about four thousand nanoseconds to go point to point. And that means that we're going to be able to do If you go ahead and go to an eight X T P setup. You will get way less. than an eight ex improvement.
12:03 And your tokens per second per user. And what we did is built our own totally custom interconnect stack. When you took Everything above the second layer Ethernet. Built a forecast'em.
12:13 that we can go out and do far, far better latencies and advanced this way too. We can go ahead and cut this by more than a factor of five X. And that allows us to then use the memory of other chips much more effectively. As you scale the world size.
12:27 You're a time for Token. Go down proportionally. Yeah. That's not that surprising given all these architectures were built before ChatGPT.
12:34 So if we're trying to build a chip for the modern workloads, it's gonna look very different. The way we, you know, organize our flops, the way we do our voltage domains, the way we do our power planes are gonna look super different, the way we do the packaging is gonna look super different, the way we do the board design is gonna look different. And then the decode side, the way we connect everything is going to look very different. So we're Now bringing forward our first generation of this low voltage inference technology, which is running at under half the voltage of any other AI chip. If you zoom all the way out, why is this so important? Like why is
13:04 The delivery. of much higher throughput, much lower cost per token. better tokens per watt, like all of these metrics that the universe is gonna start talking about more and more. Everyone knows the supply side of the equation is a big problem right now. Why is this in the bigger picture looking at a decade? the bottleneck in the technology world. Well, I think it comes down to productivity.
13:23 Where We are At this extremely interesting moment in the history of civilization. Where there is real artificial intelligence. Найси фі став, бо лайт з мод.
13:33 Can solve problems that Most humans can't, and like it's going to create new scientific discoveries, it's going to create instant access to medical care, instant access to education. And now it's just about how many people can use this at the same time, how many products can serve this at the same time. And also the speed of doing different tasks. So When you think about wall clock time. If we can take an agent that, you know, can run at a certain model quality.
13:55 And it could take a year to solve a certain task using inference time compute. If you have way faster di code sped. You can compress that into a month. So the amount of scientific and the amount of actual proliferation of technology will happen much faster. And then the second part is concurrency.
14:10 Where today It's just not possible for a billion people to use these models concurrently. Ultimately, some people are gonna get downgraded, some people's models are gonna be slower, some people just won't be able to access the hardware. A few years from now, there's going to be giant models serving billions of users. We're very much in the early innings of AI today, where the paid plans there's only a few million users in the world using paid plans at AI models. We're at One one thousandth of the global population actually using this stuff.
14:33 So if you want to serve at giant scale. A lot of things change and one of them is the number of chips that communicate together. Where people usually think about this in the context of training. You know, you have these giant training clusters, you have Colossus with over 100,000 GPUs that are all networked together. And the inference side, today, people usually think about it as an eight chip cluster, or maybe just NVL 72 as a scale up domain. But very quickly this is going to become thousands of chips and tens of thousands of chips. And the way to get the most performance there. The time between sending data from one chip to another
15:03 That primitive matters way more than is getting credit right now. So when we think about optimizing memory bandwidth for the system. you have to think about how fast these chip can communicate together, because if they can only communicate really quickly with themselves and very slowly with other chips, you're not going to actually be able to serve giant models that ten thousand, twenty thousand tokens per second. So we need multiple orders of magnitude of infrastructure built out.
15:25 throughout the entire stack, from the wafer to the watt, transistor to the token. To actually bring this stuff to the world. Yeah, I think that's You look at most other goods. Like the iPhone for example.
15:35 They've gotten this economies of scale. Or as a result. More money does not really buy a better iPhone. That if you're a billionaire or if you're Just The average American.
15:44 You buy the same phone. And tokens aren't like that yet. We're still in the very early days. Where A general purpose system.
15:52 relatively small one, is kinda handcrafting each tokens. Like they made screws. Back in late the Renaissance. And I wanted to live in the world. Where you have the same economies of scale for token making. that you do for making, say, iPhones or cars or anything else.
16:06 I think that is one of the huge unlocks. It allows A huge group of people to go use The best quality models. economies of scale have all made capitalism very
16:15 I don't know. Fair? I think that allows you to go ahead and have the same product in many, many different hands. And You're able to go then serve. Way more used on a single scale up cluster.
16:24 Allows you to get closer to that point for cocoa serving, too. Yeah. And also just Certain products aren't usable if they're slow. So if you want to serve coding models and you want people to actually use them, like there's a certain number of tokens per second you need to hit. So the question is, while maintaining that
16:39 Per token speed. How many users can I serve at the same time? And you can basically decide, I'm gonna shut off a bunch of the world from using this stuff, or everyone's gonna get a worse experience. So fundamentally you need to find ways to push out the curve and that's why there's such a pressure for new hardware. I'd like to take some time to step back and hear
16:56 Both of your stories for how you came to this idea in this company. And then kind of walk through what it's been like to build it. Because I think in so doing will understand the system that you've built for for the company itself. that will then be able to power subsequent generations of products like this one. for this crazy inference future that that we're staring down.
17:15 Rob, maybe starting with you, just Take it however far back you want. But but what what I'm curious about And your personal story was Actually the very first thing I ever heard from either one of you was your personal story. many years ago now, which really kinda blew me away. I'm most interested in your motivation ultimately for being here doing this thing.
17:33 It starts back actually in high school for me. I've been very unlucky and lucky at different points in life. This was one of the tougher times. At the end of my sophomore year of high school, I got injured at a martial arts tournament. Next day couldn't walk for some reason. I thought it was something wrong with like my SI join or something. I went through physical therapy, I did different types of scans, I couldn't figure it out.
17:52 And eventually they found this big bump on my back. And an MRI and you know told me it was a tumor. Stage four bone cancer was told I had under thirty percent chance of survival. It was like a two year crazy chemotherapy, surgery, learning to walk again, experience. I mean you go through something like that, it changes the overton window of human experience and makes you appreciate like what actually matters.
18:13 And you also ask yourself What are you gonna do if you have the chance to live? If you actually want to get through something like that, you need to be Hoping for something. And I always knew I wanted to do something very impactful if I had the chance to get through it.
18:24 And it took me a couple years to figure out what that was going to be. And at the same time as I got to college and Met a bunch of other people building cool tech. I got extremely excited by AI models, especially once GPT three came out. And I was like, wow, this is the first model that can kind of speak English. And these things are gonna get really smart. What happened was when GPT four came out.
18:43 There was GPT four V, which was the first model with image uploading. So I went through my camera roll. And I found a picture of my back with this bump on it before I was diagnosed. And I said, Hey chat GPT, pretend you're an expert doctor, a patient comes in. And it says they have this bump on their back. What could it be? And it immediately says like
19:01 This could be a tumor, you should get an MRI immediately, go to the doctor. And I just kinda sat there still. It was like That took me six months. And yeah Yesterday this feature wasn't there. Today it's here.
19:11 I go to show my parents And I got this like notification being like You're all out of image credits today. Like you need to get a pro plan. And I was like, Holy crap, like this is gonna change everything. And We clearly don't have the infrastructure to serve it. And there's very few things you can work on that can actually bring this technology
19:28 At scale to the world faster. I mean there's like plenty of people that are super smart working on models and the fabs seem maybe unreachable to work on, but It seemed like the hardware was all designed before Chat GPT. Every GPU, every TPU, every AI chip that was serving these models
19:45 before this and are retrofit to serve these modern models There's gonna be an entire new wave of architectures that came out and What a more exciting thing to work on than bringing this to everybody. Very different angle at the same time. I was running a startup incubator called PROD. Which has incubated a bunch of different companies. Some of the earliest ones being Cursor and Any Sphere, which merged and Mercore and Etch went through it, a handful of others. At the time, as these models were getting smarter, it was yeah, twenty twenty two.
20:09 I was realising all of these companies are spending all the money they raise on compute. And I had this realization as I was working on some of my own stuff that like, oh my God, all the products I want to build. are gonna cost tens of millions of dollars a year in inference. Hey, this is not going to be tenable, like the cost structure of every software company And the COGS is not gonna be like zero anymore for an incremental user. It's gonna be like quite high and it's gonna be a function of inference. And then the opex of every business is gonna also be inference as people use more and more coding agents.
20:36 So fundamentally, it seems like inference is gonna be really important. And it feels like we are on a decade and march for inference to become, you know, the biggest market in the world. So when you think about that, ten years from now There's gonna be these giant projects.
20:48 where everything in that data center fundamentally hasn't been designed today. You should go pick something and work on it. That's kinda how it got started. Yeah, but I'm really excited for you to go back about as far, probably early in high school, maybe even earlier. and tell your favorite hash marks on the timeline that ultimately led to your ambition to drop out of Harvard and start this company.
21:06 My first job ever. Was that a company called XNOR? Or I did kernel development. I was seventeen. And the seventeen year old can't sign illegally landed contracts.
21:15 So rather than go ahead and do a traditional CIA. They went ahead and sat me down and said Gavin. Don't share this information. Only companies that saw hey, maybe this is a good trade and
21:26 Yeah, we're building kernels. And Number of other companies since. Two hundred million. Did the same thing at Octo where they got bought by NVIDIA for
21:35 hundreds of millions of dollars. But when you do this sort of kernel s work. What you realised is that The math is relatively easy. But to get high speed decode, the thing that matters is data movement.
21:45 Almost all the work that you do. Optimizing. How do you move data? Around a single chip? Or across multiple chips.
21:52 Yeah, that's why we went ahead and built this cluster scale memory tech. We bring that. Interconnect time way, way lower. You can go do way more movement. And as a result, get a much faster time.
22:01 to generate each subsequent token. And build these crazy things Rob's talking about. for doing a year's worth of work in a month or more than that in the future. Can you talk about the competitive drive that's evident in some of the high school competitions? That you participated in and won.
22:17 We did a couple. For example, I was uh very active in the FTC Robotics. I'd like you to have a very talented partner, Sanford. For a long time.
22:26 Or was that? Tony Guides. All working together as often typical of first tech challenge. And the goal is to go out and build a robot. That scores the most points.
22:34 And a bunch of other things too. And first. They put a lot of emphasis around collaborating with other teams. around trying to do really good documentation. around trying to go ahead and get others inspired to go do the same thing.
22:45 Same for nitocidin. Rather than go ahead and do it this way, we're gonna win. And we did nothing else besides Bulldog that scored the most points. As a two person team. Rather than a twenty person team. That we were much, much smaller than almost every other team in the competition. We figured that if we were gonna go specialize. for to go out and do this win the damn games really well.
23:04 We wouldn't need to go ahead and advance based on the quality of our documentation or of our outreach. You're just gonna go win. And so we did. We embrusd off built a two person team. uh built a robot and decided we were going to go ahead and redesign it. Every three months.
23:17 And we did. We actually had the war record for the highest score. During this competition at one point. We were rated it by OPR a third in the world. For software development.
23:26 And uh it was a damn good machine. What from that episode can I translate as an analogy onto how you built That's the company. We're thinking about how you want to go ahead and do A full rack scale product like this.
23:38 There's a couple of key ideas. One of them is like velocity, velocity, velocity. That you win by shipping. You're not gonna go out and win by halfhing The best outreach.
23:48 Or the best communications. You were building the best product. Mm. And similarly, we think we can go do it with a lot fewer folks. That
23:56 If you're willing to go ahead and just focus on product, product, product. And parallelized relentlessly. You don't need. twenty thousand people. Like the big companies have.
24:04 You can do The best product in the world. There's far fewer people. Yeah, there's a saying of the best part is no part. I think for us it's also the best vendor is no vendor. As much as possible, we want to vertically integrate the entire product. Both because we get more performance, but we can move way faster.
24:18 So Everything from the chips to the boards to the cold plates to the interconnects to even the production. We want to do all of it as in house as possible. I think we're the only startup right now that's building its own rack as well as its own chips. And we did it all at the same time. A couple years ago is the last time we were public.
24:33 At that point, we just started building our rack team and we brought over Brian Loiler. Who built all of NVIDIA's HGX and DGX systems, which is like eighty percent of their revenue. And we said we're gonna build the rack at the same time. We actually went through multiple iterations of the rack before the chips even came back. Before the chips came back, we made thermal chips that had the exact same hotspots as we expected our chips to have. So we could build the cold plates, we could overpressurize them and blow them up. We haven't had a single leak since our chips came back with the cold plates because we already validated them.
25:01 We have a factory in Taiwan, maybe a few dozen people out there. We built a clone of a bunch of the test stations in our office. We have a two megawatt data center on this floor. And we did twenty four seven development cycles, people are doing day shifts and night shifts to actually get the hardware up and running as quickly as possible. But it's that extreme vertical integration and extreme parallelization of the schedule that lets you get products to market way faster. If you think about the building of the early team. and what it required as two young guys building this company. There's lots of very talented young entrepreneurs out there
25:31 Maybe for the first time of this scope or magnitude in a long time. All of whom probably could benefit from the lessons that you've learned. getting very sophisticated, talented people to come join you, even after careers at the other great companies. If you were teaching this as a class, like here's how to get lead talent when you're young and inexperienced and naive. What would be the syllabus?
25:49 We have a pretty bimodal Talent philosophy. It starts with We call the legends. Which is when we're trying to solve An incredibly hard technical problem.
25:58 And generally do something that hasn't been done before. We need to find the very best person in the world. And often the number one guy in the world versus the number ten guy versus the number hundred guy, huge difference in whether it's actually possible to solve the problem. We created the system we call project based recruiting. Will we map out? All of the hardest technical problems across all industries that anyone has ever had to solve.
26:17 We look at temporality, so who are the people who did the zero to one? Who is in charge, quote unquote, who actually did the work. We talk to as many people as possible and then we just track it. And you'd be surprised by The amount of people who say yes after the first conversation is pretty low.
26:31 But the amount of people who say yes after the twentieth conversation is surprisingly high. You only gotta keep at them. When you hear no from somebody who really is the best in the world. Then that really means. Hey. Should go ahead and come back when you have a few more milestones for given out.
26:44 Convincing things to see. Is hey we make And when you go ahead and hit those again and again and again. That is really belief inspiring.
26:53 When we decided we wanted to build a rack and not just a chip. We were looking at this and we're saying Yeah, how many products have actually shipped that scale? for a rack scale system that actually have the power density that we're trying to solve. And we just kind of said if we were gonna wave a magic wand.
27:08 What would the best possible person in the world look like? And be like well. If we could find somebody who like started at NVIDIA And built the entire rack team. Through all their different generations, learn all this different stuff, but is still scrappy, still understands the start of culture, but has seen scale, like that would be the best possible person.
27:26 We've mapped all of the different teams that are related to all of the different rack scale products of media. And we found three people. That we thought Could fit the bill.
27:34 And we talk to all of them. And two of them have just retired and one of them was planning to do one more generation for Nvidia and and then retire. His name's Brian. And over time we convince them to join. Brian started the HGX and DGX team at NVIDIA.
27:48 Which was yeah, a majority of NVIDIA's revenue. Tens of billions of dollars a quarter. And the other two guys ended up investing, by the way. But when you have somebody like that, they just know what good looks like, and'cause they've they've seen it. And There's so many times where we would talk to Brian and he'd just point to us and be like That's a billion dollar like a billion dollar lesson I learned. Billion dollar lesson I learned.
28:06 Like, you know, that just saves us cycles. And you pair someone like Brian With somebody like Sanford. Do you have name for them? So Brian's a legend. What's what's standard? Yeah, we say chips on shoulders put uh chips in data centers. Yeah. Sanford and Gavin in in high school were world robotics champions and Sanford was finishing his senior year of college. And we called them up a couple of years ago and we said, Hey, can you come check out what we're doing? We need some help on the platform side.
28:31 He comes for a week and we say Can you build a cold plate this week? And if you asked like any thermal engineer anything like that, they they would think you're just like totally naive, right? I mean these things take months to do. And like to be clear, they do. But you can make real progress in a week if you if you put your mind to it and you think it's possible. And he built a contraption in a week that like
28:51 actually de-risked like a pretty key power question we had. And you put those two together and they've done incredible things. One is not possible without the other because you need The extremely driven people that just keep asking why and don't know. Where the bodies are buried. to like take tons of aggressive risks and then you need the people who've seen scale and still have the startup scrappy mentality to help them along the way. So it's really the legends plus some naivete raw first principles type talent. It's not just that you have both in the company, it's that they're working together.
29:19 That's right. If I think about that funnel, anything else more interesting to say about how much better you've gotten at recruiting and like why those metrics keep getting better. One of the shocking things is I'm sure being such a contrarian bet. Kinda self selects, right? But like you're the kind of person who is Some arbitraristic. Shouldn't you go join whatever the hot company is?
29:39 Or I think go ahead and do due diligence. You will not come work here. And it's one of the things I worry about as we announce more and more of the product in the specs. We may lose some of this if we're not very careful. You kinda have to be sick in the head to join our company. When you think about it on paper, it's like
29:53 You A person who is probably a very accomplished engineer, making a good amount of money. It's liquid, it's predictable somewhere else. You're going to convince your family to move to San Jose and live in this apartment on this housing program. For the semi conductor company run by two what twenty four year olds now
30:09 That's pre product. That is going against the biggest companies in the world. in the most supply constrained environment ever created, with a design that they're saying is not going to be like ten percent better, but it's gonna be ten x better. Something must be wrong with you to do that. People are just wired differently here that they like really want to not prove people wrong who don't believe, but prove people right who do believe. They just take it personally. And you know, that's really fun to find those people and
30:35 frankly, just the nature of the company makes it very easy to whittle out the people who aren't like that. One of the very first things you and I talked about, Rob, was I started asking about Sohu, which is the name of the first product here. And You said we can talk about that in great detail, but The thing you should know is that what we're really focused on is building a machine.
30:52 that can At scale. produce these things and generations of them as efficiently at the highest possible quality level. So we want to build like the company or the machine that is the company is the thing that will produce this thing and then subsequent things. So I'd like to talk about a few principles or cornerstones of the company.
31:10 We've alluded to some of them. You've said velocity, you've said vertical integration. have become more popular topics. Parallelization is something maybe that we should talk about. But I'm especially interested in your guys' willingness to take huge risk to go faster. Maybe tell your favorite story about
31:26 Why this is the philosophy. what it's allowed you to do that maybe other companies haven't done There's a number of stories here. But uh one of my favorites.
31:34 Is There was a time. Where we were getting close to taping out the chip. We realized. Wait a minute.
31:40 One of our vendors. is a way way behind schedule. And we have two very bad options. One option. Just to keep the current vendor. And push up downlines out by on the order of a year.
31:50 Another option. Switch vendors. Start over. And also push time lines up by a year. Neither of these was a good option.
31:57 So we had to go look for option number three. And what that was was we we figured out they're all in Bangalore. I think they're actually going and doing the work. We went out and shipped a dozen of our top engineers across the world. to Bangalore.
32:10 For six months. I was there as well. I lived in Bangalore four and a half months personally. And every morning. We'd go ahead and walk across the crazy busy Bangalore streets into the office.
32:19 We'd be the first ones in. We go out and built a wide variety of tools. Well things like hey. Auditing a huge amount of the Code that was going in.
32:27 building a bunch of tools as well to make this go even faster. Making sure you're making the right design decisions on the spot right there. No twelve hour back and forth. Go ahead and decide immediately. And then at one a.m we'd go walk back.
32:39 Through the now empty Bangalore streets. And that do it all again the next day. We ran these twelve hour on each side handoffs where we had twenty four hour development cycle. We're at eight AM and eight PM every day. We'd all get on the Zoom. We'd share all the data and we'd say, When I wake up, like this must be done. Like we must get this chip out.
32:59 And it was extremely intense. At the same time we saw other chips at the same stage as us with that same vendor that ended up taking years that still aren't out today. Uh still have even taped out today. And it's that level of extreme urgency that's required. To bring products to market. What is the key to doing this? Well, this has become a trope because of Elon mostly, that like His special skill and others that seek to emulate him.
33:21 Would try to do this too is figure out like what the binding constraint is and just flo the zone personally on that thing, which is kind of like going to Bangalore or something. It seems like this is a central tenet of the business and of any business that's gonna do this kind of vertical integration. What's the key to doing that well? Like Again, what have you learned about That specific act.
33:41 For me, I think there are two key tricks to this. The first one is that You can't build a tip alone. It's gotta be a tea problem. And your most important job is to go get great people to go with you.
33:52 And right people to go ahead and be inspired and excited to go ahead and do crazy things like this. uh it is a huge ask to go say, Hey guys. Uproot your lives for Six months or in one case, twelve months. Uh we had sent one guy out well ahead.
34:05 It sucks. But we're lucky to have team members who are Enough for the right reasons. I think the second big thing too is being able to Make decisions very fast.
34:14 That One of the worst items. is when there's a factory or there's a vendor. Who is waiting for the For you to go ahead and make some call. And it's been just stalled.
34:23 And this happens all the time. Even for very small things. Sure. Sun Folks. delegate a big amount of responsibility to them and say
34:32 Make a reasonable call. Okay if you're wrong every now and then. But I would much, much rather be right most of the time and give an answer immediately. Then wait every time for the perfect response.
34:43 Speed wins. What about spending money? to go faster. There's this learn by doing thing which has become so interesting and as the world has gone away from software and towards More hardware again in the in the world of technology. that we've outsourced so much of the learn by doing.
34:58 By shipping stuff overseas and effectively just being the ideal guys here in the US. Seems like that obviously is reversing and you've adopted this way of learning by doing like you want to be in that iteration learning loop. Absolutely. And and part of that is willingness to spend and take risk with dollars. Yeah can you talk about that a little bit? I think there's a great quote of like the biggest risk is not taking risk. Very similar here, which is like Every day there's over a billion dollars of revenue in this category and a lot of it's inference.
35:24 So every day we don't ship, we're just leaving tons of opportunity on the table. So your willingness to spend money should be extremely high. If you can get a very clear ROI out of it. So we have this concept that we call prefetching. Which is when you're waiting for one thing to get done.
35:39 When you know you're going to do other things once you have it, is there ways that you can parallelize the entire schedule? So for example, like we know our chip is going to come back on a certain date. We want it to be that everything possible. that could be done without the chip is done before the chip lands. And this costs a lot of money. This means that like we want to build our entire software stack beforehand. Like we shipped racks to customer data centers without our chips in them, with all the networking, all the CPUs, all the storage all set up so we could bring all that data center software up before the chips came back. And meant that we took over seven hundred FPGAs and put the entire full reticle chip
36:13 on an FPGA cluster and ran a dozen different models with our full inference stack on them before the chips came back. Means that we built a thermal chip to mock the thermal profile out of our chip. And built cold plates based on that before the chip came back. It means we had the pr entire production line ready. It means we did many revs of the circuit board. It means the entire product was ready to go before the chips came back. And this is what it gives you.
36:35 There is another very famous AI chip company. That took ten months. To go from Getting their silicon back. To having them running inference in Iraq.
36:44 And this was like publicly announced to their investors and it was a really big deal. We were able to do it in forty days. And it's because By the time the chip came back. Everything was boring.
36:54 The software was already written, the rack was already there, the production line was already set up. We were just Go, go, go, go, go. Get everything together, you don't always catch everything, you make some tweaks on the fly, and then off you go. In that particular case too, that was a big part of it. But also like the shift I think made a big difference too. Totally.
37:10 We went out and literally had a day shift and a night shift. There are team members who would come in at ten A. M. And leave at around midnight. And leave at ten AM. You're running around the clock to get to that forty days.
37:23 Yeah, I mean over half the company lives next to the office. So it makes it easier to do that type of thing. You pay them to do that, right? Pay them extra. Do you still do that? The invisible hand does wonders. I mean hey, it works for me too. We're both there. I'd love to take one big step back and talk a bit about just the broader ecosystem here. The amount of shortages On the supply side. The exposure of
37:47 risks in the global system and the supply chain around this stuff. has become like everyday Wall Street Journal front page news. Like the stocks that people are watching and investing in and excited about. About the memory stocks. These were, you know, boring commodity like nothing burgers five years ago. And now they're the center of global attention.
38:05 If you just assess because you've been building in it. the global connected supply chain that's required to make stuff like this possible. Just riff on it. Like what scares you? What's working well? What needs to change? What do you hope you change by virtue of how you build this thing? What's your assessment of this story right now? I think that one of the most
38:24 undervalued pieces of the supply chain story. Yeah. Almost none of these things are You buy them And they don't talk to the vendor again.
38:33 You have to go collaborate. That is the most important part. Being successful, I think, in ships. With. P S M C or with memory vendors. You need that partnership.
38:42 I think that for T SMC in particular. People don't understand why it is so valuable. You look at the tech, and the tech is the best in the world. But For me, the real value is all in the service.
38:53 T SMC customer service. There's way. Way better. And I have seen uh any other company in any other industry. That's the kind of thing where if you say, Hey
39:02 But you gonna prove your yield by making this change. You can go make them a recommendation. And then we'll go run an experiment. On their own time in our case.
39:11 See if they could actually Get the higher yield. And When we found that we were right and the CM worked. They moved over the rest of the line.
39:18 And that kind of thing just doesn't happen in most industries. If I go to like the steel works plan. Say, Hey, I want you to change the composition of the steel. And they'll say Screw you. Not T SMC. It is why they are the number one and why they're going to win.
39:31 One of the things that matters a ton is power availability. And time to power. And the problem is the more power you want. the more shortage there is. It's actually very similar to chip clusters, where it's just like
39:42 Why is Colossus charging twelve dollars an hour for black holes? It's because they're the only place you can buy twenty thousand of them at once, right? Why is the 500 megawatt data centers are hard to find. It's the exact same reason. One of the things we need to think about is like how do we get way more juice. out of each megawatt. People are looking throughout the entire stack, whether it's just improving the PUE. but also entirely new hardware to get the most tokens per megawatt to solve this problem.
40:06 Building new buildings is hard. It's much easier to go from a hundred megawatt to a gigawatt th from a gigawatt to ten and ten to a hundred and We are pushing the limits of what's possible on these timelines. So there's a lot of people trying to scale in their data centers as much as trying to scale them out.
40:21 One of the interesting things about a system like this is what it replaces. Yeah. If I think about Iraq like this versus I don't know, a set of black wells or something or Rubens or or whatever's coming next. How should I conceptualize that? It's not just watts, it's also physical space. The rebers talked about this in their recent reviews call that literally this is a a problem that there's literally no space to put the systems. How should I conceptualize what this represents or replaces in terms of other units of compute.
40:46 Here's how customers think about deploying models generally. Which is When I'm building a data center or I'm building a cluster, it's not like in the abstract of like, oh, I like these chips and like yeah, this is the power footprint and so forth. It's like I have a real production workload I'm trying to surf. And for my product to be useful, there's a certain speed I need to serve for that.
41:05 And for certain products it's really fast and certain products it's really slow. Whatever the speed is. This is my speed. The question is In a given amount of power. How many users can I serve while guaranteeing that speed?
41:16 So another way to put it is ISO what's called interactivity. What is my throughput? We are just finishing kind of the early innings of of the AI infrastructure boom where people really just cared about speed. You know, GPUs were not able to reach a lot of the speeds of other types of chips, like all these S RAM chips. thousands of tokens per second, and that enabled
41:33 Tons of new use cases that got people very excited. There's an entirely new wave of AI chips. us being one of them, they're all going to be able to hit these speeds. The question then is, if you're hitting these speeds What is the number of users you can serve at the same time? And by proxy, if I have a hundred megawatt data center, how many software agents can I run at the same time?
41:51 So when people are doing that evaluation. our hardware is going to generally be able to get you an order of magnitude, more concurrency. at a given level of interactivity. So that directly translates into Tokens per watt. Tokens per dollar, all the things people care about when they're actually serving these giant mixture of expert models at scale.
42:08 There's these now famous interactivity curves. Not many people publish them. But you could see a Blackwell curve, you could see an A and B curve, which is A little bit worse. than Blackwells and it's still an eight hundred billion dollar company. So if you think about what then the impacts are of shifting that curve, not just a little bit further out, but much further out. Yeah. What are the things that most excite you about what this will enable?
42:31 I mean, I wanna go out and solve some of the hardest problems and I wanna go solve these in much less time. There were things growing up that I was not sure I'd be able to live to see. For example, the unit is uh conjecture. One of the things that we talked about in college. And I was not sure I'd see that proven in my life.
42:46 And this was done by an AI model. And it was done over a long period of time. But if you're able to then run the same model Ten times faster. You can go shrink the time to go have these breakthroughs.
42:56 And there's a huge number of other problems in math like this as well. That I worry it will take a thousand years to go prove a thing like this. You can either have a much smarter model. or a model of the same intelligence running much faster.
43:08 You can then shrink that and I can see it. It's so cool. Seeing these breakthroughs get made. I am so so excited to see much more of this happen. I think too often people think about
43:19 Tasks and applications and stuff in these very short time horizons. Doing a chat and it's like fifty percent faster is nice. But it's not like game changing. As these agents go longer and longer time horizon and the models get more and more capable, you're going to see Gigantic bodies of work that would take months of compute. And we think about this in wall clock time. Like if you talk to a pre-training researcher at a lab,
43:42 They'll tell you that wall clock time often is one of the most important things that matters. And what wall clock time means is the time from starting your run to finishing it to actually get data back. If you can shrink this time from a six month run to a two month experiment, you're gonna be able to do many more iterations and people will make changes on the model architectures to actually improve the wall clock time. Very similar here in terms of how we think about the use cases, which is
44:05 The exciting part about super low latency decode is wall clock time on long horizon tasks becomes much shorter. So a year long compute build would now take a month, and that month long compute build will now take three days, and that three day compute build will now take seven hours, and so forth and so forth. So that's the thing that I think is really hard to internalize because the models are just getting capable enough to do this stuff. I thought it was really cool months ago. When
44:31 cursor publish that they had a bunch of coding agents build an entire browser from scratch in a week. Totally nuts. And that that will soon happen in under an hour. And there's going to be many of those types of things. that are going to happen with these massive parallel agents all working on a given task. What are the ultimate limitations of these systems? Is it just like a physics question, like How many
44:51 times faster, cheaper Can we get it? Theoretically. There's a lot of room at the bottom, as they say. If you think about chip and chip latencies. On NVIDIA product.
45:01 you're looking at four thousand nanoseconds to go from one tip to another. We'll be able to do much better than that. What's the mathematical limit? You can do it in Just a handful, like two, three nanoseconds.
45:12 And they're four thousand. Four thousand today. There is A lot of room at the bottom. And same thing with things like power efficiency.
45:19 That's Sure, we're able to go and save a huge amount by bringing the voltage down by so much. But you could go lower. You could go much, much lower. It's very challenging.
45:28 But when I think about it twenty, thirty years in the future. Then I think it's inevitable. And also for uh economies of scale, for cluster scale up.
45:37 For a long time. Eight chips was the biggest uh scale up domain. And then even had to in L seventy two, bringing it to Well seventy two. But you can be way, way bigger. You look at like a fab, for example.
45:48 You have a forty billion dollar single monolithic building. With Only a handful of lines running through it. You could have the same kind of thing for some futuristic mega cluster.
45:59 Forty billion dollars, hundred billion dollars. As a giant megatope factory. Serving one or a handful of models. For a massive number of users. Do get that saying the economy is a scale thing.
46:10 Same model. Massive number of people. You you mentioned kernels engineering and that being your first job. That has emerged as a thing that nobody had ever heard of in their lives to now something they hear about all the time, the importance of it to eke more performance out of the bare raw metal. When will that just be something that AI doesn't entirely as well? Are humans still the best
46:31 Kernel engineers, are they doing it? with the assistance of AI systems like How far down will humans still be in the loop of designing these things? Like when will that go away? Today it's all very hybrid.
46:44 And the best kernels are still written by human AI collaborations. Well, so any AML has built up these fundamental primitives, like Matmoles, like convolutions. Like a chip the chip operations collectives. And making these overlap and making these really fast matters enormously.
47:00 And it's the kernel designer's job to go ahead and figure out Or can I overlap? How do I allocate memory? How do I verify you? If there's some issue like a retransmitted.
47:08 It doesn't stall the whole pipeline. These things are very challenging. But they can go and make your allot performance be say three, four percent better per optimization. You can do so many.
47:18 And we thought about our software stack. We want to negotiate it where the puck is going to be. And Three years ago. There were kind of two ways you could build software.
47:25 One of them was to invest heavily in graph compilers. These things are not very performant. But they work out of the box. I don't require a human to go come in and tweak all the kernels. But we went the opposite direction.
47:37 We are Colonel's first programming. It for a long time did not work out of the box. But if you were Colonel's expert. you could get to incredibly, incredibly high performance.
47:47 And the thing about this is that now How do the coding models get better and better? They're doing more and more of the kernel generation task. And when the models keep getting smarter. It'll eventually do all of it. It will become superhuman.
47:59 So we're gonna build where the world is going. And even today. think about our profiling tools. Or a debugging stack. We think about it from the first time.
48:07 How will the model use these tools? more than me think about how will humans use these tools. We sometimes run experiments internally. And we had codex actually get GPT OSS running from scratch just based off of our docs. Completely by itself.
48:21 And they did it I think overnight. We think about game selection a lot, and what we mean by that is Making sure we're investing our energy in the right bets. Because regardless of what you choose to work on, it will take tremendous effort. And one of the things that we started with was, you know, the decision explicitly not to build an arbitrary graph compiler.
48:39 Not to support arbitrary PyTorch, not to support arbitrary CUDA, not to support arbitrary onyx graphs. But instead We envisioned a world where there was going to be under a hundred models. That actually mattered. And they were all going to look
48:52 Very similar from the underlying mathematical perspective. And that we were going to build primitives using physics that we're going to accelerate these as much as humanly possible. And we're going to allow the most sophisticated customers to have direct access to the hardware and do whatever they want. And that has saved us a tremendous amount of time not having to build a compiler. And that has allowed us to actually get much more performance. And funnily enough. When we started
49:13 A lot of people dismiss this idea. And the only people that took us seriously were in high frequency trading. They all hate compilers too. They all write their own kernels. And we've had dozens of people from high frequency trading join the team because they saw this philosophy too. What are the limits to vertical integration? Like how do you know where to draw
49:30 the line. And I'm starting with this question to talk a bit about the broader market. The circumstances of the broader market are really interesting to me where the vast majority of chips of AI chips get bought by a very Small set of customers. Yeah. Many of those customers are themselves trying to design their own AI chips. OpenAI announced jalapeno. It seems like this very funny circumstance for like the most valuable thing in the world.
49:53 All kind of flows through a couple Chip makers, a couple chip buyers, they all seem to be thinking about doing each other's job. And then you've got the circumstance where like, okay, then these things go in a data center and you've got neo clouds and inference providers and this other part of the stack. You've got model builders and providers. Like I can imagine a world where because you have the best hardware. You design models and you build data centers. unique outside of your current Vertical. So like how do you think about where to draw the lines for the business?
50:21 We have a saying that production is the product. Ultimately what matters here is we know inference is gonna be the biggest market in the world. Whoever produces the most tokens is going to be the most valuable company in the world. So all the decisions we make
50:34 is how do we get the most token capacity online as possible. And part of that is Building a really good product that's has way more throughput, that can run at way better latencies and so forth, so we can per chip we make get way more tokens online. Another part of it is like
50:49 Not doing parts of the stack unless we absolutely have to. To get the giant scale. So there are parts that we decided to do because it was absolutely required to get the scale. Like building the rack instead of just building the chips. Like doing a CM model instead of doing a JDM model. But there are parts of it that are kind of noise to us right now. Like we're not going and building our own data centers today. That doesn't actually help us get more capacity online.
51:11 In general, our customers are actually making power and moving their clusters around to get our chips online because they're such high throughput. If there was a world where other things were a constraint, we would totally go and integrate with them, but the reality is we're just purely focused on getting as many tokens online as possible. I think this comes down to economies of scale again. A certain parts of the stack and there are huge economies to scale. And others there aren't.
51:31 For example, on designing models. Huge economies of scale there. For chip fabrication. Same story. But for example, if you think about building some small metal part inside of that rack.
51:42 There's not that same effect. We think the natural boundaries are On the chip side, on the bottom. As a model layer at the top. And we'll fill the whole gap between.
51:50 A few weeks ago there was a guy who was running A next generation AI chip for one of the frontier companies. And he's trying to recruit one of our architects. and actually kinda did an U reverse card and started recruiting the person trying to recruit our guy. And within a week we hired him. And I was going on a walk as we were kind of finalizing the offer. I was like, You're leading this super important project, why are you deciding to join?
52:11 And his answer is super interesting, which was It fundamentally is not existential for my company. for this product to win. For Google with TPUs, the revenue comes from search. Google won't fail if TPUs fail. That's right. Meta won't fail if MTIA fails, Microsoft won't fail if Maya fails, and OpenAI won't fail if Jalapeno fails. Ultimately, this is our product.
52:32 It is like completely unsurprising that the best chip in the world is built by a company that only builds that chip. It's NVIDIA, right? And like for us, like it is completely existential for us to get as much token capacity online as possible. It recruits a set of talent and it recruits a support from suppliers and from customers. View it with the level of intensity that we do. Look at the raw flops. You compare. Any of the chips built by the labs or by the hyperscalers.
52:57 They flapped in Steak for say F eight times FP eight. Is lower. And the black will be three hundred. And that makes sense because they don't have to go take the risk. They just have to go and build a similar enough product.
53:07 And not pay the NVIDIA tax. As I think about you guys building the solution, the process of doing so is solving a sequence of really hard challenges. What has been the single episode that was the hardest to overcome? We were designing the chip. We built this massive, massive FPH cluster. And verify?
53:23 The full chip work asses. And FPGs are digital entities. You can go test digital logic. But not analog laughter. And it turns out.
53:32 When the chip came back. We began to go see issues in our tensor application. Increment results. Wait a minute.
53:40 There's a problem with a back pressuring logic across a a clock domain crossing. Is failing. And this is going to do costship to produce wrong results. And it is Very, very hard to solve.
53:50 We realized. There was one? And only one way to solve it. As we had to line up two clock signals on our chip. Two within the
53:57 Fifty picoseconds. That is Literally. Fifty trillionths of a second. And we had to go get the signals aligned to this super small granularity.
54:06 And do it. On every chip. Two billion times a second. A lot of people said this was impossible. We had people quit. Yeah. That people literally were like, This problem is unsolvable.
54:17 And that's the luck, guys. Well when you have a problem like that. Step one is okay. Let's assume the problem is solvable. How would it be solved?
54:24 Well first we realized What we have to be able to do is find a way to go ahead and move our clock phase. Bye. I think a second.
54:33 And we had an idea. What if we had these two clocks? We figured out that Hey. If you give up in
54:41 Figure out the face. If we then go out and use the drifting mechanism. They go ahead and wait for just the right amount of time. To get that always lined up. We can do this extremely reliably.
54:52 And then lock the phases exactly where they have to be. We can guarantee this never happens. Mm. People were I think. Blown away that this worked.
55:00 But We made it work. How long did that take? Those are actually about two weeks.
55:07 It was a dark two weeks. It was a very scary two weeks. when that kind of thing happens, that is the most important time to go ahead and be investing effort. that that is the hardest time to go do it when you feel like things are hopeless. But like the sooner you solve that problem, the sooner you can get back to building and scaling up production into mass values. I think a lot of our story is like as Gavin says assume it is possible. Like
55:30 Assume it is possible to have a chip. With way more flops on it. Assume it is possible to have a system with way lower latency between chips. Assume it is possible to create a shared memory pool that can run at way higher bandwidth, like how would one do it? A lot of the time when we do experiments, we will do dozens of experiments, and all of them will fail. But
55:47 We only need one to work. There's multiple times. I mean Gavin, I think was leading the charge in our chip bring up with I think thirty different board experiments and three of them worked. And all three of them are worth their weight in gold. This is one of the things that people quote me and say, Gavin. Almost none of your experiments works.
56:02 I only gotta get lucky once. So one idea for one of these stories that I'm asking about, you know, difficult moments in the company's history is around the ability to raise capital to fund the thing. I think when you started it you knew you'd need capital, but you did not know you'd need the quantum of capital that you've ultimately raised and are spending to build the solution. And you hadn't raised money before. These are all new things, right? There was moments where it was really, really difficult'cause I was I was there, I saw it.
56:28 Moments where it was extremely difficult to raise the money that you did. that without which the company would not exist. It would have died. And like many great stories, there are many near death moments, but Money specifically in this new world, this this isn't software. You don't just need a little bit of money. Maybe you could tell the story about The true hardest part about raising money on, before you had something that you could show people and be so proud of and performance that you could show them and blow their socks off. It was just you guys talking about an idea. But talk about the early difficulties raising money because it was pretty hardcore.
56:59 We've had some intense moments. It reminds me of probably early twenty twenty four. before we raised our series A. And we were at this point Where
57:08 We had done enough of the architecture, done enough of the design. That we knew that the chip architecture was sound. We had to go build it, there was a lot more to do. We were ready to go into what's called the physical design stage. We needed to sign an agreement with a physical design vendor, which will cost you at least forty, fifty million dollars. And then we had this realization as the models were getting bigger and bigger. And you're seeing these giant emo models come out. That
57:32 We were going to need to build the entire cluster. Not just the chip, but we were gonna need to build boards, we're gonna need to build interconnects, we need to build cold plates, we're gonna need to figure out all of the networking and everything. And that this was gonna cost a lot more than the fifteen million dollars we had in the bank. And like, man, that was scary. Yeah. You're sitting in that moment. They think, holy crap.
57:52 We can't afford this. And I began looking at like Huh. How hard is it to go back to Harvard? At the end of twenty three. We put together
58:01 this like memo. I mean we spent like a hundred hours on this'cause we're like We have no idea how people are gonna believe us when we ask for the amount of money we're about to ask for. And it was like thirty pages, it was like extremely technical and in depth of all the different things we needed to build and like all the milestones we needed to hit and how the market was gonna evolve and all the new use cases and the cost per token and all this modeling. And then we went and talked to investors. And Every major investor in the Valley passed immediately.
58:27 They were just like Okay, two kids that just finished Harvard. Haven taped out a chip, no test chip. Inference, who knows if this is gonna be a big market, everything's gonna be training, the models still hallucinate, this could all be a bubble. You know, at the time. The biggest semiconductor fundraisers for a series was around like forty fifty million dollars. We were looking at this and we were just like tallying the bill.
58:49 We're like, We think we're gonna spend a hundred million dollars in the next twelve months. If we really want to do this, like if we want to actually get to scale And like actually get the performance we're talking about, like this is going to be extremely capital intensive. How the hell are we going to pull this off? I think that one of our key ways we Got started with this process.
59:05 We thought to ourselves What is the cheapest? Possible way. We go do this. Decided, well.
59:11 made almost nothing. Yeah. And if I ate nothing but ramen. then we would go ahead and spend basically just the money for the mass the the front meter tape out. And that'll be that. Right. And if so we've got to do it on Thirty million dollars. An obscenely low number.
59:25 And that we actually went out and got A debt provider. They're willing to go ahead and lend us the money we need to go across the chip threshold. There. I think it was first to go catalyze a series of uh other Hey, maybe we can go do one more thing, one more thing, one more thing.
59:39 Yeah, so We're at this moment where we're like If we really want to build this company, cause like we're not gonna half ass it. We're not gonna go do a test chip and spend years on it and like let the entire AI market boom while we could be building the product. Like if we're gonna do it, we're gonna go all the way. We're gonna need to find a way to get a hundred million dollars. I mean I remember Gavin and I were like sitting down in the office in Cupertino late at night
1:00:01 Just like looking at each other and we're like Could we cut five hundred K here? Can we cut a hundred K here? How long can we convince everyone not to take a salary? And we're like, holy fuck, the math is not going to close. We we we really need to solve this.
1:00:15 And there was a period of a few weeks where You kinda just go into survival mode and you call every person that could possibly know an investor and you're like We need a hundred million dollars to do this. If we do this, we think this could be one of the most important companies of all time. Do you know somebody that wants to take an aggressive bet that wants to like believe in us? Like here's all the information, like we're an open book, here's the team, it's great people, we've been working super hard, we've done these things in record time, but we have these you know a hundred things to go. Like do you wanna do this? And like
1:00:45 The snowball starts and you get A million here and two million here and you're like okay. We're not gonna run out of money this month. You get a five million dollar check, ten million dollar check, and you're like, okay, maybe I can buy those FPGAs. And you know, this snowball happened where you know we were very lucky. that we ended up putting it all together. We had a board meeting. I show you the spreadsheet.
1:01:04 And we look at it and it's like Hundred and three million. It's like these are all like soft commits. And We all look at each other and we say We're gonna take it. And that was a series A, and it's been much easier since then and we've
1:01:16 raised almost a half dozen rounds since then many of them from those investors just doubling and tripling down. that's allowed us to get to market so quickly. Like this rack would not be possible How do we not have been so aggressive? I think deserve a little bit of a Commendation here.
1:01:31 Yeah. T T SMC was willing to work with us. Any of the hundred million dollars. Back when it was still really, really scary. Synopsis.
1:01:40 Actually went ahead and let us Get some of their emulators. on extremely favorable terms where we pay over many years. Mm-hmm. It takes a lot of belief uh from your partners to go do this.
1:01:52 But I didn't know that you come out with this very strong team in All the folks who back you are not in it. Just out of pure financial incentive. They believe.
1:02:00 Why did TSM T believe, do you think? Oh, this is a great story. Even before you joined in. Every conference. A semi event. And I was one of the only young CEOs had conductors.
1:02:11 I think it's kind of a novelty. Ask me to come in there and speak. And the semi event. And I am the only Speaker they're under forty.
1:02:20 And only person there under thirty. That was at the time of twenty two. I go up and I speak. There was a speaker's dinner afterwards. And by peerlock. Happy to know if this. very senior T SM C VP.
1:02:30 It's a very nice dinner. There's like the uh former CEO of Arm there. It's very bougie, everyone's in a suit. And I'm there with this VP. And it turns out We both studied math in college. We go out and get a little piece of paper. We begin talking in great detail about
1:02:44 How do modern AI models work? Are they actual per tensor by tensor level? And the guy just gets it. And we begin talking about, hey, how do you run this very effectively? Why is Lovelton such a critical technology to make this work? And
1:02:57 The following day I get an email from T SMC saying Gavin. Wanna work with Etched? Find a way to make it happen. They've been a great partner ever since. It's amazing to think about Some of the tropes and obviously like
1:03:09 Should break the fourth wall here, like I'm a big edge investor. I've been involved for a long time. I think the absolute world of you guys. So like I'm incredibly biased. in this conversation, I'm trying to ask questions that are broader and interesting and could be objections to what you're doing and we'll we'll keep doing that.
1:03:23 But it's so interesting to me that like when you read about investing. Everyone cites this idea of like contrarian and right as the quadrant that makes all the money. And it sounds really nice. But contrarian means like everyone else thinks you're stupid. And so when you go and you you get the media knows from literally everybody, it is a fascinating quadrant to exist in before you become consensus. What was it like for you? I'm I'm super curious.
1:03:48 Well, it's interesting at the time It was the largest by a lot. first check that I had written. If I say I'm not a a math expert or a semiconductor expert or an AI expert really at the time. And so it was much more of a believed in the concept of
1:04:02 This market potentially being huge. You having made very, very clear bets on how the future was gonna look, having positioned the company in order to attack those things in a hardcore way. And then just the two of you and what I felt about you was the majority of the reason why we made the bet when we did in twenty twenty three or whatever it was. But at the time it was the biggest.
1:04:21 And I think the same thing you said about naivete. applies to investing as does to maybe building a semiconductor startup. Which is like I didn't know what I didn't know. And When I called experts, they were basically like this is stupid. They laid out in very logical terms like why this wasn't gonna work and why it was such a low probability bet.
1:04:41 And I think one of the things I've learned from it is just like you kinda have to damn the base rate. Like if you invest it on base rates. You should do something other than What? We and I do. There's always the index fund. Yeah, there's always an index fund, exactly. So It's actually never been scary for me. I think probably most of that is'cause there's a lot I don't know.
1:04:58 And if I knew more. about what you guys have done and the difficulty, I probably wouldn't have done it. I don't know what that says about like maturing as an investor. Like maybe I don't want to know, you know, a lot more and ha and have some of that healthy naivete. I don't know. Funny, I mean a lot of the traditional semiconductor funds missed the entire AI chip, like all the AI chip companies. And like all the coding experts missed all the coding companies. I think it's very hard to realize the constraints have changed, and when you've looked at tape outs for twenty years.
1:05:26 And you've seen so many of them not work on the first try or the second try or the third try, like you couldn't even run a workload. You totally forget that EDA tools are way better, and that FPGAs w exist today in a way that they didn't exist before. And all the types of validation you can do today. just wasn't possible for us. A lot of our believers either they were kind of on two sides of it. They were just believers in the market and the team. Or they were building chips today in extremely technical, like the high fifty trading firms, where they would literally audit everything from the microarchitecture and the RTL to like the board designs and like the schedule and the software stack. And we would sit down with like 10 of their people who build their own chips.
1:06:04 And they're asking us such detailed questions that we're wondering, are they gonna build the chip? It was really on either of those sides. And if you were anywhere in the middle, like you just wouldn't understand it. In the investing world, they all often talk about variant perception, something that you see or believe that others don't, right? Then that perception creates the opportunity. I think I've invested I don't know, five or so times a matched. And every time when you do it, things are getting bigger and bigger. It does get a little scarier and scarier. And and Because you guys have been so quiet in the marketplace.
1:06:32 I think it's very easy to dismiss you. As the stakes get bigger and bigger, those dismissals are harder to hear. I do think betting On something that you see when what you hear from the outside world is very different. That perception gap.
1:06:48 Equals opportunity. Exactly. The last thing I would say is The Accumulated evidence. of your guys' and your team's ability to solve seemingly impossible problems. Is one of the most interesting
1:07:00 things a company can have. It's like a binary, like companies do this or they don't. I don't think it's a big advantage. People who have been here for a long time. Yeah. Some new joiners who are scared shitless. You see a thing like this and there are uh old
1:07:12 Timers. Another one. Yeah, there's definitely a find a way mentality. If you're here You're here because you assume it's possible. So like we we can't be saying it's impossible. Everything is solvable.
1:07:28 And we're just going to work at it until we figure it out. A favorite story I have, this guy who's kind of a legend in Silicon Validation, who joined our team. And we were doing the early stages of what's called Wafer Sort. when your chips are coming out of Fab, they go out on these wafers and you have this thing called a probe card that attaches to the wafer before you dice it with these probe pads and you send these electrical signals to basically test which chips are good and bad. So when I slice the wafer into a bunch of chips, I can package it and only package the good ones. We go through our first wafer. It's like
1:07:56 Two three AM because we're doing it with T SM C over the phone in Taiwan. We have the screen. With the wafer it's all gray and each ship is gray. And then as you start running the patterns, the squares are supposed to turn green or red. And The all turn red.
1:08:11 Fuck. Like, this is really bad. Like everybody's like Guys, take a breath. The puzzle begins. You have to have the attitude of like, yes, you will go out and you will go stare into the abyss, and you will go see scary things. And
1:08:29 We'll solve them. When did you see the first green square? Within a day of that. But in the moment it it's extremely scary. And there's a certain type of person who just like is addicted to that feeling of just feeling the fear and solving it and we are lucky to have a lot of those people here. If you think about applying all of this earned know how from this last several years and now thinking ahead to Gen two, gen three, and beyond. Sure.
1:08:52 What will you be doing most differently? As a result of everything that you've learned. Just like from a conceptual standpoint, like the way that you will attack designing and producing this next one based on what you learned doing it the first time. It took us a while to get to the primitives that we think are really what matters for scaling inference.
1:09:09 We tried a bunch of things early on from compilers that would turn different models into FPGAs. to burning weights in silicon to splitting your HPM to KV cache and weights and all these different things. And there was a lot of cycles of learning'til we got to the point that we realized that like Fundamentally, if you want to run a majority of tokens in the world You need to do three things.
1:09:29 Need to build a chip with the most flops. And given power budget. You need to build a chip. that has the lowest latency between other chips, so the biggest scale up domain possible. And you need to produce as much of it as possible.
1:09:40 And I think Probably in the first half of our journey so far, we learned the first two, and that informed the design a lot, and that informs a lot about the bets we're making in the future. with the low voltage inference and the cluster scale memory. But the production part, I think, in the past year has become extremely Obvious how much people want to deploy this stuff.
1:09:57 If you can have it available today. The best ability is availability. If I have a thousand chips today, someone's gonna use them. And you know, we need to build a chip that's not just like way better than what's been built before, but it needs to be available at many gigawatt scale. We need to be able to be building a product that is producible at gigawatts per month from the limit. As we think about that, a lot of the design decisions we're making with our next gen, which you've seen already.
1:10:21 It's just about simplicity removing tons of parts. trying to assemble and disassemble the thing again and again and learning how to make it as quick in the cycle times as possible in production, making sure it's gonna be reliable, making sure it's gonna be serviceable. And making sure it's gonna be producible at gigantic scales. What about other problems in the ecosystem that are outside of your control, such as capacity at the leading nanometer at T S M C or availability of HBM four
1:10:46 memory or you know, some of these other things were like Everyone is fighting for uh scarce unit of capacity or whatever. How do you face up against th those realities when you're trying to produce as much as humanly possible.
1:10:59 The people deploying the most compute in the world. Do think about supply a bit zero sum. Which is There's only so many wafers being produced on a given nanometer node. On a given fab.
1:11:09 Right. And there's only so much memory being produced. And that's why actually for our first gen product We built it on a different supply chain than the Rubens. But we're on four nanometer, Rubens on three nanometer. Yeah, we're on a different HPM than Rubin's and so forth. So it actually is not a zero something. It's a positive something where more is more.
1:11:26 So often when we're talking to people deploying at scale. It's not a decision between a gigawatt of a GPU and a gigawatt of us. It's Two gigawatts. And I think as much as possible thinking about supply chain early in the design decisions, because if you have the most performant product and you can't produce it,
1:11:42 Then you're just a podcast. That's the other big thing about vertical integration too. Is Well certain things like for the chips and for the memory, you have to go ahead and uh partner. For most of the other stuff.
1:11:52 Those are also very highly in demand components. And the more that you build yourself. The more stuff you can go do on top of what the world can currently build. It is not all you're taking availability with somebody else. You're adding way, way more I think that's how you win.
1:12:05 One of the things I realize we haven't talked at all about is the models themselves. Which is kinda crazy, the things behind all of this demand. Anything interesting that you would say about the way that You see models progressing.
1:12:20 How? You as thinkers about hardware. think hardware might impact where the models themselves go in the future. One of the most important ideas That we believe in.
1:12:30 Is that machines don't think like people think. You look at airplanes, for example. Airplanes don't fly, birds fly. That when you think about how mechanical Devices have to work.
1:12:40 It's often very different. And in much the same way. Four people. Storing data? And loading memory is very cheap for neurons.
1:12:48 And doing math is relatively expensive. It is the exact opposite. For tips. Generally one data's very expensive. And doing math is very cheap.
1:12:59 And as time goes on. Then you end up finding that Math gets cheaper. at a rate that is faster. Then uh
1:13:05 Memory gift cheaper. J due to this fundamental limit on any kind of D RAM device. You shouldn't think about. How can I make my model use a huge Huge amount of compute.
1:13:16 What if I had, for example, many copies drawing at the same time? But it's activated a huge number of experts. But if I had gigantic experts, I could go ahead and run on multiple server racks at the same time. That is how I think you'll build models that are The next generation of intelligence.
1:13:31 And context too. There's been a lot of work on a very efficient inference. What if I don't load the full context into memory? And most of the time I think that makes a lot of sense. You don't even build a superintelligence. Why can't they go look at
1:13:44 A billion tokens of context. Why can't it spend a huge amount of compute to go ahead and read all in a super fast? I would love to g be able to talk to a machine that was able to go attend to every book ever written and short term memory. And uh I think that's gonna get to a point where you can't.
1:14:00 A theme in models right now is this focus on something called dynamism. But just this ability to control the level of computation and memory spent at a per token or per user level when doing attention. as well as this ability to dynamically in your chip on the fly send data to other chips for different MOE models doing
1:14:20 Certain types of operations. And you know, the reason is fundamentally As we are scaling context length, as we're scaling model size, as we're scaling the amount of computation per user, we're looking for ways to be more efficient. So you know the first thing is like Gavin says, you know, a mixture of experts, yeah, architectures where you know, maybe we don't need every parameter being used for every token. But maybe there's things where even at a token level we can say, well, this token needs this context from this other token, they can share that memory so we don't have to have overhead of of using in the memory as much.
1:14:49 Maybe this token is really important, so we should spend more compute. We should have longer context on that token. So hardware that really accelerates these types of Very dynamic computations extremely important. And you can imagine current hardware that was designed before those types of architectures. have lots of overheads in doing them. So you basically end up in these really bad worlds where you have
1:15:09 inefficient hardware at doing this dynamism. So therefore you can't run it very well. Or you have these very blocky architectures that are kind of applying blunt force to many different tokens that all need more or less computation. I have two questions about the future. We've talked a lot about what you built so far and how you built it. The first is about the new ways that people might start using these systems.
1:15:31 The raw technology Longer run times. Things of this nature. When inference gets much cheaper, faster, more accessible, there's more total supply and it's better. What are the things that you think people will use that? capacity to do that are the most interesting exciting to both of you.
1:15:46 Yeah, there was a viral tweet by Noam Brown. Where said that as these models are having longer and longer time horizons. They can do tasks that take, say, six months. And there's often not enough time to go and evaluate them. For such a long period of time.
1:16:01 Because by that point you'll have a new model app, you'll want to go evaluate instead. And with tech like We built. or cluster scale memory. You can go ahead and run that six month job much faster.
1:16:11 But there's a second piece of this too. Talking to him about it. He's now an angel as well. Where It's not just the time. It's also the number of people or agents who are working on this.
1:16:21 If you're trying to go and evaluate Can a human build a rocket? You will find out the answer is no. No, well first you can go build a rocket. Instead of you have to go put a team together.
1:16:31 And I believe the same thing will be true of ages too. You want to go ask can an agent go out and build Some crazy futuristic Piece of software. you'll probably need a very large team.
1:16:42 Uh maybe that's ten. Maybe that's a million. You have to go and have this enormous amount of both. Cluster scale memory. To go ahead and have that very short time.
1:16:51 Per token. And a huge amount of flops. To be able to go run that whole fleet. I'm gonna be a little futuristic. I firmly believe we are on a global march. of inference becoming majority of global GDP.
1:17:04 It may take more than ten years, but it's gonna happen. And right now we measure productivity as a society. uh GDP per capita. But really it's going to look much more like agents per megawatt or it maybe agents per gigawatt by then. And while we're being futuristic.
1:17:18 I think this is the second to last year Or a majority of the workforce is going to be human. I think in twenty twenty seven you're going to see there's going to be more agents doing knowledge work And humans. And it's going to be extremely interesting to see what happens.
1:17:31 You could imagine a world where For countries a m majority of their energy ends up going into data centers doing inference. And the energy efficiency of those data centers basically governs how many agents and therefore, you know, how big their workforce is. So you're going to see, like as Gavin is saying one agent or a team of five to ten agents working on group projects for a couple days.
1:17:52 So you can do pretty cool stuff'cause they're smart, but Not going to be civilization scale. What happens when you have countries that can have literally a billion concurrent agents Like a billion people in the workforce working twenty four seven concurrently. On the same stuff.
1:18:07 It's just kind of unfathomable what's going to happen. But it's going to be the biggest proliferation of technology humanity's ever seen. I think as well, like When you have these Huge, huge amounts of demand, you get this idea of economies of scale again. Yep.
1:18:20 Or think about people. Have a brink. I'm not using the whole thing all at the same time. That's only part of it's gonna be active and this is the way healthy brains work. M models, it worked much the same way.
1:18:30 On MOE model, only a small fraction of the parameters is being used. For any given token at any given moment. But If you have a large number of users. on a piece of hardware, you can go kinda take that brain.
1:18:41 experts on many different servers and run a huge amount of volume through it. So you'll have a bunch of different pieces of traffic. You'll have many of them using each part of the brain any given in time. And you also make the cost per
1:18:54 Thought costs per token way way lower. You're gonna end up with these. Giant. scale distributed brainess. The form factor of this is a big data center.
1:19:02 With a bunch of chips, with a huge amount of flops and a huge amount of scale up interconnect. You think we'll see a trillion dollar individual data center. Absolutely. It is a matter of time. It's like asking
1:19:12 Well you say a billion dollar fab or ten billion dollar fab Or a hundred billion dollar fab. It is inevitable. that the economies of scale don't stop at oh forty billion dollars is the magic number for fabs. No.
1:19:25 The cost wafer keeps going down. And you keep spending more money. And the same thing will be true of Plants to go out and make steel. More plans to go out and make tokens.
1:19:34 A very smart alien lands on Earth. And wants to know from each of you How you would frame up this opportunity. That that you guys are tackling. What do you say to them?
1:19:44 Try me to us. Thinking is really valuable. But Every company in the world. Run this on thinking.
1:19:49 And we are entering this really unique moment in time. Where Yeah, the machines. I can go think. Almost as good.
1:19:56 And as soon as good and assumed better. than the best humans can. building these machines is gonna go be a huge opportunity. More important than that. The way in which you go ahead and run
1:20:06 This kind of thinking is going to be very, very different. As demand goes higher and higher and higher and higher. There's a unique moment right now to build A new set of solutions, a new roadmap. For how do you run this?
1:20:17 The future quadrillion parameter models. For milion people. All at the same time. On a gigantic scale up cluster. We are in a new era of intelligence.
1:20:26 Where The cost of producing intelligence Is dramatically so much cheaper. then the value of the intelligence But we are in a many year, probably many decade
1:20:38 Supply shortage. Of these tokens. And basically any chip or any system that can produce tokens Is likely to be extremely valuable. And you should find some part of the supply chain of the token.
1:20:51 I can be everything from model training down to what we're doing in the silicon and otherwise. To spend time on and push the frontier. And that the companies that are the largest are frankly going to be the companies That produce most of the global supply of tokens. and own majority of the supply chain of that token.
1:21:07 And importantly. It's people who build systems. That As they get more and more Ships put together.
1:21:13 Get cheaper. The way you want this to scale is not that oh if I want to go serve ten times more tokens, I buy ten times more servers. It must be some solution where Only those are ten times more tokens. Then
1:21:24 I get some economies of scale benefit. with MySech cluster scale memory tech. It allows me to then Not charge as much. as ten times more for those back tokens. What a ridiculously exciting future that you guys are
1:21:36 Building to enable. When I did this with Gavin last time I asked him my traditional closing question, so this time I'll ask you. What is the kindest thing that anyone's ever done for you? During my cancer treatment there was a big decision I had to make. The doctors came to me and said.
1:21:49 It's time for you to decide. Do you want to get surgery or do you want to get radiation? Here's the trade off. If you get surgery, you're more likely to live. But you have to assume you'll never be able to walk again. If you get radiation, you'll be able to walk again.
1:22:01 But It's not the same probability that you'll live. You may die. What do you want to do? And I was sixteen.
1:22:08 And My parents said you have to make this decision for yourself. I thought a lot about it for a long time and decided I'm gonna do the surgery. I get the surgery, one of the things you do when you get a tumor resection is they do something called a necrosis analysis. Where they look at all the different cell and say is this cell dead or alive?
1:22:25 Because if you have a bunch of cancer cells that are live You have a problem. And they looked at it and they said, You know You usually want ninety eight, ninety nine percent necrosis. For us to say you're in the clear. You're below that.
1:22:36 You should go get radiation. And there was only a few machines in the world that actually could do the type of radiation I needed. One of them was in Boston. I was in a wheelchair and I needed to move to Boston for multiple months. And both of my parents decided to move out and drop everything they were doing and and live with me. And I'm eternally grateful. She's beautiful. Thanks guys. If you enjoyed this episode, visit Colossus.com. You'll find every episode of this podcast complete with hand edited transcripts. You can also subscribe to Colossus, our quarterly print, digital, and private audio publication featuring in-depth profiles of the founders, investors, and companies that we admire most. Learn at Colossus.com slash subscribe.
What you see above is a preview of the first minutes. One unlock costs 10 credits and covers this episode forever: full segment and word-level timestamps on this page, plus .txt, .srt, .vtt and word-level JSON downloads, as many times as you like.