Transcript
Chetan Puttagunta and Modest Proposal - Capital, Compute & AI Scaling - [Invest Like the Best, EP.400]
0:00 I know firsthand how complex the tech stack is for asset management firms. And seemingly every new tool and data source makes the problem even worse, adding more complexity, more headcount, and more risk. Ridge line offers a better way forward, one unified platform that automates away the complexity across portfolio accounting. Reconciliation, reporting, trading, compliance, and more, all at scale. Ridge line is revolutionizing investment management, helping ambitious firms scale faster.
0:25 Operate smarter and stay ahead of the curve. See what Ridgeline can unlock for your firm. Schedule a demo at ridgeline.ai. Hello and welcome everyone. I'm Patrick O'Shaughnessy and this is Invest Like the Best. This show is an open ended exploration of markets, ideas, stories, and strategies that will help you better invest both your time and your money. Invest like the best is part of the Colossus family of podcasts, and you can access all our podcasts, including edited transcripts, show notes, and other resources to keep learning at joincolosis.com.
1:00 Patrick O'Shaughnessy is the CEO of Positive Sum. All opinions expressed by Patrick and podcast guests are solely their own opinions and do not reflect the opinion of positive some. This podcast is for informational purposes only and should not be relied upon as a basis for investment decisions. Clients of Positive Sum may maintain positions in the securities discussed in this podcast. To learn more, visit psum.vc.
1:29 My guests today are Chaithan Pudigunta and Modest Proposal. If you're as obsessed as I am about the frontier in AI and the business and investing implications, you will love this conversation. Chaythan is a general partner and investor at Benchmark, while Modest Proposal is an anonymous investor who manages a large pool of capital in public markets. Both are good friends and frequent guests on the show, but this is the first time that they have appeared together. The timing could not be better. We might be witnessing a pivotal shift in AI development as leading labs hit scaling limits and transition from pre-training to test time compute. Together we explore how this change could democratize AI development while reshaping the investment landscape across both public and private markets.
2:07 Please enjoy this great discussion with my friends Chaith and Pudagunta and Modest Proposal. So Chain, maybe you can start by just telling us From your perspective, what is going on right now that is most Interesting. in the technology part.
2:23 of the story of LMs and their scaling. Yeah, I think We're now at a point where it's Either consensus or universally known. that all the labs have hit some kind of
2:37 plateauing effect on How we perceive scaling for the last years, which was specifically In the pre training world. And the Power laws of scaling.
2:49 Stipul that The more you could increase compute. in pre-training, the better model you were gonna get. And Everything was thought of in orders of magnitude, so
2:59 Throw ten X, more compute. At the problem. And you're gonna step function and model performance and intelligence and This certainly led to incredible breakthroughs here. And we saw from all of the labs
3:14 Really terrific models. The Overhang on all of this Even starting in late twenty twenty two. Was at some point we were gonna run out of text data that was generated by human beings.
3:27 And we're gonna enter the world of synthetic data fairly quickly. All of the world's knowledge effectively had been tokenized and had been digested by these models. And sure there were niche data and private data and all these little repositories that hadn't been tokenized, but In
3:45 terms of orders of magnitude, it wasn't gonna increase the amount of available data for these models particularly significantly. As we looked out in twenty twenty two You saw this Big question of was synthetic data gonna enable
4:02 These models to continue to scale. Everybody assumed, as you saw that line, this problem was gonna really come to the forefront in twenty twenty four. And here we are. We're here and Where All trying to train on synthetic data, the large model providers.
4:18 And now As it's been reported in the press and has all these AI lab leaders have Gone on the record. We're now hitting limits.
4:29 Because of synthetic data. The synthetic data as generated by the LMs themselves are not enabling the scaling and pre training to continue. And so we're now shifting to a new paradigm called test time compute. And what test time compute is in a very basic way.
4:46 Is you actually asked the LM To look at the problem. Come up with A set of potential solutions to it. And pursue multiple solutions in parallel, you create this thing called a verifier.
4:58 And you pass through the solution over and over again iteratively. And The new paradigm of scaling, if you will, the X axis Is time measured in logarithmic scale.
5:11 And intelligence is on the Y scale. And that's where we are today where It seems that Almost everybody is moving to a world where we're scaling
5:21 On Pre training and training to scaling on what's now being called reasoning, or that is Inference time, test time. However you want to call it.
5:31 And that's where we are as of Q four twenty twenty four. Just a follow up question on the overall picture. So setting aside CapEx and all this other stuff that we'll talk about with the big public tech companies. In just a moment. Is it reasonable to say based on what you know now
5:49 That The switch to test time scaling. Where time is the variable. This V likea Who Cares? As long as these things keep getting more and more capable
5:59 Isn't that all that matters and the fact that we're doing it In a different way than just based on pre training. Does anyone really care? Does it matter? Two things that come up. Pretty quickly.
6:11 In test time Or reasoning paradigm, which is As LMs explore the Space for potential solutions. Very quickly.
6:22 As a model developer or somebody working on models, you quickly realize that algorithms used for test time compute. might exhaust the useful search space for solutions quite quickly. That's number one. Number two, you have this thing called a verifier that's looking at What's potentially a good solution, what's potentially a bad solution, what should you pursue?
6:45 And The ability to figure out what's a good solution and what's a bad solution or what's a optimal path and not An optimal path. It's unclear that that scales linearly with infinite compute.
7:00 And then finally Task themselves. Can be complex. Ambiguous And the limiting factor there may or may not be compute. So
7:12 It's always really interesting to think of these problems as if you were to have infinite compute To solve this problem, could you go faster? And certainly there's gonna be a number of problems. in reasoning where you could go faster if you just scaled compute. But oftentimes we're starting to see evidence that
7:30 It's not necessarily something that scales with compu linearly with the technology we have today. Now can we solve all that? Of course. There's gonna be algorithmic improvement. There's gonna be data improvement, there's gonna be hardware improvement, there's gonna be all sorts of optimization improvements here. The other thing we're still finding is
7:52 The Inherent knowledge. Or data available to the underlying model. That you're using for a reasoning? still continues to be limited. And just because you're pursuing test time
8:05 It doesn't mean that you can break through all previous data limitations. By just scaling compute a test time. So It's not that We're hitting walls on reasoning or we're hitting walls on test time.
8:19 This is the problem set And the challenges and The computer science problems. are starting to evolve. And As a venture capitalist, I'm very optimistic that we're gonna be able to solve all of them, but they're
8:31 To be solved. So if that's the sort of research lab's outview. Modest, I'm curious for you to give us the big public tech companies down view, because so much of this story has been the spend capex, the strategic positioning, the quote unquote ROI on all this spend. And how they're gonna earn a return on this insane outlay of capital.
8:53 Do you think that Everything Jathan just said is well reflected in the stance and the pricing and the valuations of the public tech companies. I think you have to start
9:04 At a macro level. and then get to a micro level. At a macro level, why this is so important is Everyone knows the Mag Seven. a large a percent of the S P five hundred they represent today.
9:19 But beyond that I think thematically AI has permeated Or broader into industrials.
9:28 into utilities. And really makes up I would argue somewhere between forty and forty five percent of the market cap is a direct
9:38 And if you even abstract to the rest of the world. You start bringing in ASML, you bring in TSM C. You bring in the entire Japanese chip sector. And so if you look at the cumulative market cap.
9:52 That is a direct play on Artificial intelligence right now, it's enormous. And so I think As you
10:01 look across the investment landscape you almost are forced to have an opinion. On this. because most people in some form or another are benchmarked against an index that is going to be a derivative play on artificial intelligence. At the micro level
10:18 I think that This is a fascinating time because all of public market investing. is scenario analysis and probability weighting different path. And if you go back to when we talked probably four months ago.
10:31 I would say that the distribution of outcomes has shifted. And at that point in time. Pre-training. And scaling on that axis. What
10:42 Definitely the way. And we talked about what the implications were at the time. We've talked about Pascal's wager, or we've talked about prisoners to laugh. And in my mind
10:54 It was easy to talk about that when the cost of anteing up was a billion or five billion dollars. But we were rapidly approaching the point in time. where the ante was gonna be twenty billion dollars or fifty billion dollars. And You can look at the cash flow statements of these companies.
11:13 It's hard to sneak in a thirty billion dollar Training. And so The success of GPT five class broadly. Let's apply that to all the various slabs. I think was going to be a big proof point.
11:27 As to whether or not The amount of capital was committed. Because these are three four year commitments. If you go back to when the article was written on Stargate, which is the Hypothesised hundred billion dollar
11:41 data center that OpenAI and Microsoft were talking about. That was a 2028 delivery. But at some point here in the next six to nine months it's a go no. We already know that the three to four hundred thousand chip. Super cluster.
11:56 is gonna be delivered end of next year, early twenty twenty six. But we probably need to see some evidence. of success on this next model in order to get the next round of commitment. So
12:08 All that as a backdrop, I think at the micro level, this is a really powerful shift if we move. From pre training to inference time. And there are a couple big ramifications. One Yeah.
12:20 Better aligns revenue generation. and expenditure. I think that is a really, really beneficial outcome for the industry at large. which is in the pre training world. You were gonna spend twenty, thirty, forty billion dollars on Cephex.
12:34 Train them all over nine to twelve months. Do post training. Then rule it out. Then hope to generate revenue off of that in entrance. In a test time compute scaling world
12:46 You are now aligning your expenditures with the underlying usage of the model. So just from a pure efficiency and scalability On a Financial side This is much, much better for the hyperscalers.
13:02 I think a second big implication again, we have to say We don't know that free trading scaling is going to Stop. But if you do see this shift towards entrance time. I think that you need to start to think about
13:17 How do you re architecture the network design? Do you need millian chip superclusters? in energy low cost land locations. Or do you need smaller Lower latency, more efficient.
13:33 Inference time data centers. Scattered throughout the country. And as you re architect the network. The implications on power utilization. Grid design.
13:46 A lot of the I would say narratives that have underpinned Huge swaths of the investment world, I think have to be rethought. And I would say To date,'cause this is a relatively new phenomenon.
14:00 I don't believe that the public markets have started to grapple with What that potential new architecture looks like. and how that may impact some of the underlying specs. Chief, I'm curious. Maybe to tell the story of Deep Seek.
14:14 And other things like it. Where You see new models. being built by small teams For relatively small dollars.
14:23 that are competing in performance with some of the Leading edge models. Can you talk about that phenomenon and What it makes you think about or the implications for the landscape. It's really amazing in the last Call it six weeks of time.
14:40 The number of teams we've met. Here at Benchmark. That are two to five people. And Modest has talked about this in your podcast before, which is that The story of technology innovation has been there's always been two to three people in a garage somewhere in Palo Alto doing something.
14:57 to catch up to incumbents very, very quickly. I think we're seeing that now in the model layer in a way that we haven't seen, frankly, in two years. Specifically, I think We still don't know one hundred percent that pre training and training scaling isn't coming back. We don't know that yet. But at the moment at this plateauing
15:18 time we're starting to see these small teams catch up to the frontier and what I mean by frontier is where are the State of the art models. Especially around text performing. We're seeing these small teams of
15:32 Quite literally two to five people. Jumping to the frontier with Spend that is Not one order, but multiple orders of magnitude less than what these large Labs we're spending to get there.
15:46 I think part of what's happened is the incredible proliferation of open source models. Specifically what Meta's been doing with Lama. Has been an extraordinary force here. Llama three point one comes in three flavors, four oh five billion, seventy billion, eight billion.
16:04 And then Lama Three dot two comes in one billion, three billion, eleven billion, and ninety billion. And You can take these models. Download them.
16:15 Put them on Local machine, you can put them in a cloud, you can put them on a server. And you can use these models to distill fine tune, train on top of, modify, et cetera, et cetera, and catch up to the frontier. With
16:30 pretty interesting algorithmic techniques. And because you don't need massive amounts of compute, or you don't need massive amounts of data. You could be particularly clever and innovative. About a specific vertical space or a specific technique or particular use case. To jump to the frontier.
16:49 Very very quickly. I think that is largely changing how I personally think about the monolayer and potential early stage investments in the monolayer. There's a lot of ifs here and a lot of dependent variables and literally in six weeks none of this could be true anymore, but If the state holds, which is that
17:09 Pre training isn't scaling because of synthetic data. It just means that you can now Do a lot more. Jump to the Frontier very quickly with the minimum amount of capital. Find your use case, find where you're most powerful. And then from that point onwards. The hyperscalers
17:27 Frankly become Best friends. Because Today, if you are at the frontier, you're powering your use case. You're not particularly GPU constrained anymore.
17:38 Especially if you're gonna pursue test time inference or test time compute or something like that. And you're serving let's say ten enterprise customers. Or maybe it's a consumer solution that's optimized for a particular use case. The compute side of it. Just doesn't become as challenging.
17:54 as it was in twenty twenty two. In twenty twenty two you would talk to these developers and it just became a question of Well, could you get a hundred thousand cluster together'cause we need to go train and then we have to go buy all these data and then Even if you knew all the techniques. All of a sudden you would pencil it out and say, like, I need a billion dollars to get the first training run to go. And
18:15 That just is not a model historically that's been The venture capital model. The venture capital model has been could you get together a team of extraordinary people. Have a technology breakthrough. B Capital Light. and jump way ahead of incumbents very quickly and then somehow get a distribution foothold and go.
18:34 At the model area for the last two years That certainly didn't seem like it was possible. In literally in the last six Eight weeks. That's
18:44 Definitively changed. I think it's important. The point about meta open source and the hyper scalers. Open source pushing the frontier, smaller models being able to scale.
18:58 to very successful points. Is enormously beneficial, particularly for AWS who doesn't have uh native LLM. But if you just take a step back and think about what historically cloud computing was. It was providing a set of tooling.
19:14 to developers and builders. AWS first articulated this vision, I heard it publicly. In September when Matt Garman was at Goldman Sachs conference. But Their view clearly has been
19:27 That LLMs are just another tool. That generative AI is another tool that they can provide their enterprise customers and their developer customers. to build the next generation of products. The risk to that vision was
19:41 An all powerful generalizable mob. And so again, this is where you sort of have to rethink. If we're not going to build these massive pre-trained Entities where you drive training loss down to near nothing. And that in some form or another builds
19:59 The metaphorical God. If instead The focus of the industry is at test time at entrance time And trying to solve real problems.
20:10 At the point of need for a customer. I think that again re-engineers and re architects the entire vision. of how this technology rolls out. And I think we need to be humble that We don't know what Lama Four is gonna come out with. We don't know what Rock Three is gonna come out with. Those are the two models.
20:30 that are currently being trained on the largest cluster ever. So Everything we're saying right now. may be wrong in three months. But I think the entire
20:40 job right now is to ingest all the available information and replot the various scenario paths. With what we know today. I do not feel like People have updated their priors.
20:54 As to how these pads may go forward. If This is correct. I'm curious. Chan, how
21:02 the idea that maybe now you would invest in a model company because of this change. I remember you tell me over dinner two years ago That As a firm, you just decided we're not investing in these companies like you said, it's just not the model that we do. We don't write billion dollar checks for the first training run. And so we're not investing in that part of the stack. We're investing more in the application layer, which we'll come back to
21:22 In a little bit in this discussion. But Maybe say a bit more about this updated view on how that could work out, what a sample investment could look like.
21:31 And whether even if Lama four is the pre training scaling loss hold If that even changes that,'cause it would just seem like something like Deep Seek just only benefits from Okay, now instead of three point two, it's four and like we still doing our thing and It's still better and cheaper and faster and whatever.
21:47 So yeah, what do you think about this new view you have Potentially investing in model companies, not just application companies. In Meta's last earnings call Mark Zuckerberg talked about them starting long before development and he said that
22:03 Lama 4 is being trained on a bigger cluster than Anything he's ever seen out there. The number that was quoted was it's bigger than a hundred thousand H one hundreds. Or bigger than anything I've seen reported for what others are doing. And he also said, you know, the smaller Lama 4 models should be ready in early twenty twenty five.
22:23 What's really interesting about that is that Regardless of whether Lama 4 is a step function From Lama three. Kinda doesn't matter. If they push the boundaries of efficiency and get to a point where Even if it's incrementally better, what it does to the developer landscape is
22:43 Pretty profound because The force of Lama today. has been two things and I think this has been very beneficial to meta is one The transformer architecture that Lama is using Is it
22:55 sort of standard architecture, but It has its own nuances. And if the entire developer ecosystem that's building on top of Lama It's starting just assume that that Lama three transformer architecture is the foundational and sort of standard way of doing things.
23:12 It's sort of standardizing the entire stack. Towards this llama way of thinking. All the way from The hardware vendors will support your training runs to the hyperscalers and on and on and on. And so Standardizing on Llama itself.
23:29 It's starting to become more and more prevalent. And so If you were to start a new model company. What ends up happening is starting with Llama today is not only great because Llama is open source, it's also extraordinarily efficient because the entire ecosystem It's standardizing on that architecture.
23:47 And so You're right. As an early stage fund with five hundred million dollars of capital And we're trying to make thirty investments every fun cycle. A billion dollar training run Is essentially you're committing two funds
24:00 To do one frame run, that may or may not work. And so that's an ex Extraordinarily capital. Intensive business And by the way, the depreciation schedule for these
24:12 models is frightening. Distillation as a technique. makes defensibility of these models and these notes of these models Extraordinarily challenging. And it really comes down to what's your application on top of it, what's your network effects, how are you capturing your economics there and all of that. And I think what has now The case as of today.
24:32 Is If you're a two to five person team You can take something like coding as an example. And you could Push your way into a model that generates better coding answers faster.
24:45 By fine tuning and training on top of Llama. And then offer an application with your own custom models. that really produces extraordinary results for your customers, whether it's developers or something like that. And so Our particular approach and strategy here has been To invest heavily in applications, starting when we saw OpenAI APIs start to take off, we started to see developers talk.
25:12 About these open AI APIs. The summer of twenty twenty two. And A lot of our effort starting then was to just find entrepreneurs that were thinking about leveraging these APIs to go after The application layer and really start thinking about
25:28 What are applications that simply could not exist? before this current wave of AI. Obviously we've seen some really incredible successful companies. come out of that. They're still early, but the kind of traction they're seeing, the kind of customer experience they're providing, the kind of
25:45 Well metrics, all of that has been extraordinary. You had Brett Taylor. On your podcast a couple of weeks ago. So Sierra is an example of this in procurement. We have a thing called level path. Many other examples across the portfolio at the application layer. Where you can just go through Every single large SAS market.
26:06 And go after it with an application layer investment and start to really think about what's now possible that wasn't possible. Two, three, four years ago. I'm curious to talk a little bit about The big
26:19 foundation model players that we've talked about Llama, but less about XAI, Anthropic, and OpenAI. maybe Mata starting with you. I'm curious for l just your thoughts on their strategic positioning and the things that are important for each And Maybe open AI as an example.
26:35 Maybe the story here is just what a great brand they built and that they have so much distribution and they have all these great partnerships and people know it and use it and they have lots of people paying them twenty bucks or whatever. Maybe the distribution's more important than the product in the model. I'm curious what you think about these three players that have So far dominated, but seemingly. through this analysis so far, like
26:56 It's important that they keep innovating. So I think the interesting part for open AI was because they just raised the recent round And there was some fairly public commentary around what the investment case was. You're right. A lot of it oriented around the idea. That they had
27:13 to escape velocity on the consumer side. And that chat GPT was now the cognitive referent. And that over time they would be able to aggregate an enormous consumer demand side. and charge appropriately for that.
27:29 And that it was much less a play on the enterprise API and application building. And that's super interesting. If you actually play out what we've talked about. When you look at their financials
27:43 If you take out Training on the If you take out the need for this massive off front expenditure. This actually becomes a wildly profitable company quite quickly in their Projections.
27:57 And so In a sense It could be better. Now Then the question becomes, what's the defensibility of a company that is no longer step function advancing?
28:08 on the frontier. And there I think This is ultimately gonna come down to one Google. is also advancing on the frontier and they
28:17 most likely will give the product away for free. And meta, I think. We could probably spend an entire episode just talking about meta and the embedded optionality that they have on both the enterprise side and the consumer side. Well let's stick to the consumer side.
28:32 This is a business that has over three billion consumer touch points. They are clearly rolling meta AI out into various surfaces. It is not very difficult to see them building a search functionality. I joke they should buy perplexity. But you've also just had the DOJ come out and say that Google should be forced to license their search index.
28:57 I can think of no bigger beneficiary in the world than Meta having the opportunity or at marginal cost. to take on Google search index. But the point is that I think there will be two very large scale internet players. giving away what essentially looks like chat GPT for free. So it will be a fascinating
29:19 Case study in Can this product that has dominant consumer mind share. My children know what? Chat GPT is. They have no idea where Claude is.
29:28 My family knows what chat GPT is. They have no idea what Brock is. So I think for open AI the question is Can you outrun Free. And if you can
29:41 And training Becomes less of an expense. This is going to be a really profitable company really quick. If you go to anthropic I think they have
29:51 An interesting dilemma which is People think Sonnet three dot five is possibly the best model out there. They have incredible technical talent. They keep ingesting more and more of open AIs.
30:03 researchers and I think they're gonna build great models. But they're kinda stuck. They don't have The consumer mind share.
30:12 And on the enterprise side. I think that Lama is going to make things very difficult. for the frontier model builders to try to grab rate value creation there. So they're stuck in the middle, wonderful technologists, great product. But not really a viable стратегі.
30:32 And you see they raise another four billion dollars. To me that's vindicative that Free training is not scaling so well because Four billion dollars is
30:43 Not anywhere close to what they're gonna need. If the scaling vector is pre training. I don't have a good sense for what their strategic path forward is. I think they're stuck in the middle. X A I I will plead ignorance on that one.
30:59 He is a one of a kind talent and they're gonna have a two hundred thousand chip cluster and they have A consumer touch point. They're building an API. But
31:11 I think If pre training is the scaling vector They're up against the same math problem that everyone else has. Only possibly mitigated by Elon's unique ability to raise capital. But again, the numbers get so big so quickly in the next four or five years that that may even
31:31 Be greater than him. And then if it's test time compute and algorithmic improvements and reasoning. What is their differentiation, what is their go to market. When you have people who have staked their claim on the consumer side.
31:47 And then you have an open source M D on the enterprise side that's every bit as formidable. So when you look at those three, I think it's easiest to see what open AI's path forward is. One thing I will say about OpenAI though is Noan Brown, who I find to be one of the most effective communicators in the research world. He was on Sequoia's podcast recently.
32:11 And he was asked about AGR. And he said Look, I think When I was outside of Open AI. I was skeptical of the whole AGI thing that that was actually what mattered to them.
32:25 And when I got inside of OpenAI It was very clear to me that they are very serious about AGI and that that is their mission and everything else is in service. of AGI. It's easy for us to sit on the outside and articulate the strategy that we might pursue if we were in charge there. But I think we need to be cognizant of the fact
32:48 That Part of the reason they gotten to where they are today is because they are on a mission. That mission is to develop AGI. And we should be very humble about Ascribing any other
32:59 End game for them than that. In my personal. The belief is that AGI is very close by. Say more.
33:10 And and why isn't it not already here? These things are smarter than almost everyone I deal with. Yeah. I think So AGI As narrowly defined or maybe expansively defined, depending on your viewpoint. Is
33:23 A highly autonomous system that surpasses human performance in economically valuable work. In some cases, it's very easy to argue AGI is here using that lens. I think what is pretty clear is that If you look at the announcements made by OpenAI and their Execs that have given interviews in recent weeks.
33:44 An example that's brought up is end to end travel booking as something where That's something we can expect to see in twenty twenty five, where You can Prompt the system to book travel for you and it'll just go do it. And That
33:59 Is A new way of thinking, which is end to end Task completion or end to end work completion. That involves Obviously reasoning that involves agentic work, that involves using computers.
34:14 as Claude has come out with and You're combining multiple ways of these large language models interacting with the ecosystem itself. Putting into A very nice package that then is just able to do
34:29 The end to end work. And fully automated. And do it better than humans. And In my view From that lens, we're very, very close to it. And I imagine that
34:42 Well Be pretty close to or at A G I In twenty twenty five. I don't see how given the current progress and the current innovation And now moving to test time compute and reasoning.
34:55 A GI is not around the corner with that lens. And it's funny'cause we sort of become the frog boiling in water where We passed the Turing test pretty Easily. And yet
35:05 Nobody sits here anymore and talks about Holy crap, we passed the Turing. It just came and went. And so it could be that This declaration of AGI is something along the same lines where it's like Yeah, of course the model can book end to end travel. That's not actually that difficult.
35:23 Where two and a half years ago, if you had said Hey, there's an algorithm that you can tell them what you want to do. It books it end to end and sends you a receipt. You would say no. So there may be some of this boiling frog to it where all of a sudden you wake up one day
35:40 And a lab says, Hey, we've got AGI and everyone's sort of like ah, cool. Okay. There is one particular reason, though. That lab declaring AGI is interesting. In a broader sense, which obviously is the relationship with Microsoft.
35:55 And Microsoft first disclosed last summer that They have the full rights to the IP of open AI. Up until A GI is achieved.
36:06 And so If OpenAI elects to declare that AGI is achieved. I think then you have a very interesting dynamic. between them and Microsoft which will compound an already very interesting dynamic Which is at play right now. So that's something to watch next year, certainly.
36:27 for public market investors, but also for the ramifications of the broader ecosystem because I do think Again, if we're right about the path that we're pursuing now, there will be a lot of reshuffling of relationships. and business partnerships. As we go forward.
36:43 Chatham, was there anything else in Modest's assessment of the big players and we'd love to hear your thoughts on Google since we didn't talk about them as specifically. Anything that he said that you disagree with or would press further on? No, I think What we just don't know is we don't know the underlying discussions in all of these rooms. And we can speculate and understand what we might do.
37:05 But I think ultimately Every internet business or technology business ultimately has come down to either On the consumer side, distribution then combines with some kind of network effect and lock-in effect, and then you're able to just run away with that and separate from the field.
37:22 And then on enterprise, it's largely been a business that's driven by technology differentiation and Delivery of that. Technology. With Great SLAs with great service with very
37:35 unique approaches to solution delivery. And so Modest comments. on consumer and how consumer is gonna evolve. I think is exactly right. You have
37:46 Meta Google And XAI with consumer touch points. You have OpenAI with an extraordinary brand today with ChatGPT and A ton of consumer touch points already. On the enterprise side, the challenge has been that these APIs have largely to date not been as reliable as what developers expect. Developers have gotten used to
38:09 Because of the excellent work of hyperscalers that If you are out there with APIs For A product That product should be infinitely scalable.
38:20 Available twenty four seven and the only reason the API ever goes down is'cause some giant data center lost power or something. There's very few reasons why an API should fail. It has become the developer mindset to enterprise solutions. And Over the last two years The quality of AI APIs has been a huge
38:41 challenge for application developers. And so what's happened As a result is people have figured out workarounds and have solved all those problems with pure innovation. But going forward In this again, we keep going back to this if If
38:56 Pre training and scaling is not the way to do it. And it's all about test time compute. This is where again we go back to the traditional way of hyperscalers. And I think this is where AWS is extraordinarily advantaged because Azure and Google have great clouds. But
39:15 AWS has the biggest cloud. It has really built for resilience in a way. That's very, very differentiated. And even today, if you're running llama models. You want to run Lama models.
39:29 on AWS, or if for some reason you have some very specific use case and you need to support on prem customers, you can at very large Financial institutions That have complex regulatory environments or compliance reasons. You can run these models on prem if you choose to.
39:47 And AWS has even gone there with V P Cs and GovCloud and all this kind of stuff. And so If we assume that pre-training and scaling there is done, then all of a sudden AWS becomes extraordinarily powerful. And their strategy here could just be
40:06 buttons with everybody in the developer ecosystem over the last couple years and not pursue their own LLM efforts. Well they are pursuing, but not sort of like in the same way that others have. Will likely end up becoming A pretty good strategy because all of a sudden you you have the best API. Service.
40:25 The other part I think is Google. Which we haven't talked about yet is Their cloud is very good at certain things. So they have an enterprise business. That enterprise business is actually pretty scaled now, if you look at the latest earnings.
40:39 And obviously their their consumer business is Is dominant and There has been a perception that They're getting disrupted today. I think these forces are very disruptive to them.
40:53 But it's uncle that the disruption has already happened. What are they doing about it? Obviously they're trying and it's pretty clear that they're trying very hard, but I think it's an interesting One to watch. And the one that I like to watch
41:08 Because It's the class of innovative dilemma, and they're clearly trying to be on the good side of not being innovated away as an incumbent. They're trying very hard. And so there's very few cases in business history of the incumbent Preventing the innovator's attack.
41:25 And if they do defend their business through this era, that'll be an extraordinary achievement. Mm. Yeah, Google is so fascinating because You had a brilliant cell site analyst. Carlos Kardiner, who unfortunately passed away, but
41:42 In two thousand fifteen and sixteen. He spent Many many reports writing about Google's progress. towards artificial intelligence. and the underlying work that they were doing at DeepMind.
41:55 actually was so fond of it, he went to go work at Google ultimately. But sort of first expose this idea of the underlying work that they were doing there. In neural nets, in deep learning. And It's clear they were caught off guard by the brute force scaling of the transform.
42:13 That What advanced this wave of technology was literally throwing compute But if you read any of the interviews with people who foreshadowed this data wall One of the things they talked about was that self play might be A mode to overcome
42:32 the lack of data. And who is better at self play? Than deep mind. And if you look at the pieces that Deep Mind brings. from before the transform.
42:44 And what they bring together. with the transformer and scaling up compute. It seems as though they have all the pieces. To win. Now the question I have always had is not
42:57 Can Google win at AI? It's Is winning whatever that looks like ever going to possibly replicate How good winning was. in the current paradigm.
43:08 That's really the question. To Jaden's point, it would be amazing if they overcome the dilemma and win. But I think they have the pieces there. The question really is If they can build a business out of the assets that they have.
43:22 that in any way looks as good as what is arguably the greatest business model we have ever seen. Which was Internet search. So I'm equally as fascinated to follow them. I think on the enterprise side. They have incredible models and in and incredible assets. I think they have a lot of trust to earn.
43:40 I think that over time They've come and gone in that world. And so I think that's a harder access of attack for them. But certainly on the consumer side and certainly in the model building side, they have all the assets in place to win. The question is just what does that prize look like?
43:58 particularly now if it doesn't look like there may be one or two models to rule them all. Shaitan, I'm curious as an investor seeking a return. What path You personally hope for.
44:11 I Personally hope for AI to continue For a really long time. You need big disruptions as a venture investor.
44:21 To unlock distribution. And If you just look at what happened in the internet or in the mobile and where value accrued Value predominantly accrued at the application layer in those two waves. Now Obviously our
44:36 hypothesis and my hypothesis was that this layer again was gonna be very receptive to distribution unlock because of innovation at the AI application layer. I think that's largely played out so far. It's still early days, but the application vendors that have Come out with production AI applications. For both consumer and for enterprise.
44:58 have found that those solutions which can now only exist because of AI are unlocking distribution in ways that was frankly not possible. In the world of SAS or pro sumer SAS or whatever. We'll give you a very specific example. With an AI powered application, we're now going to CIOs of Fortune five hundred companies.
45:18 Showing these demos. And Two years ago there were really nice demos. Today it's A really nice demo combined with five customer references of peers that are using it in production and in experiencing great success.
45:31 And What becomes very clear in that conversation is that what we're presenting is not a five percent improvement over an existing SaaS solution. It's about We can eliminate significant amounts of software spend. And human capital spend.
45:47 And move this to this AI solution in your 10x traditional ROI definition of software. is easily justified and people get it. Within thirty minutes. And so you're starting to see these what used to be a very long sales cycle for SaaS. In AI applications, it's Fifteen minutes to a yes, thirty minutes to a yes.
46:09 And then The procurement process for an enterprise completely changes. Now the CIO says something like Let's put this in as quickly as possible. We're gonna run a thirty day pilot. The minute that's successful, we're signing a contract and we're deploying right away. These are things like three, four years ago in SaaS was just completely out of the realm of possibility because you were competing against incumbents, you were competing against their distribution advantage, their service advantage. And all this kind of stuff and it was very hard to prove
46:38 why your particular product It was unique. And so Since twenty twenty two And I'll call it since Chat TPT, November twenty twenty-two. That seems like a really good line of pre and post In this world.
46:53 We've made twenty-five investments in AI companies, and for a s five hundred million dollar fund with five partners, that's an extraordinary pace. The last time we had that kind of a pace was Surprise when the app store came out in two thousand nine. And then the pace that we had That kind of pace was again in ninety five, ninety six with the internet.
47:15 And in between those you see us and our pace being pretty slow. We average around maybe five to seven investments a year in non-disruptive times. And clearly now. Our pace is dramatically increased. And if you just look at Of those twenty five companies.
47:31 Four have been infrastructure companies and the rest have been application companies. And we just invested in our first model company, which hasn't been announced yet, but it's two people, two extraordinary brilliant people. that are jumping to the frontier with very little capital. And so
47:49 We've clearly bat and anticipated that there's dramatic innovation and distribution unlock happening at the application layer. We've seen that happen already. These products are truly as a software investor are Absolutely amazing. They require a total rethinking from first principles on how these things are architected.
48:10 You need unified data layers, you need new infrastructure. You need new UI and all this kind of stuff. And it's clear that the startups Are significantly advantaged. Against incumbent software vendors. And it's not that the incumbent software vendors are
48:26 Standing still. It's just that innovators dilemma in enterprise software is playing out Much more aggressively in front of our eyes today. than it is in consumer. I think in consumer the consumer players recognize it and are moving it and are doing stuff about it. Whereas I think in enterprise it's just
48:44 Even if you recognize it, even if you have the desire to do something. The solutions are just not built in a way that is Responsive to Dramatic re architecture. Now could we see this happening? Could a giant SaaS company just pause selling for two years and completely re-architect their application stack?
49:04 Sure, but I just don't see that happening. And so If you just look at any sort of analysis on what's happening on AI software spend. Something like it's eight ex year over year growth between twenty twenty three and twenty twenty four. On just pure spend. It's gone from a couple of hundred million dollars to well over a billion in just a year's time. And you can see this pull, you can feel this pull if you're in any one of these AI application companies. It's like,
49:33 More of these companies are supply constrained than demand constrained. We talk to the CEOs of these application companies And they just say things like, Well, I see demand as far as I can look out. I just don't have the capacity to like go service all the people that are saying yes to me. So I'm gonna segment it and go to where they are. And My hope as an investor is that it continues to play out this way and that we have
49:59 stability to just pursue these angles. And frankly, the model layer stabilizing is a huge boon for this application layer. Primarily because As an application developer, you were sitting there watching the model layer take step function leaps every year. And you kinda didn't know what to build and what you should just wait on building, because obviously you want to be completely aligned with a model layer. Because the model layers are now moving to reasoning.
50:27 This is a great place for an application developer. One thing You know as an application developer is humans are not patient. And so You need to Always build solutions that optimize on performance and quality.
50:40 You cannot go to a user as an application developer and say, like, I'm gonna deliver a high quality response. It's just gonna take longer. Well that's not been a winning argument. Now for certain use cases is that Possible, could you have it run in the background for 24 hours? Absolutely, but those use cases are not widespread. And predominant and people aren't gonna
51:01 be willing to buy that kind of stuff. And so If As an application developer today, all of my board meetings over the last couple of weeks have been these companies saying In this new reasoning paradigm.
51:13 We're really confident that we can invest in these four things that we've been super hesitant to. In the last year and a half. But now we're gonna go all in on these bets. And the kind of performance gains you're gonna see out of our systems is gonna be huge. Sorry, why is that the case? Why does the reasoning thing make it so that their confidence goes up? Just like spell that out.
51:33 Well, if you are an application developer and you're Looking at The models today. And you're saying I can see clear efficiencies for my use case, but I have to invest in these five infrastructure layer things and these UI things. But if a new model comes out in six months and blows all that investment away, just because the model itself can do it.
51:54 Then why would I ever invest in these things? I'm just gonna wait for the model to do it and then bet on that. But in this reasoning paradigm if All the labs pursue reasoning. And reasoning is intelligence on the y axis, time on the x axis. And that's where we're going.
52:10 Then any improvement that I can make in my own tool to make either that reasoning time dramatically compressed because of the way algorithmically I'm feeding reasoning and I'm able to take the data and manipulate it and all that kind of stuff. I should invest in it now if reasoning is now the new paradigm. And the last mile delivery at the application layer against these reasoning models means that I'm building technology And tooling that
52:35 model companies are very, very unlikely to build. And as those using systems continue to get better, my last mileage and last mile delivery systems understand are still advantaged and defensible. Do you both have favorite examples of this beyond coding and customer service, which seem to be the two dominant And incredibly exciting and cool
52:56 Use cases with lots of companies chasing after versions of that. Do you have other favorite examples that you know would fit the CIO of the Fortune Whatever company saying, like, we need this in our company now? I can give you twenty of them. Maybe like categorically is my question. There's coding, there's support.
53:19 Essentially top down. Look at the biggest spends of enterprise software. And you could attack that with An AI powered AI first solution. And so We've got a great company called Alove and X that's going after sales automation. We've got a great company called Leia that's being used by lawyers to dramatically increase the efficiency of their work. I think legal has been a very interesting question.
53:42 Because people assume that lawyers work on billable hours. If you're automating billable hours, aren't their economics gonna change. Well, now the evidence two years into this is that actually lawyers end up becoming way more profitable by using AI. And the reason is that a lot of the work That was ro repetitive. And hard. And done by junior people inside of law firms.
54:04 Law firms weren't able to bill for that stuff anyway. And so if you can take down the time to do document analysis from three or four days to twenty four hours. All of a sudden you free up all your lawyers to do all the strategic work that they can bow for. And stuff that is extremely valuable for clients.
54:22 We've got a company that's automating Accounting is an example and financial modeling. We've got a company that's Changing how game development is working. We've got somebody that's going after circuit board design. It's been a hugely manual and human intensive thing and and computer systems are particularly really good at. And we recently
54:41 Invested in something going after an ad network. Now that's been something that's not been touched for a long time from startups, but it turns out Matting The people that have Inventory. with the people that want to do advertising in the AI world is just way more efficient. And so
54:59 We invested in a company that's got a new document processing model. And they're going after open text. When was the last time a startup Part about open text. It's been a long time.
55:11 Where these huge incumbent SaaS markets were thought to be open to new startups. So you had to pursue. More niche, more vertical. And I often joke because like I saw this, it was payroll for field workers working in eastern Europe. was the SaaS company that you like legitimately had to think about in 2019. And now we're back to like large swaths of horizontal spending in to say like
55:38 Hey, there's an incumbent here that's worth ten billion plus. The market here is ten billion of annual spend. AI makes a product here Easily ten X better. Faster
55:50 And all the things that users want when they See it and you need a new platform. to come out with that kind of an advantage and that's what this is. And Patrick, you asked back at the beginning about the big debate on RLI and CapEx and all this and
56:07 When you listen to Chip and when you listen to other investors in the application layer, when you listen to the hyperscalers. The big takeaway over the last three months is The use cases are coming. Yes, everybody knows about coding. Everybody knows about customer support. But this is really starting to permeate and get out.
56:26 into the broader ecosystem and the revenues are becoming real. The challenge on the ROI question always was. Okay, you put the capital in here. You then amortize it over the entrance. But meanwhile, you're then stacking the next
56:43 quantum of capital for the next model. And so everybody could draw those extrapolations and say Oh my God, it's not just that Microsoft is gonna spend eighty five billion dollars. In cash capex. Inclusive of the leases in twenty twenty five.
56:59 It's what does it mean for twenty six, twenty seven, twenty eight. Because the pre trained models were getting so big. If and again it's an F. We are plateauing and we're spending less money on pre-training. and moving that capital towards inferencing.
57:16 We know that spend is coming. We know the revenue generation of the customer is coming. And so it becomes much more easy to say This spend is warrant. I think it is important that people remember the underlying clouds of these companies, meaning just the normal storage and compute, are still growing. high teens to low twenties.
57:39 So there's some capital that needs to be allocated towards that when your hundred billion dollar business growing eighteen percent. You're a sixty billion dollar business growing twenty five percent. It's the Incremental capital above that, everybody was Very
57:53 concerned about six, nine months ago. My personal takeaway coming out of Q three was Okay, I see it. There's use cases here. The inferencing's happening. Technology is doing what it's supposed to. The cost of inferencing is plummeting.
58:08 The utilization is soaring. You put that together, you get a nice rising pot of revenue and everything's good. Sachin and Adela talked about this. The challenge is you spend the money for the model. You get it on inferencing, but then we're spending on the next model. If we can start to say
58:24 Hey. Maybe we're not going to spend the next fifty billion dollars on the model. the ROI calculation looks a lot better. And something that you asked Chasen was Why is stability of the model layer?
58:37 Important. I think Sam Alton gave the right answer on this, which was six months ago. He was on a podcast and said If you're scared of our next model being released. We're gonna run you over.
58:48 If you're looking forward to our next model coming out. Then you're in a good position. Well, if the actual Reality is The next model is gonna be at inference time and not
58:59 Retraining. You probably have less worry about them steamrolling. So I think everything that we're talking about in this one pat is very conducive. to a favorable economic reality.
59:14 For the entire ecosystem, which is all the Attention capital being put towards infrastructure. The real concern was Do we need to spend fifty hundred two hundred billon dollars? to build these ever more accurate models in pre training.
59:30 Where do prices most reflect Extreme optimism or hype still. I've certainly seen my fair share of private markets companies, let's say series A type companies, that price at Extremely high valuations. They're often incredible teams and very exciting.
59:46 But they're also playing in spaces where if something works, you could imagine lots of other very smart investors funding some competitor. So you see these scenarios where it's like, Great team, high price, high potential competition, really exciting. Everything's moving fast. I'm curious what signals you both read from valuations and or multiples. Right now.
1:00:06 In the private markets One of the things that's happening is just the dramatic drop in prices. Uh Just compute Whether it's inference or training or whatever, because
1:00:19 It's just becoming way more available. If you're sitting here today as an application developer Versus Two years ago. The cost of inference of these models is down a hundred X, two hundred X. It's Frankly, outrageous. You've never seen cost curves that look this steep.
1:00:36 That fast. And this is coming off of Fifteen years of cloud cost curves, which were amazing and mind blowing by themselves. The cost curves on AI are just a completely different level. We were looking at cost curves.
1:00:51 And the first wave of application companies that we funded in twenty twenty two. You look at the inference cost and it would be like fifteen to twenty dollars per million tokens. on the latest frontier models. And today Most companies don't even think about
1:01:06 inference costs. Because it's just like Well, we've broken this task up and then we're using these small models for these tasks that are pretty basic, and then we're like the stuff we're hitting with the most frontier models are these like very few prompts. The rest of the stuff, we've just created this intelligent routing system. And so our cost of inference is essentially zero, and our gross margin for this task is 95%. You just look at that and you're just like, Wow, that is a
1:01:32 totally different way to think about Application gross margins. Then what we've had to do with SaaS and what we've had to do with Basically software for the last decade plus.
1:01:44 And so I think that's where you're starting to look at and saying the entire application stack. for these new AI applications, and it starts with people that provide inference. It starts with the tooling and the orchestration layer. So we have a portfolio company
1:02:00 That's extremely popular called Ling Chain and the inference layer we have fireworks. These kinds of companies are seeing extraordinary usage by developers and then all the way up the stack to the applications themselves. I think just the pace of innovation, pace of commercial success is driving a lot of excitement.
1:02:17 With private investors. What is also appealing of model stability is now we can finally assume if this sticks. That all these companies are gonna be fairly capital light. Because If you're not having to spend a lot on pre training, if you're not gonna have to spend a lot on inferencing,'cause most of the hyper scalers are now gonna present you with really reliable APIs.
1:02:38 At these kinds of costs. It's a great time to be in the application development business and it's a great time to be in the application development stack. Modest, what do you think? On valuations. I think you have to start in general with animal spirits.
1:02:53 If you go back to the week before Chat GPT was released. If you go to the fall of twenty twenty two. Tech had probably just suffered its most brutal bear market.
1:03:06 Since The dot com. Collapse. It was arguably worse for the median tech stock than even the financial crisis. You had some of the very large growth funds down sixty, seventy percent.
1:03:19 You had The hyper scalers laying off people for the first time ever. You had Capex cuts, you had Op excuses. It was a
1:03:29 very different Vibe. In the tech world. And in the public markets at large. The release of chat GPT
1:03:38 Catalyzed the reemergence of animal spirits and it's been a progressive process. But I think where you are today, you have the public markets trading at twenty four times earnings. And again, this goes beyond just the Mag seven at this point. I mean, Google trades it, I think nineteen or twenty times. So they're not one of the offenders here.
1:04:00 And so I think in general, there is a lot of optimism baked into the public markets. A lot of which is tied thematically to this idea that we're in a new platform era. And that the sky is the limit for a lot of various new concepts. So there's that global overhang.
1:04:21 If we are right. I think that what it really comes down to is understanding What does this new Path forward look like. If
1:04:32 CapEx. And hyperscaler OpEx is more closely tied to revenue generation. If you listen to AWS, one of the fascinating things they say is they call AWS a logistics business. I don't think anyone Externally it would sort of look at
1:04:49 cloud computing and say, Oh yeah, that's a logistics business. But their point is Essentially what they have to do is they have to forecast demand. And they have to build supply on a multi year basis to accommodate. And over twenty years they've gotten extraordinarily good at
1:05:05 What has happened in the last two years, and I talked about this last time is You have had An enormous surge in demand. hitting inelastic supply because you can't build data center capacity in three weeks. And so if you get back to a more predictable cadence of demand.
1:05:22 Where they can look at it and say, Okay. Війно. where the revenue generation is coming from. It's coming from test time. It's coming from Chasen and his companies rolling out. Now we know how to align supply with that.
1:05:38 Now it's back to a logistics business. Now it's not Ry Mothballed nuclear site in the country and try to bring it online.
1:05:49 And so instead of this land grab, I think you get a more reasonable, sensible, methodical roll out of It may be, and I actually would guess. That If this path is right. Yeah.
1:06:03 Inference overtakes training much faster than we thought. And gets much bigger than we may have suspected. But I think the path there. in the network design is gonna look very different. And it's going to have very big ramifications for the people
1:06:19 who were building the network. who were powering the network. Who were Sending the optical signals through the network. And all of that, I think.
1:06:30 As not really started to come up in the probability weighted. of a huge chunk of the public market. And look, I think most people overly fixate on NVIDIA. because they are sort of the poster child of this.
1:06:47 But there are a lot of people downstream from NVIDIA. That will probably suffer more because they have inferior businesses. NVIDIA is a wonderful business doing wonderful things. They just happen to have seen the largest surge in surplus. I think that there are ramifications far far beyond.
1:07:07 Who is making the bleeding edge GPU, even though I do think There will be questions about okay. Does this new paradigm of Test time compute. Алавфор кастомізацін.
1:07:21 at the chip level much more than it would have if we were only scaling on pre trade. But I think this question Whenever I have this in normal conversations, people overly fixate on NVIDIA.
1:07:33 I think people like to debate. that particular name. But I think there's a lot of other derivative plays of the AI build out. Where the distribution of outcomes has shifted and that has not been reflected yet.
1:07:46 I just think it's really important to think about in the test time and the reasoning paradigm from an application layer, how many of your prompts Actually utilize Reasoning. as a way to respond to those prompts and
1:08:02 Yes, application developers as the technology becomes more available in usable will use way more of it than they are today, but If you just look at the current techniques and the wows you're getting from the application layer already, What percent of prompts or what percent of queries are gonna use reasoning It's very hard to squint and say it's gonna be ninety percent of queries.
1:08:26 That doesn't seem like it's gonna go that way because again Your users are not gonna wait. humans are inherently impatient and you've a solution that's like just spinning And thinking your users are gone. It doesn't matter what sector they're in, they're just gone. And so yeah, you can have a certain set of tasks that take a long time and deliver great accuracy.
1:08:46 But speed is by far the most important consideration for these application developers. And so Are we just gonna have a system that continues to just go back and through and back and through? And utiliz all this compute in what market share of queries use that. It's hard to imagine that being super majority of queries. And so then the implication
1:09:08 At least From a private market. early stage investor, which take huge grains of salt on what it means for anything other than my world. But the implication there is simply that You just don't need as much
1:09:21 Compute. As you did with training. Training is just a constant exercise. You're scaling and you're just really hitting all your compute power all the time and is going.
1:09:33 And the application layer. It's extraordinarily bursty. You're gonna have certain tasks that need a lot right away. And for a lot of it, you just don't need a lot. And so this is where again, like hyperscalers and Things like E C two and S three were incredible and now in this new world
1:09:49 The solutions from hyperscalers are really terrific. I think AWS's training um and the TPUs from Google are really, really terrific and they offer a great Developer experience. I think part of What has been known for application developers is that GPUs are really tough to use. For this use case.
1:10:09 Getting max utilization out of GPUs. Chain together. Whether you're buying that from Dell or whether you're buying it from a hyper scaler is just really hard to use. But With new software innovations, that's obviously gonna get better. And then the stuff that's coming out from the hyperscalers themselves, they're really, really terrific. And you just don't need to
1:10:28 Hit him as hard as you did when you were doing training when you're doing Test time compute. I think it's a really important point. in the utilization of the GPUs. If you think about a training exercise
1:10:40 you're trying to utilize them at the highest possible percent for the uh a long period of time. So you're trying to put fifty hundred thousand chips. In a single location And
1:10:52 Utilize them at the highest rate possible for nine months. What's left behind is a hundred thousand chip cluster. that if you were to repurpose for inferencing is arguably not the most efficient Bill.
1:11:06 Because inference is Kiki. And bursty. And not consistent. And so this is what I'm talking about that
1:11:14 I just think from first principles you are going to rethink. Wow. You want to build Your infrastructure to service A much more inference focused.
1:11:27 World. than a training focus world. And Jensen has talked about the beauty of NVIDIA is that you leave behind this in place infrastructure that can then be utilized. And in a sunk cost world, you say, sure, of course, if I'm forced to build a million chip supercluster in order to train a fifty billion dollar model.
1:11:49 I might as well sweat the asset when I'm done. But from first principles. It seems clear. You would never build a three hundred fifty thousand Chip cluster.
1:12:01 Well Two and a half gigawatt of Hour. In order to service the Type of request.
1:12:07 that Chaithan's talking about. And so if you end up with much more edge computing with low latency and high efficiency. What does that mean? for optical networking. What does that mean for the grid? What does that mean for the need for on site power versus the ability to draw from the local utility.
1:12:27 I think these are the types of questions I would be very interested to read about But to date a lot of the analysis is still focusing on What's gonna happen when We light up three mile island.
1:12:42 because the new paradigm is really too soon to change. Do you think that we still need and will see though Tons of innovation in the semiconductor world and layer. Whether it's networking, whether it's optical, whether it's chips themselves, different kind of chips.
1:12:59 I would imagine this would accelerate it even more. Because It was very difficult to foresee a world where you took on Big green in Train.
1:13:11 The way I think about this. Over hundreds of years is you have a gold rush. A land grab and everybody's just doing whatever they can. But in technology, then as some stability sets in, you get an optimization period. You've already had that.
1:13:26 On The inferencing side it's what Chief and referenced is. People had time to optimize. the underlying algorithms in compute and Inference has fallen ninety nine percent.
1:13:38 It's the same thing that happened with Internet Transit. At the end of the bubble. Which was People said, No, you can never stream a movie. Do you have any idea how much that would cost? And the cost of transit has fallen twenty-five percent a year, like clockwork, for twenty years.
1:13:55 The literal profit pool of that business is static for twenty years. And so I think We've had this mammoth. Demand search. And I think that if we get a little bit of stability and everyone can take a breath.
1:14:10 There will be the two guys in the garage optimizing every single thing possible. And that's the beauty of technology over the long term is it is deflationary. Because it's an optimization. problem. But you don't have time to optimize when you're in land grab mode.
1:14:29 I quoted this to you last time. The data center industry They were power neutral. There was no demand growth in power for the entire data set of business for five years. That was because you were in the ful stage of the cloud data center build up.
1:14:43 I don't know when you'll reach that point. I mean, we know that these guys on a three four year through at least twenty six or twenty seven. are pot committed to Their build out. When in that path will everyone have time. To take a deep breath and say, Okay.
1:14:59 Now let's figure out how to run these more efficiently. That's just the nature of things. Same thing on the compute side. I just think We haven't yet gotten to the point where technologists have been able to apply their optimization. They've been in the implementation. And I'll give you uh a couple of data points from my end. So
1:15:17 My partner Eric is on the board of a great semiconductor company called Cerebros, and they recently announced that Inference on Lama three point one four or five billion for Cerebrus is it can generate Nine hundred plus tokens per second, which is A dramatic order of magnitude increase. I think it's like seventy or seventy-five times faster than GPUs for inference as an example. And so
1:15:41 As we move to the inference world The semiconductor layer, the networking layer, et cetera. There's tons of opportunities for startups to really differentiate themselves. And then the second thing I would bring up is I was just recently talking to the CIO of a large financial services institution.
1:15:58 Who said that Over the last two years they were Were you buying a lot of GPUs because they assumed that they were gonna have lots of AI workloads and who knows if maybe they needed to do some training themselves. And so those systems are now being installed into their data centers and they're now online. And we're in this world where
1:16:16 You don't need to create your own model. And even if you did, you just fine tune an open source model and it's not that heavy. And so His view is like look, if you have AI applications that run on prem, like it's essentially free. I have all this capacity. I'm not using it for anything. Inference is light.
1:16:33 And so I have at the moment infinite capacity to run AI applications. On from And it'll cost me zero marginal dollars. Because it's all s all the stuff is up and running and I'm not using it for anything.
1:16:46 So I'm ready to buy. So not only are all these application things that you're talking about hugely exciting because they unlock R and all the stuff. But the minute you can run any of this stuff on prem on our stuff. Та дramatically decreases the cost for us. And so it's just like win win win all over the place when you have something like that. And that's the current state of play. Now how long does this
1:17:07 Overcapacity last application developers are famous for using all the capacity and pushing the limits, and all of a sudden what used to be overcapacity ends up becoming Undercapacity because all of a sudden we have all this breadbended build out and we decide to stream a video on it. And so of course AI applications are gonna get more sophisticated and Swallow up all this capacity. But that is just a much more predictable world.
1:17:29 And much more sane world from an investment perspective. then scaling to infinity on Pre-training. The one thing I am curious to monitor is
1:17:43 It's important to remember that the reporting was not that the models weren't getting better. It's that the models weren't getting better relative to expectation. or the amount of compute applied to that. So I think we do need to be cautious. To conclude
1:18:01 that the labs are not gonna keep trying to figure out the unlock on the pre-training side. I think the question there is one, what should we be looking for? But then two is If they continue to push on that vector. Do we believe
1:18:18 And this was the question I always wrestled with was If scaling laws held in pre training. Would people be willing to spend a hundred billion dollars? And I know that everybody says if you're playing for the ultimate prize, you would.
1:18:34 But has enough doubt been cast. That simply brute forcing pre training. Is the path. to that ultimate unlock. Or is it now some combination of pre-training, post training, and test time compute?
1:18:51 In which case, again, I think that the world is just the map. Is much more sane. And I've seen a lot of notes coming out saying people are declaring the end of AI progress and all that. And hopefully. The takeaway from today is none of that is what I think people really in the weeds looking at this are saying. People are saying AI is full speed ahead.
1:19:14 I think the question is just what the axis of Advancement is. And from my seat The math seems much more sensible. Everything seems much more rational. Pursuing this path.
1:19:28 Rather than The Up front cost. being spend any amount you can. to build this hypothetical God.
1:19:38 So I think this is a much better outcome if this is the path that it we end up going down. I'm curious what you think, if anything, is the most under discussed part. Are there things that you find yourself thinking a lot more about than you hear discussed from your friends and colleagues? On
1:19:57 The public investor side. Just reading Southside. Reports that We haven't seen cell side reports or analysis on what this new paradigm of test time compute means and how things change. And so I'm really looking forward to way more cell side analysis on this new paradigm shift.
1:20:13 I think there's also Little coverage in the private markets. I think it's known to people that are meeting these entrepreneurs is just how capital efficiently these entrepreneurs are getting to the frontier today. And this is just a shift that's happened very, very recently. And you're seeing people just show up and having spent under a million dollars to match performance, not broadly, but in specific use cases with the frontier models. And that's just not something that we were seeing two years ago or even a year ago. And so
1:20:41 I think that's pretty dramatically undercovered. Free training is a big test of capitalism. If we pursue down this path. I feel much better with a microeconomic background. Analyzing what's gonna happen.
1:20:54 Because you don't have to put in the NPB of God. And I just think that that's much better. In terms of what I'm looking forward to reading and hearing, yeah, I I'd love to see And thoughtful analysis really wrestle with Right now I feel like it's a little defensive.
1:21:12 People are defending the fact that scaling's not done, it's just move. So that's great. But let's now work through the second order effects, the third order of facts and How does this really manifest itself? I think it's very good for the overall ecosystem, the overall economy.
1:21:30 But I think there's gonna be a lot of surplus shifting from pockets that look like winners before and pockets that look like losers. What outcome in the next six months would most disorient you? Well two Dramatic examples on the positive side, if somebody came out with the results that Po training was back on and there was a huge breakthrough on synthetic data, and all of a sudden
1:21:53 It's go go again and Ten billion dollar and hundred billion dollar cluster would be back on the table. You would go back, but all of a sudden the paradigm would be Wild. All of a sudden we would now be talking about a hundred billion dollar
1:22:07 Supercluster. That was gonna pre-train And then obviously If my expectation comes out that next year we're gonna call AGI We're gonna have AGI and we're building a hundred billion dollar cluster because we had a breakthrough on synthetic data and it all just works and we can just simulate everything.
1:22:23 That would be pretty dramatically disorienting. I think Another scenario is It's pretty clear now that While we've exhausted data on text, we are not close to exhausting data on
1:22:37 Video and audio. And I think that It's still TBD on what these models are capable of. on new forms of
1:22:47 Modes. And so we just don't know because the focus hasn't been there, but now you're starting to see Large labs talk more about audio and video. What these models will be capable of from a human interaction perspective.
1:23:01 I think it's gonna be pretty amazing. I think you've just seen already how much leaps have gone into image generation and video generation. And what does that look like? In a year's time and two years time. Could be pretty dramatically disorienting. Yeah, I think the hard part as a non technologist is
1:23:20 For the last year, year and a half, the question has been What would GPT five bring if it adhered to the scaling. And no one could really articulate because all we know is
1:23:33 Okay. Trying loss would be lower. So you'd say, okay, this thing's more accurate at next token prediction. But as far as what does that actually mean from a capability standpoint, what's the emergent capability we were unaware of before it was released? So I think it's really hard to know ex ante what you're looking for.
1:23:52 Other than the labs coming out and saying This is so good in its accuracy. That it warrants Staying on this long linear trajectory.
1:24:05 Of spec. And if someone comes out and says that, I think Irrespective of this entire conversation or what you may believe, you have to say, Okay, that's happened. Again, I just think you have to have a super open mind. And if we were
1:24:19 But it wasn't in the open. I just think you have to be updating your priors constantly. So clearly, like Jason said, I'd be looking for that. Personally, I watch Lama closely. There's clearly a risk at some point that they decide not to keep open source it.
1:24:36 And if I were other players in the ecosystem, I would be doing my damnedest to make sure that Lama stays Open. And there are certain ways you could go about doing that, but I think that's one thing because there Willingness to spend at the frontier.
1:24:56 And make those models available the way they do. I think has completely changed. the strategic dynamic in the model industry. So That's another one that I would be paying attention to.
1:25:09 I have a philosophical question as we near the end of the discussion. Which is around ASI. So if AGI is here or coming next year How the both of you would even think about I guess it builds on that point about what do we even expect from a GPT five that stays on the scaling wall.
1:25:26 What does it mean? Because There are fewer and fewer things, at least in a simple chat interaction. than I could imagine it doing a much, much better job on, or even what that would look like. And Again, we're probably just in the early innings of application development and fine tuning and improvement and algorithmic updates and blah blah blah. So
1:25:44 I'm curious just philosophically what you think the litmus test could or might be for something beyond What we have naturally as the existing models get tweaked and tuned and better. What does the ASI even mean? Does it mean it solves previously impossible math or physics challenges or something else?
1:26:01 What does that idea mean to you both? These are my words. I don't remember who originally said them, but Humans are really good at changing the goalposts on expectations and AI in the nineteen seventies meant something different than what it meant in eighties and nineties.
1:26:20 Two thousands and in twenty twenty four. And so If a computer can do it, humans have a really good way of describing that as automation. And whatever a computer can't do, that now becomes the new goalpost for AI. And so I think that these systems are already extraordinarily intelligent. And are extraordinary at Replicating human intelligence and sometimes exceeding human intelligence.
1:26:44 I think if you just look at the path that some model developers like Deep mind and several startups are pursuing with things around math and physics and biology. It's very clear that there's gonna be applications and outputs of these models that are going to be things that humans were simply not capable of doing before.
1:27:04 We already have seen that. In things like protein folding today. We're starting to see a little bit of that. As it relates to math proofs. I am confident we're gonna start seeing that as it relates to physics proofs.
1:27:15 And so my optimistic hope for humanity is that I don't know, we'll be able to open wormholes or something. We're gonna be able to Study general relativity at a scale that we haven't been able to before, or study black holes or simulate back holes in a way that we haven't been able to. All that sounds a little bit ridiculous at the moment, but certainly like the way things are progressing and the way things have progressed. We don't know what is possible and what's not possible. And to bring it back to an investor point of view.
1:27:41 When you have the unknown future where the possibility is Up to your imagination. That's a usually a really great time to be an early stage investor. Because that means that Technology has unlocked.
1:27:55 And usually when technology unlocks in a grammatic fashion, distribution also then unlocks. And you can now go get customers that were very expensive to get. And so previously, if you wanted to build a consumer application You then had to factor in the tax of the app stores, search, ad networks, and all that kind of stuff. And all of a sudden it just became a very quick Exercise in unit economics. And similarly in SAS it was like A productivity, gross margin, and infrastructure costs, and you just tried to do a spreadsheet exercise.
1:28:27 And early stage investing started to become more spreadsheet like than true technology innovation. I think when you have big breakthroughs like this. Everything sort of changes again. Like distribution is Nearly free if you have something unique and there's like a word of mouth And morality factor to it. The technology spend
1:28:45 is really again, you go back to just investing in your developers and your research scientists and the R D and the ROI and R D starts to become remarkable again. That's what's most exciting as an early stage investor is that we just don't know what the future holds and therefore it's back to human ingenuity and People able to push these boundaries. When it's exciting for a early stage investor, I think it's terrifying for a somewhat sceptical public market investor. Prices are based on vibes and not math in the spreadsheet. Yeah.
1:29:16 Oh. With ASI. I think we talked about this before. This whole concept, the reason people spend so much time on it is Because it is so profound ultimately. You have people who are invoking
1:29:30 quasi religion. in their view of what we're building. Any time that comes into play. I think the stakes are just higher.
1:29:38 It's kind of unknowable. And it's super complicated. So we all love to debate it, but I think the one thing we haven't touched on here, which is There's a pretty fervent belief amongst a group of people that There will be recursive self improvement.
1:29:52 At some point in time. And I think that that would be a big path to unlock in Whatever hypothetically ASI means is when the machines are smart enough To learn themselves and teach themselves. On a less sort of dramatic
1:30:07 You the way I think about this. There's Alpha Go, which famously did that move that no one had ever seen and can I think it's like move thirty seven everybody was super confused about and ended up winning. And Another example I love is Norm Brown'cause I like poker, talked about his poker bot.
1:30:24 confused. It was playing high stakes, no limit. And it continually overbed. Dramatically larger sizes. than Crows had ever seen before. And he thought the bot was making the mistake.
1:30:37 And ultimately it destabilized the pros so much. Think about that, a computer destabilized humans. And their approach. that they have to some extent taken on overbetting now into their game. And so those are two examples where
1:30:53 If we think about free training being bounded by the data set that we've given it, if we don't have synthetic data generation capabilities. Here you have two examples where algorithms did something outside of the bounds of human knowledge. And that's what's always been confusing to me about this idea
1:31:13 that LLMs on their own could get to superintelligences. Functionally. They're bounded by the amount of data we give them. Up five. And so if you have examples like this where algorithms are able to
1:31:29 Get outside. of what they're initially bounded by. That's super interesting. I'm not smart enough to know where that leads us. But that's the kind of thing that I feel like is the next thing to come is how do you escape the bounds of what you're given upfront? I think what's
1:31:46 Remarkable. From my perspective is how much of this innovation is happening in the United States. And how much of it is happening in Silicon Valley? We've had a rough couple of years since the pandemic and It's really amazing.
1:32:00 There was an investor friend of mine who's not based in Silicon Valley and he was just saying, I can't believe it's happening in Silicon Valley again. And it's just become this beacon where All of the labs were based here, a lot of the people that are working on these applications, these infrastructure companies, et cetera, are here. Or even if they're not here, they're somehow connected to being here and are often visiting here a lot. And
1:32:24 I would say that The focus on the innovation here Is really extraordinary on a the progress being made in the United States. In Silicon Valley specifically is extraordinary. I do think that there's
1:32:38 a level of attention that investors and entrepreneurs now have is how fragile the system is and how much we need to protect it and continue to invest in it. And I think there's a lot of focus Now that
1:32:50 Innovation is something that Needs to be protected and But I think a lot of people are now paying a lot of attention to make sure that All of this innovation that's happening in the United States. Continues to benefit.
1:33:02 Everybody. And I think that's the really optimistic and cool thing to recognize. the agglomeration of tracks are real. If the reporting is right, the way the transformer paper came to pass. is that someone was roller blading down the hall and
1:33:17 heard two guys talking about something and went in and whiteboarded it. Two more people came by and Who knows how much of that is apocryphal at all, but It is fascinating to see from an economist standpoint that like these human network effects are real. And that COVID did not destroy them and that work from home did not destroy them.
1:33:37 And that there really is something tangible to being together and The synthesis of ideas. multi disciplinary coming together to build this world changing architecture. Guys, it's always such a blast talking to you both. I'm lucky to get to do this in private. It's fun to do it in public. Thanks for your time.
1:33:54 Worse. Thank you. If you enjoyed this episode, check out Join Colossus.com. There you'll find every episode of this podcast complete with transcripts, show notes, and resources to keep learning. You can also sign up for our newsletter, Colossus Weekly, where we condense episodes to the big ideas, quotations, and more, as well as share the best content we find on the internet every week.
What you see above is a preview of the first minutes. One unlock costs 10 credits and covers this episode forever: full segment and word-level timestamps on this page, plus .txt, .srt, .vtt and word-level JSON downloads, as many times as you like.