Transcript

Sergey Levine - Building LLMs for the Physical World - [Invest Like the Best, EP.465]

Free .txt

0:02 And welcome everyone. I'm Patrick O'Shaughnessy and this is Invest Like the Best. This show is an open ended exploration of markets, ideas, stories, and strategies that will help you better invest both your time and your money. If you enjoy these conversations and want to go deeper. Check out Colossus, our quarterly publication with in-depth profiles of the people shaping business and investing. You can find Colossus along with all of our podcasts at Colossus.com. Patrick O'Shaughnessy is the CEO of Passitive Sum.

0:29 All opinions expressed by Patrick and podcast guests are solely their own opinions and do not reflect the opinion of Positive Sum. This podcast is for informational purposes only and should not be relied upon as a basis for investment decisions. Clients of Positive Sum may maintain positions in the securities discussed in this podcast. To learn more, Visit PSUM dot VC Um My guest today is Sergey Levine, one of the co-founders and researchers at physical intelligence. As a disclaimer, I'm an investor in physical intelligence because I believe it's one of the most important companies tackling the problem of robotics.

1:05 As you hear us discuss today, robotics has what I would call a scarecrow problem. All of these amazing physical devices are becoming ever more possible in all sorts of cool permutations. But what they all really need is an intelligence, a brain. And that is what they're developing at physical intelligence. They're trying to develop foundation models that can make any physical robot do any task in any environment. The nature of our conversation today is all of the problems facing robotics and all of the promise of solving these problems across the world. I hope you enjoy this great conversation with Sergei Levine. Serge, this is gonna be a real treat and a blast to learn about.

1:43 Possibly the most exciting impactful area of technology being developed. Just to set the stage before we go back in time. Maybe you could just define physical intelligence as you see it. Fundamentally the goal of physical intelligence is to develop robotic foundation models that can control

1:59 Basically any bodied system to do any task. Broadly speaking. You could imagine that In the same way that a language model It's kind of

2:07 rapidly evolving towards a system that can do any task that can be expressed in language, what we would like is to Building. a new class of models that can do any task that can be done by a physical actually device. Part of the thesis of this company is that We believe that doing it at the full level of generality

2:22 might actually in the long run be easier than trying to special case, very specific narrow application domains. Again, in much the same way that for language models, it turned out to be Easier in some ways. Uh Solve.

2:34 natural language task in their full generality. than to narrowly target like machine translation or sentiment analysis or whatever. That may not be obvious why you would make that that versus A robot that just does your dishes or something? What are the key trade offs to understand and why make the decision that he made?

2:50 In the world of natural language we saw that there were a lot of efforts to develop domain specific solutions that tackled specific problems. Somebody would spend a lot of time thinking about how like English differs from French and then build a machine translation system. The reason that language models

3:05 took over for all of those different application domains is because they can leverage much broader sources of data. It's Not even as simple as saying like, Oh, we have this data for this application, this data for this application, but like merge everything. It's actually more than that. It's when you can leverage Weekly label data.

3:21 In the case of language models, they're just mine from the web. you actually learn more about the world. So you've established like foundation of world understanding. And then on top of foundation turns out to be much More effective to build out different applications. To bring this into robotics. the calculus doesn't look quite the same because in in robotics we don't have like a intranet size data set that we can just draw on.

3:39 But this notion of understanding the world, if anything's actually more important robotics because If you have many different tasks, maybe even many different physical systems.

3:49 Then you can go from training individual dishwashing specialists or laundry folding specialist. And instead train a model that actually understands physical interaction. People can master new skills very, very rapidly.

4:00 Because we understand physical interaction, we can intuitively grasp what's gonna happen in this new unfamiliar situation. They'll just like bootstrap things really, really quickly. If we can draw on data from many sources, many applications, many robots. then we can have a model that has a physical understanding and it'll be much, much easier to put new applications on top of that platform. What is the hardest part about

4:21 Building in this way. For you When you see other approaches that are more may be legible to the average person. Oh, there's a robot moving around doing this one specific thing. It looks a certain way.

4:33 What's the hardest part about This approach. As you're doing it. I think this has actually been kind of an issue in my whole career because When you work on robotic learning

4:42 The more general the more th this becomes important is Effective robotic learning, effective generalization isn't actually the optimal way to have like a really exciting demo. The way to have a really exciting demo is to pick a really cool task. control everything else in the environment, like set it up so that it's perfectly clean.

4:59 Perfectly. pristine and just make it work in that one setting. That's the way you make a robot demo. And generalization. You can't just show it.

5:07 In one spot. The point of generalization is that it does something relatively mundane that any human could do, but it does it in any situation. So we we had some demos that we released last April. Where we showed our robot cleaning kitchens. It's cool, but if you watch an individual video out of context, it's just like okay, it's like picking up plates, like anybody can pick up plates.

5:24 Except that we just put it into that home. just for that demo and it never had training data from that setting. So obviously you kinda have to like understand what's going on to appreciate why this is actually pushing the frontier. What is your model for The stakes. of what you're doing.

5:38 If you are successful. I'm curious for you to define what that would mean successful other than We cross this chasm of general physical intelligence. But if you cross that line. Then what?

5:48 One of the things that I think would be really, really exciting that would be enabled by a general purpose embody foundation model. is the ability to unlock people's imagination in how they build robots and other embodied systems. Personal computers were a really big deal.

6:03 In my mind because it it made it possible for lots of people to hack together all sorts of like really cool stuff. Cambrian explosion of like amazing applications that started in the nineties and so on. And then was further accelerated by the internet.

6:15 And I think something like that might happen in the world of robotics. But it can't happen today because if you want to put together some cool new robotics application, some cool new robotics idea, you kinda have to build a monstrous stack and you you need to basically solve the intelligence problem. But if there is

6:30 a solution that someone can build on top of. There's a foundation model that you can prompt that'll provide like basic functionality and then you can fine tune it a little bit or adjust it in some way to your application. Now it actually makes it a lot more tractable for lots of people, lots of companies, lots of individuals to try all sorts of different things. Sometimes we think that robots are gonna be

6:48 One thing. There's people and now we're gonna make like metal people and that'll be robots. But I don't think that's how it's gonna be because No technology has been like that. It's gonna be more like kind of a tool kit where you can put together all sorts of like really cool applications, get really creative with it. You know, maybe I'm gonna make a robot with like five arms and this one is gonna look like that. It's gonna move, this one's gonna hang from the ceiling.

7:06 and figure out kind of the right thing to tackle your domain, maybe also experiment with software. But you need the right. platform on top of which to do that. And I think the foundation model can be that thing. What are the in your mind the pros and cons of the humanoid

7:18 approach to robotics. There's a lot of value to that. There's a lot of value to capturing the imagination and there's a lot of value getting people to think about what the future might look like. In a way, that's understandable. In my mind it's one of many

7:31 possible kinds of robots that we're likely to have. The challenge of intelligence. Looks very similar. For all these different robots. I don't think we should be tackling intelligence in the context of

7:41 one specific body. I think we should to handle it in a general way. Because otherwise it's just really hard to get a handle on this, we need lots of data. The cool thing about being able to build robots is that ultimately they don't have to be Constrained to look like humans at all.

7:55 You can build the right tool for the job. You could imagine that you're building a house with a robot that is a swarm of one thousand quadcopters. And I think that in the future we'll have to do it. A robotic foundation model? Which can then

8:08 be adapted to all sorts of applications. And am I really run the gamut from Bulldozers or something. Two Humanoids, two robotic arms.

8:15 Maybe it would need to be adapted to each one, maybe it would need to be fine tuned, maybe we would need something in context to understand how that body works. But the fundamentals of how you interact with objects, how things move in the world, how causality works, like that's all conserved. For all these different systems. Do you have a favorite example of what might be possible with true general intelligence that might not be possible with A humanoid only intelligence or something.

8:37 There are a few things that I think are worth thinking about. One is that we can make machines that are very big and machines that are very small. This is not by any means a short term thing, but in the long run. I think there's lots of really exciting applications in medicine and surgery. Where we not only might in the long run not be limited to robots that

8:53 We might not be limited to robots that can even be controlled by humans. currently, for example, in robotic surgery. It's done entirely with the trillion operations, so you need something that are personal control in real time.

9:04 With the right level of dexterity. And of course That limitation holds for current learning enabled systems too, but in the long run. We could imagine addressing that. Think about the

9:14 most important hash marks on the timeline of robotics research that have gotten us to here. I always think it's super helpful to set the historical context before we talk about What state is today and where we're going. Can you walk us through that? at some level doing end to end control for robotic systems is a very, very old idea.

9:30 The first, for example. autonomous driving systems that used end to end. Learning. They existed in the nineteen eighties. Alvin was

9:38 Nineteen eighty six or eighty seven. And that was a driving system that was demonstrated to drive on highways controlled by a neural network. And then from a camera. The neural network was Tiny. There are some.

9:49 very venerable concepts, but Historically, what has been really difficult in robotic learning is that You need a system that handles the application you want to address. That is

9:59 cost effective to train for the application, meaning that you don't need like a huge amount of data for every single application you want to tackle. handles long tail scenarios with common sense, so if something weird takes place in the world, it needs to like have a reasonable response to it. And then also for the thing that it's actually supposed to do, it needs to be robust. Fast and reliable.

10:16 And getting all those things together is very, very hard because With machine learning. It works best when there's a lot of data. So if you sort of naively approach a robotic problem and say like I want to do washing dishes. the all big thing to do is to collect like an enormous amount of data washing dishes.

10:30 But that's not cost effective because then you go on to the next application and you go through that process all over again. Being able to train general purpose models that can handle many tasks is essential to this, because now you need a lot less data for each new task. But then Even further, and this is the thing that has probably changed the most in the last few years. You also then need to handle

10:48 The Unusual scenarios. For the unusual scenarios You are probably not going to have experience. What you need to rely on is

10:57 knowledge. that you've acquired from other sources. They can ground in a new situation. And people are extremely good at this. If you're driving a car and there is something going on in the middle of the road or someone put up a sign saying

11:09 Don't go here. There's the gas leak or something. You've probably never experienced that before, but you can put these things together and figure out what you're supposed to do in that unusual situation because you have common sense. This has been like a huge mystery. In robotic learning world.

11:21 Where do you get the common sense? And this is what's changed in the last few years because Turns out that Multi modal language models are really good. at pulling in knowledge.

11:31 and trying to articulate that knowledge. They're not very good at grounding that knowledge in physical situations, but they know stuff. There's a path to get that common sense by essentially leveraging the knowledge. That's contained in multimodal LMs. But all there's also a challenge because you have to somehow plug into that knowledge in the right way.

11:48 You can't just like show it a picture and say, What would you do here? because it doesn't have the context. It doesn't know that you're a robot. This is what you look like. This is what's going on. That's a technological challenge and we made some headway on addressing that technological challenge, the research community in Charlotte. But most importantly it's kinda the light at the end of the tunnel though.

12:03 Now we have this way of pulling in lots of knowledge, which can help us handle those long tail scenarios. Are there hash mark equivalents on the timeline of the AlexNet or the Transformer, are there big major events that you think everyone will point to when writing the history books about this? I think it's very early on right now.

12:20 To like answer that definitively. Probably the first time. And learning systems which were in the eighties. That's definitely a milestone. the first deep reinforcement learning systems, which were in in the early twenty tens, those are probably a milestone because deep reinforcement learning gives us a way to go beyond human level performance, which I think will be essential for robotic systems. And then there's the more recent stuff, but

12:40 I don't know how that's gonna shake out as far as Whether that's something that's People will point to but I do think that the advent of Multimodal LLMs. that can be adapted to robotic control to bring in that common sense. I do think that's a really important advance.

12:51 I think we're probably gonna see quite a few important events in the next few years. And Maybe those will be the things people point to. Can you tell us your own personal history of approaching the problem? Maybe the origin of when you first became interested in why

13:04 And then How you've decided what to spend your personal time and attention on ever since then. So I started working in robotics in Two thousand fourteen.

13:13 after I finished my graduate degree and started a postdoc with Professor Peter Beal at U C Berkeley. I hadn't worked on on on robots before, but I figured I should get a little bit more education. after finishing my degree and his lab worked on robots. So I tried to apply what I had learned previously. Before that I worked on computer graphics.

13:32 The thing that I've always wanted to really figure out is How to get AI systems. They get better and better the more they do things.

13:40 Because I think that's tremendously powerful. If you can have a system that gets better and better the more it does something And it just keeps getting better in this no limit that it can master all the skills he wanted to do. Initially I tried to approach it. In a very blank slight way.

13:54 You start with nothing, you practice a particular skill, then you get better at that skill. You can do that. limited setting and you get something that works. But it's very hard to turn that into Mm.

14:04 General. system that can work in open world settings because if I practice something over here and then it goes over there, now something is different, it needs to practice all over again. When I worked at Google. Afterwards. I tried to see if

14:17 We can do that. But now parallelize it across many robots. So collective learning. Can you put twenty robots in a room and have them all learned together? And that works. And it generalizes. But it's very hard for that to handle these tailcases, these edge cases.

14:31 savant of this particular task, and that's all it knows in the world. Next episode I mentioned before is combining this ability to practice skills with lots of prior knowledge. And that's actually a really, really hard problem.

14:43 It's not just in robotics where it's a hard problem. I think it's a hard problem in all of AI because arguably The two big impressive results in AI over the last few decades have been gener of AI and deep reinforcement learning. Like if you want to single example to evitemize this train of AI. That's like LMs, deep reinforcement learning, AlphaGo.

15:00 They're both very, very impressive, and they're very impressive for very different reasons. The general the eye is impressive because it can reproduce some of the things that humans can do. Like it can draw pictures that look like human pictures. Right text. Deep R L is impressive for the opposite reason. It does things that humans hadn't thought of.

15:16 The big challenge. And This is what I'm leaning up to and what I hope to figure out here our physical intelligence is has to combine those threads.

15:24 how to bring in all of that knowledge that you get with genre of AI. but also go beyond just human level performance. with reinforcement learning. What literally have you done and are you doing to Make that happen.

15:36 In a Past few years, we started off first by Developing the basic foundations. The basic foundation is what's called a vision language action model. The vision language action model you can think of as An L M?

15:49 That has been adapted for robotic control. So the way these things are trained is they're first trained on text data. Then they're adapted with lots of image data from the web to understand images, and then they're adapted to robots with lots of very diverse robot data. That's a a starting point. That's a way to take all of that web knowledge, get it into a model that can control robots. And get some interesting behaviors out of it.

16:10 And then from there. We studied Two threads. How to get this thing to handle unusual situations with common sense.

16:17 and how to get it to improve with reinforcement learning. The way uh you get common sense is by essentially using chain of thought. The robot enters a scene and instead of directly Starting to move. it thin about what it was asked to do. So if it was told clean up the kitchen, looks at the scene and says Based on this, I should pick up the plate.

16:33 And then it goes and does it. So that Unlocks all this prior knowledge because those Intermediate inferences benefit from the web scale pre-training.

16:44 That handles edge cases. And then the reinforcement learning part comes in after you've practiced it a few times and keep getting better and better at the task. directly through your experience. For example, we had this demo on making espresso. That system practiced making those espresso.

16:58 many, many times and use that to improve robustness, improve speed, improve throughput. And we're not done with that. Like I think there's a lot more to do there, but we have the starting point. The robot data itself. Is the right way to think about it, I'm looking at the Gen one of these things.

17:12 I see a camera here, maybe there's some sensor somewhere else. Is effectively the data being gathered by various sensors strategically placed on the robot at different parts? Yeah. Something I'll say about sensors is that I think you can actually get away with less than one might think and still do quite a lot.

17:27 This platform here has Three cameras, one on each wrist, and a base camera. It doesn't have touch sensing, it doesn't have force sensing, it's very bare bones and very low cost. I'm sure that more sensors could make it better. But a good learning method. can actually compensate for deficient sensing fairly well.

17:41 The wrist cameras are essentially a touch sensor in disguise because you can see Local deformations when you touch something. If I think about the analogy to the expert systems of the eighties and nineties in Basic AI. to the lesson that scales all you need and the sort of counterintuitive nature of that, that you're not teaching it any specific thing, just blasting it with data.

17:59 And there's this reservoir of internet data. Talk about how to create the reservoir of data needed for this. So I don't think anybody really knows how much robot data is needed to have truly generalizable and powerful embodied AI. My sense was that We actually don't need to know.

18:15 What we need to do is get to the point where these systems are useful enough that they can go into the world. And gather more data themselves. Tesla doesn't worry about how much data their cars can collect. If anything, it's the other way around. That's a little too much data. The key is not so much to quantify here is exactly the price tag of getting the ultimate robot data set. The key is to get a system that can go into the world that's useful enough.

18:38 that does a wide variety of different things and that can keep pulling in more data. You brought up the example of Tesla the beautiful system of a thing that's useful without the AI to begin with,'cause the human drives it and it gathers data. Why then not start with your best guess at something that's useful as a single robot? to to have the same sort of flywheel thing happen. I think it's a good idea.

18:58 And do you think that's an approach that You'll pursue. I don't think that there's like one right answer, right? So I think there are some domains where deploying a system under human control makes a lot of sense. There's some domains where Deploying a partially autonomous system. Is very reasonable.

19:11 That's kind of domain dependent, because robots aren't just one thing. Some people might not want a robot in their home that is constantly being controlled by A person offside? But maybe for some applications that doesn't matter. if you mark the start of physical intelligence through today, what has been the most surprising thing to you that you've

19:28 Discovered or the nature of how the Research has gone. One of the things that's been surprising to me is that I think we've made a lot more progress on dexterity than I thought we would. We had good reason to believe that if we just collect more and more data, that just steadily gets better.

19:41 What was surprising is that we could also get these systems to perform very dexterous behaviors. Without really doing anything particularly special for that. The same, by the way, also applied to Getting systems to work on different embodiments. Where

19:55 We could get our models to work on all sorts of other robots, including robots with multi fingered hands. Robots with different numbers of degrees of freedom. And obviously we need to get data and we need to fine tune the model. But the model itself didn't need to change. It didn't even need to be told.

20:09 Through any kind of prompt. what the robot was. And that was also surprising to me because I would have thought that We would need some. fancy techniques to adapt the system to faster, more dexterous, more complex tasks and also to different kinds of embodiments. But

20:22 It actually seems to generalize pretty well. I'm always interested in like the spectrum of capabilities and especially Where the systems today are more advanced than you think people would probably expect and where they're less advanced than people might expect. This is something that's always been very Tricky.

20:38 to understand in robotics. There's this idea that roboticists always talk about called Morovic's paradox. That's actually true in all areas of AI, but especially in robotics, this is a big deal. We kinda have a cognitive bias. to think that things that are easy for us will be

20:52 Easy for the machine. solving calculus problems is difficult for most people. Picking up a cup was easy for most people. So we think, Oh, machines should be able to do this. But it's actually the other way around. There are things that are easy for us. Because they have to be, otherwise we wouldn't survive.

21:05 We're very good at spotting the tiger in the jungle because the people that weren't so good at it got eaten by the tiger and they're not around anymore. Because of that we have this cognitive bias and we think that there are Things that should be very easy. But they're actually very difficult engineering challenges. However

21:20 something that is changing is that machine learning slightly changes that equation. programming something by hand to pick up any cup anywhere, that's difficult. Getting a machine learning system to do it if you have data for it. It's actually not that difficult. And I think increasingly what we'll see is a shift where

21:35 Domains where collecting data is straightforward. They actually end up falling into the easy bucket over time, even if they are physically intricate. But there will be domains where collecting data is difficult where You need to use more common sense where you need to reason at multiple levels of abstraction.

21:50 connect phys that you've learned in other areas to knowledge that you got from the web. And those will be tough and that's where we'll need more Technology accounts. What is the science of common sense? When we say that

22:01 What does that mean? For the purpose of robotic learning. We can think of it as Applying semantic inferences? Using knowledge learned from other domains.

22:12 to the current physical task at hand. You can think of common sense As The opposite of muscle memory. So muscle memory, like if you play a sport, you practice something a lot, you hardly think about it, you just kinda do it on autopilot.

22:24 Common sense. In my mind. I don't know if this is a conventional definition, but I think it's a reasonable definition. Is when you know something to be true. Because You saw it, or you read about it, or you heard it.

22:35 And now you are in a situation where that fact is highly pertinent. And you are able to make that connection, apply it to your situation, grounded it in the environment that you're in and make the right decision. One of the other differences that's so interesting to me is people that have used chat everyone's used chatbot now. You query it, you get an answer, query, get an answer.

22:52 We're now seeing what happens with cloud code and other things where you give it something complicated and it's able to do a very long Without failing. What's the similar thing long range? In robotics. It's something that we're working on quite a bit right now and the methodology is not that different at some level.

23:09 the way that our models work now, as I mentioned, is they use this kind of chain of thought process to reason about the task. When you have that, you can actually do very long horizon tasks. You can have uh a robot that goes and Takes out all the dishes from the dishwasher, puts them in the correct cabinets, wipes down the counter, all that kind of stuff. The interesting thing here is that

23:27 We found Maybe about six months ago. That Our models had gotten to the point that Where they could be improved.

23:36 Just from supervising them with high level instructions. You take a a robot, you put it in a new kitchen, you ask it to clean the kitchen. It gets to work. And then it fails somewhere.

23:46 So now okay. What do you do? Well you you add more data. Traditionally what you what we would do in that situation is add more teleoperation data to cover a wider range of kitchens. But what we tried kind of on a whim is to see, okay.

23:56 Well what if we don't add more teleoperation data? What if we just add more data Labelled with the semantic command. So basically just take whatever the robot experience and just label it with some semantic command, but don't add any more Ball of Lactions. And that actually helps.

24:10 It actually improves its ability to generalize. So what that means is that the bottleneck had actually shifted from the lowest level meaning the l the robot's ability to physically do the task, to this like middle level, where now the system is more bottlenecked by its ability to interpret the scene and select the correct next step. Which can be supervised with language. That's a big deal because now that means that someone can literally talk to the robots. Coaching, basically. Yeah, exactly.

24:31 And make it better just by talking to it. We're in twenty fifty. And there's no robot in my kitchen doing my dishes for me. What do you think the most likely explanation is for it not having gotten there by that point? I can think of a few reasons.

24:46 My suspicion is that There is A long tail of challenges that has to do with the interaction of technology and People.

24:55 Autonomous cars aren't that different in this regard where Getting to a level of comfort. with deploying autonomous vehicles on the road. was a significant challenge. No

25:05 ran in peril with getting the technology to a level. early Tesla self driving was a bit controversial because it wasn't perfect. There was a question like are people comfortable with this? level of imperfection. Probably there are some tasks for robots.

25:18 Where People will be comfortable with something that's not perfect, something that needs to learn from its mistakes. There are some areas where we will not be comfortable. Are you comfortable with it? occasionally breaking your dishals. Maybe in a few years it will stop breaking those dishes, but maybe in the meantime it's not quite there.

25:31 Are you comfortable with a robot like that? in a home where there's like small children. Maybe not. That's okay. I think that figuring out how those factors interact And what that means for the timeline and for how these systems get better with experience. I think that's a tricky question.

25:45 I think it needs to be approached. Very carefully, with a lot of sensitivity. There may be some domains where It makes a lot more sense. For these systems to be deployed and shrunk.

25:53 and collect more data. And maybe other domains require more care. Can you imagine a purely technical explanation? For why something might not work. I think the place where I would see the biggest technical risk is

26:04 Dealing with the breadth of different situations. If we're talking about A well defined But slightly chaotic environment like

26:14 Cleaning hotel rooms or Assisting human. Cooks. In a restaurant. I have like a very good sense for how I get that under control.

26:21 If you're imagining A robot going into a home, one place where I can anticipate a challenge is that There are a lot of other unexpected things that can happen. And you need a system that's Very good at

26:32 Inferring what's going on. And adapting to it. Or acting intelligently. And I think we have a lot of ideas for how we can approach it. But that is the hardest part of the problem because

26:42 When you're in a situation where Just about anything could happen. And You're controlling like a physical device. Affects the world around it.

26:50 Then you really need to get things right, at least at some level. Pretty much in every case. Like it doesn't mean that you always have to succeed, but it does mean that you always have to do something Sensible. That people are okay with.

27:01 And I think there are a lot of really good ideas for how to do that. But that is Probably the most challenging. Or the equation. If I go back to thinking about the right model to think about the physical intelligence approach to doing this whole exercise.

27:13 Help me make it as simple as possible. So one might be we're gonna build a whole variety of different kinds of form factors to do a whole variety of different kinds of things and mash all this data together. And start to, you know, experiment with how we can On evals make it better. Is that just the simplest way of doing it? Is there an even Simpler way. And I'm asking because I'd love to then contrast it with some

27:31 other approaches that you're interested in that you're not doing that others are doing. In my mind, the most important thing to get right is To get the system to be general, in particular, to get it to be general with respect to How it can be improved. For example

27:45 hand design robotic controllers. are not very general aspect to how they can be improved because it requires like a human engineer to go in and improve it. A learning based perception system is more general because All it requires is human labelers to go in and label more data.

27:58 A system that learns autonomously from data that it gathers through its own experience is even more general because you don't need the human labelers. The key is this generality particular with respect to improvement. And the decisions we make are to a very large extent center around that.

28:13 I don't know if the correct design for a robot is to have three cameras. I don't know if it needs like a touch sensor, I think. We're very agnostic to that. I think we'll try a lot of those. Different. Choices I'm not even sure if in the long run it's gonna have a language model. Maybe it'll have some other kind of model that's trained on very diverse data.

28:28 The key is this level of Generality. What other approaches are the most interesting to you? One thing that's like a very Important question.

28:37 In this area. And something that I think the research community and the tech community is not fulfilled is The dichotomy between different data sources, particularly with respect to real data and simulation. It's a very controversial topic. I have a very strong opinion about it.

28:52 But I think that's It's worth acknowledging that If we look for example at Humanoids. We've seen videos of humanoids doing all these acrobatics.

29:02 There's a particular pipeline that makes that work, which is very heavily reliant on Simulation. And very light on real world data. Off country zero real world data. And then there are the

29:12 approaches that work well for robotic manipulation. That often are the opposite. They often use very little simulated data. often use large amounts of real world data and very large foundation models. And it is surprising.

29:25 That in these two robotic domains, the dominant approaches look so different. It may be that. one will win out and there's a particular approach that can handle everything in the long run, or maybe there's some sort of synthesis of of these ideas. That's important. I don't know the answer that I have.

29:41 Subjective opinions, I think the approach we're taking is a very good one. But I think that it's interesting to look at that and see why is it that these things are so different. Can you talk about the contrast between cool and useful? The Boston Dynamics robot is very cool. The back flip is Super cool.

29:56 Inverting the body. It all looks really good. I don't know what I need. That requires a robot to do a backflip. So I'm curious how you think about optimizing around cool versus useful.

30:08 The strategy we've taken is subject to the constraint that it's useful, make it as cool as possible. We make decisions first and foremost based on assessment of what will drive the tech forward towards this truly general, broadly applicable Robotic Foundation model.

30:23 But In doing that. We try to stress test it against the toughest challenges we can throw at it. The toughest challenges are often ones that look cool.

30:32 We didn't set out for example to build a robot that can make espresso or can Full laundry. But in the process of building these general systems, we figure like these would be particularly challenging, particularly exciting. Things to try with them.

30:45 To see how far we can push them. Can you talk about the robot Olympics? There was A gentleman named Benji Holson who used to work at Everyday Robots.

30:54 He spends a lot of time thinking about. Tasks that robots can do. So he wrote a really interesting blog post. A while back. There was this

31:01 Robot Olympics that was held in China where robots would like run around on a on a track and jump and so on. But maybe these aren't the real challenges we should worry about. How about a robot Olympics centered around Essentially everyday tasks that people do. That's kind of more of X Paradox thing where tasks that people find really easy but that robots struggle with.

31:19 And he had things like opening a door. Washing a frying pan. With grease on it. Using a plastic bag to pick up dog poop.

31:27 Things that people don't find particularly challenging, but that no current robotic system can do. And he listed Maybe a dozen of these things. This wasn't part of a concerted like research project. We had developed processes and

31:38 Systems for just ingesting new tasks. No. we want to use for all sorts of tasks and we figured, okay, like a good way to test this is to say, like, hey, here's like a big list of tasks. Let's just go through this process that we've developed. And see

31:49 It works basically. So it's Almost like a test of like our internal operations and model training system. When we tried these things and actually turned out that we could solve almost all of them. Well there's one we couldn't do which was turning a dress shirt inside out because the grippers on this thing wouldn't fit inside the sleeve.

32:04 So we probably need to change the gripper. I think on a technicality, we didn't succeed at peeling an orange. Because he said, Do it with the fingers and our fingers weren't strong enough so we had to use like a little tool, like a little knife, basically. Everything else we could do. If anybody watches those videos, one thing that I think is important to keep in mind is We didn't like develop anything special for this. We literally use this as a test of our

32:25 Task onboarding process. There's something interesting there because it suggests the power of generality that when you have this general system you can really just like onboard all these crazy tasks. Without really

32:36 Do anything particularly sophisticated. I was curious before when you said Superhuman ability. Dexterity or something like that, where we're limited by what we can do, or maybe by what we can control, even if it gets smaller. What are some of the other dimensions?

32:50 That we might surpass human ability on in terms of physical ability. What are the other trend lines? So here's a fun one. We were Working on a task where

33:00 Our robot had to plug in Things like power cables or Ethernet cables or something like that. When a person does this Obviously if you practice it a lot, you'll get really good at it, but when a person does this without having practiced a lot. You pause frequently, right? Because

33:13 It's not a physical thing. You just have to parse what's going on. You have to make sure that's like all a line and all that stuff. So you do it very slowly. And if you're Telly operating a robot, you do it even more slowly because there's this level of indirection. It turns out to be like pretty straightforward to go in and find all those pauses and remove them. you can speed things up further so you can get to a task where a person demonstrates what it means to succeed, and then you can have the robot practice the task and succeed in the same way, but a lot more quickly, a lot more efficiently.

33:39 The most general way to do this is with reinforcement learning. But there are also like some simple tricks you can do that to if you just want speed. So that's like one example of something where You can have a machine that does it a lot better. You know, at some level we have like a processing bottleneck, like that's why the person does it slowly because they have to process what's going on.

33:53 But Speeding up processing is something that People understand quite well in computer science. There's this amazing Michael Crichton novel called Pray. Where it seems like for a given problem there may be an optimal or set of optimal

34:06 Shapes of the robot. to perform the task and that what you should do is analyze the problem then Have something that can almost like morph or transform into the right Form factor. How do you think about

34:17 That the Innovation on the form factor side rather than the data and model side. I think that in general in robotics The ability to innovate on form factors has been very constrained.

34:30 Because of the AI challenge. If you have a traditional AI pipeline, like you know, you're doing some motion planning and stuff like that. It's hard to just go and cobble together some new robot because when you do that you have to like characterize the dynamics of a system. You have to do society. You have to build up all the stuff.

34:46 If you could just put together a robot In your garage. Load up. a robotic foundation model. and tell it to do a bunch of stuff. Like maybe it won't be perfect now. Maybe it needs more data to really perfect it. But you can at least like get the thing moving.

34:57 I think that can be a really powerful engine just to get everybody to experiment with this stuff. I don't think that I'm like the right person to design the perfect robot. There are people here, of course, who are a lot better at that, but in general I think that It's just like with personal computers, so I think the key is to

35:10 let people experiment and play around with it and just radically lower the barrier to entry for that. Then we'll see a lot of creativity. When we first started using computers. There's A limited number of form factors. Now you can have a computer in your phone, a computer in your car, embed a computer in your refrigerator.

35:25 They're everywhere and they're very different. Generality. good software, good foundation on top of which you can build applications. Those are key to enabling that. Your co founder Lockheed once described to me the feeling of physical intelligence for a human

35:36 is like learning how to ride a bike. Like there's that moment when you didn't know how to do it and then you do know how to do it. And that feeling is physical intelligence, that snap of understanding. There's actually a physiological explanation for this. There were studies that were done in monkeys. using tools. And you can actually find Where in the brain.

35:54 Which neurons activate. for the monkey to figure out where its hand is. It turns out that if it's using a tool. They activate. Based on the location of the tool tip, not based on the location of the hand.

36:03 The tool of being an extension of your body is a real physiological thing. Like your brain literally does that. Knowing that, what does that do to impact The approach to your research. It says that physical intelligence.

36:15 should be at some level agnostic to embodiment. The the a good foundation model should figure out How to manipulate whatever body It's controlling.

36:25 whatever tools it has at at hand. There's basically one problem, not many different problems. There isn't like a humanoid problem. and a car problem and a bulldozer problem and a robot bolted to the table problem. There's one problem. And if you solve it as full level of generality. That's really, really powerful. We're in the early stages of seeing some of the job and other sorts of transformation.

36:45 in businesses in the economy, et cetera, that LLMs Make possible. Certainly we've seen it in Engineering. How do you think about what might happen or what you hope will happen?

36:54 when we're at a similar stage, whenever that happens to be for robotics, where all of a sudden we have this thing that's general, that's useful. The world's very efficient at deploying these things. People are creative. Where do you expect to see the world start to change most in the early days? I really don't know. I don't think anybody would have been able to predict how the L M stuff evolves and people would have guessed, but

37:14 This is why I keep coming back to this idea that Maybe the key is to let people try lots of things. one of the really amazing things about applications of LMs is that they are really accessible and somebody could put together a really cool new prototype that under the hood is just prompting Cha chicketee or something.

37:30 But they can experiment with it, they can try it out. See what it does. And there's an amazing Power to having a Lots of smart people rapidly iterating and prototyping lots of things.

37:39 That's a lot of why physical intelligence has really put a premium on engagement, like we've open sourced. Our models. We Would like to

37:48 engage with lots of other companies that are Building robots. Because We all see a lot of power in this. effect of having many people.

37:55 Trying out lots of things. What are the major controversies in the robotics community? To me a controversy someone gets in in an argument with me at a conference, but I can tell you that the kind of arguments that I found myself in, and it's kind of an interesting trajectory that In the early days The main argument I would have with people is

38:13 Does learning have a place? Yeah. Robotic AI. I think part of why that was often A controversial point is that

38:23 In a traditional engineering pipeline, robots do look very different. Than software R Fox. They're physical. They can affect stuff around them. They're safety considerations.

38:33 There are A lot of weird situations they can get into and it for the robotics research community. to really internalize that you don't necessarily need to program in

38:44 things like knowledge of physics. You don't necessarily need a physics simulator inside your robot when it's planning. We can actually have a learning system. Figure all that stuff out. That was a very controversial thing for a very long time. I think at this point There's a lot of acceptance that

38:56 learning is a really important part of robotics, but I don't think there's still universal acceptance that End to end learning. is the right way to go. Basically I don't think there's universal acceptance of the bitterless. The better lesson says that You should not program the machine

39:09 To think the way you think it should think. but you should let it learn from data. And that is not a universally accepted idea. I think there's good arguments against it. But I think that's a good thing.

39:18 in the long run, if we want that generality, especially generality of the machine's ability to improve. then we need it to primarily be learning from data. What is the good argument against? My best attempt at steel manning this.

39:30 Is that If you want something reliable in a really complicated open world setting then you can't afford not to use what you already know. about the physical world and we've got textbooks full of this stuff. So why don't we just plug in what we know from the textbooks?

39:43 What is compositional learning? Can you describe that? One of my students. He had this idea where he Ask a language model. To provide a recipe for how to make

39:53 A sandwich. in international phonetic alphabet. International phonetic alphabet is these symbols that they use in a dictionary to explain how to pronounce a word. And it's very peculiar because it only ever appears for individual words in a dictionary.

40:06 Do you never see freeform text written in her astronomic alphabet? But if you ask a good language model It will write. paragraphs in IPA for you. And that is compositional translation. That means that you have never seen this particular language, this particular alphabet used to write paragraphs, but you understand paragraphs, you understand that it's compositional with different alphabets. So you can solve the problem.

40:25 You can imagine the same thing coming up in robotics, that you've learned a repertoire of skills. And now you can combine and mix those skills. and apply them to solve new problems. It makes me wonder what the last type of tasks You think will be possible.

40:40 For a robotic system. To achieve. I think changing a child's diaper will be really, really hard. This really is just Morvik's paradox all over again, that people are extremely good at certain things. We're very good at physical things, we're also very good at interacting with other people.

40:56 And that makes sense. We have to be. That's a lot of our existence. So things that involve behaviors to interact with other people. Where you have to like

41:05 Help somebody. I think that's a lot harder than people appreciate. Elderly care. Taking care of small children. I think those things are gonna be hard and they're probably gonna be harder than people think.

41:14 And the stakes are very high. It's not just that, the stakes are high in many places, it's just that It's probably the pinnacle of something that fools us into thinking that it's Easier than it really is. We are so

41:26 evolved for interacting with people and doing things physically. If you're Helping somebody Get up the stairs or get out of bed.

41:35 You don't have to think very carefully about how you're gonna do that. So I think it's really the pinnacle of Morvic's paradox. If I think about an LLM as a brain And now it's effectively studied everything. I don't know how else to put it. And then I think about a robotics model's brain instead. What are the dark parts of the brain?

41:50 What has it not been able to study? What are the areas that have just been really difficult? That matter but have been hard. For us to get into. One of the things that people are remarkably good at

42:02 Is Using physical analogies to understand other situations. I don't know whether this is something that LMs can or can't do, but It is something that people use a lot. They use it in everyday life and they also use it for very sophisticated problems. So for example, you could say, That company has a lot of momentum.

42:18 That's a physical analogy, know exactly what it means. I don't have to explain that statement to you. But If you actually think about that, it is quite a complex thing. There's like a lot riding on that word momentum. There was an interview with Richard Feynman when he talks about

42:31 Teaching that he talks about. analogies that he makes. In regard to subatomic particles. We use like the word spin. The thing is not really spinning, like it's not like a spinning top. But

42:40 All those kind of analogies. Help us make sense of it. And not just in a way that allows explaining concepts, but it actually leads to conclusions. It actually leads to inferences and those inferences actually make sense. We're so primed to interact with the physical world, so primed to have physical intelligence that you can use it in everyday speech by saying that company has a lot of momentum. And you can use it when

42:59 Advancing fundamental Theoretical physics. That's kind of remarkable. I don't know if LLMs can do that. Maybe they can, but I think that's a great idea.

43:08 really understanding physical interactions, causal structures, all that kind of stuff. There is something special about that and it's clearly something that people get a lot of mileage out of. I love talking about the role of researchers and the actual people doing the research. In L L M world it's fairly shocking how few people are at the global scale responsible for basically all the progress.

43:26 In LLMs, someone like Ilia as an example. What is that like in robotics? How many people in the world are truly impacting this trajectory and then I want to ask what good research means. I think those questions are often very hard to answer about science because I think that

43:42 We sometimes have a tendency. Especially when we look at history. To Underline milestones. And certainly in machine learning this is the case.

43:51 Alex that was a big step for them. That's true. But I think it's also important to remember that These Advances

43:58 They happen because Lots of people are trying lots of things and even some of the failures are actually very instructive. I complained before a little bit in a low key way about the controversy around end robotic learning, but I don't know if robotic learning would have advanced the same way if it were not for the controversy, so to speak. It is true that you can look through the list of successes and mark down that like oh like these folks

44:20 have a a history of repeatedly hitting home runs. But I think in reality in the scientific community It's not just the home runs that are responsible for progress. And even some of the failures and even some of the bad ideas. are very instructive in pushing towards the good ideas.

44:34 That's fascinating to think about. The example you gave before is so interesting where the research insight was like just give it some coaching and it gets better. It seems like that sort of insight. can be very powerful and High leverage.

44:45 Which makes me wonder like what have you learned about what makes for a great researcher? Research is definitely a different From engineering because In research the important thing is to Get an answer to a question.

44:56 Which often requires cutting some corners. One of the most delicate decisions in research Is When do you try new things, or is when do you stick with what you're already trying?

45:07 That's very, very delicate. It's very, very hard to figure that out. And if you get it wrong. then you can miss something really remarkable. If you get it wrong and you don't stick with something for long enough. You might be like right there, you might be about to get to the answer, and then you stop just short of it. That's terrible.

45:21 Or you could get stuck. hammering against something that's never gonna give way for years. Deciding when to Turn a little bit and look this way and that's the one. To open yourself up to more opportunities.

45:31 versus when do you should you keep hammering on the thing because you're about to get the solution. That's often the most important decision and Some people have an instinct for getting that right. That counts for a lot. You've obviously been in and around and are

45:41 Great researchers. What are these people like? as people how do they tend to be distinctive from The average person. I think they're just the same.

45:50 I have a very hard time thinking of a single set of personality traits. There is no constant, basically. There might be a commonality in that To do effective science, you have to be very passionate about that.

46:03 But even that passion can come from many different places. I've worked with people that were remarkably effective that are just driven purely by the desire for novelty. They don't give a damn about what their technology does. They don't give a damn about whether it's useful. They just want like cool new ideas. I've also worked with other people. that really want to solve a particular thing and they're just as happy building stuff as they are testing out experiments, as they are hammering away at things. Whatever it takes.

46:26 You mentioned The difference between research and engineering, which also makes me think of manufacturing. Elon would be fond of saying that the factory is the product. The hardest part of this whole equation is actually the scale up of Whatever this thing ends up looking like. Making, you know, a hundred million of those.

46:41 How do you think about that part of the equation, or is it too Remote at this stage. I think it's an important part of the equation. I'm not sure it's it's like the part of the equation. That

46:50 We most need to figure out right now. But it's certainly part of it. A lot of how I prefer to think about this. is to

46:58 Figure out the hard part. And then enable A lot of experimentation on the other parts. Making a robot at scale is difficult. Making a robot at scale is even more difficult.

47:09 if you don't know what kind of software's gonna run on afterwards and you're not even sure whether it's the right kind of robot. One of the really valuable things we can get of general purpose AI tools like robotic foundation models. is the ability to like get a lot of the other stuff figured out. So that at least

47:22 Some of the uncertainty goes away. So that when you scale things up. You have some confidence that this is like really gonna work. A lot of people that listen to this are entrepreneurs, people that run companies. a very popular question has become how should a traditional company begin to think about using LLMs or preparing itself for

47:40 the ongoing improvement of these models. How would you answer the same question for robotics? How would you encourage companies to think about This. The technology is changing so rapidly.

47:51 I want to illustrate why this question is difficult with an example. Here is a Particular uncertainty about the tech. Will the robots rely more on demonstrations? Or on

48:00 Reinforced learning from a dominous data. We're working on both of those things and they're clearly both important. But how somebody should prepare for the technology will be pretty different. If they're expecting that they need

48:11 lots of teleoperation to produce lots of demonstrations. A little bit of autonomous experience. versus the opposite, like a tiny number of demonstration that huge amounts of autonomous experience. Like is it ninety ten or ten ninety? That's something that we're hopefully gonna learn about over the next few years.

48:24 But it does change. the correct approach. Pretty dramatically. That's kind of a case study of how Changes in technology will

48:31 Dramatically alter this. From a business standpoint is the right way to think about it. get really clear on the economics of the labor in your business. I'm curious how you think about that. The way that this will change

48:43 The nature of labor itself. coding tools are like a a really nice example to look at for a template of how this might work. It's not like coding tools came on the scene and suddenly We don't need software engineers anymore. It's that

48:57 The coding tools. increase the productivity of individual software engineers. There's some amount of work that needs to be done to make sure that people are able to use them. There's some amount of technology development that needs to be done to make them useful for the appropriate use case and these things are co-evolving and they're also still changing. coding agents are different than code completion tools and so on. But I think it's like a nice template for us to look at to see how

49:17 AI tools combine with people doing a job. increase their productivity and also raise new challenges. And I think we'll actually see something like that with robotics too that. A more realistic template is not like the humanoid goes in and the people just leave.

49:32 There are some aspects of the job. that can be done by a robot. Something can be done with a robot working together with a person. Some of that can be where the person needs to like do something special to make the robot more productive. Some where it's the other way around, where the robot does something that makes the human more productive, and it'll be this kind of dance that we've seen with coding tools. Do you have a favorite robot?

49:50 It's not part of what physical intelligence is doing. And if so, why it could be anything could be a Factory robot could be a Optimism. Boss and dynamics.

49:59 I do really like the Boston Dynamics robot. Especially the new version of the Atlas, because It is In some ways very human like and in some ways very not human like. They made some interesting decisions about how they want more range of motion on the joints, so it can do some pretty cool things.

50:13 It's also a very agile robot, which is really cool. It makes well those awesome demos. So I'm a big fan of that. I'm generally a big fan of like everything that Boston Dynamics has done. Should or could anything be read into the fact that Boston Dynamics has been doing very cool demos for a very long time and don't actually do anything useful for customers? I think it's also a fair question for lots of robotics companies, to be fair. There is a lot of value. in demos that serve to illustrate challenges on the road to something useful and productive.

50:41 Obviously you can also do a demo without being on the road to something useful and productive. There is value in demos, I think that Demos that are used correctly in service to a mission can Provide people with an illustration.

50:53 Of what to expect. And they also provide a challenge. You just have to be like honest and setting up the challenge. How much do you think about the business endpoints? To this point, Roomba is like the best selling robot of all time.

51:05 In the consumer category, which is surprising. And of course we might be on the edge of some sort of Cambrian explosion. But how much of your cycles do you spend thinking about This is the shape of a product that might result from this that maybe is the way we bootstrap our way to all this data. It's just something that's very hard to reduce to like a very concrete answer right now.

51:23 It's not too bad to like think about a space of possibilities. A lot of what we're doing when we develop our models, when we experiment with different Tasks when we Do demos like the Robot Olympics. Underneath we're kind of prototyping. What does it look like?

51:37 when we try to do something real with this, to different degrees of real. And what goes wrong. It is something we think about a lot. It's not something that I have even close to like a concrete answer to, but There's a space of possibilities and a lot of what we actually are planning to do in twenty twenty six is also experiment with

51:52 Different. Things in that space. When you study the history of general purpose technologies, which certainly This would be a major one if it comes to fruition. You often find this constellation of things happening around that thing that enable it.

52:05 L's are a direct compliment to what you're doing. Are there any other surprising technology areas or trends that help you do what you do, but are different. Robotics hardware has become dramatically more affordable over the last few years.

52:21 When I started working in robotics about a decade ago I worked with a robot called a PR two, which I believe had a cost of about four hundred thousand dollars. when I started my lab at UC Berkeley I used a robot that was in the ballpark of thirty thousand dollars.

52:35 Now e charm on this thing. Is maybe a temple that. We think that can be even less. That's not due to like any one single technology. involves both hardware and software. So the kind of

52:45 low cost arms that we have here. they wouldn't be useful in an industrial setting because traditional control methods that rely on a great deal of precision wouldn't be able to use them. And I think that does make it a lot more practical to think about general purpose robotics today. For people that would want to be fairly technical about following

53:01 major milestones that are happening in this field. Where does that information show up? What do you read to stay informed about what's going on or watch? So a lot of it shows up in research papers. Research papers unfortunately are not a very accessible

53:15 source of information because It takes a bit of care to like sort through everything and figure out What is the signal and what does something really mean? Research results are sort of intended.

53:26 Four. An audience that already understands the starting point from all the past research results. Robotics and I think technology in general is one of those things where The

53:35 public facing artifacts the demos and the videos that somebody might post on social media. Are often actually not very good for providing a sense for the true underlying state of things because they're sort of meant more as a demonstration at the edge of capability. And grounding them.

53:50 What does the demo really mean requires digging deeper. probably research papers are the way to go. Sometimes even worse than that, you have to actually go talk to the individual people and find out what the inside story really is. And maybe that's not a great situation to be in, but that's kinda how Science works. As we look forward to the future in your mission, what feels the most uncertain?

54:08 I do think the timeline is uncertain. If anything, my sense of the time has gotten more optimistic since we started, but It's uncertain. Because of the nature of the technology. This is something.

54:18 Where there's a bootstrap challenge. getting to a particular level of usefulness so that robots can be deployed so they can do useful tasks so they can start collecting Data. From open world settings.

54:28 At scale. Because that's social. Sudden Vent getting past the activation energy, I think there was a lot of uncertaint about the timing of that.

54:37 That's exasperated by the fact that The timeline looks different depending on what kind of technology it is deployed. Example I gave before about whether it's a It should be data collection through teleoperation. or data collection with autonomous systems or something in between, maybe shared autonomous, maybe like this this coaching kind of thing.

54:52 those all sort of change the picture in terms of how deployments work and how in the wild data collection works. So because of that I do think there's quite a bit of uncertainty. You're in such an interesting position because You're at the center of research, lots of different kinds of people

55:05 are talking to you, asking you questions. What are Questions that you're surprised people don't ask you. Well I think the question you asked earlier actually about How somebody should prepare.

55:15 There's a variant of that question which would be something like if I want to start using Autonomous robots for a thing. What should I start setting up? Should I set up operations?

55:26 Should I modify my task in some way so it's more accessible? Should I design new hardware? Maybe I should design new so I can plug your software into it. And I think people make a lot of assumptions about that. For example, one assumption is Machine learning requires data, so let me just figure out something that will collect data. That's not often the best assumption because you need the right kind of data.

55:44 Maybe it some data is easy. It's easy to get like videos of people doing something. That doesn't mean that's the right kind of data, and it might be domain dependent, it might be Depending on the thesis about the technology that will Succeed. So I think that People do make a lot of assumptions about that. Not that I necessarily have a better answer for them, even if they asked me, but

55:59 It's something where there's a big space of possibilities. We talked about these like big uncertain long term timelines. What is the very next thing you are trying to solve? a big focus for us right now is actually Better understanding

56:13 this mid level reasoning part of the problem. Because we think that we have a pretty good sense for how to acquire But Getting those little physical behaviors to generalize.

56:23 requires bringing to bear a lot of this common sense knowledge. The representation of that might be really important. So LLMs make certain kinds of representations very convenient. They make it very convenient to basically turn text into other text.

56:36 But that's not necessarily the best representation. for what an embodied system needs to do. Sometimes it needs to think about things more spatially. sometimes semantically, sometimes other representations, and trying to figure out exactly how to structure that internal thinking process. Might be a very important question.

56:52 The answer to that question might be different. in the world of embodied foundation models than it is in the world of LLMs. So that's like a concrete thing that we're working on now. If I could somehow get the Hundred most.

57:03 informed and active robotics researchers in the room at once and pull them. On how certain they are that Things will have to be a little bit more. unlimited capabilities and how soon that might happen.

57:13 Where do you fall in that distribution? Probably I'm on the optimistic end when it comes to Established robotics researchers and on the pessimistic end. Relative to Robotics entrepreneurs.

57:25 I understand the entrepreneur part for sure. They're optimistic by nature. Why are you on the optimistic end of the researcher community? Robotics has a very long history. Which Has precious few successes.

57:37 Especially when it comes to robotic AI. So I think if we're being honest about it. most robots that are out there doing useful work are still running Say the art technology from the nineteen eighties. Because the robotics problem is hard. Not our fault. It's just a difficult problem.

57:50 Because of that. I do think that there is good reason for caution. Maybe we've made a lot of headway on this part of the problem, but there's like many other problems that still remain. Part of why I'm optimistic about this is that

58:01 I kinda have like a sense of what has proven tough for me before. And I can see a lot of the puzzle pieces that I'm imagining could be slotted in. to address many of those things. As my co founder, Carol likes to say it.

58:12 when you've climbed a mountain. Only then do you see if there's another mountain after it. In robotics, there's been a lot of experience of lots of mountains. Some caution is justified. Given that endurance is required.

58:22 Who or what most inspires you? Boston dynamics. I think there's like a lot of things that we can debate on the technology side. But there's A lot of value

58:33 And Repeatedly Showing Something That people wouldn't have thought

58:39 Possible. Even if There's all all sorts of caveats and assumptions and so on. And certainly in robotics. Whatever we might say about demos and whatnot, like I think it's very fair to say that

58:49 People have revised their thoughts about what's possible from seeing Some of that stuff. I think I'm also Inspired by органіzons that create an atmosphere

58:59 For Experimentation. There are some research labs that have done a very good job of this. Open AI has historically done a great job of this, of

59:07 creating an atmosphere where individual. Researchers can experiment with things. And be empowered. To see things through.

59:14 Chad GPT was Basically John Sholman's That experiment for a while. It wasn't a concerted corporate strategy with lots of spreadsheets and

59:23 Pie charts. It was A pet project. I think there's something pretty inspiring about organizations. that empower people to have pet projects turn into

59:31 World changing. successes. Certainly one of the aspirations that I and my co founders have here at Physical Intelligence is to Provide. Some of that, to the best of our ability.

59:40 It's hard to do. I feel like Google used to have that one day, you can do whatever you want. Thing. Is that the spirit of it? I was absolutely shocked when I started working at Google.

59:51 That the level of leverage that I felt I could have. One of the projects that I did with many of my colleagues there in twenty fifteen was

59:59 colloquely referred to as the arm farm. So we Chuck. A couple dozen robots, put them In a lab and have them collect data. I found out from somebody that they had a warehouse full of robots that nobody was using.

1:00:09 I asked Jeb Dean and Ms. Mount Hook if we could stick them in a lab, and I was just thinking like okay. They're not gonna take me seriously I was A level four research scientists. Jeff was like, Yeah, let's do it. What do you need?

1:00:19 I just remember feeling like wow, I had never in my life Thought that I'd have that leverage. I mean, I was very young at the time. That's very special. And I think getting to a place where people can unlock their creativity and have that kind of agency.

1:00:32 Can make for a very remarkable place. My friend Jesse has this great question, which is For companies that you're not involved with. Which one do you most hope succeeds and why? People just say boom a lot because they want to fly places faster.

1:00:45 Increasingly as I've asked this question, people have said buy because the sheer impact that it might have If you're successful is massive on such a global scale. And it's been really fun just to hear about all the ins and outs of how you're thinking about the problem and attacking it. When I do these interviews I have the same traditional last question for everyone.

1:01:03 What is the kindest thing that anyone's ever done for you? It's a tough question to answer because I do think there are many moments in my career where I got a leg up on something. I think I have the kind of personality where I sometimes don't appreciate in the moment and only reflect on it afterwards.

1:01:17 the three moments in my career that stand out. Actually one of them I'd already mentioned to you, which was the arm farm thing. I'm especially grateful to Jeff and to Vincent for Willing to take that. That's on me and my colleagues. And then there are a couple other moments.

1:01:30 When I started my Post talk with Peter Beale at Berkeley. I had zero robotics experience. I had done virtual character animation and computer graphics. I felt like that was

1:01:39 I bet on My potential more so than my actual accomplishments. And there was another moment even earlier on. I got an internship at Nvidia that got me to like experience some cool

1:01:48 Stuff. When I was just like a sophomore and I think the hiring manager for that also took a bet on me and I think that these kinds of things. They really matter in a person's career and I think that At the moment I should have been more grateful, but certainly in hindsight it's something that made a big difference and hopefully I can make that difference in other people's careers as well. Well, I've learned so much from you and your co founders and so much today. Thank you so much for your time.

1:02:09 Okay. If you enjoyed this episode, visit Colossus.com. You'll find every episode of this podcast complete with hand edited transcripts. You can also subscribe to Colossus, our quarterly print, digital, and private audio publication, featuring in-depth profiles of the founders, investors, and companies that we admire most. Learn at Colossus.com/slash subscribe.