Al Engineering 101 with Chip Huyen (Nvidia, Stanford, Netflix) Transcript from https://podmenti.com/t/d4f134932e0ca39a The question that got asked a lot and a lot is how do we keep up to date with the latest AI news? Why why do you need to keep up to date with the latest AI news? If you talk to the users who understand what they want or they don't want look into the feedback, then it can actually improve the application way way way more. A lot of companies are building AI products. We are even ideal crisis. Now we have all this really cool tool to have do everything from scratch. It can have your write code, it can have your website. So in theory we should see a lot more. But at the same time people like somehow stuff. They don't know what to build. All is AI hype, the data is actually showing most companies try it, doesn't do a lot, they stop. What do you think is the gap here? It's really hard to measure productivity. So I do ask people to ask their managers would you rather have give everyone out of team very expensive coding Asian subscriptions or you get an extra head cow. Almost everyone that managers could say head cow. But if you ask VP level or someone who managed a lot of teams, they could say one AI. assistant because as managers you are still growing. So for you having one extra head cap is big. Whereas for executive maybe you have more business metrics that you you care about. So you actually think about what actually drive productivity metrics for you. Today my guest is Chip Huen. Unlike a lot of people who share insights into building great AI products and where things are heading. Chip has built multiple successful AI products, platforms, tools. Chip was a core developer on NVIDIA's Nemo platform. An AI researcher at Netflix. She taught machine learning at Stanford. She's also a two time founder and the author of two of the most popular books in the world of AI. Including her most recent book called AI Engineering, which has been the most read book on the O'Reilly platform since its launch. She's also gotten to work with a lot of enterprises on their AI strategies, and so she gets to see what's actually happening on the ground inside a lot of different companies. In our conversation, Chip explains a lot of the basics, like What exactly does pre training and post training look like? What is RAG? What is reinforcement learning? What is RLHF? We also get into everything she's learned about how to build great AI products, including what people think it takes and what it actually takes. We talk about the most common pitfalls that companies run into, where she's seeing the most productivity gains, and so much more. This episode is quite technical, more technical than most conversations I've had. And is meant for anyone looking for a more in depth conversation about AI. If you enjoy this podcast, don't forget to subscribe and follow it in your favorite podcasting app or YouTube. And if you become an annual subscriber of my newsletter, you get a year free of 16 incredible products. including Devin, Lovable, Replic, Bolt, N8M, Linear, Superhuman, D Scrip, Whisperflow Gamma, Perplexity, Warp Granola, Magic Patterns, Recast, JPRD, and Mobbin. Head on over to Lenny's Newsletter dot com and click Product Pass. With that I bring you Chip. When after a short word from our sponsors. This episode is brought to you by D Scout. Design teams today are expected to move fast, but also to get it right. That's where D Scout comes in. DScout is the all-in-one research platform built for modern product and design teams. Whether you're running usability tests, interviews, surveys, or in the wild field work. D Scap makes it easy to connect with real users and get real insights fast. You can even test your Figma prototypes directly inside the platform. No juggling tools, no chasing ghost participants. And with the industry's most trusted panel, plus AI powered analysis, your team gets clarity and confidence to build better without slowing down. So if you're ready to streamline your research, speed of decisions, and design with impact, head to dScout.com to learn more. That's D S C O U T dot com. The answers you need to move confidently. Did you know that I have a whole team that helps me with my podcast and with my newsletter? I want everyone on that team to be super happy and thrive in their roles. JustWorks knows that your employees are more than just your employees. They're your people. My team is spread out across Colorado, Australia, Nepal, West Africa, and San Francisco. My life would be so incredibly complicated to hire people internationally, to pay people on time and in their local currencies, and to answer their HR questions 24-7. But with just works, it's super easy. Whether you're setting up your own automated payroll, offering premium benefits, or hiring internationally, JustWorks offers simple software and 24-7 human support from small business experts for you and your people. They do your human resources right, so that you can do right by your people. Just works. For your people. Chip, thank you so much for being here and welcome to the podcast. Hi, Lenny. I've been a big fan of the podcast for a while, so I'm really excited to be here. Thank you for having me. I wanna start with this. Table slash chart. Are you shared on LinkedIn? A while ago that went super viral and I think it went Super viral'cause it hit a nerve with a lot of people. And let me just read this and we'll show this on YouTube for people that are watching. So it's this very simple table you share it of What people think will improve AI apps and what actually improves AI apps. What people think will improve AI apps. Staying up to date with latest AI news. adopting the newest agency framework. Agonizing what vector databases to use. Constantly evaluating what model is smarter. Fine tuning a model. And then you have what actually improves AI apps. Talking to users, building more reliable platforms, preparing Better data. Optimizing end to end workflows. Writing better prompts. Why do you think that's a such a nerve with people and just w if you have to boil it down, what do you think is what do you think people are missing about building successful AI apps? What automation they get us a lot and a lot is that How do we keep up to date with the latest AI news? And I like Why why do you need to keep up to date with the latest AI news? I know it's a very cut here. Got into it, but I just so much news out there. A lot of people also ask me Questions like How do they choose between two different technologies? Like maybe like recently like MCP versus like Asian to Asians, right? Like protocol. And it was like which one is better or like this or that. And I think it's a a simple question that you should ask them. It's like first, Like if uh how much of the improvement could you get, like from like optimal solutions versus non optimal solutions, right? And sometimes it was like actually it's not much. Right? And I'm going to say, okay, if it's not much improvement, why do you want to spend so much time debating something that doesn't uh makes a much difference to your performance. And another question they ask is like if you adopt a new technology, like how hard it could be. Should I switch? That I'll show another. And sometimes it will like, oh I think it could be like A lot of work stitching it out. And I was just like, Hm. Let's say he's a new technology. It hasn't been tested by a lot of people. And if you adopt it, you would be like stuck with it forever. Like, do you actually may want to adopt it? Right, and maybe you want to think twice about. About like Overcommit. to like uh new technology that hasn't been Battle test it. I love your just broader advice is just simple. Like talk to to build successful AI apps, talk to users. Build better data. uh write better prompts, optimize the user experience. Well versus just like what is the latest and greatest, what's the best model to use right now, what's happening in AI. Let me follow the thread of this idea of fine tuning and basically post training. There's all these terms that people hear in AI, and I think this is gonna be a really good opportunity for people to learn what we're actually talking about. Since you actually do these things, you build these things, you work with companies doing these things. And um there's a few terms I want to sprinkle in through the conversation, but let's start with this one. What what's the simplest way for someone to understand? What is the difference between pre-training and post training and then just how fine tuning fits into that, just what fine tuning actually is. Shop disclaimer, I don't have like full visibility into like on what like this big secretive like frontier labs are doing. Uh but right from what I heard, right. So so I think it's like one is um like supervised file tuning when you have demonstration data and you have like a bunch of like l um experts like okay, here's a prop, right, and here is uh what the answer should be like. And I you you you just train it like on like to like stim like uh simulate um like emulate what The human export would be like And that's also like what a lot of people would like, um the the open source. models are doing as they do it by distillation. So instead of having human experts to like write Really good starting up grid. answers to my prompts that get like very popular Famous good models. to like generates a respond to it and like getting this chin smaller model to emulate. So so it's something that you see people just like So that's because I some I I I really appreciate open source community, by the way, but like going from like have being able to train the models that can emulate a existing good model. It's very different from like being a TG trained a good model is like an out before. Existing. good model. So it's a big step there. Uh so yeah, so like we have my supervised five tuning. And another thing is that's like very big and great, uh I'm not sure you have guests talking about it already, but like reinforcement learning is like everywhere. Let's pause on that because I would definitely want to spend time on that. And that's such a cool topic that I want that's merging more and more in my conversations, but just to even summarize the things you just shared, which I think is really, really important stuff. So the idea here is A model essentially this algorithm piece of code that someone writes. It's say the frontier models are feeding it just like the entire internet of content. And basically it's Trying to test itself on predicting In all of the in all across all the data, the next Word. Essentially the token is a simpler way is the correct way of thinking about it, but a simpler way to think about it is like the next word in a in text. And As it gets it. wrong, it adjusts these things called weights, essentially. Uh. Just like is that a simple way to think about it, even though that's oh that even that's just like very surface level? So I think of language modelling as a way of encoding statistical information about language. Right. So so let's say that um you we you we both speak English. So we kinda get a sense of like what is more statistically likely. Like if I say My favorite colour is then you would tell okay, that should be another color, like the word blue would be much more likely to appear than the word like uh in a table. Right, because statistically So so it's a sign get is this um is is It's a way of encoding statistical information. So like when language modeling when you train a large amount of data, like it got you see a lot of languages, a lot of domains. So it can tell, like, okay. Ubic size is standard. then it user do the prompts and it could come like with the next uh most likely Token. Uh so by the way, it's not a new idea. I shouldn't have video so it's the idea comes very, very old, like from the nineteen fifty-one papers. Um um the English entropy. I think it's like called Shannon. It's a great paper. And I think it researches a story I really like. Um is from did you read Shello Home, by the way? Uh yeah, I read a few Sherlock Holmes books, yeah. Yeah, so so this is story of like when Sherlock Holmes was using this statical information to like have sewn a case. So he was getting um so this is this story, uh there is Uh somebody left a message uh with a lot of stick figures. So Sherholm was like okay. He knows that in English. Це мус. Come on letter is e. Then the Morse come on stick figure. Must be Mm. Right, and then he goes he stuck like that, he was really so Uh so the uh the code. So I think there's language so In a way that's like simple language modelling, right? But instead of like at a work level he does it as like Top like character level. And token is something in between, right? A token is not quite a word, uh, but it's bigger than a character. So let's say uh we say token because uh it helps us like read would help us reduce vocabulary because which character is like smallest amount of like vocabulary right now. So the unpaper has a turn is the character, but words can have like millions and millions, right? Uh whereas um Um Tokens you can like Be able to like get Like the sweet spot maybe the two. So let's say that we have a uh the new word, like uh Uh how do I say like Podcasting, right? I say is a is a new word, but it can divide in a podcast and ing. So people understand, okay, podcast, we know the meaning, we know that ing is like uh like a verve like gerund whatever it is. So we know the word like podcasting. So that's why it's a token comes in. But yeah, uh the pre-tuning is basically like encoding statistical information of language to have you predict what is most likely. Um I think that most likely is the most simple way of doing it. Uh because it's more like building a distribution of like okay, so there's a next token could be like more like like ninety percent of the channel it could be like a color. Not doing like ten percent of the time could be something else, right? So you're based on distribution so language could like pick. Like depending on your sampling strategy. Like do you want it to always pick the most likely token or do you want it to pick something more creative? You know, so so so I think my sampling strategy I think is something extremely important. That's it can have you boost up performance in a in a huge way and very, very underrated. Okay, awesome. So essentially Yeah, a model is just code with this whole Uh set of weights, essentially the statistical model that has learned to predict what comes next after certain words and phrases. Yeah. And then post training and fine tuning specifically is doing that same thing. So Pre training, you get like GPT five. Fine tuning is someone taking GPT five and doing the same sort of thing. uh adjusting these weights a little bit for specific use cases on data that they Uh find as necessary to do their very specific use case. Is that a simple way to think about it? Yeah, I think you're playing with just like functions, right? So let's say just like you you have uh Maybe as a functions of like Maybe Lenny's height is maybe like One X Like one X plus something, like two X, like one and s plus something is is a weight, right? So you change it until you fit. the uh the correct data, which is like my height and your height, right? So you can think of the width is just like a weight like they function. So you so so you like chain adjust the weight so they can fit the data, which is the training data. Awesome. Okay so we're talking about pre training, post training, fine tuning. Is there anything else here that's important to share about just like what this is ex exactly? What people need to understand about These parts of training. So the vast majority of time we don't touch on like Christian or like as users, we don't. Right. Already done for us. Yeah. So so I think my actually it's a bit of fun like uh process like when my friends trading models, they try to play with the pre trading model and they're horrendous. So they're like saying things like it was like Oh my gosh, this is like Yeah, it's crazy. Um so so is this very interesting to look at like how much of like Post training can change. The mode behavior Um, yeah, and I think that's where like a lot of time as that a lot of people are spending energy on nowadays, they're frontier lap is on like push training. Because uh pre training uh I think um so pre training have been usually increase the general cab capacity of of of a model capabilities of a model. And it depends on it. You need a lot of data and like motor size, like you increase um to increase the model abilities. And at some point we are actually like have quite maxed out on the internet data. Right, and then people like text it up. I think a lot of audioing like but other data like audios and videos, and everyone's trying to think of like what is the new source of data. But if we're like portraiting, uh but like metal quasi is just more of like Everyone can have very similar. Pre-shooting data. Is that post trading is where they make a big difference nowadays. This is a good segue to you talked about supervised learning versus unsupervised learning. I love we're getting into this, by the way. This is super interesting. So you talk about labeled data. Basically supervised learning is AI learning on data that somebody has already labeled and told it here's correct versus incorrect. For example, this is spam versus not spam. This is Uh a good s short story. This is not a good short story. We've had uh the COs of a lot of these companies that do this for labs, uh Mercor and Scale. Handshake, uh there's micro Uh there's a few others. So is is that Essentially what these companies are doing for labs, giving them labeled data, high quality data to train on. It is in a way, but I think it's more like a part of a big equation. So there are a lot more different components than that. So that's why I was talking about reinforcement learning. I'm not sure if your CEOs that you interviewed bring up like that term. Uh so the idea is that um Yeah, once you both like So like let's say you have a model, give the model like a prop, right, and it produce an output, right? You want reinforce or anchorage the model to produce an output. That is better, right? So so like the how like now is it comes to like how do we know that the answer is good or bad? Right. So easy people realize on like um Signals. So one way to get like a first one good or bad is like human feedback, right? You have you have two responses. You can okay, this one's better than the other. Um and we do that is because like as humans, uh we tend to it's very hard to give like concrete score. But it's easier to do comparisons, right? Like if you ask me, Okay, give this song a score. I'm not a musician's, like and don't know like how hard it is. Yeah, I don't know like what like now ten and six, you know, and like if you ask me again a month from now and I completely forgot it and say, Okay, maybe now seven or like four. I don't know. But then if you ask me, Okay, here are two songs and which one could you prefer to play for the birthday party. I was like, Okay, I can play it before this song. So like comparison is a lot easier. Um So yeah, so we have a human um you have human feedback uh and then you use this human feedback to trade a reward model. So I tell I wish. Like so and then the free road model will have you like okay, it's a model that produces this response. Is the reward model can score? Is this good or bad? And you try to bias toward like producing better model uh the better responses. Another way is like you can instead of using a humans, you can use like AI. Right, like the rest born and say ye yes or good good or bad, right? Or in fact the thing is that people are very big on now and it's like ver verifiable rewards, but it's like natural. Um so basically they give it a math problems and then math solutions, uh Like it's a model up with a solution. It's you know that okay, it's uh expected response in it for you two. And it doesn't provide for it to and it's then it's wrong, right? It's not a good response. Um so so yeah so like uh a lot of time people like using this, um human labor, like human Um Human laborers should like produce like ma like um how say expert Questions? And I say expected answers. And in the ways that like desired systems that like verifiable. So is that the The models can can be trained on, yeah. Okay, I'm really glad you went there. This is essentially R L H F Reinforcement learning with human feedback, which is exactly what I wanted to also talk about. Right. Yeah, so um I think it's like it's general. It's like it's a way of learning. It's like training is convertible learning and whether it learned from human feedback or like AI feedback or like verifiable rewards, uh I think they say you say it's just different way of like Um Clipping signals. Awesome. Yeah. That's uh I we had that C of Anthropic on the podcast and he talked about their version of RLHF, which is AI driven. Reinforcement learning. I love the way you phrase it, where you basically You want to help the model you want to reinforce correct behavior and correct answers. And this is the method to do it, whether it's Say an engineer. Seeing an output. from a model being like no, here's how I would code it differently. And then training And it's training a different model that the original model Works with to tell it am I correct or not correct is that right? Yeah. I I think I think that's a way of of looking into it. And I think that's a space is so exciting nowadays because He has so many like domain exports. Task? That the model. Like some more developers want models to do well on, right? Let's say uh like accountant, right? Like maybe I want to use a model to have an accounting task. So I need a lot of like accounting data, like examples for my accountants. So you need to hire a lot of them, should I do it? Or maybe it wants a physics program, or you want to do um I know, like legal questions and stuff or like engineering questions or like somebody was telling me they want to do like uh using like coding for to source scientific problems and not just like coding to be a product, which is another different whole realm of things. And I also like using very specific toolings like uh Yeah, like uh I'm not sure like what apps you use, but maybe like um for eighting app or uh like QuickBooks or like Google X Excel. Like they're very specific, like tool specific um expert expertise. So you want the model to what you learn. So I they need a lot of like humans expert. in this area should like create data. To trade them. Um and it's a massive thing. Because uh everyone wants a lot of data and like one slabs have like unlimited budget. Uh, but uh where there I think this is also like a little bit of low key. Interesting economics. I'm not sure you've talked to like So guess about I thought it's very interesting to think about. Because it's very lopsided, right? Because like is it only like a very small Numbers of Frontier Labs, right? And they want a lot of data. And there's like a massive amount of like startups or companies that providing later data. So like you can see this companies like this startup like doing later labeling, but they have like maybe they have like massive AR. But you ask them like okay, so how many customers you have? And they could be like uh very small numbers. I'm not sure I'm not sure you you you it's all you're smiling. Yeah, we chat we chat about that. Yeah, so so I'm like a bit like look me uneasy, right? I hear have like a company's growing like crazy, but it's like heavily dependent. On Like two or three. companies. And at the same time I if I if I was this company from TLS. What would be the right? economical things. For me to do, right? Now I want a lot of startups. I want to have a lot of providers. So it can pick and choose. And then this providers can also like to compete each other. So so I feel like yeah, so So yeah, just it kinda makes a whole economics is very interesting to me and I'm curious to see how it plays out. What I'm hearing is you're uh you're bearish on the future of these data labeling companies. Because as you said, they don't don't have a lot of leverage over pricing'cause they have so few customers. And there's so many people getting into the space. So basically Even though they're some of the fastest growing companies in the world you're feeling like There's there's a challenge up ahead. I'm not sure if it's some bearish on it. Um, I think um Curious. Because I think Things have has a way of work out in ways that I Don't expect So I think that maybe these companies they have a lot of data, maybe they wouldn't be able to use that to like have some insight that helps them like stay ahead of the curve. You know? So I don't know. Oh. A very fair answer. Okay, while we're on this topic, I wanna chat about evals, which is a very recurring topic in this podcast. This is the other piece of data content these companies share that AI labs really need. Can you just talk about what And eval is the simplest way to understand it and then how this helps models get smarter. So I think the people approach Eva, I think they're like two very different problems. Once is a app Right. And like can I say have an app, uh do like a Maybe A chatbot. Very simple, and I was the first thing that came to my mind. Um And I want you to know if Chebo is good or bad, right? So so it needs to come a little way with like if I'm checkboard. Um another thing is uh I think of this as a uh Task specific Avant Design. So let's say I'm a motor developer and I want to make my model better at curb writing. Right. And it was like okay. But how how how do you even imagine current writing, right? So як ви ні сам вачули. Okay, understand curve writing and think about like what makes good story like what makes a story good and then design the whole data set. And then criteria to evaluate creep writing. Um so yeah, so so I think there's that I think it's like more like eval. Design. That is very interesting. Uh come up with crit criteria. Um And then also like Trained people. Like how to do it effectively. Um so I guess uh in a quiz I think Ivar is really, really fun. Because it's extremely creative. Uh I was looking at like different air on building and was like, wow. Like Is it's not dry at all. It's just like super super super fun. With whole podcast any valves with Hammel. Hammel and Shreya and Uh And that's exactly what they talked about. It's just it's actually really fun to create evals for for companies, especially. So let's still dig into that one a little bit more. There's this kind of debate online that I don't know how big of a deal this debate is, but it feels like people uh Spent a lot of time thinking about this. This idea of Do we need evals for AI products? Some of the best companies say they don't really do evals, they just Go on vibes. They're just like, Is this working well? Can I feel it? Or not. What's your take on just the importance of building evals and the skill of evals for app AI apps, not the model companies. You don't have to be like Absolutely perfect I think to win. You just need to be like good enough and being consistent about it. Okay, this is not a philosophy I follow. But like I have worked with enough companies to see that play out. So when I say like why come you don't evaluate, let's say you are like an executive, right? And you want to have a new use case. So here's a use case you you started out, we built, and it's like it works well, right? The customers are somewhat happy, don't you don't have the exact measure for it, but like so the traffic keeps increasing, like people seem happy, people keep buying stuff, right? And then here's our engineer coming like okay. We need Ivar for it. And so I and it's not saying it was like, okay, how much effort do we need to put into it up? And they were like, Okay, uh maybe like two engineers as much as much. And it could maybe would improve sort and would say, okay, so how much expect gain can I get from it? And the engineer would be like, Oh, maybe you can improve it from like Eighty percent should like eighty two percent, eighty five percent, right. And it was like okay, but it'd be taken that two engineers and be gonna launch a new feature. Then it could give me like so much more, like improvement, right? So I think it's like one of them is like Eva. Sometimes people think of Eva. It's like okay, this is good enough to st touch it and do spell of energy on Eva. I would like only incremental improvement where it spends the energy. On like another use case. And maybe in a scale good enough that you does a vibe check it, right? So so I I do think it's like maybe like that's a debate it's about, um And do things that's like a lot of time people just like get things to the to some place when it's like okay. Good enough. People run. But and then but of course it's like there's a lot of risk associated with it because if we don't have a clear Match it uh you have good feasibility to hustle application is a model is a performing, it might do something very dumb or it it can cost you like, I know it's some something like crazy. can happen. So so yeah, so um So so I do think AVA is very, very important if you have if operate a scale. And where like failures can have like catastrophic consequences. then you do need to be very tyrannical about like what you put in in front of the users, understand different failure modes, like what could go wrong. And also maybe in a space when that like it's it's a feature as of the product is as a competitive vantage, right? You want to be the best at it. So you want to have like a very strong understanding of like where you are. and like way or with our competitors. But it's just something that's like more like a low key. Okay, it's like something it's like okay that's not the core and like it helps with our users. Then maybe you don't need to be so so obsessed or about it is okay, that's good enough for now. And if it fails, then it fails. Okay, I know it's like it's so terrifying, but like, yeah. Um Yeah, I think it's on about the question of like regional regional investment. Um I'm a big fan of Avar, I love writing Avar. And the size like I understand why some people would choose to not focus on ever right away and choose like bringing on new functionalities instead. Awesome. That is a really pragmatic answer. What I'm hearing is Evals are great and very important, especially if you're operating at scale. But pick your battles. You don't need to write any battles for every little feature. Something that Hamel and Shreya shared is that People need just like I don't know, five or seven evals for the most important elements of their product. Is that is that what you see, or do you see a lot more in production that people Building. Mm. I I don't Think of like Just a fixed number on like the Evars. Like what was a go to Evar, right? The go to Evar is to guise the product development. Um so so like you see Avar, um because I think I'm a big fan of Avar is is that it has you uncover opportunities where the progress of doing well. So I sometimes see in a very obvious okay, so okay, we look at the Avar and we realise it's like, okay, it performed really poorly. On just like specific segment of users. And then we're looking to it's like, okay, what what what what's what's wrong with it. And it turns out it's like we just like don't have a good messaging to it. So like maybe we'll show like just focus on the things of reading pool can improve significantly. Yeah, so I can guess the number of eval is Like we have the same product but like hundreds of different metrics, right? Crazy. This is because like set product is like general, right? Have different things have like on Avar for like I don't know, like uh verbosity, have like one Avar for like user sensitive data. Um and like another is like um for length. But like has a number of like um Okay, that's just example, compared simple. Like diff research. So so you have the application, you have like be as a model. She like do deep research for you, right? Like okay, like have a prompt and I may say, okay. do me a com comprehensive research. on old Nenny's podcast. And have me like sell like uh propose like a s show me report on what kinda topics he's interested in, what kind of videos get the most views, or like what topics that he's missing on that he should be covering, right? Like have a kinda like prompt. Then how do you evaluate? The as a resort. Right. I don't think there's like one like metric that would help. Maybe it's just like maybe you have like a hundred I don't somebody has a benchmark and is a get like a hundred expert, like write a bunch of prompts and then go through like all the on the answers on a yeah and I do it and it's like it's extremely costly and slow. Right. But if I might have something else, for some more like uh one way was thinking about it, um I was talking to a friend about it and and one way is just like how to produce the result of the of the summary, right? At first you need to be able to do like gather information. And to gather information you need to do a lot of search queries. Uh you like gather the graph the social results and then from the social results you like uh aggregate and then maybe say, Okay, I'm still missing out on this. You have to go another route and like another route. And it's the another summary. So every step of the way. You need evaluations. Right, it don't need to N to end. So maybe if it was a search query, it might first think about like Okay, now I write five such queries. And I look into like how good are the search queries? Like do they like are they like similar to each other? Because in the five search queries that are very similar, like okay, then it podcasts, then it then it podcast uh last month then it podcast like two months ago, right? It's not It's not very exciting. But like if the quality is a podcast like the the keywords are like more Um More diverse, right? And then look at the results of the of the search query, and they say you enter the search query, like and then it pass cat Data labeling. And then they come up with like ten pictures, uh, ten results. And then you come up with like oh Lenny podcasts on uh I don't know, mo um I don't know, like uh Frontier laps and have like 10 resorts. And my look was a different web page, like how much of them overlapping. Like are we are we doing both like the breath, like getting a lot of But also like do we have death and also like relevance because if we come up with a search queries are completely irrelevant to the to to the original prompt. So I feel like every aspect of it, it would need a way of evaluating, right? So I don't think it's just like How many Eva should I get, but like How many Ivar should Do I need to get a good coverage, a high confidence. in my application's performance and also to help me understand like where it is not performing well so that I can fix it. Awesome. And I'm hearing also just especially for the very core use case, like the most common path people take in your product is where you want to focus. Yeah, so yeah. Okay. Let me uh there's one more term I wanna cover and then I wanna go to Somewhat different direction. Rag. People see this term a lot, R A G What does it mean? So Vragus then for ritual augmented generations in a so not A specific shoe Genip AI? So um the idea is just like for a lot of questions we need contact to answer. So I think it came pretty oh, I think it's from the paper twenty seventeen. So so someone was like Um so they realize it's like for a bunch of like benchmark when the question answering benchmarks. They realize it's like, okay, if we give the model information about the questions, then it's the answer can be much, much better. So what it does is try to retrieve information from W Wikipedia. So for for questionable topics, it's like reshrift that and then put it into the context and like answer it does much better. So I feel like it sounds like a no brinner, right? I mean, like obviously. So so I think that's what Rackett as a simplest sense is just like providing the model with a relevant context so so that it can answer the questions. And that's why like things get like uh really um more more interesting. Because traditionally when it started out, uh RAC is mostly like text. Um so so we we we talk about like a lot of way of like how to prepare data. So that the model can retrieve uh effectively. Let's say there's like not everything is a Wikipedia page, right? Wikiped Wikipedia page is pretty contained and you know that okay, everything about it is about a topic. Uh, but a lot of time have documents extremely a lot. Right, and like they have a weird way of like structuring the documents. Let's say that um you had documents about Lenny uh podcast. Right. And in the f in the future in the beginning of documents like From now on. Podcast wouldn't refer to Lenny's podcast, right? So let's say somebody in the future is like okay, tell me about Lenny. Right, money's work. And because Many You just don't know, uh, you might not retreat it. And the document is long enough that it chunk into a different part so like the second part happen s doesn't have the the word many, so you cannot reach it. So I have to find a way to process data. So that makes sure that's like It can retrieve. the information is as relevant to the query even though it might not immediately like obvious that is related. So if we'll come up with like only thing if I think like contextual visual. Um like uh giving extra angles of data. uh relevant like maybe in a summary metadata so that it knows um or a sumable use of like as um hypothetical questions is very interesting. Like for a given channel of like documents, I must get it a bunch of questions. such a chunks can help answer. So that when I have a a query, it was like, okay, does it mesh any of the like hypothetical questions? So it can it can fetch it. So it's very interesting approach. Okay, so maybe Before I go to the next thing, I just want to say this like data preparations for RAC is extremely important. And I will say that's like in the a lot of the companies that I've have seen that's like the biggest performance. in their solutions coming from like better data data preparations. not agonizing over like what red databases to use. Uh we've got very data base. Uh a coin is very important if you care about like things like latency or like if you have like very specific S patterns, like we heavy or right heavy. Of course it's like it matters. But in terms of like pure quality answers, right? I think they did a preparation and just like hands out. When you say data preparation, what's an example to make that real and concrete for us to understand. So like one way is like uh um smashing as in like um you have like chunks of data. So we can think about like how big of H chunk should be, right? Because um If it's like sort of think about like it's it's a context you want to maximize, maybe you can I'm very simple example, right? Now you want to retrieve like a thousand words, right? So if H chump's data is too long. Um then so if if a minute champ is long then it's more likely to contain more relevant meter data, so we can retrieve more. But If It's too long, like then he had thousands of wood. And so Chang is like a thousand words to get a rich one chang. So it's not very useful. But it's too short. Then you can can retrieve more relevant information. Like oh some it can retrieve uh wider range of like documents and chunk. But at the same time, each chang is too small. She contained relevant information. So we have like very nice like Chunk design, like how big each um should be. Uh, you add like contextual information, like summary, metadata, hypothetical questions. Uh somebody was telling me just like um a very big performance it got is that from um rewriting their data in the question answering format. So like instead of having like of so they have a podcast, right? Instead of just chunking of the podcast, you just like reframe rewrite it into like Here's a question, here's answers, um, like and and produce a lot of them. It can use AI for that as well. So that's pretty simple of data processing. A lot of example we I see is like for people helping like using AI to have like specific uh tool news and documentations, right? And a lot and we write documentation usually to our document. occumentation today is written for human. um reading. And AI reading is different because it's different because humans we have a common sense. And we kinda know what it is. Um so so or on things that are like uh human experts they have the context that AI doesn't quite have. So somebody told me is that like what's a big change they have is like Let's say that um you have a you have a function, a document um documentation for this maybe Z library. And Z library says okay, the output of this one is like maybe talking for like uh I know some crazy term. Maybe there's some uh temperature or something under grab. It should be like one Zero or minus one. And as a human expert, maybe understand the scale and what one in the scale mean. But like for I just really doesn't understand what I mean. So so actually have like another allotation annotation layer layer. For AI. I say, okay, good. Temperatures. you've got one mint like that. It's not like it's assurant temperature. It's more like associated with a scale. But there. So is that saving that or just data processing to make it easier for AI to retrieve the relevant information to answer the questions? This episode is brought to you by Persona, the verified identity platform helping organizations onboard users, fight fraud, and build trust. We talk a lot on this podcast about the amazing advances in AI, but this can be a double edged sword. For every wow moment, there are fraudsters using the same tech to wreak havoc. Laundering money, taking over employee identities, and impersonating businesses. Persona helps combat these threats with automated user business and employee verification. Whether you're looking to catch candidate fraud, meet age restrictions, or keep your platform safe. Persona helps you verify users in a way that's tailored to your specific needs. Best of all, Persona makes it easy to know who you're dealing with without adding friction for good users. This is why leading platforms like Etsy, LinkedIn, Square, and Lyft trust persona to secure their platform. Persona is also offering my listeners 500 free services per month for one full year. Just head to with persona.com slash Lenny to get started. That's with Persona dot com slash Lenny. Thanks again to Persona for sponsoring this episode. Awesome. Okay. So you've talked a bit about how you work with companies on these sorts of things, on their AI strategies, on their AI products, how they build. Which tools they build, all these things. I wanna spend a little time here. 'Cause a lot of companies are building AI products. A lot of companies are not having a good time building AI products. Let me ask a few questions along these lines of what you've learned working with companies that are doing this well. One is just, I guess, in terms of AI tool adoption and adoption generally in companies, there's all this talk recently of just like all this AI hype. The data's actually showing most companies try it. Doesn't do a lot, they stop. And so there's all this just like maybe this isn't going anywhere. So in terms of just adoption of tools and AI within companies, what are you seeing there? for Gen AI in company, I think there are two type of Gen AI. toolings that I've been uh I've seen s like ones is to like um internal productivity. Right, and I have coding tools, Slack chat bot, um Uh internal knowledge like A lot of big enterprises have some kind like a rubber irrelic, um model so then but like with access like maybe some different kind of rack. So we should I think we'd talk about that up um okay like text based rack. I haven't talked about an Asian tech rack or like having like multi mode rack yet, but this like yes as a whole very exciting area or rather. Um yeah, so like b basically to allow the employee to like Аксес. internal. document uh some way someone ask like okay um I'm having a baby what it could be is a maternal paternal policy, right? Or like am I having these operations good hell benefit like cover that or like I want you like interview or I want you like refer my friend, but could be the process for that. So a lot of this like having chatbot, internal chatbot to have with internal operations. Um, and um another things um another category is more like Customer facing? Um so or like partner facing. Um so so chatbot is a big one. You have a hotel chain, you might have like a booking chatbot, which is like somehow massive, like a lot of booking chatbot because I guess it's it's it's I do have this theory of like a lot of applications uh companies pursue because they can measure the concrete outcome. And I feel like booking or a sales chatbot is very clear, right? Like what's a conversion rate right now with a chatbot uh with human operators, and what could be conversion rate with a chatbot. And it's something somehow I think it's like very clear outcomes in companies are easier to buy into This um this solution. So a lot of companies have that like customer uh facing chatbot. Uh so yeah, so so that is um Another category of two. Um and I think that Um I don't know for customers or external facing tools, um because People are driven to Um people are driven sho choose applications with clear outcomes. So so the questions of uh adopting them is really based on like whether they see the outcome or not. Of course it's not perfect because sometimes uh the outcome can be bad not because the idea or like the application's idea sell is bad just because a So I know the the the process of building it is like not that great. Um, yeah, so so it's tricky. For the internal adoptions of like toolings or internal productivity, that's where it gets tricky. I would say like a lot of companies, uh was it think of a sharing. Like I think of a sh have like usually have very Um have like two key aspects, right? It's like use cases. And the second is talent. You may have a great data for great use cases, but you don't have talents and you cannot do it. So a lot of time in the beginning with Gen AI and still and something that I'm really admired a lot of companies for that is like exactly it was like, Okay, we need our employees to be very Gen AI aware. Like very AI literary, right? So what it does is I start like Maybe like Adopting a bunch of tools. for for the team to use. They have a lot of upskill uh upskilling uh workshops, like they encourage learning. I think it's like a really, really good thing. And it's also a willing to spend a lot of money into like adapting, like um giving people like Chester B D subscriptions, cursor subscriptions, uh Cloud Code subscriptions like to get the employees. She like to to be more AI literate. Um and then the thing is like a lot of the secondary no comes from me say, Okay, we spent a ton of money on this tooling. But then we don't See because you can see the usage. But people don't seem to use them as much. And what is the issue? So so yeah, so I think there's that that is um That is tricky, yeah. What do you think is the issue? Is it just They're not they're like they don't know how to use them. Like what do you think is the gap here? Do you think we'll get to a place of just like wow, work is completely different because of AI for a lot of companies. The main thing is like it's really hard to measure productivity. Okay. So I tell you a lot of people and they was like first of all on a sample is coding, right? A lot of companies not using coding agents, uh or the coding AI said it's coding. Uh and um I was asking it was like And once I could do do you think that like it helps with your productivity? And a lot of times the question is very hand wavy. It's like Okay, it was like okay, I feel like it's paid better, right? And I said okay, because we have more PRs Uh we see more code and then immediate correctness. Okay, but of course quote number of life code is not a good metric for that. Right. So so it's it's really, really tricky. And it's something uh funny. Uh so So so I do ask um people to ask their managers because I work with like either they've VP level so they have like multiple teams under them. So I asked them like okay, do you ask a manager um like okay, would you rather have access, uh good you rather have give everyone on the team, like very expensive. Coding Asians. Um subscriptions. Or you get an extra headcount, right? And let's say it's like maybe like um and and Almost everyone. Good say the measures could say head cow. But if you ask V P level or like someone who managed a lot of teams, they would say just like they could one AI. A system. as it's tools. And the reason is that we could say like, okay, because as a manager is right, because you are still growing. Like you're not as a level when you you manage hundreds of thousands of people. So for you like having one Hascount. Is like is big. So you want that not for productivity, but because you just want to have more people working for you. Whereas for executive you care more about like the Um the maybe we have more like business metrics that you you care about. So so you actually think about like what actually drive drive productivity uh metrics for you. Uh so so yeah, so it it's tricky. Um and I think it's like the question of like productivity Um is not I'm not sure it's like fundamentally is the some people more productive, but it's just like we don't have a good way of measuring Productivity improvement. Uh another thing is also very wily. Um and I think that people Do tell me does they notice different buckets of of employees, like different reactions? to AI as the tools. Like first of all, I I'm I keep going back to coding because it's for it is big and it's like easier to show my rhythms about. Um so This is like um I haven't even report. Like One team would tell me is that like, um one of the people would tell me, Okay, amongst one his engineers You think it's like senior engineers would get the most output. I would be more productive because Like, okay, so that person's very interesting. So so he actually divided his team. to like three bucket. But he didn't tell them obviously. He was like okay, here's more like currently like uh best performing, average performing, and lowest performing. And then it does a randomized trial. So let's give like half of eight. of H proof like access to like to like cursor. And I was noticed like over time, I was like, Okay, so Something funny, like the the group that get the biggest performing boost, like in his opinion, so it was very close in his team. The biggest boom boost by the senior engineer, the the highest performing. So the highest performing engineer get the biggest boost out of it. And then the second group is just like the Um The Irish performing. So so he he so his opinion is like okay The highest performing engineers are also not more proactive. They always say no has a solid problem. So I I have some sold problem better. Whereas the people who always have my lowest performing, they only don't care much about work, right? So like it's easier to just like go on autopilot, get it to like generate like that code and just like do it. And I just don't know how to do it. As another company, however. They tell me that's like actually senior engineers are the one most resistant. She like using AI. As is tooling. Because I said it's like okay, but AI Because they are more upionated and they have very high standard. It was like okay, but AI cool, Jericho just sucks. So just like very, very resistant in using it. So I don't know. I I haven't Fight. be able to reconcile like very different Report on that yet. This is so interesting. So just to make sure I'm hearing what you're the story, so there's a company you work with That did a three bucket test with their engineering team. Where they Created three sorts of groups. the highest performing engineers, mid performing engineers, lowest performing engineers. Uh and gave some of them so they gave some of them access to say cursor. Was it cursor or what did they give them access to it was cursor? I think they say it was cursor. Okay, cool. And so within I didn't work with them just more like a friend company. Okay, it's a friend's company. So did they give like half of the higher performing engineers cursor and half not, or how did they do the split? Yeah, so like they give like half of the entire company, but like half for each bucket. Yeah. And then it observes a difference in like production. I see. Yeah. So how do they even do that? They're just like, Okay, you get cursor, you don't get cursors. How did they do that? That's so strange. Yeah, I I do again just the mechanics of it. Uh, but but I was like at rest back here for doing a randomized trial. That is so cool. Yeah. Okay, wow. How large was this engineering team? Was it like hundreds of people? Um it's it's not that large. It's about like maybe the um Three, twenty forty? Yeah. Thirty to forty. Okay. Yeah. Wow. Okay. So they found that the highest performing engineers had the most benefit. from using AI tools and then behind them was the middle Tier engineers and the worst performers. Yeah. Yeah. But it's so not the same everywhere. Um, yeah, yeah, different. Right. There's other example you shared. of just senior engineers in this one example are most resistant to changing the way they work, which I get. I I do feel like the most valuable people right now, other than ML researchers. Uh uh and AI researchers like yourself. uh are senior engineers because it feels like junior engineers Just like so much of this is now done by AI, but it's but an engine that knows what they're doing, that understands how things work. At a large scale. with AI tools. Just basically like infinite junior engineers doing their bidding. Feels like an extremely valuable and powerful asset. Yeah, uh I definitely like really appreciate uh as you see companies like we appreciate engineers who are Um have a good understanding of the whole systems and be able to have good problem solving skill, like thinking holistically instead of like local I uh locally. Um or when our company FCS the way they work as I told me is it's like we're completely different now. And like so they actually restructure the engineering orc so that like they get more senior engineers should be more in the peer review. because I like to get like uh sort of writing guidelines on like what is a good engineering practice is um what is a process would be like? Maybe like okay, so is it right, like a lot of like processes are on how to work well. And then they um and then they have Mm Junior engineer just like Produce court in and like some APR, but senior engineer more in the reviewing case. So I think is it might be P PR for the future. So another company actually told me something very similar. So they can't preparing for future once they only need a very small. group of like very, very strong engineers. to like create like Processes. And like reviewing. call to get into production, but I get like AI. or like junior engineers to I produce code. But then the question becomes just like How does one become? A very strong. That's right. That's right. I feel like the development. Yeah, so so I don't know what's the process. I was thinking about like yeah, um No one's thinking about it. It's just it's a problem. We won't have any more in ten, twenty years, there'll be no more. Engineers'cause no one's hiring GD engineers. Although I could make the case student engineers, people just getting into computer science right now are just Native. AI native. And in theory, you could argue they will become really good really fast if they're curious. Aren't just you know. The delegating learning and thinking to AI, but learning how to actually using it to learn how to code well and architect correctly. Like you could argue they will be the most successful Engineers in the future. I do think that what I mentioned is I loading to architect, um I think I I group that in my system thinking. I do think it's very important skill. Because I think AI can help automate a lot of like Um Destroyed it. Skill. But like knowing how to like utilise the skills together to solve a problem. Is is very uh it's it's It's hard. So there's a a webinar between um Mean so I mean. was my one of my favorite professors. He was a chair of the curriculum of the C department at Snapford, so he spent a lot of time thinking about CS education is right. Like what what what should students learn nowadays in the era of like AI coding. And then the other person is like Andrew, which is of course is like a a legend in the AI space. And Nira Sammy present like Sami said something very interesting. Is it like he said like A lot more thing as C S is about coding, but it's not. My coding is just a means to an end. Like CS is about system thinking, like using like coding to solve actual problem and problem solving will never go away. Because I what like I can automate more stuff, the promises get bigger. But as a process of understanding what causes the issue and like how to like design step by step solution to it will always be there. Um so I've taken an example of um I have like I actually have a lot of issues with like AI for like um In the wave I is debugging. So I'm not sure you use a lot of air for quitting, but like a uh something I noticed and also seeing for my friends, it's like It is pretty good when you have very clear well defined tasks. Maybe the right documentation is s fix the specific features or like build an app from scratch, right? Like doesn't have to interact with a large existing code base. But it added something like a little bit more complicated. Uh maybe require interacting with other components and stuff. It's usually like not that Good. Um and and for simple I was using AI to like use um to deep line applications. Um and it was testing out a new uh posting service. It was not for me to know it. It was like okay, like usually they for me so what AI does give me is like confidence is trying new tool. Like before what AI is like trying new tools history, not documentation for the beginning, but AI was like, Okay, just try it out and s and and learn. So I was testing as a new hosting service. And it kept getting a buck. It was like very, very annoying. And it was like Okay, I asked uh Clark was like fix it. And it keep giving like it keep changing the way, like maybe change the environment variable fix the code, maybe not change from the function to this function, maybe change the l language, maybe maybe it doesn't process Jemis Well, I don't know, whatever. And then right. And it was like Okay, that's it. I'm just gonna read the uh documentation myself and see what's wrong. And it turns out it's like I'm on another tier. As a fish I want did not is not available in this tier, right? So I feel like okay, so the issue with Clark Code is trying to focus on fixing things from a very a different component, whereas the issue is from a different component. So it's in a I think of like okay be understanding a how different components um work together and where the source of the issue might come from. You need to you need to give a holistic view of it. And it's made me think it's like okay, how do we teach AI like system thinking? Like that right and I have all the human experts like having like right like very much people scaffold Oh it's just like okay, Fox this guy problem, look into this, look into that, look into that and then stuff. So so I think that could be one way. But that's what made me think it's like how do we teach humans, like system thinking. Um yeah. So so yeah, so I think it's very interesting. Um Skill. I I do think it's very important. That's exactly the same insight Brett Taylor shared on the podcast. He's the Covener Sierra, he created Google Maps, he was CO of Salesforce, Quip, a few other things, and I asked him just like should people learn to code? And His point is exactly what you said, which is Learning taking computer science classes is not about Learning Java. And Python, it's learning how systems work and how Code operates and how Software works broadly, not just Here's like a function to do a thing. One thing that I wanted to Help people understand you you wrote this book called AI Engineering, which is essentially helping people understand this new Genre of engineer. And you have this really simple way of thinking about the difference between an ML engineer and an AI engineer. Which has a really good corollary to product managers now. of just like an AI product manager versus uh non AI product manager, the way you describe it in ML engineers built models themselves. AI engineers use existing models to build products. Anything you want to add there? One thing I really dislike about writing books is that you have to define like this. And and I think it's like no definitions can be perfect because they're always really edge cases. Um but yeah, in general, I think it's like gen uh AI as a service, like more as a service. Like when somebody build the models for you and the base. model performances are pretty shock. So so it's like it's enabled people to just like okay, now I want you Integrate. A I ensure my product. I don't need to learn what green design is. Even though knowing that would really help. Uh but but yeah, it's like it's makes the entry barrier really low for people who want to use. A I should be a parecch. And at the same time AI could Go abilities are like so strong. It's like it's also like increased like the possibilities, like the type applications that AI can be used for. So I think like yeah, so it was entry bearers like super low and I said demand for like A applications like a lot bigger. So it feels it's very very exciting. It's opened up like a whole new Ball of possibilities. Yeah. It's like now you don't have the time You don't even have to spend time building this AI brain. Now you can just use it to do stuff. uh such a such an unlock. Okay, maybe just a final question. you get to see a lot of where what's working, what's not working, where things are heading. I'm curious just if you had to think about it. In the next two or three years, just where things are heading. What do you think? What do you think? How do you think? Building product. will be different? How do you think companies working will be different if you had to think of Maybe the biggest change we expect to see in the next few years in terms of how companies work. I Thing is a lot of organizations said don't move that fast, right? Um But at the same time. They will move faster than I expected. Uh because uh again, I think it's like bias like and don't work with a dinosaur companies don't get I think a lot of Exactly who come to me are like very forward looking. So maybe for me I'm very biased, uhwards the word like organizations is like Move fast. Um so so yeah, so I think one big change I see is just like in organizational structure. Um, I think it's a a lot of value plays um in like um So before it we have like a lot of disjoint team. Like we have very clear like engineering team. Good up to him. But then there's a question of like Who should write it about? Right, like who should own the metrics. And it turns out it's like if all it's not a um it's not a it's not a separate problem. It's a system problem, right? Because you you need to look into different components, how to interest each other, you need to use the behaviors, because you need to know what users care about so that you can So that you can like write it off. reflect what users care about. So so all of that like you can swan it from like you look into different component architectures, uh place cardrails and stuff. So it's just engineering, but understanding users is like what product, right? So so because of like a lot of things and of our extremely important. So like that kind of bringed product team and like engineering team, even like marketing team, like user acquisition, like very close to each other. So so yes, in so in a way so I think we were structuring so that's more communications between like previously very distinct functions. Another thing is I also see as teams, um Of course I Think about like what can be automated. in the next few years and what we're cannot be automated. And I see that people already like shedding like Actually is it's a little bit like Scary to think about it, but I saw things it's like the team Web Tobies it's like okay, it's been you and me, but we have we like got rid of these functions, right? Like for a lot of things like uh previously outsourced, for example. Like traditionally is a visa that's outsourced thing that's not core to them and like can be done with like not um can be then more Um System Mike, uh um systematized. Um so so with that you can actually like use AI to like automate a lot of that. And so as a separation people think more about like what is the value of like junior engineers or senior engineers, how to restructure engineering. Okay. For that? Um so yeah, so I do definitely think that um is one thing to success organization people just moving pieces around and I think about like use cases, um Whether we wanted to like spin out new use cases and who would lead the new effort and I yeah. Um That is one Big. Uh change. Another thing's in terms of like AI I think is there's um I'm not sure how true this is. Um, I guess I'm I'm also like on the cam of like thinking that it's Has merit. is it's a kind of like, okay, um base models we have probably like not quite maxile, but we What We are unlikely to see like really, really Strong, like crazy strong. model. So like is you remember like when we have like GPT, right? And GPT two, which is a big step up, like on automatic to like. like better than like G P D and then GP three, which like much, much bigger. G four, much, much bigger. And I of course I'm going to G P D five, but like is GP five like that Scale of like Much bigger. Like a step jump compared to like the previous. I think it's a debateable, right? So so I think that it's like we had reached a point with like the base model um performance improvements is not gonna be like Mind blowing. in the last three years. Uh is it so so I think it's like a lot of like improvements we're gonna see in the post training phase, in the application building phase. And um And yes, that's what I think is that's where I feel I was See a lot of improvement there. I also very like interest in like multi modality. Um so we've seen a lot of uh text based uh but I think there's a lot of um audio, videos, use cases, uh that is very, very exciting. And I think Rio it's not quite as soft as it was thinking because I do work with like uh with with like a a a couple of like voice startups. And I'm gonna talk to a vo thing about voice is a Entirely different. Yes. Uh so let's say her chatbot, right. We go from a text chatbot to voice chatbot. It's like the console are completely different. Because now with voice shot, right? We need to think about like latency. Because having multiple steps are first like have like tech like voice to text. text to text and text question into text answer and then like and then text to voice answer, right? So it's like manageable hops. And like latency becomes very important. And it's a question like what does it make you sound natural? So for example like people think it like um In in in AI and humans, when when humans talk to each other. Like if I say I say You were trying to interrupt me and I say, um chip. That's right. I would click pause. and I try to hear you up, right? But sometimes I just even say say some word like acknowledge when I mm Mm. I shouldn't stop. I just continue. So the question of like force interruption and where there is like I should should I stop or not? Like it's it's a big in the what perceived as like natural conversations. And that's also regulations, right? Because like A lot of time people want to build AI chatbot voice chatbots that sound like humans. Try to like trick users into thinking they're talking to humans. But it's a very maybe potential regulation saying like okay, you have to disclose to users. When to top If If the bot is talking to his human or um AI. So so I think just like um there's a whole space. I think it's not quite as sole as as you think uh is it but it's all it's all not Quite like a AI foundation model problem, right? Because like a human interruption detection is actually a classical machine problem. a framing of like you can view classifier. for that. Oh like the question of like let us see uh actually it was a massive engineering challenge, not an AI challenge. Oh, because they can be an AI challenge because we're trying to build a voice to voice model. So instead of having like Having to first transcribe the voice from me into text and then get a model Jared text answer and get another model to like to information speak, you can send your voice your voice directly. So that is something we're working on, but it's like very hard. Um yeah. So so yeah, so like Even audio, I think of it is like the easier than video. Right, because if you do have like For image in voice. Uh it's already like pretty hard. So I think it's a lot of challenges in that space. That was an awesome list of things. Let me mirror them back real quick. So what you're predicting in the next few years, things that will Change in the way we work. And these actually resonate. with so many conversations I've had on this podcast. So this is just kind of doubling doubling down on where things are heading. One is The blurring of lines between different functions instead of just like eng design engineering. Everyone's gonna be doing a lot of different things now. Uh two is just more of work being automated with agents and all these AI tools. And just In theory, productivity going up. Third is a shifting from pre-training models to post training, fine tuning and things like that, because To your point, model. Models maybe are slowing down and how smart they're getting. Although I'll point folks to the chat with the co-founder of Anthropic, he made a really good point here. He's like We're really bad at Understanding what exponentials feel like when we're in the middle of that. And also models are being released more often. So the difference between them we may not notice because they're just happening more often versus GPT three came out like a year I don't know. af before after JPT two. So uh Maybe true, maybe not. And then the fourth point you made is This idea of multimodal investing in multi modal experiences. I cannot wait for Chat GPT voice mode to get better at interruption. Like exactly what you're saying. I'm just like talking to it and then someone makes a little sound. It's like Okay. And then you have to And then it's like and then it stops talking. It's so annoying. I'm shocked that we don't have better voice assistants at home yet. I think I have been testing out a bunch, honestly. I keep hoping oh my God. Zach could be the one isn't I don't know how many of them I just like had to get away because they're not that good. I think it's coming. I hear it's coming. Anthropic's working with someone. Uh that I don't know if it's launched or not yet. Yeah. I was really want to bring back to what it mentions about like the um so guests like from Anthropic mentioned about the performance uh improvement. I think there's a big change. Um I think like um there's difference between um a model based capability. So I'm talking about like the pre-trained model, right? Versus a perceived performance. So so let's say it's like um I'm actually uh thought about like are you familiar with the term test time compute? Uh I don't think so. Yeah, so um so the idea is it's like okay, like um do you have some Or fix them out computer, right? So you're gonna spend a lot of compute on Appreciating our training the model. And then I'm Then a lot of uh some computer only five children and the ratio of like pre tuning and posturing computer is like crazy very different different map. Um Um, and also like since then it has just been comp uh on like generate inference. When I have a trend and five tender model now it want to like serve it to users. So I might type a question as a prom and it's like generate like do inference like and zap requires a compute. And I guess I feel more discussion of like, uh should I spend more compute on like precinct and Or fight tuning. Or inference, right? Because like inference and people five hours just like test and compute. So like Spending more compute on inference is like calling like Test them like uh compute. Uh like as a strategy of like just allocating more resources, computer resource to share it. Uh In front. when I shouldn't bring better performers and how does that do it? Like let's say um let's say you have a math questions, right? And maybe instead of just generic one answer, I can share like four different answers and say, Okay, whichever is uh the best according to some standard. Uh all I okay have four answers and then maybe like Three of them say four A two and one of them says like twenty eight. You say, Okay, three of them in the in in agreement. So the answer should be four A two. Right. So like just people shouldn't generate a bunch of it. Or another thing is like a lot of time like reasoning, uh thinking, it's just like people should like generate more thinking tokens and spend more time thinking before showing the final answers. Uh it's like require more compute, whereas like give more uh more and more So so yeah, so so I think it's like From the user perspective, right? Like When is a model spend more time exploring. different potential answers, thinking longer. It can give you much better final answers. But the base model itself. Does not change. Does it make sense? Yes, that does. Absolutely. Yeah. Uh that is a good corollary to uh To Ben Man's. Point. Yeah. We covered a lot of ground. I've gone through everything I was hoping to learn and more. Before we get to our very exciting lightning round, is there anything else that you wanted to share, anything else you want to leave listeners with? So I do work at a few companies that do this things of like they want employees to like come up with ideas. So there's a big debate on like what is a better way for a strategy of I should it be top down or like bottom up? Right, should like executive come up at like Why not should like kill a use case and like everyone like allocate resources to that? Or should you give engineers and PMs and smart people like come up with ideas? And I think it's a mixture of both. So some companies it was like okay, we hire a bunch of smart people. Like let's see like what they come up with. And they they organize like more hackathons or like in internal challenge to get people to to build product. And one thing's that um I noticed it's like a lot of people just like don't know what you built. Uh and it shocked me, like why I feel like we are in some kind of like an idea crisis, right? Now we have all this really cool tools. You have you like do everything from scratch. I can have you like design, it can have you like write code, it can have your website. So in theory we should see a lot more. But at the same time people are like somehow stuck, like they don't know. What to build. And and I think it's like and maybe you see a lot of had to do with like maybe like um society expectations. Because I we have gone through uh we have gone into this phase of like specialization. It's like people like uh very highly uh specialized and people are supposed to do like focus on one thing. really well instead of like a big picture then we don't have a big picture of you it's hard to come up with like ideas of what you build. So so I don't want like uh when when I work with this company, I just hike that on like we do work out like a how come up with a guideline, like how to come up with ideas. And usually what we think of is like okay like one TB is like Go Look from the last week, right. Like for all weeks just like pay attention to what you're doing, what frustrate you. And when something frustrates you think I sne It's yeah, like can it be done a different way, so it's not frustrating. And it can talk like people can swap accept subnotes or teams, and I even see it come on frustrations. Maybe it's just something you can think about is just to build something around that. So yeah, so I feel like um just like notice like how we work, uh thinking of like way, so like constantly ask questions like how can it be better? And then I just build something to like address the frustration. I think it's a good way to just like run and adopt AI. I think people have felt exactly what you're describing every time they open up one of these vibe coding tools where they could just describe anything you want. I'm like, I don't know, what do I want? And and I love this very tactical piece of advice, just like what frustrates you, just pay attention to where you're frustrated. For example I just built a very cool little vibe coded app. I was working on a Newsletter post inside Google Docs. And I I pasted all these images into the Google Doc from screenshots and stuff. And and then I forgot, oh yeah, you can't take images out of Google Docs. It's like this Hotel California experience where you can paste stuff into it. Very hard to get images back out. So I I just went to all the vibe coded tools and just build an app that I can give you a Google Doc U and Let me download all the images automatically. And it worked amazingly well and it made it really cute and I'll I'll link to it in the show notes. Oh, I would love to see that. I do I'm very bullish on using A I just create like micro tools. But it's just something just like Make your life a bit easier. In a hundred a hundred percent. I feel like that's one of the main ways people are using these tools just like A little niche problem they have. With that. Chip, we've reached our very exciting lightning round. I've got five questions for you. Are you ready? Yeah, always. No, uh it depends on how hard the questions are. They're very uh consistent across every guest, so Uh I imagine you've heard them before. First question, what are two or three books that you find yourself recommending most to other people? Oof, I'm really terrified of like book recommendations because I feel like what books a person should read really depends on what they want and where they're in life and where they want to get to. Um but I just several books that I do think is have she really changed the way I So one thing is just the fish gene. This I should understand, uh it actually changed. Uh it actually helped me with the question like whether I want to have kids or not. Uh, because it's like uh understanding more of like um Yeah, a lot of our functions of way we operate is the functions of our genes. Uh and gins once in the one thing was like to procreate us. So so yes, in a little way. But I still like So but also proposed another thing is like, um so everyone wants to live forever, right? And maybe it's not like Consciously, but subconsciously. We we do we do want that. And and I said two ways, like one is like by genes. My genes one is just like Once it could do forever, but it's also the two ideas. Um, I think there's something gonna meme. Uh it's like being able to if you have some ideas out there and then this like last for a long time, this is one boy should like live on. I know it's like it's a little bit like um Uh shreck? But I thought it's very interesting. So yeah, the books I really, really like is like from like um the book from um Singaporean um previous um I think he says a Like uh Liu Kong Yo, I'm not sure what's the title of it, but like he did so he was the one who left Singapore from uh he's changed uh Singapore from a three world country to a fourth world country within twenty five years. And I have never seen any country leaders have spent so much effort into like putting down his thought on like how To build a country. of like like that. Um yeah, and I say talk a lot about like public policy, uh like how to like create policies of anchorage. People should do the right things. That is good for the nations. And also talking about like uh foreign affairs, foreign policies, like So it's a really good book to think about. For me it's like system thinking, but like it's a different kind of system which a country which a lot of us don't get a chance to like ever experiment in our life. So it's good to learn about that. What was the name of that second book? Uh It's calling from third to first world fashion. I think we have it somewhere here. Yeah. There it is. That's awesome. I definitely want to read that. That's a really good tip. Uh I've heard a lot about just the impact he's had and I've seen all these videos on Twitter of just his really wise insights into how to build a thriving Society. And clearly it works. How do you see if Tancho rises such a thick book? Insane. That is Claude, please summarize. I'm just joking. Uh, by the way, Selfish Dean, I also absolutely love that book. That is such a good choice. It's such an under the radar kind of book that really Change the way I see the world as well. Uh, so really good pick. Okay, next question. Do you have a favorite recent movie or TV show you really enjoyed? So I watch a lot of movie and TV shows as a research. Uh because I I working on my first novel and I recently uh sold it. So I'm interested in like what makes it it's a drama, it's not a science fiction or uh anything that like tech people usually read. So it's it's very like I know it's a very um out of the left fell Of left fill and like very um so it's almost like reading Watching T V to see like what kind of stories become popular trying to understand the trope and and stuff like that. So I'm not sure if the audience are like Well what's one? What's one that taught you something about writing? I think you could like JMC Pelez? Is a Chinese TV show? Cool. Okay. Haven't heard how that went on the podcast before. Okay, cool. Yeah. Next question. Do you have a life motto? that you often think about Come back to when you're dealing with something hard. Whether it's in work or in life. This sauce we're nihilist. I think society like in the end nothing really matters. Uh easy thing like in the grand scale thing like in a billionaires. Nothing will like no one will ever be there. I think okay. Someone will argue with me about that. So I'm going to say like I still might theories like in a billion years, like none of us will never exist. So like so like wherever like messy things, like crazy things we do, or like how bad we do it. I mean no one would be remember We've been there to remember it. And I think in a way it's like it's self scary, but it's very liberating. Because it's just allows me okay, let's just try things out, right? Like why does it matter? And then it's a story of like recently, um so have uh some family member who passed away recently. And uh I was talking to my dad. Because I couldn't be home for that. I was asking my dad like, Okay, is there anything I can do? To make the person l like Oh s something like comfort, so anything that you can get that person's. And my dad was just like What can he possibly want? at this moment. Like and it's made me real like as the end of life, like there's nothing that can bring you like Like much, no like money, no product, nothing. And in a way to make okay, what really do really care about. That's the end of the day. Um, so I guess it's like a thing about it. It's just like okay, maybe I fail it, maybe I don't get a contract, maybe do things like is it but is it at the end of life, like I don't think that I should really Matters. So in a way it's like it's kinda liberating. Uh, I know you said it might be nihilistic. This is what Steve Jobs shared too in one of his most famous speeches just We will all die some day, so don't take things so seriously. And it is freeing, absolutely. It just makes you appreciate every moment every day you have, just like yeah. Let's just do something hard and scary. Okay, final question. You talked about how you're writing a novel. Most people in tech uh have never written Something creative and fiction. What's just like one thing you learned in the process about how to write better stories, better fiction? A lot of time when we read, uh we get tripped up by some small things. So m I think like I I wanna should do curve writing because I just want to become a better writer. And it passes like maybe try my w I a different audience could have made like become better like anticipating what this different type of audience would want you here and like are able to care about. So it's a way for me to get a so I think if I write it. Oh, it even like any kind of like content creations it's about like Predicting. The user's reactions, right? Just kidding. Yeah. So so like you do a podcast it's like okay, what kind of the user is gonna find engaging, right? And and I find it's like a little bit like uh in a lot of companies like you have like launch a product, you have a narrative coming out. So okay, what kind how do we positions this product in a way that like users would want, right? So I feel like I have done technical writing for a while and I felt like I have Has some experience like trying to predict what Engineers would want you here. Uh all I care about. But then I don't have an experience like this. completely different type of audience. So that's what I want you to like career writing, writing a story. And that's why I was doing a lot of research on my question. I mean, going research, but I should enjoy a lot, like watching a lot of drama. So just see like what what it will like Um so the one thing that I care about is just like I think I learned it's like what like emotional journey was from my editor. Right. So like when we write something where we we care about like how users would feel. Like across the the story. Like we want something in the beginning, right? We want something just like We we need to have a hook so that we will continue reading. But we also don't want too much of like drama. Because we'll get like too tired, right? Like uh because like you're the emotionally exhausted. Like because it's like you're being like emotionally manipulated. Like a lot of time. So if you have like emotional emotional journey, maybe like some some climax or like some something more chill. And maybe like I also care about another thing that I I didn't realise like for me So for technical writing. You entirely focus on the content. like the argument it's very impersonal, right? Like it it like for example, like people like ML compilers. Like doesn't matter if See like the person telling them about compiler or not, right? Because it's just like objective like I but like for for novel, people care about like character likability. So so like in in the first version is my story and makes the characters like a little bit more like uh very uh very logical, very rational. And just does everything just like very rationally. And then the feedback I got is I have a very good friend read it and he was he's a amazing person, he's a great person. And he was like shit, I'll be honest with you, I hate that person. So it doesn't matter as a story. It's just like the person is so unlikable. So he's a second version and makes a person uh second a lot more likable. Like what How she made it? character more likable is that you put in some vulnerability. Like some of the time that okay maybe it's a person like has set back because somehow we can relate to it. See in a lot of ways like It's very interesting. It's like a lot of it is like um Yeah, a lot of it is it's about like understand the emotional bit. Uh like how's the users feel. Not just about the story, but also about the characters. That is so interesting. Wow. I learned a lot more there than I thought. That was awesome. Really good example. Chip, two final questions. Where can folks find you online if they wanna reach out and maybe work with you or Uh, maybe even just share the stuff that you offer if folks want to reach out. And then How can listeners be useful to you? I'm like a master social media, LinkedIn, Twitter. I don't post a lot, but I keep telling myself that I should do more because I quite like the conversation with uh with um Bit readers. Uh so I'm should be about to start a a sli um uh a sub spec. Um So I have like a placeholder for SubSpect right now and I'm thinking of doing it for more system thinking because I think it's a very interesting skill. Um, and so like thinking of doing a YouTube channel on book reviews and basically books that have you thing. Better so I think it's the first book I'm gonna review is probably like this book. because it's like my favorite book growing up, uh and I have been like keep on reading it. Uh so yeah, so how can it be helpful? Like Send me books that you like. Uh books that Help You have change. So what you think? Or change you the way you do anything. So I would appreciate it. Amazing. I'm I'm excited to read that book. Uh Chip, thank you so much for being here. Thank you so much, Lenny, for having me. By everyone. Thank you so much for listening. If you found this valuable, you can subscribe to the show on Apple Podcasts, Spotify, or your favorite podcast app. Also, please consider giving us a rating or leaving a review, as that really helps other listeners find the podcast. You can find all past episodes or learn more about the show. at Lenny's podcast dot com. See you in the next episode.