Why AI evals are the hottest new skill for product builders | Hamel Husain & Shreya Shankar (creators of the #1 eval course) Transcript from https://podmenti.com/t/f88289fbb2a15505 To build great AI products, you need to be really good at building evals. It's the highest ROI activity you can engage in. This process is a lot of fun. Everyone that does this. immediately gets addicted to it when you're building an AI application. You just learn a lot. What's cool about this is you don't need to do this many, many times. For most products you do this process once and then you build on it. The goal is not to do evals perfectly, it's to actionably improve your product. I did not realize how much controversy and drama there is around evals. There's a lot of people with very strong opinions. People have been burned by evals in the past. People have done evals badly, and then they didn't trust it anymore, and then they're like, Oh, I'm anti eval. What are a couple of the most common misconceptions people have with evals? The top one is we live in the age of AI. Can't the AI just eval it? But It doesn't work. A term that you used in your post that I love is this idea of a benevolent dictator. When you're doing this open coding, a lot of teams get bogged down in having a committee do this. For a lot of situations, that's wholly unnecessary. You don't want to make this process so expensive that you can't do it. You can appoint one person whose taste that you Trust. It should be the person with domain expertise. Oftentimes it is the product manager. Today my guests are Hamal Hussein and Shreya Shankar. One of the most trending topics on this podcast over the past year. has been the rise of Evals. Both the chief product officers of Anthropic and OpenAI. Share that evals are becoming the most important new skill for product builders. And sin, this has been a recurring theme across many of the top AI builders I've had on. Two years ago I had never heard the term e bells, now it's coming up constantly. When was the last time that a new skill emerged that product builders had to get good at to be successful? Hamel and Shreya have played a major role in shifting Evals from being an obscure, mysterious subject To one of the most necessary skills for AI product builders. They teach the definitive online course on Evals, which happens to be the number one course on Maven. They've now taught over two thousand pms and engineers across five hundred companies. including large swaths of the OpenAI and Anthropic teams along with every other major AI lab. In this conversation, we do a lot of show versus tell. We walk through the process of developing an effective eval. Explain what the heck Evalves are and what they look like. Address many of the major misconceptions with evals, give you the first few steps you can take to start building evals for your product. And also share just the ton of best practices that Hamil and Trey have developed over the past few years. This episode is the deepest yet most understandable primer. You'll find on the world of e bells. And honestly got me excited to write evals, even though I have nothing to write evals for. I think you'll feel the same way as you watch this. If this conversation gets you excited, definitely check out Hamill and Shreas's course on Maven, we'll link to it in the show notes. If you use the code Lenny's List when you purchase the course, you'll get thirty-five percent off the price of the course. With that, I bring you Hamel Hussein and Shreya Shankar. This episode is brought to you by Finn, the number one AI agent for customer service. If your customer support tickets are piling up, then you need Finn. Ben is the highest performing AI agent on the market with a 65% average resolution rate. Fin resolves even the most complex customer queries. No other AI agent performs better. In head-to-head bake-offs with competitors, Finn wins every time. Yes, switching to a new tool can be scary, but Fin works on any help desk with no migration needed, which means you don't have to overhaul your current system or deal with delays in service for your customers. And Fin is trusted by over 5,000 customer service leaders and top AI companies like Anthropic and Synthesia. And because Fin is powered by the Fin AI engine, which is a continuously improving system that allows you to analyze, train, test, and deploy with ease. Fin can continuously improve your results too. So if you're ready to transform your customer service and scale your support, give Fin a try. For only 99 cents per resolution. Plus, Fin comes with a 90-day money back guarantee. Find out how Finn can work for your team at FIN.ai slash Lenny. That's Finn.ai slash Lenny. This episode is brought to you by D Scout. Design teams today are expected to move fast, but also to get it right. That's where D Scout comes in. DScout is the all-in-one research platform built for modern product and design teams. Whether you're writing usability tests, interviews, surveys, or in the wild field work. D Scat makes it easy to connect with real users and get real insights fast. You can even test your Figma prototypes directly inside the platform. No juggling tools, no chasing ghost participants. And with the industry's most trusted panel. Plus AI powered analysis, your team gets clarity and confidence to build better without slowing down. So if you're ready to streamline your research, speed of decisions, and design with impact, head to dscout.com to learn more. That's D S C O U T dot com. The answers you need to move confidently. Hamil and Shreya, thank you so much for being here and welcome to the podcast. Thank you for having us. Yeah, super excited. I'm even more excited. Okay, so A couple years ago. I had never Heard the term Evals. Now it's one of the most trending topics on my podcast. Essentially that to build great AI products, you need to be really good at building evals. Uh also turns out some of the fastest growing companies in the world are basically building and selling and creating evals for AI labs. I just had the CF Mercor on the podcast. So there's something really big happening here. Uh I wanna use this conversation to basically help people understand the space deeply. But let's start with the basics. Just what What the heck are evals for folks that have no idea what we're talking about, give us just a quick understanding of what an eval is. And let's start with with Hamilton. Sure. Evals is a way to systematically measure and improve an AI application. It really Doesn't have to be scary. Or unapproachable at all? It Really is at its core. Data analytics. On your LLM application. In a systematic way of Looking at that data. And Where necessary creating metrics around Things so you can measure. What's happening and then so you can iterate And do experiments and improve. So that's a that's a really good broad way of thinking about it. If you go one level deeper just to give people a very even more concrete way of imagining and visualizing what we're talking about, even if you have a example to show would be even better. W what's a what's an even deeper way of understanding what an eval is? Лецею A real estate Assistant. You know, application. And It's It's not working the way you want. It's not writing emails to customers the way you want. Or it's not. Uh You know Calling the right tools. Or any number of errors. And Before evals. You would Be luck with guessing. You would maybe fix a prompt. In hope that you're not breaking anything else with that prompt. And You might rely on vibe checks, which is totally fine. And vibe checks are good, and you should do vibe checks. Initially. But it can become very unmanageable. Very fast because as your application grows. It's really hard to rely on vibe checks, you just feel lost. And so Evals help you create metrics that you can use to measure How your application is doing. In kind of give you a way to improve your app your application with confidence. The you you have a feedback signal. So just to make it very real. So imagining this uh real estate agent, maybe they're helping you book a listing or go s op see an open house. The idea here is you have this agent talking to people, it's answering questions, pointing them to things. as a builder of that agent, how do you know if it's giving them Good advice, good answers. Is it Telling him things that Completely wrong. So the idea of evaluation essentially is to build a set of tests. That tell you Is how often are Is this agent doing something wrong that you don't want it to do? And there's a bunch of ways wrong. You could define wrong. It could be Uh just making up stuff. It could be Uh just answering in a really strange way. Uh the way I think about evil isn't tell me if this is wrong, just simply is like unit tests. For it. For code and then they you're smiling, you're like, No, you idiot. Oh, that's not what I was thinking. Okay, okay, tell me, tell me. Does how does that feel as a metaphor? So okay. I like what you said first, which is we had a very broad definition. Evals is a big spectrum of ways to measure application quality. Now unit tests are one way of doing this. Maybe there are some non negotiable functionalities. that you want your AI assistant to have. And unit tests are going to be able to check that. Now, maybe you also because these AI assistants are doing such open ended tasks. You kind of also want to measure how good are they at very vague or ambiguous things like responding to new types of user requests or You know, figuring out if there's new distributions of data, like new users are coming and using your real estate agent that you didn't even know would use your product and then all of a sudden you think like oh there's a different way you want to kind of accommodate this new group of people. So evolves could also be Yeah, a way of looking at your data regularly to find these new cohorts of people. Evals could also be like metrics. that you know, you just want to track. Over time. Like you want to Track. People Saying yes, thumbs up. I liked your message. Um you want to tr Very, very basic things that are not necessarily AI related. but can go back into this flywheel of improving your product. So I would say on the end on Overall, right, unit tests are a very small part. Of that very big puzzle. Awesome. You guys actually brought an example of in Evo just to show us exactly what the hell we're talking about. We're talking in these big ideas. So How about let's pull one up and show people here's here's what an eval is. Yeah. Let me just set the stage for it a little bit. So To echo what Shreya said. It's really important that we don't Think of Evals As just tests. The common trap that a lot of people fall into. Because they jump straight to the test, like let me write some tests. And usually that's not what you want to do. You should start with some kind of data analysis. to ground what you should even test. And that's a little bit different than Software engineering where You have a lot more Expectations of how the system Is going to work. With LMs it's a lot more service area. It's very stochastic. So We kinda have a different flavor here. And so the example I'm gonna show you today It's actually a Real estate. example. It's a different kind of real estate example. It's uh from a company called Nurture Boss. I can share my screen to show you their website. Just to help you understand this uh use case a little bit. Let me share my screen. So this is a company that I worked with called Nurture Boss. And it is a AI assistant for property managers who are managing apartments. And it helps with various tasks such as Inbound leads. Customer service, booking appointments, so on and so forth. Like All the different sort of operations you might be doing as a property manager, it helps you with that. And so you know, you can see kind of What they do. It's a very good example because It has a lot of the complexities of a modern AI application. So There's Lots of different channels. That you can interact. Do the AI with like? Chat. Text, voice. But also there's tool calls. Lots of Tool calls for like booking appointments. Getting uh information about availability, so on and so forth. There's also Rag retrieval. getting information about customers and properties and things like that. So it's pretty fully fledged. In terms of an AI application. And so They have been really generous with me and uh allowing me to use their data as a teaching example. And so we have anonymized Yeah. But what I'm gonna walk through today is Okay. Let's create let's do the first part. Of How we would start to build evals for nurture boss. Like why would we even want to do that? So let's go through the very beginning stage what we call error analysis. Which is Let's look at the data. Of their application. And First start with What's going wrong? So I'm gonna jump to that next. And I'm gonna open an observability tool. And you can use whatever you want here. I just happen to have this data loaded. In a tool called Brain Trust. But You can Loaded in anything. You know, it's not we don't have a favorite tool or anything. In the blog post that we wrote with you Uh we had the same example, but in Phoenix. Arise. Um and I think Amon on your blog post use Phoenix or I've as well. And there's also Langsmith, so These are like kind of like different tools that you can use. So what you see here On the screen. This is Logs from The application. And Let me just show you how it looks. So What you see here Is and let me make it full screen. So this is one particular interaction that a customer had with the nurture boss a application. And What it is, it's a detailed log of everything that happened. So it's it's a It's called a trace. And it's just the engineering term for logs of a sequence of events. The concept of a trace has been around for a really long time. But it's especially really important when it comes to AI applications. So we have all the different components and pieces In information. that the AI needs to do its job. And we are logged all of it. And we're looking at A view of that. And so you see here a system prompt. The system prompt says you are an AI assistant working as a leasing team member. at retreat at Acme apartments. Remember I said this is anonymized, so that's why the name is ACme apartments. Your primary role is to respond to text messages from both residents and Perspective. Uh Both current residents and prospective residents. Your goal is to provide accurate, helpful information, yada yada yada. And then there's a lot of detail around guidelines of how we want this thing. to Behave. Is this their actual system prompt, by the way, for this company? It is. Yes. It's a real system prompt. That's amazing, because that's really how it's rare you see actual company product system prompt. That's like their crown jewels a lot of times. So this is actually very cool on its own. Yeah, yeah, it's really cool. And You know, you see all these different sort of features that they want to or different use cases. So things about tour scheduling, handling applications. guidance on how to Talk to different personas, so on and so forth. And you can see the user just kind of jumps here. It says Ask, okay, do you have a one bedroom with study available? I saw it on virtual tours. And then you can see that The L um Call some tools. It calls us Get Individuals Information Tool. And it pulls back That person's information And then It gets the community's availability. So it's you know Mm-hmm. It's querying a database with The availability for that apartment complex. And then Finally the AI responds, Hey, we we have several one bedroom apartments available. Um but none specifically listed with the study. Here are a few options. Uh then it says can you let me know when one with the study is available? Then it says I currently don't have Specific information on the availability of a one bedroom apartment. User says thank you. And the AI says you're welcome. If you have any more questions, feel free to reach out. No. This is An example of A trace. And this is We're looking at v one specific data point. And so one thing that's really important to do When you're doing data analysis. Of Your LM application. is to look at data. Now You might wonder There's a lot of these logs. It's kind of messy. There's a lot of things going on here. How in the hell are you supposed to Look at this data. Do you want to just drown in this data? How do you even analyze this data? So it turns out there is a way to do it that is completely manageable. And it's not something that we invented. It's been around in machine learning and data science for a really long time. And it's called error analysis. And what you do Is The first step in conquering data like this It's just to write notes. Okay. So you gotta put your product hat on. Which is why we're talking to you. Because Product people have to be In the room. Um, and they have to be involved in sort of doing this. You know, usually a developer is not suited to do this. Especially if it's not a coding application. And I'm just uh just to mirror back why I think you're saying that is because this is the user experience of your product. People talking to this agent is the entire product, essentially. And so it makes sense for the product person to be involved, super involved in this. Yeah. So let's let's reflect on this conversation. Okay, a user asked about Availability The AI Said, Oh, we don't really have that. Have a nice day. Now for a product that is helping you with Lean Management. Is that good? Like, do you feel Like this is The way we want it. To to go. Not ideal. Yes, not ideal. And I'm glad you said that. A lot of people would say, Oh, it's great, like the AI did the right thing. It said we don't It looked, said we didn't have available, and it's not available. But with your product hat on, you know that's not correct. And so what you would do Is you would just write a quick note here. You would say, Okay. Um You know, you might Pop in here. Let me just And you can write a note. So every observability ha application has ability to write notes. And you wouldn't try to figure out If something is wrong in this applic you know, in this case It's kind of not doing the right thing. Um will you just write a quick note? Um should You know, should have Can Handed off. To a human. And as we watch this happening it's Like you mentioned this and you'll explain more. You're doing this this feels very ma manual and unscalable. But uh as you said, this is just one step of the process and there's a system to this. And it's just the first part. And you have to do it for all of your Data. You can you sample your data and just take a look. And It's surprising how much you learn. When you do this. Everyone that does this Immediately. gets addicted to it and they say this is the greatest thing that you can do. when you're building an AI application, you just learn a lot. You're like hmm. This is not how I want it to to work. Okay. And so um that's just an example. So you write this note. And then we can go along to the next trace. So this is the next trace. I just pushed a hotkey on my keyboard. Let me go back to Uh looking at it. And these tools make it easy to go through a bunch and add these notes quickly. Yes. And so this is another one. Similar system prompt. We don't need to go through all of it again. We'll just jump right into the user question. Okay, I've been texting you all day. Maybe it's Is it funny? Um Um And Uh The user says please Okay, yeah, this one is Yeah, d this one is just like an error in the application where You know Um Th this is a text message application. And so You know, it's a tech the sorry, the channel through which the customer is communicating is through text message and you're just getting like really garbled. And You can see here that it kind of doesn't make sense. You know, like the words are being cut off like in the meantime. And then the system doesn't know how to respond.'Cause you know how people text message. They like write short phrases. Yeah. you know, split split their sentence across four or five different turns. Yeah, so this is uh This is a different kind of error. This is more of Hey, we're not handling this interaction correctly. It is more of a technical problem. Um, rather than hey, the AI is not doing exactly what we want. So we would write that down too. Like it's amazing you're catching that too here. Otherwise you'd have no idea this is happening. Yeah, you might not know this is happening, right? And so you would just say, Okay. Um you'll write a note like Oh. Conversation flow. Is Janky. Because of text message. And I like yeah, I like that. I like that you're using the word janky. It shows you just how informal This can be at this stage. Yeah. It's supposed to be chill. Like just don't overthink it. And there's some. There's a way to do this. So The question always comes up. How do you do this? Do you look at do you try to find all the different problems in this trace? What what do you write a note about? And the answer is Just Write down the first thing that you see that's wrong, the most upstream error. Don't worry about all the errors. Just capture The most the first thing that you see that's wrong. And stop and move on. And you can get really good at this. The first two or three can be very painful. But You know, It doesn't We can You know, do a bunch of them really fast. So here's another one. And Uh let's skip the system prompt again. And the user asked, Hey I'm looking for a two to three bedroom. With either one or two bats. Do you provide virtual tours? And A bunch of tools are called And it says Hi Sarah, currently we have Three bedroom, two and a half bathroom apartment available for Two thousand one hundred and seventy five dollars. Um, unfortunately, we don't have any two bedroom options at the moment. We do offer virtual tool tours, you can schedule a tour. Blah blah. It just so happens that There's no virtual tour. Mm. Right. Nice. You know, it is Hallucinating something that doesn't exist, then you would you kinda have to bring your context as an engineer. Or even, you know, product content and say, Hey, this is kinda weird, like You know, we shouldn't be telling person about virtual tour when it's not offered. So you would say, Okay Uh, you know. Offered virtual tour tour. And you just, you know. You just write the note. So you can see there's a diversity of different kinds of errors that we're seeing. And we're actually learning a lot. about your application. Um In a very short amount of time. One common question that we get from people at this stage is Okay, I understand what's going on. Can I ask an L M to do this process for me. Great question. And I loved Hamill's most recent example. Because What we usually find when we try to ask an L M to do this error analysis is it just says the trace looks good. Because it doesn't have the context needed. to understand whether something might be, you know, bad Product smell. Or Yeah. Not. For example, the hallucination about scheduling the tour, right? I can guarantee you I would bet money on this if I put that into chat GPT and asked, is there an error? it would say no. Did a great job. But Hamill had the context of knowing, Oh, we don't actually have this virtual Tor functionality. Right. So I think in these cases it's so important to make sure you are manually doing this yourself. Um and we'll talk a little we can talk a little bit more about when to use LLMs in the process later. But like number one pitfall right here is people are like, Let me automate this with an LLO. Do you think they'll we'll get to a place where where an agent can do this? Oh no no no sorry, there are parts of error analysis that an LLM is suited for. Which we can talk about later in this podcast. But right now in this stage of free form note taking It's not the place for an LLM. And this is something you call open coding. Yes, absolutely. Uh another uh t term that you used in your post that I love and that's fits into the step is this idea of a benevolent dictator. Maybe just talk about what that is and maybe Sharia cover that. Yeah. So Hamill actually came up with this term. Okay, maybe Hamill covered that. No problem. And we'll actually show uh the LM automation in this example. Because we're gonna take this example. We're gonna go all the way through. Amazing. And so And so um Benevolent dictator is just A catchy Term. For the fact that When you're doing this open coding. A lot of teams get bogged down. in having a committee do this. And for a lot of situations that's wholly unnecessary. Like you know, people get really uncomfortable with oh okay, you know We want everybody on board, we want everybody involved, so on and so forth. You need to cut through the noise. Um in a lot of organizations If you look really deeply Especially small and medium sized companies, there's really Like you can appoint one person whose taste that you trust. Um And you can you can do this with a small number of people and often one person. And that's it's really important to make this Tractable. You don't wanna make this process so expensive that you can't do it. You're gonna lose out. So That's the idea behind Benevolent Dictator is Hey. You need to simplify this. Across as many dimensions as you can. Another thing that we'll talk about later is When you go to building An LM as a judge, you need a binary score. You don't want to think about is this like a one, two, three, four, five, like assign a score to it. You can't that's gonna slow it down. Just to make sure this benevolent dictator point is is really clear, basically this is the person that does this note taking. And Ideally they're the expert on the stuff. So if it's law stuff. Maybe there's like a legal person that owns this. It could be a product manager. Give us advice on who this person should be. Yeah, it should be the person with domain expertise. So in this case Yeah, it would be The person who understands The business of leasing? apartment leasing and has context to understand if this makes sense. It it's always a domain expert, like you said. Okay, for legal it would be a law person, for mental health it would be the mental health expert, whether that's like a psychiatrist or You know, someone else. Cool. Um Oftentimes it is the product manager. Cool. So the advice there is pick that person. may not feel so super fair that they're the one in charge and they're the dictator. But they're benevolent. It's gonna go be okay. Yeah, it's gonna be okay. You're just trying to it's not perfection. You're just trying to make progress. And In Git signal quickly, so you have an idea of what to work on. Because it can become infinitely expensive if you're not careful. Yeah. Okay, cool. Let's go back to your examples. Yeah, no problem. So This is another example. Where We have Someone Saying, okay, do you have any specials? And This is the Or the AI responds, Hey, we have a five percent military discount. User response can you in the switch is a subject. Can you tell me how many floors there are? Do you have any one bedrooms available or one bedrooms on the first floor? And The AI responds, Yeah, okay, we have several one bedroom apartments available. And then the user wants to confirm any of those on the first floor. And how much are the one bedrooms? And then also. Is is a current resident. So it's all they're also asking I need a maintenance request. This is actually pretty like you could see the messiness of the real world in here. And the assistant just calls a tool that says transfer call. But it doesn't Say anything. It just abruptly does transfer call. So it's pretty jank, I would say. Like it's just not Yeah. Another kind of jank. A different kind of jank. So you don't wanna when you write the open note, you don't wanna say jank because what we wanna do is we wanna Understand what And when we look at the notes later on, we'll understand like what happened. So we're just sort of say Um you know, did not confirm Call transfer. With Uh With user. It doesn't have to be perfect. You just have to have a general idea of what's going on. Cool. So Let's say we do and we U Treya and I We recommend doing at least a hundred of these. The question is always like how many of this do you do? And so there's not a magic number. We say one hundred is because we know Не а сунз ю старт дуюйнис. Once you do twenty of these You will automatically Find it so useful. But you will continue doing it. So we just say one hundred to uh mentally unblock you so it's not Intimidating. Like, don't worry, you're only gonna do a hundred. In there is a A term for that. Oh. So so the right answer is Keep looking uh traces until You feel like you're not learning anything new. Maybe Shreya should talk about it Yeah, so there's actually a term. in data analysis and quanti qualitative analysis called theoretical saturation. So what this means is when you do all of these processes of looking at your data. When do you stop? It's when you are theoretically saturating. Or you're not uncovering any new types of notes. New types of concepts. Or nothing that will like materially change the next part of your process. Um, and this kind of takes a little bit of intuition to develop. So typically people don't really know when they've reached theoretical saturation yet. That's totally fine. when you do two or three examples or rounds of this, like you will develop the intuition. A lot of people realize like, Oh, okay, like I only need to do forty. I only need to do sixty. Actually I only need to do like fifteen. I don't know, like depends on the application and develops like how s depends on how savvy you are with error analysis for sure. And your point about we probably want to You're gonna wanna do a bunch. I imagine it's because you're just like, Oh, I'm discovering all these problems. I gotta see what else is going on here. Exactly. And uh promise at some point you're like not gonna discover new types of problems. Yeah. Awesome. So let's say you did a hundred of these, what's the next step? Yeah, okay. So you did a hundred of these. Now you have all these notes. So this is where you can start using AI to help you. Um you so the part where you looked at this data Is important. Like we Discuss you don't want to automate this part too much. Humans will still have jobs. This is a takeaway here. That's great. Yes. Just reviewing traces. At least there's one job left for now. Yeah. So o yeah, exactly. Um And so Okay, you have all these notes. Now To turn this into something Useful. You can do basic counting. So basic counting is the most powerful Analytical technique in data science. Because it's so simple. And it's Kind of undervalued. Um in many cases. And so it's very approachable for people. And so the first thing you want you wanna do is take these notes. And you can categorize them with an LM. And so there's a lot of different ways to do that. Right before this podcast. I took three different Uh Coding agents are You know. uh AI tools. In how it categorized These notes. So one is okay, I uploaded into a cloud project. I uploaded a CSV of these notes. And I just exported them directly from This interface. Um, there's a lot of different ways to do this, but I'm sh I'm showing you The simple, stupid way. The most basic way. Of doing things. And so Dump the CSV in here. And I said, Please analyze the following CSV file. There's a and I told it there's a metadata field that has a note. In it? But what I said is I used the word open codes. And I said, Hey, I have different open codes. In that's a term of art. That's um LMs know what open codes are and they know what axial codes are. Because It is a T it is a concept that's been around for a really long time. So Those words help me shortcut. What I'm trying to do. That's awesome. And the end of the end of the prompt is telling it to create it. Excel codes. Yes. Creating acyl codes. So What it does is So maybe it's worth talking about what are axial codes or like what's the point here. But you have a mess of open codes. Right. And You don't have one hundred distinct problems. Actually mo many of them are repeats. But because you phrase them differently. Right, and and that put you shouldn't have tried to create your taxonomy of failures as you're open coding. You just wanna get down what's wrong and then organize, okay, what's the most common failure mode. So the purpose axial code basically is just a failure mode. It's like the label or category. And what our goal is is to get to this clusters of failure modes and figure out what is the most prevalent. So then you can go and run an attack. That problem. That is really helpful. Basically you're just synthesizing all these categories into categories. And them. Super cool. And we'll Uh include this prompt in our show notes for folks so they don't have to like sit there and Screenshot it and try to type it up themselves. Yeah, great idea. Um And so Claude. You know, went ahead and analyzed the C S V file, decided how to parse it, blah, blah. We don't need to worry about all that stuff. But it came up with a bunch of axial codes. Basically axial codes are categories. Like Shreya said. So one is okay, capability limitations misrepresentation. Process of protocol violations, human handoff issues. Communication quality. It created these categories. Now Do I like all the categories? Not really. I like some of them. It's a good first like Stab at it. I would probably rename it a little bit because some of them are a bit too generic. Like what is Capability limitation. That's a little bit too broad. It's not actionable. I wanna get like a little bit more actionable with it. So that If I do decide it's a problem, I know what to do with it. But we'll discuss that in a little bit. Um so you can do this like with anything, and this is the dumbest way to do it. But Dumb sometimes is a good way to get started. So and and this is what LMs are really good at, taking a bunch of information and synthesizing. Absolutely. Synthesizing for us to make sense of, right? Note that you know it's not tell us it's not automatically proposing fixes or anything. That's our job. But You know, now we can wade through this mess of open codes a lot easier. Another thing that's interesting here in this prompt to generate the axial codes is you can be very detailed. If you want. I want each axial code to actually The You know, some actionable failure mode. And maybe the LLM will understand that and propose it. Or I want you to group these open codes by, you know, what Stage. Of the user story. That it's in. So this is where you can, you know Be creative or do what's best for you as a product manager or engineer working on this, and that will help you do the improvement later. So there's no definitive prompt of here's the one way to do it. You're saying there's You can iterate, see what works for you. Absolutely. It's interesting the tools don't do this, or or do they try and they just don't do a great job. No, I don't think they do it. We've been screaming from the rooftops. Please, please do this. I do think it's a little bit hard, right? Like part of this whole um experience with the Evals course Hamill and I are teaching are like a lot of people don't actually know this. So maybe it's that people don't know this and they don't know how to build tools for it. Um And Hopefully we can demystify some of this Magic. And just to double click on this point, like this is not a thing everyone Does or knows this is something you two developed. Based on your experience doing data analysis and data science and At other companies. Well, I want to caveat we didn't invent error analysis. We don't actually want to invent things. That's a bad that's bad signal. If somebody is coming to you with a way to do something that's like entirely new and not grounded in hundreds of years of theory and literature. Then. You should. I don't know, be a little bit wary of that. But what we tried to do was distill okay, what are the new tools and techniques that you need to No, make sense of the LM error out analysis, and then we created a curriculum or structured way of doing this. So this is all very tailored to LLMs. But you know, the terms open coding, axial coding. are grounded in Um social science. Amazing. Okay. Like what's funny about you do guys doing this is I just want to go do this somewhere. I don't have I don't have any product to do this on. But it's just like oh this would be so fun. Just sit there and Find all the problems I'm running into and categorize them and then try to fix them. I love that. Hamble pulled up a video. What do you got going on here? Yeah. So I pulled up a video just to drive home Shreya's point. Like we are not inventing anything. So what you see on the screen here is Andrew Eng. One of the famous machine learning uh researchers In the world who have taught Lot of people, frankly, machine learning. And um you can see this is a eight year old video. So And he's talking about error analysis. And so this is a technique that's been used To analyze stochastic systems. For ages. Um and it's some it's something that if you're just using the same machine learning Ideas and principles is bringing them in into here. Because again, these are stochastic systems. Awesome. Well, one thing we're working on getting Andrew on the podcast. We're chatting, so that'll be really fun. Uh two, I love that my other my podcast episode just came out today is in your feed there and it's standing out really well in that feed, so I'm really happy about that thumbnail. Very nice. Yeah, the recommendation algorithm is. Don't screw my algorithm. Okay, cool. So we've done some synthesis. What's I know we're not gonna go through the entire step. This is like you have a whole course that takes many days to learn this whole process. Okay, so you can you can do this through anything and you know, I've used I the same thing works just fine in chat GPT, the same exact prompt. You can see it It made axial codes. I really like using Julius AI. Um it's one of my favorite tools. Julius is a is a kind of a third party tool that uses notebooks. I personally like Jupiter notebooks a lot. And so It's more of a data science thing, but A lot of product managers that are kind of learning notebooks nowadays and is kind of cool. It's like a fun playground where you can like write code and look at data. We don't have to go deeply into that. Just wanted to mention you can use a lot you know, AI is really good at this. So let's go into the fun part. Here we go. So now we have all the Ax we have these axial codes. So the first thing I like to do I Have these open codes, right? And I have the axial codes that Let's say You know the m Like that we assigned from The cloud project or the chat GPT. And so what I do Is I collect them. First and then take a look, like does these axial codes make sense? And I look at the correspondence between the different axial codes and the open codes. And I and I go through an exercise and I say, Hm, do I like these? These codes. Like can I make them better? Can I refine them? Can I make them more specific? Um you know Instead of like being generic, I make the very specific inactionable. So you see the ones that I came up with here or tour scheduling, rescheduling issues, human handoff or transfer issue. Formatting error with an output. Conversational flow. We saw the conversational flow issue with the text messages. Uh making follow up promises not kept. And And so basically what I can do what you can do now Is like you have these Axial codes. And Um so I just collect them into a list. So this is an Excel formula. just collects these codes into a list and now we have a comma separated list of these codes. And then what you can simply do is you could take your notes. That you have those open codes. And you can tell an AI And this is using Gemini and AI. Just for simplicity, this is like The you know, again, we try to keep it simple. Categorized Uh the following note into one of the following categories. That's the way this for folks watching, there's like I like all these different prompts and formulas you're sharing. This is like the Uh Google Sheets AI. AI prompt. Yeah. Yeah. Mm. And so basically what you can do is you can then have You can categorize your faces into one of the buckets. And that's what we have here. We have categorized all those problems that we encountered into one of these things. And this is automatic, which is very exciting. I mean the AI is doing it. So this also drives home the point that your open codes have to be detailed, right? You can't just say janky. Because If the AI is reading janky, it's not gonna be able to categorize it. Even a human wouldn't, right? It would have to go and Remember why, said Jenky. Mm-hmm. So it's important to be, you know, somewhat detailed. In your open code. Okay. So avoid the word janky is a good rule of thumb. Yeah, avoid the word janky out of the way. Okay. I was being funny. Yeah. Okay. What are some of those other words just uh that come that people often use that you think are not good. I don't think it's specific words. I think it's just people are not detailed enough in the open code, so it's hard to do the categorization. Great. And by the way, the reason you have to map them back is because say Claud or J P T J P T gave you suggestions and you change them and Iterate it on them. So it doesn't just go back and say, Cool, whatever in each bucket. Yeah, yeah. Great. That's a really good question actually. It's good to iterate and think about it a little bit, like Do I like these open codes, or do these actually make sense to me? Just like anything that AI does. It's really good to kind of put yourself in the middle. Yeah. Just a little bit of loop. Still space. Yes. Great. Yeah. One of the things that I like to do with this step if I'm trying to use AI to do this labeling is also have a new category called none of the above. So An AI can actually say none of the above in the axial code. And that informs me, okay, my Axial closed are not complete. Like let's go look at those open codes. Let's figure out what some new categories are, or figure out how to reword my other axial codes. Awesome. And what's cool about this is you don't need to do this many, many times. Like for most products, you do this process once and then you build on it, I imagine, and you just make it over time. Absolutely. And it gets so fast. Like people People do this like once a week and you can do all of this in like thirty minutes. And like suddenly your product is like so much better than if you were never aware of any of these problems. Yeah, it's absurd to feel like you don't You wouldn't know this is happening. Like watching this happening I'm like how could you not Have no idea. Most people. Yeah. We'll we'll talk about that. There's a whole debate around this stuff that we want to talk about. Uh okay, cool. So you have this you have the sheet. Okay, so here's the big unveil. This is the magic moment there we go. Right now. So We have All these codes we that you know we applied. The ones that we like. On our traces. Now you can do The ta-da, you can count them. So Here's a pivot table. And we just can do pivot table on those. And we can count how many times those different things occurred. So what do we find? Fan on this on these like traces that we categorize. We found seventeen conversational flow issues. And I really like pivot tables'cause you can do cool things. You can like double click on these, you can say, Oh, okay, let me let me take a look at those. But That's going into an aside about Pivot tables, how cool they are. Um You know. W now we have Just a nice rough cut. Oh What are our problems? And now we have gone from chaos. To some kind of Thinking around oh, you know what? These are my biggest problems. I need to fix conversational issues. You know, maybe these human handoff issues. Not necessarily the count. Is the most important thing. You know, that might be something that's just really bad and you want to fix that, but Okay, now you have Some way of looking at your problem and now you can think about whether you need evals. Uh for for some of these. So You know, with the Yeah, there might be some of these things that might be Just dumb engineering errors. that you don't need to write an eval for because it's very obvious on how to fix them. Um maybe the formatting error with output. Maybe you just forgot to tell the LM how you want it to be formatted. In like you didn't even say that in the prompt, so like just go ahead and fix the prompt. Maybe. You know? And we can decide like, okay, do you want in uh do you want to write an email for that? You might be you might still want to write an evolve for that because you might be able to test that We just code. You could just test the string, does it have the right formatting, potentially. Uh without running an L L M. So there's a cost benefit trade off? Two evals. You don't want to Get carried away with it? Um But you want to start you want to usually ground yourself in your actual errors. You don't want to skip this step. And so the reason I'm kind of spending so much time on this is like This is where people get lost. They go straight into evals like let me Let me just write some tests. And that is where things go off the rails. Um So let's let's Okay, so let's say we want to tackle one of these things. Yeah. So for example Uh let's say we want to Tackle this human handoff issue. And we're like hmm. I'm not really sure how to fix this, like That's a kind of subjective sort of judgment call. On you know, should we be handing off to human And I don't know immediately how to fix it. It's not super obvious. per se, yeah, I can like change my prompt, but I'm not like sure. I'm not a hundred percent sure. Well, that might be sort of an interesting Um For example. So there's different kinds of evals. One is code based. Which you should Try to do If you can. Because they're cheaper. You don't have to you know, LM as a judge is something it's like a meta eval. You have to evo that eval to make sure the LM The judging is doing the right thing. Which we'll talk about in a second. So Okay. Elm as a judge, that's one thing. Okay, how do you build an element as a judge? Before we get into that, actually, just to make sure people know exactly what you're describing there. these two types of evals. One is you said it's code based and one is uh L M as judge. Maybe Shreya, just help us understand what the what code base eval even is. It's just like it's like essentially a unit test. Is that a simple way to think about it? Maybe eval is not the right term here, but think like automated evaluator. So when we find these failure modes, one of the things we want is like, okay, can we now like go check the prevalence of that failure mode in an automated way without me manually labeling and doing all the coding and the grouping and I wanna run it on thousands and thousands of traces. I wanna run it every week. That is okay. You should probably build an an automated evaluator to check for that failure mode. Now when we're saying code based versus L M based, we're saying Okay, so maybe I could write like a Python function or a piece of code. to check whether that failure mode is present in a trace or not. And that's possible to do for certain things like Yeah, checking. The output is JSON. Um, or you know, checking that it's marked down or checking that it's short. Like these are all things you can capture and code or you could approximately capture and quote. Uh when we're talking about L L M Judge here, we're saying This is a complex failure mode. We don't know how to Evaluate in an automated way. So maybe we will try to use an LM. to evaluate this very, very narrow, specific failure mode of Hand offs. So just uh Try to mirror back how what you're describing. You wanna test what your say agent or AI product is doing. You ask it a question, it gets back with something. One way to test if it's giving you the right answer is if it's consistently doing the same thing that you could write a code to t to tell you this is true or false. For example Will it ever say there's a virtual tour? You could ask it. Is do you provide virtual tours? It says yes or no, and then you could write code to Tell you if it's correct. Based on that specific answer. But if you're asking about something more complicated and it's not binary, you almost need Like in a In a one world you need a human to tell you this is correct. the solution to avoid humans having to review all this every time automatically is L's replacing human judgment. And you'd call it a L M as judge. The L M as being the judge if this is correct or not. Absolutely, you nailed it. Um So people always think like oh This is at least as hard of my problem of creating the original agent. And it's not. Because you're asking the judge to do one thing. Evaluate. One failure mode. So the scope of the problem is very small. And The output. of this LLM judge is like pass or fail. So it is a very, very tightly scoped thing that L M judges are very capable of doing very reliably. And the goal here is just to have a suite of tests. Now, Ryan, before you ship to production That tell you things are going. The way you want them to. The way your agent is interacting with you. The beautiful thing about LLM judge is you can use them in unit tests or CI, sure. But you could also use it online. For monitoring. Right. Like I can thousand traces every day run by LLM judge. Real production traces. And see what the failure rate is there. This is not a unit test, right? But still now we get like a extremely specific measure of application quality. Cool. That's a really great point,'cause a lot of people dis evaluate this like not real life thing. It's a thing that you test before it's actually in the real world and What's actually happening in the real world. You're saying you could actually you should actually do exactly that. Yeah. That's your real thing running in production. And it's like a daily hourly sort of thing you could be running. Totally. Awesome. Okay. Uh Hamil's got a a example of an actual LM as Judge Eval here, so let's take a look. I love how Shreya really teed it up. Um For me. So thank you so much. So what we have is a Ellen as a judge prompt for this one specific failure, like Shreya said. You would want to do one specific failure. And you want to make it binary. Because we want to simplify things. We don't want, hey, like score this on a rating of one to five, like how good is it? That's just mostly In in most cases that's a weasel way of like not making a decision. Like no, you need to make a decision. Is this good enough or not? Yes or no? Can be painful to think about what that is. But you should absolutely do it. Otherwise this thing becomes very untractable. And then when you report these metrics, no one knows what three point two versus Three point seven meats. This is yeah, we see this all the time also, and even with like expert curated content on the internet. Where it's like Oh, here's your LLM judge evaluator prompt. You're the one to seven scale. And I always think I s always text Hamill, like, Oh no, like now we have to fight the misinformation again because we know somebody is going to try it out and then come back to us and say, Oh, I have four point two Average and we're gonna be like Okay. Oh It's wild how much drama there is in Eval's space. We're gonna get to that. Oh man. This episode is brought to you by Mercury. I've been banking with Mercury for years. And honestly, I can't imagine banking any other way at this point. I switched from chase, and holy moly, what a difference. Sending wires, tracking spend. Giving people on my team access to move money around so freaking easy. Where most traditional banking websites and apps are clunky and hard to use, Mercury is meticulously designed to be an intuitive and simple experience. And Mercury brings all the ways that you use money into a single product, including credit cards, invoicing, bill pay, reimbursements for your teammates, and capital. Whether you're a funded tech startup looking for ways to pay contractors and earn yield on your idle cash, Or an agency that needs to invoice customers and keep them current. Or an e-commerce brand that needs to stay on top of cash flow and excess capital, Mercury can be tailored to help your business perform at its highest level. See what over 200,000 entrepreneurs Mercury. Visit Mercury.com to apply online in ten minutes. Mercury is a fintech, not a bank, banking services provided through Mercury's FDIC Insured Partner Banks. For more details, check out the show notes. Okay, so This is your judge prompt. There's no one way to do it. It's okay to use an LM to help you create it, but again, put yourself in a loop. Don't just blindly accept what the LM does. And in all of these cases, that's what we did. Like with the axial codes, we Kind of iterated on this. You can use an LM to like help you create this prompt, but make sure you read it. Make sure you Edit it, whatever. This is not necessarily the perfect prompt. This is just the stupid like very keeping it very simple just to show you the idea. is like okay for this handoff failure Um, you know, I said, Okay, I want you to output true or false of is binary. It's a binary judge. That's what we recommend. And then we then I just go through and say, Okay, like when should you be doing a handoff. Yeah, I just list them out. Like okay, ex explicit human request ignored or looped. Mm. Uh some policy Mandated transfer Sensitive resonant issues, tool data unavailability. same day walk in or tour requests. You know, you need to talk to a human for that. So on and so forth, right? And so the idea is like now that I know that this is a failure from my data, I'm interested in iterating on it because I know this is actually happening. All the time? And like Shreya said, like it would be nice to have a way not only to Evaluate this on Ex like the data I have, but also on production data. Just to get a sense of like well what scale is happening, let me find more traces, let me have a w you know, a way to iterate on this. And so we can take this prompt. I'm gonna use the spreadsheet again. So The first step is Okay, uh when I'm doing this judge, I wrote the prompt. Now, a lot of people stop there and they say, Okay, I have my judge prompt or done. Good. Like let's just let's just ship it. And that's uh The prompt says if the judge says it's wrong, it's wrong. They just like accept it as the gospel be like, Okay. No L Msa is wrong. It's It must be wrong. Don't do that. Because that's the fastest way that you can have evals that don't match What's going on? And when people l lose trust in your evals, they'll lose trust in you. So It's really important that you don't do that. And so one before you release your LM as a judge, you want to make sure it's aligned to the human. So how do you do that? Is you actually you have Those axial codes. You wanna like measure your judge against the axial code. And say like hey, d does it agree with me? Does my own judge does it agree with me? Just measure it. And so what we have here is okay, I say assess this L L M trace. Again I'm using Just spreadsheets here. Assess this LM trace. According to these rules. In the in the rules are just the prompt that I just showed you. And I Ask it, okay. Is there a Hand off error. True or false? So then this column Me just zoom in a bit. Column H I have Okay. Is did this error occur? Column G is whether I thought the error occurred or not. You can see you going through it manually, you do that. Yeah, yeah. And which we already did. We we already went through it manually. So we don't it's not like we have to do it again, because we kind of have that. Gee. code from the axial coding, we already did it. Um you might have to go through it again if you need more data. And there's a lot of details to this on like how to do this correctly. Um you want to split your data and do all these things so that you're not cheating. But I just want to show you the concept. And basically um what you can do is measure the agreement. No. Mm-hmm. One thing you should know as a product manager Is A lot of people go straight to this like agreement. They say okay, my judge agrees with The human It's some percentage of the time. Now that Sounds appealing. But it's a very dangerous metric to use because A lot of times errors Have Um you know. They only happen on the on the long tail. And they don't happen as frequently. So like if you only have the error Ten percent of the time. Then you can easily have ninety percent agreement. I just Having A judge. Say Uh it passes all the time. Does that make sense? So like Ninety percent agreement might look good on paper, but it might be misleading. And that's rare. It's a rare. Yeah. So You know, as a product manager or someone Even if you're not doing this calculation yourself. If someone ever reports to you Agreement. You should immediately ask, Okay, tell me more. Like you need you know you know need to look into it. They give you more intuition. Here is like a matrix. Okay. Of this specific judge. In the Google sheet. And this is again a pivot table. Just keeping it dumb and simple. Is okay. On on the uh rows I have What did the human think? What did I think? Did it. Have an error, true or false. And then Did my judge have an error? True or false. The intuition here is exactly what Hamil said, right? You need to look at each type of error. So when the human said false but the judge said true or vice versa. So those non green Diagonals here. And if they're too large Then go iterate on your prompt, make it more clear to the LLM judge so that you can reduce that misalignment. You wanna get to a point where most you're gonna have some misalignment. That's okay. We talk about in our course also how to code correct that misalignment. But in this stage If you're a product manager And the person who's building the L M Judge Evil Has not Done this. They're saying like oh it agrees Seventy five percent of the time we're good. They don't like have this matrix and they haven't iterated to make sure that These two types of errors. Have gone down to zero, then it's a bad smell. Go and ask them to go fix that. Awesome. That's a really good tip is Is what to look for when someone's doing this wrong. Yeah. Actually, can you take us back to the L M As judge prompt. I just wanna highlight something really interesting here. I've had some guests on the podcast recently who've been saying Evals are the new PRDs. And if you look at this, this is exactly what this is. Like product managers, product teams are right. Here's what the product should be, here's all the requirements, here's like the how it should work. They built the thing and then they test it manually often. What's cool about this is this is exactly that same thing. And it's running constantly. It's telling you Here's how this agent should respond in very specific ways. If it's this, this this is do that. If it's this, do that. And so it's exactly what I've been hearing again and again. You could see it right here. This is Like the purest sense of what a product requirements document should be is this Eval judge that's telling you exactly what it should be, and it's automatic and running constantly. Yeah, absolutely. And it's kind of derived from our own data. So of course it's a product manager's expectations. What I find that a lot of people miss is they just put in what their expectations are before looking at their data. But as we look at our data, we uncover more expectations. that we couldn't have dreamed up in the first place, and that ends up going into this prompt. So that is interesting. So it's not so your advice is not Skip straight to Evals and LLM as judge prompts before you build the product. still write traditional one pagers PRDs to tell your team what we're doing, why we're doing it, what success looks like. But then at the end you could probably pull from that and even improve that original PRD if you're evolving the product. Uh using this process. I would go even further to say you're going to improve It's going to change. You're never gonna know What the failure modes are going to be up front. And you're always going to uncover New No. Vibes that you think that your product should have, where you don't really know what you want until you see it with these LLMs. So you've got You gotta be kind of flexible, have to look at your data, have to PRDs are a great abstraction for thinking about This That's not the end all be all. It's going to change. I love that. And Hamill's pulling up some cool research report. What's this about? Oh, this is one of the coolest research reports. You can possibly read if you wanna know about evals. So it was authored by someone named Shreya Shankar. Oh my God. And Her collaborators. And so it's called Who Validates the Validation. That's the best name for a research. So I sh I should let Shreya talk about this. I think the One of the most important things to pay attention to in this paper are the criteria drift. Yeah. So we did the super fun study when we were kind of doing user studies. with people who were trying to write LM judges. Or just validate their own Lol M outputs. And we were this was I think this was before Evals was like extremely popular, I feel like, on the internet. This was we did this project like late twenty twenty three, was when we started it. But then I the thing that really was burning in my mind as a researcher is like, why is this problem so hard? We've been having machine learning and AI. For so long. It's not new. But suddenly this time around everything is really difficult. So we just did this user study with a bunch of developers and we realized Okay. What's new here is that you can't Figure out your rubrics up front. People's opinions They think of failure modes only after seeing ten outputs. they would never have dreamed of in the first place. And these are experts, right? These are people who have built many LLM pipelines and now agents. Before. And just You can't Ever dream up everything in the first place. Um and I think that's so key. in today's world of AI development. Okay, that is a really good point. That's very much reinforcing what we were just talking about and that's why Ham will pull this up. Yeah, okay, great. You still gotta do product the same way, but now you have this really powerful tool that Make helps you make sure what you've built is Correct. uh it's not gonna replace the PR D process. Cool. How many Evals of these, how many say, I don't know, L Animus judge prompts do you end up with usually say I don't know. Like I know obviously depends the complexity of the product, but what's like a Number in your experience. For me, like between four and seven. Oh, that's it. It's not that many'cause a lot of the failure modes, as Hamill said earlier, can be Fixed by just fixing your prompt. You just didn't think to put it in your prompts and now you put it in your You shouldn't do an eval like this for everything. Just the the pesky ones that Um the oh you've described your ideal behavior in your agent prompt, but it's still failing. Got it. So say you found a problem, you fixed it. in traditional software development, you'd write a unit test to make sure it doesn't happen again. Is your insight here is don't be even bother writing an eval around that if it's just gone. I think you can if you want to, but the whole game here is about prioritizing. You have finite resources and finite time. You can't write an eval for everything, so prioritize the ones that are the more pesky areas. And probably the ones that are most risky to your business if they say something like Mecha Hitler Scrock and Cool. Okay. So that's that's very uh relieving that this'cause this is this was prompt us like a lot of work to really think through all these details. But it's a lot of one time cost. But now forever you can run this. On your application. Right. And I wanna say okay. Data analysis? is super powerful. Is going to drive lots of improvements very quickly to your application. We showed the most basic kind of data analysis, which is counting. Which is accessible to everyone. You can get You know. More Sophisticated with the data analysis. There's lots of different ways to sample Look at data. We kind of made it look easy. In a sense. But There's a lot of skills here. To do to it well. Um You know. building an intuition and a nose for how to sort through this data. For example Let's say I find conversational issues, this like conversational flow issues. Maybe If I was Trying to chase down this problem Further I would think about ways to find other conversational flows flow issues that I didn't Code. You know, I would maybe dig through the data in several ways. Um And there's you know, different ways to go about this. It kind of Yeah. It's a very Similar, if not almost exactly similar as kind of traditional analytics techniques that you would do on any product. Give us just a quick sense of what comes next. And then let's talk about the debate around evals and a couple more things. So what comes next after you've Built your L M judge. Well, we find that people just try to use that everywhere they can. So they will put the L M judge in unit tests. As you and they will know like Oh, here are some example traces where we saw that failure because we labeled it. Now we're gonna make those part of unit tests and make sure that every time we push a change. To our code. These tests are gonna pass. They also use it for online monitoring. People are making dashboards on this. And I think that's incredible. I think like the products that are doing this, right, they have a very sharp sense of how well their application is performing. Um And people don't talk about it because this is their moat. Right, so people are not gonna go and share all of these things because Makes sense, right? If you are an email writing assistant and you're doing this and you're doing it well. You don't want somebody else to go and build an email writing assistant and then kind of get you out of business. So I really want to stress the point that it's like try to use these artifacts that you're building wherever possible, online, repeatedly. Um use them to drive improvements to your product. Oftentimes Hamill and I will Kind of We'll tell people how to do this up to this very point. And it clicks for people and then they like never come back again. So either they have I don't know, quit their jobs, they're not doing AI development anymore. Or They know what to do from here on out. Um I think it's the latter. But um I think it's very powerful. Like just watching you do this. Really open my eyes to what this is and how systematic the process is. Yes. I always imagine you just sit on a computer, okay, what are the things I need to make sure Work correctly. And what you're showing us here is here's it's a very simple step by step based on real things that are happening in your product, how to Catch them, identify them, prioritize them, and then Catch them if they happen again and fix them. Yeah, it's it's not magic. Like anyone can do this. You're gonna have to practice the skill, like any new skill you have to practice. But you can do it. Um and I think what's very empowering now is that product managers Are doing this and can do this and can really build very, very profitable products with this skill set. Okay. Great segue to a debate that we kinda got pulled into that was happening on on X the other day. Uh I did not realize how how much controversy and drama there is around evals. There's a lot of people with very strong opinions. Uh so about Shreya, give us just a sense of the two sides of the debate around The importance and value of e bills. And then give us your perspective. Yeah. So all right, I'll be a little bit placading and I say I think everyone is on the same side. I think the misconception Is that people have very rigid definitions of what Evals Yes. For example, they might think that Evals is just unit tests. Or they might think that Evals is just the data analysis part. And no online monitoring or any no monitoring of product specific metrics, like Actually number of chats. Engaged in or whatnot. Um, so I think everyone has a different mindset of evals going in. And the other thing I will say is that people have been burned by evals in the past. So I think people have done evals badly. One concrete example of this is they've tried to do an LLM judge, but it has not aligned with their expectations. They only uncovered this later on and then they didn't trust it anymore. And then they're like, Oh, I'm anti Evals. And I hundred percent empathize with that because Okay. You should be anti Likert scale LLM judge, I absolutely agree with you. We are anti that as well. So a lot of the misconception stems from two things, right? Like People having a narrow definition of evals and then people Not doing it well and then getting burned and then Wanting to avoid other people making that mistake. And then unfortunately X or Twitter is like a medium where you know People are misinterpreting what everybody is saying all the time. And you just get all these strong opinions of like don't do evals. It's bad. We tried it. It doesn't work. We're Clawed code or You know Whatever other than famous product and we don't do emails. And there's just so much nuance behind all of it because 'Cause a lot of these applications are standing on the shoulders of evals. Coding agents is a great example of that. Claude Code. Right they are standing on the shoulders of Claude. Baseball not baseball the The fine tune Claude models. have been evaluated on many coding benchmarks. Yeah. Can't argue against that. And just to double d just to make clear exactly what you're talking about there. one of the heads uh I think maybe the head engineer of Cloud Code went on a podcast and he's like Oh, we don't do evals, we just vibe. We just look at vibes and vibes meaning they just use it and feel if it's right or wrong. And I think that kind of works. So there's two things to that. Right. One is they're standing on the shoulders of the evals that their colleagues are doing for coding. Of the Cloud Foundational model. Absolutely, right. We know that they report those numbers because we see the benchmarks. We know who's doing well on those. The other thing is they are actually probably very systematic. about the error analysis to some extent. I bet you. that they're monitoring. Who is using Claude? How many people are using Claude, how many chats are being created, how long these chats are. They're also probably monitoring in their internal team. They're dog fooding anytime something is off, they maybe have a cue. Or they send it to the person developing Claude Code, and this person is implicitly doing some form of hair error analysis that Hamill talked about. All of this is evils. Right. There's no world in which they are just Being like I made Claude Code, I'm never looking at anything. Um And unfortunately, right, when you Don't Think about that or talk about that. I think that the community most of the community is beginners, right? Or people who don't know about evals and want to learn about it. Um and it sends the wrong message there. Now I don't know what Claude Code is doing, obviously. Um But I would be willing to bet money that they're doing something. In the form of emails. We'll also say that Coding agents are fundamentally very different. than other AI products because the Developer is the domain expert. So you can short circuit a lot of things. The and also the developer is using it all day long. So there's a type of dog fooding and type of dom domain expertise. That is You know can collapse the activities. You don't need As much data. You don't need as much feedback or exploration because you know. So Your evil process You know, should look Different. But because you're seeing the code. Like you see the code it's generating, you can tell this is great. This is terrible. Yeah. Yeah, and so and so I think a lot of people had generalized Coding agents because coding agents are the first AI product. Released. Into the wild. And I think it's a mistake to try to generalize that. At large. The other thing is yeah Engineers have a dog fooding personality. But there are plenty of applications where people are trying to build AI in certain domains and And they don't have dog fooding for like doctors, for example, are not out there trying to get all the most incorrect advice from AI. And be tolerant and receptive to that. So it's Very important to keep, I think, these nuanced things in mind. What I'm hearing from you Shreya is interestingly is that If You if humans on the team are doing very close uh data analysis, error analysis. Dog fooding it like crazy. And essentially they're The human evals. And you're describing that as that's within the umbrella of evals. So you could do it that way if you're very If you have time and motivation to do that, or you could set these things up to be automatic. Absolutely. Uh it's also about the skills, right? People who work at anthropic Are very, very highly skilled. Um They've been trained. the data analysis or software engineering or AI and whatnot, right? Yeah. You know, you can get there. Anyone can get there, of course, by like learning the concepts, but Most people don't have that skill right now. Dog fooding is w is a dangerous one. Only because A lot of people will say they're dog fooding. They're like, Yeah, we dog fooded. But Are they really? And A lot of people aren't really dog fooding it. At that visceral level. that you would need to to have to close that feedback loop. So that's the only caveat I would add. There's also this kind of feels like straw man argument of evals versus A B tests. Talk about your thoughts there, because that feels like a big part of this debate people are having. Like do you need evals if you have A B test that are testing production level metrics. So A B tests are again another form of evals, I imagine, right? Like when you're doing an A B test, right, you have two different experimental conditions and then you have a metric that quantifies the success. of something and you're comparing the metric. And again, right, an eval in our mind is systematic measurement of quality. Some metric. Um you can't really do an A V test without The eval to compare. Um So maybe maybe we just have a different Weird take on it. Yeah. Okay. So what I'm hearing is like you consider A B test as part of the suite of evals that you do. I think when people think A B test it's like we're changing something in the product. We're gonna see if this improves some metric we care about. Is that Is that enough? Why do we need to test every little feature? Like if it's Impacting a metric we care about as a business. We have a bunch of A V tests that are just constantly running. This is now a great point. Um so I think a lot of people prematurely Do A B tests. because you know, they've never done any in error error analysis in the first place. They just have hypothetically come up with their product requirements and they like believe that You don't we should test these things. Um, but it turns out, right, when you get into the data, as Hamill showed, the like the errors that you're seeing are like Not what you thought what the errors might be. They were these like weird handoff issues or like I don't know, like the text message thing was strange. Um, so I would say that like if you're going to do A B tests and they're powered by actual error analysis, as we've shown today. Then that's great. Go do it. Um But if you're just going to Do them, which we find that people try to do. Just try to do them based on like what you hypothetically think is Why is important. then I would encourage people to go and like rethink that and kind of ground your hypotheses. Do you have thoughts on what Stat Sigs gonna do at OpenAI? Is there anything there that's interesting? Just like that was a big deal, a huge acquisition. A B test company, people are like, Oh, B test of the future. Uh Thoughts. You know, just w to add to the previous question. A little bit. Is Why is there this debate A B testing versus Evals? I think fundamentally Evel's Is People are trying to wrap their head around What how to improve their applications. Fundamentally. Um You need to do data science. You need data science is useful in products. Like Looking at data, doing data analytics, there's many s different suite of tools. And um you don't need to invent anything new. Sure, you don't need like necessarily the whole breadth of data science and it looks slightly different. Just slightly. With LMs. Um You know, you might your tactics might be different. And so Really what it is is like Using Analytic tools. Uh to understand your product. Now people say the word eval is trying to kind of like carve out this new thing. And saying no evals and then A B testing, but if you zoom out. It's it's the same data science as before. And I think that's what's causing the confusion is hey We need data science thinking. In AI products is you know, it's helpful to have that thinking in AI products like it is in any product. Uh is my take. On that. So yeah. That's a really good take. Like I think just the word evals triggers people now. Yeah and if you just call it we're just doing air analysis using doing data science to understand where our problem break our product breaks and just setting up tests to make sure we know. It sounds boring. No no no we need a mysterious term like evals to to really Get the momentum going. Your question about static, um, I think it's very exciting. To be honest, I don't know much about it because, you know, I just imagine that they're this company that Many there's a tool that many people use and maybe it just so happened that Open AI acquired them. I'm sure they've been using them in the past. Um I'm sure open AI's competitors. Are you using stat Sig as well? So maybe there is something strategic in that acquisition. I have no idea. I don't know any think there. But I think those are really the bigger questions for me than you know, is this fundamentally changing A V testing or making evals more of a priority? I I think they've always been a priority. I think open AI has always been doing some form of them. And open AI has gone so far. Oh, yeah. historically speaking, as to like Go and look at all the Twitter sentiment. And try to do Some sort of retrospective on that. And then tie that back to their products. Like they're sh certainly they're doing some amount of evals before they ship their new foundation models, but they're going so much beyond and being like, Okay, let's find all the tweets that are complaining about it, all the Reddit threats that are complaining about it, that go try to like figure out what's going on. So it goes to show that like evals are very, very important. No one has really figured it out yet. People are using all the available sources signal that they can to improve their products. What I will say is I'm really hopeful that It might shift. The Or create a focus. Within open AI? Hopefully. Up until now, a lot of the big labs Understandably the focused on general benchmarks. Mike. M MLU score, human eval, things like that, which are very important for foundation models. And you know, those Not very related to product specific evals like the ones we talked about today, but like handoff and stuff like that. Like those You know, they tend not to correlate. Yeah, I think it's not a lot of it. Sorry to say. Hm. Exactly. And so Um You know, If you look at the eval products, let's say the one Up until recently that some of the big labs have, they don't have error analysis. They have genera a suite of generic tools, cosine similarity a hallucination score, whatever. And that doesn't work. It's a good first stab at it. It's okay. You know, at least you're doing something getting people maybe it's like getting people to look at data. But Oh Eventually what we hope to see is okay. Some a bit more data science thinking in s this like eval process. Which hopefully the tools will get to. Hamel and I should not be the only two people on the planet that are promoting like a structured way of thinking about application specific evals. It's like mind boggling to me. Why are we the only two people doing this? The whole world. We're not the only people and that more people catch on. Well, the fact that your course on Maven is the number one highest scoresting course on Maven, clearly there's demand and interest. And there's more people, I think, on your side. Interestingly, uh just a s an example you've been sharing on Twitter that was I think is informative. Everyone's been saying how Claude Code doesn't care about it. Evals they're all about vibes and everyone's like with and they're the best coding agent out there. So clearly this is right. More recently there's all this talk about codex, open AI codex. being better and everyone switching and they're so pro E balls. I would Yeah. So Gets me every time. The internet's so inconsistent. My favorite thing was um Like yesterday, I believe, like a couple of labmates and I Where uh getting like dessert or something. And somebody said like oh um Do you like codex or clod better or whatever? And the other person said Oh I like Claude. And then someone else said but the new version of codex is better. And then the first person said, Oh, but the last I checked was two days ago, so maybe my The thoughts. Maybe I'm not up to date. And I was like, Oh my God. So true. This is the world we live in. No, my dad. Uh huh. I wanna ask about just top misconceptions people have with evals and top tips and tricks for being successful. So maybe just share one or two each of each. So let me just start with misconceptions. And maybe I'll go to the Hamill first. Just what are a couple of the most common misconceptions people have with Eval still? The top one is hay. I can just buy a tool. Plug it in. And it'll do the eval for you. Why do I have to Worry about this. We live in the age of AI. Can't the AI just eval it? That's the most common misconception. And People want that so much that people do sell it. But It doesn't work. That's the first one. Shoot. We need humans still, great. I think that's great news. The second one. That You know, I see a lot is Hey, um Just not looking at the data. You know, so In my consulting People come to me with problems all the time. And the first thing I'll say is Let's go look at your traces. And you can see The kind of their eyes. pop open and be like, What do you mean? Like, yeah, let's look at it right now. And they're surprised that I am going I'm gonna go look at individual traces. Um And we always it always one hundred percent of the time Learn a lot in Figure out what the problem is. And so Um I think people Just don't know how powerful looking at the data is like we showed on this podcast. I would agree with that. Those are the top two, okay. Is there anything else or those are those are the ones like solve those problems? Oh those are definitely And then I guess the th one I would add is there's no one correct way to do evals. There are many incorrect ways of doing evals. But There are also many correct ways of doing it. And you gotta think about where you are at. With your product. How many r how much resources you have? Um And figure out the plan that works best for you. It'll always involve some form of error analysis, as we showed today, but how you operationalize. those metrics is going to change based on where you're at. Amazing. Okay. What are a couple of just tips and tricks you wanna leave people with? as they start on their evil journey or just try to get better at something they're already doing. So tip number one. Is Just Don't be alarmed. Or don't You know? Be scared. Of looking at your data. The process we try to make it as structured as possible. There are inevitably, you know, questions that are going to come up. That's totally fine. you might feel like you're not doing it. Perfectly, that's also fine. The goal is not to do evals perfectly, it's to actionably improve your product. And we guarantee you, no matter what you do. Yeah. doing parts of these process, you're going to find ways of actionable improvement. And then you're going to iterate on your own process from there. The other tip that I would say is we're very pro AI. Use L's to help you organization Any thoughts that you have. throughout this entire process. So this could be everything ranging from like initial product requirements. Right. figure out how to organize them. uh for yourself figure out how to improve on that product requirement stock based on the open codes that you've created. Right. Like don't be afraid to use AI in ways that You know, present information better. for you. Sweet. So don't be scared. Uh use LMs as much as you can throughout the process. But not to replace yourself. Right. Okay, great. Still jobs. Great. How. Yeah, let me actually share my screen so when I show something. So to piggyback off what Shreya said. Is If you heard any phrase in this podcast You've probably heard look at your data more than anything else. And so it's so important. That we teach that you should create your own tools. to make it as easy as possible. So I showed you some tools when we're going through the live example of like how to annotate data. Most of the people I work with They Realize how important this is. And they vibe code their own tools. Are they We shouldn't say vibe code. We there's just They make their own tools. And it's it's cheaper than ever before. Because you have AI that can help you and AI is really good at creating simple web applications. That can show you data. They have You know. I can write to a database. This is very simple. And so for the nurture boss use case? We wanted to remove all the friction. Of Looking at data. And so what you see here is just some screenshots of Uh what The application that they created. Looks like it's just Okay, they have the different channels, voice, email, text. Um they have the different threads. They Hid the system prompt by default. Little quality of life improvements. And then they all actually have this axial coating part. Here we can see. Okay, in red the count of different errors. They automated that part. In a nice way. And they just they created this within a few hours. Um And so It's really hard to have a one size fits all thing for looking at your data. You don't have to go here immediately. But Something to think about. is make it as easy as possible because again, it's the most powerful activity that you can Engage in. It's the highest ROI activity you can engage in. And so um You know, with AI Yeah, just remove all the friction. That's amazing. And again, I think the ROI piece is so important. We haven't even touched on this enough. The goal here is to make your product better. Which will make your business more successful. Like this isn't just a little exercise to catch bugs and things like that, like This is the way to make AI products. better because the experience is how users interact with your AI. Absolutely. If any you know, we teach our students, hey, when you're doing these evals, if you see something that's wrong. Just go fix it. Like the whole point is not to have Evals a beautiful Eval suite. Where you can point a Edited it and say, Oh look at my evals. No. Just fix your application, make it better. Do you know, if it's obvious, do it. So totally agree with you. Amazing. How long a question I asked, but this is I think something people are thinking about. How long Do you spend on this? Like how long does it usually take to do the first time? I can answer for myself for applications that I work with. Usually I'll spend three to four days really working with whoever to do initial rounds of error analysis. Like lot of labeling, feel like we're in a good place to create the spread sheet that Hamill had and everyone's Kind of. On board and convinced. And even like a few LM judge evaluators. Uh, but this is a one time cost. once I figured out how to integrate that in unit tests or I have like a script that automatically runs it on samples and I will create a cron job to just do this every week. I would say it's like I don't know. I find myself probably spending more time looking at data because I'm just data hungry like that. I'm so curious. I'm like I've gained so much from this process and it's like put me above and beyond in any of my, you know, collaborations with folks. So I wanna keep doing it, but I don't have to. I would say like maybe Thirty minutes a week. After that. So it's a week, essentially. A week essentially up front and then like thirty minutes to keep improving on uh adding to your suite. Yeah, it's really not that much time. I think people just get overwhelmed by how much time they spend up front and then thinking But they have to keep doing this all the time. Amazing. Is there anything else that you wanted to share or leave listeners with, anything else you wanted to kinda double down on as a point? Before we get to our very exciting lightning round. So I would say uh this process is a lot of fun. Actually. So It can it's like okay, you're looking at data, oh it sound like you're annotating things. Okay. Actually like so I was just looking at a client's data yesterday. The same exact process. It's a application that sends emails. uh recruiting emails. To try to get candidates to apply for a job. And We decided to start looking at traces. We jumped right into it. Like, hey, let's look at your traces. The Uh, we looked at a trace. The first thing I saw was This like email that Is worded like Given your background. Blah blah blah blah blah. So ask the person right away. And this is where putting your product hat on. And just D Critical And this is where the fun part is. I said, You know what? I hate this email. Like, do you like the email? Given your background when And when I receive a message given your background comma, I just delete that. So I'm like what is this Given your background with machine learning and blah blah. I mean, this is a generic thing. Like So I asked the person like hey You know, can we do better than this? Like this is kind of like a This is like a Sounds like Generic recruiting. And they're like oh Yeah, maybe Yeah, like It's the AI the'cause they were like they were proud of it. They're like the AI is doing the right thing, is sending this email with the right information with the right link. With the right name, everything. And so that's where the fun part is. Is like put your product hat on. And get into like is this really good? Something I wanna make sure we cover before we get to our very exciting lighting round. is this is just scratching the surface of all the things you need to know to do this well. Uh I think this is the best primer I've ever seen on how to do this well. Nice. But You I think we did it. But you guys teach a course that goes much, much deeper for people that really want to get good at this and take this really seriously. Share what else? you teach in the course that we didn't cover and what else you get as a student being part of the the course you teach at Maven. Yeah, I can talk about the syllabus a little bit and then Hamil can talk about all the perks. Um Go through a life cycle of error analysis than automated evaluators. Then how to improve your application, like how do you create that flywheel for yourself? We also have a few special topics that we find like pretty much no one has ever heard of or taught before, which is exciting. One is how do you build your own interfaces. So we kind of Go through actual interfaces that we've built and we also live code them on the spot for new data. Um And we show kind of how we use claude code. cursor whatever we're feeling in the moment that day. To build these interfaces. Um and we also talk about kind of broadly cost optimization as well. So we've I a couple of people that I've worked with Um, they get to a point where their evals are very good, their product is very good, but it's all very expensive because they're using like state of the art models. So how can we kind of replace certain uses of the most expensive GPT five models. With Yeah, five nano For many whatnot and save a lot of money, but still maintain the same quality. So we also give some tips for that. Um Hammel, you want we also have many perks. Yeah, talk about the perks. Okay. The perks. So My favorite perk is there's a hundred and sixty page book. That's meticulously written. That we've created. Yeah. Walks through the entire process in detail of how to do evals. That supplement the course. So you don't have to Sit there and take a All these notes. We've done all the hard work for you. And we have Like documenting it in detail. Um You know, and organized things. So that Is really useful. Another really interesting thing and something that I got the idea from you, Lenny. Is Okay, this is an AI course. Education shouldn't be this thing where you your only Watching lectures. And doing homework assignments. So students should have access to an AI that also helps them. So what we have done is We've Uh You know, just like there's the Lenny bot that you have. Yeah, Lennybot.com. Uh we have Made the same thing. With the same software that you're using. And we have put everything we've ever said about evals into that. So every single lesson, every office hours, every Discord chat. Any blogs, papers, anything that we've ever said publicly In within our course. We've put it in there. And we've tested it. With Um A bunch of students. And they've said it's helpful. Um so we're giving all students ten months Free unlimited access to that. Alongside the course. Amazing. And then you'll charge for that uh later down the road. I just think one month at a time, I don't know. Eight months and then we'll have to figure it out. Uh I was thinking we should have this whole interview should have just been our bots talking to each other. That's amazing. I would I would watch that. Only for like ten minutes. Then I don't know what they're talking about. Yeah. Maybe maybe thirty seconds. Do you guys train it on the voice mode, by the way? That's my favorite feature of this of Delphi's product. If not, you should do that. Oh I think I'm I d can't remember. Okay. Um I should I should look at it. Definitely should. Now that we have this podcast episode, you could use this uh content to train it. It's eleven laps powered. It's so good. Uh okay. So that how do they get to I guess that's okay, they get to that once they become a Uh. Sign up for the course. And then you'll get a bunch of emails. Everything will be clear. Hopefully. Amazing. of all the students who ever taken the class and that discord is so active I d I can't go on vacation without getting notified on the plane. Bitters. Incredible. Okay. Uh with that, we've reached our very exciting lightning round. I've got five questions for you. Are you ready? Yes, let's go. Let's do it. Okay, so I'm gonna bounce between you two. Share something if you want. You can pass if you want. First question, Shreya. What are two or three books that you find yourself recommending most to other people? So I like to recommend a fiction book because life is about more than evaluation so recently I Read Pachenko by Uh then Jen Lee. Uh really Great book. And then I also am currently reading Apple in China. Which And then Name of the author is slipping my mind, but this is kind of more of a exposition written by journalists on how Apple did a lot of manufacturing processes. In Asia or the last um Couple of several decades. Very eye opening. Yeah. Yeah, I have them right here. Uh so these so I'm the nerd. Okay, so I'm not as cool as Shreya is. So I actually have like textbooks, which are like my favorite. So this one is a very classic one. Machine Learning by Mitchell. Now it's kind of theoretical, but the thing I like about it is It um really Drives home the fact that the That Occam's razor. Is Prevalent. Not only in like Science But also in machine learning. And AI. So a lot of times the simplest in also engineering. So like a lot of times the simpler approach Generalizes better. And so that's the thing I kind of internalized deeply from that book. And um also really like this one. So another textbook I totally have a nerd. This is like also a very old one. Mm. Well um And this is like, you know, Norveg Algorithms. And it's just like I really like it because it's just human ingenuity. And it's very like lots of clever, useful things in computing. I'm at Berkeley. Okay. The people that did that research. Yeah, uh Yeah. Textbook authors. Super cool. No man. Nerds, I love it. Okay, next question. Favorite recent movie or T V show? I'll jump to Hamill first. Okay, so I'm a dad of two parents. Oh sorry, uh two kids. So yeah, I'm a dad of two kids. And I don't really get The time to watch any T V or movies? So I watch whatever my kids are watching, so I've watched Frozen like three times in the last week. Only three. Oh okay, in the last week. Okay. Yeah. So That's not great. That's Mahammel. Frozen. I love it. Okay. Sure. Yeah, I don't have kids, so I can give all these amazing answers. Actually, so my husband and I have been watching The Wire recently. We never actually saw it. Growing up. So We started watching it and it's great. I feel like everyone goes through that. Eventually in their life they decide I will watch the wire. I know your life. I it's great. I It's such a great show. Oh man, and but it's so many episodes and everyone's an hour long and such a good moment. We get through like two or three a week. Mm-hmm. So we're very smart. Okay. Next question. Do you have a favorite product you recently discovered that you really love? And we'll start with Strea. Yeah, I really like using cursor. Honestly now, clawed code. Um Oh I'll say why. So I think as a I'm a researcher More so than anything else. I write papers, I write code, I build systems, everything. And I I find that like a tool I'm so bullish on AI assisted coding because like I have to wear a lot of hats all the time. Um and now I can be more ambitious. With the things that I build. Um and write papers about. So I'm super excited about those. Cursor, it was my entry point. Into this. Uh, but I'm starting to find myself trying always trying to keep up with all these AI assisted coding tools. How? Yeah, I really like cloud code. And I like it because I feel like the U X is outstanding. Um there's a lot of love that went into that. Um, it's just it's just really impressive as a terminal application that is that nice. Ironic that you two both love clock code when it's just built on vibes. I think it's false. It's not just built on vibes. Yeah. There we go. Okay, two more questions. Uh, Hamil, do you have a favorite life motto that you find yourself using And coming back to you in work or in life. Keep learning. And think like a beginner. Mm. Beautiful. Sure, yeah. I like that. Uh, for me it's to always try to think about the other side's argument I find myself sometimes just Encountering arguments on the Internet like this. Recent emails debates. And like Really think okay. put myself in their shoes. There's probably a generous take, generous interpretation and I think we're all much stronger together. than if we start picking fights. My vision for Evals is not that Hanela and I become billionaires. It is that Everyone can build AI products. And we're all on the same page. Slash everyone becomes billionaires. Yes. Yeah. Amazing. Final question. When I have two guests on, I always like to ask this question. And I'll start with Hamill. What's something about Shreya that you like most? What do you like most about Shreya, and I'm gonna ask her the same question in reverse. Yeah. Sharia is One of the wisest people that I know. Especially for being so young. Relative to me. I feel like she's like much wiser than I am. Honestly, seriously. Um very grounded. And has Like a very even perspective on things. And so I'm just really impressed by that. All the time. Yeah. Sure, yeah. Yeah, my favorite thing about Hamill is his energy. I don't know anybody who consistently maintains momentum and energy like Hamill does. Um, I often think that like I Would start carrying. Much less about e balls. If not for Hamill. And Uh everyone needs a hammer in their life. For sure. No. Well, we all have a Hammel in our life now. Uh this was incredible. This was everything I'd hoped it'd be. I feel like this is the the most in Interesting in depth. uh consumable uh primer on evals that I've ever seen. I'm really thankful YouTube made time for this. Uh two final questions. Where can folks find you or can they find the course? And how can listeners be useful to you? I'll start with Shreya. Yeah, uh you can reach me via email. It's on my website. If you Google my name, that is the easiest way to get to my website. You can find the course if you Google AI Evals for engineers and product managers, or just AI Evals course. You'll find it. Um we'll send some blinks, hopefully, after this, so it's easy. And how to be helpful? Two things always for me. One is ask me questions when you have them. I will try to get to the res Respond as soon as I can. The other one is tell us your successes. One of the things that keeps us going. Is somebody tells us Like what they implemented or what they did. A real case study and Hamill and I get so excited from these um And it it really keeps us going. So Please share. Yeah, uh it's pretty easy to find me. I'm my website is hammel dot dev. No, it can give you the The link? Um you can find me on social media, LinkedIn Twitter Um thing that's most helpful is To echo what Shreya said, we would be delighted. We're not the only people teaching Evals. We would love other people to teach Evals. And so Any kind of blog posts Writing especially Uh As you go through this and learn this, that you want to share We would be delighted to help reshare that or amplify that. Amazing. Very generous. Thank you two so much for being here. Uh I really appreciate it. And you guys have a lot going on, so so thank you. Thanks, Lenny, for having us and for all the compliments. Bye, everyone. Thank you so much for listening. If you found this valuable, you can subscribe to the show on Apple Podcasts, Spotify, or your favorite podcast app. Also, please consider giving us a rating or leaving a review, as that really helps other listeners find the podcast. You can find all past episodes or learn more about the show at Lenny's Podcast.com. See you in the next episode.