OpenAI researcher on why soft skills are the future of work | Karina Nguyen (Research at OpenAI, ex-Anthropic) Transcript from https://podmenti.com/t/e43e0066df5b05c9 Not only are you working at the cutting edge of AI and LMs, you're actually building the cutting edge. When I first came to Andarbine, I was like, Oh no, I really love front and engineering. And then the reason why I switched to research is because I realized, oh my god, Cloud is getting better at front ends. Cloud is getting better at like coding. I think Cobb can like What skills do you think will be most valuable? going forward for product teams in particular. Creative, thinking. And you kind of want to like generate a bunch of ideas and like filter through them and just build the best product experience. I think it's actually really, really hard to teach the model how to be aesthetic, a really good visual design, or like how to be extremely creative. in the way they write. What do you think people m most misunderstand about how models are created? When you taught the model some of the self-knowledge of you actually don't have a physical to operate in the physical world. The model would get like extremely confused. Today my guest is Karina Nguyen. Karina is an AI researcher at OpenAI, where she helped build Canvas, tasks, the O1 chain of thought model, and more. Prior to OpenAI, she was at Anthropic, where she led work on post training and evaluation for the Clot 3 models, built a document upload feature with 100k context windows, and so much more. She was also an engineer at New York Times, was a designer at Dropbox and at Square, It's very rare to get a glimpse into how someone working on the bleeding edge of AI and LLMs operates, and how they think about where things are heading. In our conversation, we talk about how teams at OpenAI operate and build product, what skills she thinks you should be building as AI gets smarter, how models are created, why synthetic data will allow models to keep getting smarter, and why she moved from engineering to research after realizing how good LMs are gonna be at coding. If you enjoy this podcast, don't forget to subscribe and follow it in your favorite podcasting app or YouTube. It's the best way to avoid missing future episodes, and it helps the podcast tremendously. With that, I bring you Parina. Nueva. This episode is brought to you by Interpret. Interpret unifies all your customer interactions, from gone calls to Zendesk tickets, to Twitter threads, to app store reviews. And it makes it available for analysis. It's trusted by leading product orgs like Canva, Notion, Loom, Linear, Monday.com, and Straba. To bring the voice of the customer into the product development process, helping you build best in class products faster. What makes Interpret special is its ability to build and update customer specific AI models that provide the most granular and accurate insights into your business. Connect customer insights to revenue and operational data in your CRM or data warehouse to map the business impact of each customer need and prioritize confidently. and empower your entire team to easily take action on use cases like win loss analysis. Critical bug detection, and identifying drivers of churn, with interpret's AI assistant wisdom. Looking to automate your feedback loops and prioritize your roadmap with confidence, like Notion, Canva, and Linear, visit ENTERPRET.com slash Lenny to connect with the team and to get two free months. When you sign up for an annual plan. This is a limited time offer. That's interpret.com slash Lenny. This episode is brought to you by Vanta, and I am very excited to have Christina Cassiope, CEO and co-founder of Vanta, joining me for this. Very short conversation. Great to be here. Big fan of the podcast and the newsletter. Vanta is a longtime sponsor of the show. But for some of our newer listeners What does Vanta do and who is it for? Sure. So we started Vanta in 2018 focused on founders. Helping them start to build out their security programs and get credit for all of that hard security work with compliance certifications like Stock two or ISO twenty seven oh one. Today, we currently help over 9,000 companies, including some startup household names like Atlassian, Ramp, and Langchain, start and scale their security programs, and ultimately build trust by automating compliance, centralizing GRC, and accelerating security reviews. That is awesome. I know from experience that these things take a lot of time and a lot of resources. And nobody wants to spend time doing this. That is very much our experience, but before the company to some extent during it. But the idea is with automation, with AI, with software, we are helping customers build trust with prospects and customers in an efficient way. And you know our joke, we started this compliance company, so you don't have to. We appreciate you for doing that, and you have a special discount for listeners, they can get$1,000 off Vamta. at banta.com slash Lenny, that's V A N T A dot com slash Lenny. For one thousand dollars off Anta. Thanks for that, Christina. Thank you. Karina, thank you so much for being here. Welcome to the podcast. Thank you so much, Slenny, for inviting me. I'm very excited to have you here because Not only are you working at the cutting edge of AI and LMs, you're actually Building the cutting edge of AI and LMs. We recently launched this feature which basically uh the first agent feature of open AI. I also just did this survey. I don't know if you know about this. I did this I did a survey with my readers and asked them what tools do you use every day in your work and most use. And chat GPT was number one above Gmail, above Slack, above anything else. Ninety percent of people said they use Chat GPT regularly. It's so it's absurd. It wasn't around two years ago. Yeah. Uh also we're recording this the week that OpenAI announced Stargate, which is this half trillion dollar investment in AI infrastructure. So there's just like a lot happening. uh constantly in AI and you have a really unique glimpse into How things are working, where things are going, how thing how work gets done. So I have a lot of questions for you I want to talk about. how you operate and how you work at OpenAI. Where do you think things are going? What skills are gonna matter more and less in the future. And also just where things are going broadly. So how does that sound? Sounds great. Thank you so much. Um Yeah. I was extremely lucky. To join early days on topic and kinda learned a lot of things. Uh there and and I joined OpenAI around like eight months ago. So yeah, I'm excited to have more. Okay, I'm gonna definitely ask you about the differences between those, but I wanna start More technical. And just dive right in. I want to talk about model training. People always hear about models being trained, this these big models, how much data it takes, how long it takes, how much money it tosses. It takes how uh how we're Running out of data, which I want to talk about. Let me just ask you this question. What do you think people most misunderstand about how models are created. Moral training is more an art than a science. And in a lot of Ways like V S like model trainers. Think a lot about like data quality is I like it's one of the most important things in model team is like Uh, how do you ensure the highest quality data for certain like interaction uh model behavior that you want to create. But the way you debug models is actually very similar the way you debug software. Um so one of the things that I've learned early days and ontopic was like We've discovered, especially with like Clock three training. When you taught the model some of the self knowledge of like, hey, like, you actually don't have a physical body to operate like in the physical world. But then at the same time we had theta They kinda taught the model. Um some of the function calls. Which is like This is how you set the alarm. As a model would get like extremely confused. Uh about like Whether It can set an alarm in the f but it doesn't have a body in the physical b uh role. So it's like the model gets confused and sometimes it'll like overrefuse. So sometimes it says like I don't know like uh sorry, I cannot help you. And so there is always like a balanced trade off between Uh, how do you make the model to be more helpful for users, but also Um not being harmful. uh in other scenarios. So it's always about like How do you make the model like more robust and like operate across like variety of diverse scenarios. That is so funny. I never thought about that. Most of the data that it's trained on is kind of like assuming It's like a human describing the world and how they operate and there's Well assumes there's a body and you could do things and the model told you don't have a body. Yeah. Uh Okay. I wanna talk a little bit about data. While we're on this topic, I know you have strong opinions here, there's kinda this meme that Models are gonna stop getting smarter because they're running out of data. They're trained in um large part on the internet. And there's only one internet. And they've already been trained on it. What more can you Show them about the world. And there's this trend of synthetic data, this term synthetic data. What is synthetic data? Why do you think that's important? Do you think it's gonna work? I think there are two questions here. Um we can unpack. Or one other time. But uh people say we're hitting the data wall. Uh, I think people think more in the terms of like pre trained large models that are trained on the bi well on the entire internet. To protect the next Talking. But what actually the model is learning during that process is actually How do you compress? The compression algorithm here. The model learns to compress a lot of knowledge. And it learns how to model the world. Um So the next prediction of the word like Teach me how to Drive. Basically, and you only have like a few words that will match that, a car. So the model actually learns Um about the world a in in itself. So it's like it's modeling human behavior sometimes. It's modeling and when you talk to like pre train models. Which are very, very large. They're actually extremely diverse and extremely creative. Because you can Talk to almost any Reddit user through Progene model. But I think what's happening right now is like new uh paradigm of like oh one serious is that like The scaling And the chaining itself. Is not hitting the wall. And that's because Basically we we went from like Raw data sets. From from featuring models. To Infinite amount of tasks. That you can teach the model in the post training world. via reinforcement learning. So any task, for example, like How to search the web. how to use the computer, how to write uh well, like All sorts of tasks that you like trying to teach the model. All the different skills. And that's why I'm be saying like there's no data wall or whatever, because there will be infinite amount of tasks. And that's how the model becomes extremely super intelligent. And we are absolutely getting saturated on all benchmarks. So I think the bottleneck is actually in evaluations. Uh, but you don't have Um All the frontier Like Eva's like I don't know, um, GPTA. Which is like A Google proof. Fashion answering like PhD level. And Dollar J the benchmark is like getting to like, I don't know, more than like Sixty, seventy percent, which is what G H D gets. Um so it's like they're h literally hitting the wall and like evolve. I wanna follow both those threads. So the first is on this idea of synthetic data. Is a simple way to understand it that the models are generating the data. That future models are trained on. And you ask it to generate all these Ways of doing stuff, all these tasks as you described, and then the newer models trained on this data that the previous generated. Some tasks are synthetically curated. So this is like an active like research area is like how do can you constru uh synthetically construct like new tasks model to like learn. Sometimes, you know, like when you develop products You get a lot of like data from the product and like user feedback and you can use that data too in like uh this like post-changing world. Um Sometimes you still want to like use like human Um human data because Uh actually s some of the tasks can be like really, really hard. uh to a teach. Um Like like like experts like Only no like sort of knowledge about like some chemicals. like biological knowledge, so like you actually need to tap into uh the expert uh knowledge a lot so Yeah, I think To me like synthetic data training is more um Well like Product is like a rapid model iteration for similar product outcomes. And more into but uh The way we made canvas and tasks and like new like product features for G Biki was Mostly done with by synthetic chaining. Let's actually get into that. That's really interesting. I want to talk about evals, but let's follow that thread. So talk about how this helped you create Canvas. So when the first came to open AI I really How this idea of like Okay, like it would be real cool. for Chapity to Actually like change the Visual interface but also sh Change like the way It is with people, so Going from like In a chad boss. to more of a collaborative agent. And the collaborator. It is like a is like a stop towards like Мой чай такси сенс. Um they become like innovators ultimately. And so the the entire team of like applied engineers, designers, product Like research kind of like God. Like formed. uh in the air. Oh almost out of like nothing. It's just like a collection of people who just like got together. And we rapidly started each waiting with each other. Actually like caves is like like one of the I would say like the first project of Open AI were researchers and applied engineers started working together from the very beginning of the Product development cycle. And I I think like there's A lot of things that we have learned. On the way? But um I definitely came to with the mindset of like we need to do like a really rapid model iteration such that like it would be much easier Four. Engineers to you know ripe with the latest model. possible, but also learn from like User feedback or like early like internal dog food. uh how do we improve the model very rapidly. And Yeah, it's really hard to like um kind of like figure out like how people when you deploy a product how people would be able to like use it. And so like The way you synthetically train the model is basically figuring out like what are the most core behaviors That you want that this product feature. To do. And for Canvas, for example Uh it was If Game Donzo likes to be mean. He versus it was How do you trigger canvas? Full promise like Write me a long essay. when the user intention is mostly like Iterating of our long documents. Or write me a piece of code. Or when to not trigger canvas for prompts like Can you tell me more about President like I don't know. Um so some of the general questions. So you don't want to like trigger campus because the user intention is mostly getting answered, not necessarily like iterate uh always a long document. The second behavior is Um How do you how do we teach the model to update the documents when the user ask? So one of the behaviors of the Uh the model is actually have like the a t some agency and autonomy to literally go to the document and like select s uh specific sections and Either deleted or edited. So highlighted and rewrite certain sections. So sometimes the model Sometimes the user would just like say Change the second paragraph to be something friendlier. And you would have to like teach the model to Literally find the second paragraph in the document and change it. To a friendly tone. So basically you teach Both like how to trigger like Uh edit itself. But also how do you teach the model to get higher quality at it? uh for the document. In case of like coding, for example. Uh there's also like the question of like how good the bottle is that like completely rewriting the document. versus like v having like very specific targeted ads. So that's like another like layer of decision boundary within like edit itself. It's like select. the entire document then like rewrite completely, or you want to like have like very talking custom keyboard. And you know, like when we first launched the model, We would bias the model towards like more rewrites? Because we saw the quality of the much higher. But over time you're like kinda shifting based on like user feedback and what they're learning from iterative deployment. Lastly, this the third behavior that we taught synthetically the model is how to make comments on any document. So The way we use that is like We would use a one model. to produce to like seem away of like music conversation. Let's say like Write me a document about XYZ. But then we used O One to like produce the document. And then the kind of injective like user prompt to be like Oh make some comments, critique my piece of writing. uh or critique this piece of writing that you just named Um And then the tell the model to like make comments on the document. On like very specific talking to the document. And so it's like also like what kind of comments you want the model to make, like do do they make sense or not? Like how do you teach The quality of that, um And It all came down to like measuring progress. Why very robust evolves. But yeah, this is how you like used like a wan and the like kinda like synthetic data condition for like the staining. Okay. This is so interesting. Uh, so you talk about this idea of teaching the model and you mention how it's Using synthetic data to teach the model different behaviors. Is a simple way to think about it. Basically that's where you you do that by uh showing it what success looks like using basically evals. Is that the simple way to think about it? Like Here's what you doing this successfully would look like and that teaches it Okay, I see this is what I should do. Yeah. Amazing. Yeah, you got it. Okay, got it. Um I wanna start unpacking what your day to day looks like as you're building these sort of things. Is it like you sitting there Uh talking to Some version of chat GPT. Uh crafting these evals. Sometimes I do that, but sometimes I do set I think I learned this so much. From It's like People spend so much time just like prompting models and like quality delivery by bash all the time, and you actually get A lot of new ideas. Oh, how do you make the model? uh better. It was like oh like this is this response was kinda weird. Like why is it doing this? And you start like debugging or something or like You you start like thinking out like new methods or like how do you teach? the model to respond in a different way, like have better personality, let's say. So it's the same thing of like How Personality. mm is made like in the models of the goals? Like very similar methods. But yes, I I think my time I don't think I have changed. I think when they first came, I was like mostly like research I see work. Uh so I was like Building a lot of like I was like writing codes, like you know, changing models. Writing evolves, working with PMs and like designers to like learn teach them how to like even think about like evaluations. I think it was like Really cool experience. And I think it was just like an a a an adoption of like How do we like do this like product management of like yeah features or like um AI models. Um Yeah, but now it's like mostly like you know, like management and like mentorship, um, I'm still like doing I see like research codes. after like four PM although but um Yeah, I just kinda like changed. All right, don't talk too much about being a manager'cause everyone's firing their managers. Who needs managers anymore? That's the what I hear now. Just kidding. It's interesting that so much of your time was spent on teaching product teams how evals integrate and how important That is and I've heard this a few times and I Haven't personally experienced it yet, so I think it's an important thread to follow is just How writing These evaluations is gonna become increasingly an important part of the job. of product teams, especially when they're building AI features and working well on So Can you just talk a bit more about what that looks like? Is it like sitting there with an Excel spreadsheet? Basically showing like here's the input, here's the output, here's how good the result was. Talk about what that actually looks like very practically. It certainly depends a little on the wood you're developing. But uh there are various types of like immunities. So sometimes I do ask product managers or uh there's also like new role that we have like model designers. to um Kinda like go through some of the user feedback maybe or like think of like various like user conversations. That should have triggered Like under this circumstances, it should trigger canvas. And then you have this like ground truth label of like, okay, with this conversation it should look trigger candles, and with this conversation it should not trigger candless. And you have this like very bind deterministic kind of like evolution. That for like This is modern behaviors is like this. uh when we were launching tasks, for example, like How do you make correct schedules? is like actually really hard for the model. But we both out like some of the deterministic evaluations. Then it's like, okay, like if the USA says like seven PM it's like the model should say seven PM. So if you can like have a defeministic evaluation, whether it's like Pass or fail. Um so yeah, and like the way it dresses, I was like Sometimes I ask Pragmatist just like go create like a rule sheet, like have different tabs and like Um What's the current behavior, what's like the ideal behavior. And like why like some notes. And sometimes we usually use it for eval, sometimes uh we use it for training. Because like if you give the spreadsheet to like a one model, it can probably figure out like how to uh teach itself. A good behavior. And I think there are second type of like evolves. There is kind of more prevalent is like uh human dominations. And you can have Specific trainers, or you can have like internal people. to um When when you have like a conversation of the prompt and then you have like various completion of models, you kinda choose the win rate. Which model is the best? Which model produced the highest quality? comment or Edit. And then you can have like continuous win rates. And as you develop new models, it should always like win over the previous models. So Um it depends on what you want to measure. So interesting. Like basically what I'm hearing in This is something I'm learning about as I talk to people. Is product development start might move from this like Here's a spec PRD. Let's build it together. And then cool, let's review it. Are we happy with this two from that to Hey, AI Build this thing for me. And here's what correct looks like, and I'm spending all my time on what does correct look like. Any valves essentially. You definitely want to like measure progress uh the model and this is where evasive is because like you can have Prom Third's model as a baseline. Already and of the most robust evolves is the one where Pronto baselines uh get the lowest score or something. And then because then you know like, oh if you're trained a good model, then it should like just like hill climb and that you out. all the time while not like also like regressing on like other intelligence evolves. So it's like I think it's more what that's That's what I'm saying, like it's more of an R than a science. It's like, okay, like if you optimize the model for this behavior, like you kinda don't want to like bring damage in like other areas of intelligence or This is happening like all the time in every lab and every like research team. Um I would say Like prompting is like also a way to like prototype like new On the good news? Oh, the early days of Andar wina. was working like file uploads feature. I remember I was just like Yeah, prompting the model. Um To just like I mean when we were like launching like Honda Ky contacts. I was just like prototyping this in the local local parson where sh I did the demo and like people really, really loved it. And they just like wanted like API for like file uploads or something. And then that's when it clicked to me like I also like run the blog post. On temphe, like it clearly to me like prompting is like a new way of like product development or prototyping for designers. And for like privilegers. For example, one of the features that I want to do is like have like personalized uh recommended personalized starter prompts. So when ever you come to like Cloud. Like It should like recommend you like starter prompts based on what your interests are. And so like you can Literally do it like Haunting. For that. Another feature was like generating titles for the conversations. It's a s very small like micro experience, but I'm really proud of Uh The way we did that was because we we took like five later conversation from the user and like asked the model like what's the style of the user. And then like for the next kind of new conversation, the generated title will be of the same like style. Micro experience is like those. Oh, that's so cool. Did you do that at Ethropic or at OpenAI? One topic. Okay, cool. I love the file upload feature that Claude has, by the way. Oh chat GPT doesn't have that yet, is that right? Um, I think it has. I think like the way it's implemented is like very different though. Okay, maybe it's the PDF feature because I use it all the time with colour. Okay. That's cool. Someone needs to get on that. Uh man, it's wild how many features you built that I use every day and that many people use every day. This prototyping point you made is really important. It's something that comes up a ton on this podcast also of how that is maybe the way that AI has most impacted the job of product builders. Recently is just prototyping. Instead of going from Showing just like here's a PRD, here's a design. PM's more and more just here's the prototype with the idea that I have and it's working. You can play with it. Yeah. Yeah. Okay. I wanna spend a little more time on how you operate. So you talked about You built this in launch this task. Feel is that is that the way you describe your tasks? Yeah. So talk about how that emerged and let's better understand just how you collaborate with product teams and how open A works in that way. Whatever you can share there. I think Canvas and tasks are Uh going into the bucket of Ducks where it's like more like Short like medium terms? And um Actually the way cameras and tasks About to be it was like It started was like one person prototyping. And Creating like A spec? It's kinda like PRD. It's like creating a spec of like the behavior of the model. I don't think like Tasks is like extremely like Grand baby. B ground breaking feature necessarily. What makes it like Really cool is Because the models are so general model can now search, they can like write sci fi stories, they can like Search for stocks they can like summarize your news every day. Because the models are so general Like giving something familiar to people that like, you know, notifications like very familiar. Like having reminders is like very familiar. So like Feeling like a Form Focta for the people who like Very familiar. Same as like Camus Moses like a bulldogs are very familiar. Oh, but then you add like magical ammo and then it's like it becomes like very powerful. But the way comes like operationally, like Yeah, it says this like a prototype, like literally prompted. prototype or like how you would want like the model to behave. For like tasks, for example, like You kinda like need to design. A little bit like design design systems design thinking is like Okay, like well Is the more is this the user says like um Remind me to go to lunch like eight a M tomorrow. Okay, what kind of information does a model need to extract from that prompt in order to create a reminder. And so this is how you like Like design like a spuck. For um a new feature. Like a tool. Canvas and tasks are all tools. So it's like how do you like create the tool stock? And then it's like mostly mostly like like Uh developing JSON schema is like Okay, like from this problem, maybe the model should extract like The time To the user request it. And then you're thinking about like which which form right do you want the time to be? And then like How do you One the model to like Not a five you is like Basically I the user should give instruction to the model. Uh, and then this instructions would like fire off like every day or something at the that particular time. Uh, so for example if you say like Search like every day I want to like learn No about the um Latest AI news. Um The model should derive into like Okay, like search for the latest AI news. And this will will this task will get fired at that go type that the model that the user requested. And then you know, your design was like Tull Spark and then Actually I don't know, like I feel like sometimes like It's like through conversations I Like I thought like people asked me to like join the listen like Genuinely like, Oh my god, like we need to be so charged. And we need like some support, like You've like to train the models, or sometimes like Tango was like most of like I just pitched the idea of like It got staffed quite immediately during the break. Um, so I feel like it it's like depending on the project. And usually with staffing it's like mostly like a product manager. Um model designer um Actual product designer. A couple of researchers and buy to like applied engineers. Depends on the complexity of the project. And then like Yeah, it takes it to for tasks it took like. I don't like Two months or so? To go from like zero to one, basically. Oh wow. Um For canvases was like Four five months, I guess. Um To go find out the each one. But uh yeah, and then like you know, you teach Part of mine is just how to like build evals and like maybe You know. How how do we Not really like ship. uh the better feature, but how do we think like more longer term? Like what kind of cool features that you want tasks to have? Like I think it would be nice for tasks to be like extreme a little bit more personalized It'd be nice to have like to create tasks via voice in on a mobile, right? Like so you kinda need to like this is how you get like research mo roadmap right here is like thinking like how the feature will be developed in the future. And then from there on it's like You like start? Creating data sets like Uh with e was you want to make sure that goes Well, and then like You Mm, need to have like a trade off between like what methods you want to use. And the reason why I really love like synthetic, like relying purely and synthetic data instead of like collecting. data from humans is because it's like much more scalable. It's cheap, let's have like you literally sample from the model. And you teach the core behaviors of the model, then that will generalize um to all sorts of the the worse coverage. And when you launch the better feature, you learn so much from the users. That you can like All your synthetic. That's Can be can be shifted in the distribution of how the users behave in the on the private behavior, and this is how you improve. Uh this is what happens kind of stuff. When we learn from better to Sh. This episode is brought to you by Loom. Loom lets you record your screen, your camera. And your voice to share video messages easily. Record a loom and send it out with just a link to gather feedback, add context, or share an update. So now you can delete that novel linked email that you were writing. Instead, you can record your screen and share your message faster. Loom can help you have fewer meetings and make the meetings that you do have much more productive. Meetings start with everyone on the same page and end early. Problem solved. Time saved. We know that everyone isn't a one-take wonder when it comes to recording videos. Saloon comes with easy editing and AI features to help you record once and get back to the work that counts. Save time. Alig your team, stay connected, and get more done with Loom. Now part of Atlassian, the makers of Jira. Try a Loom for free today. at loom.com slash Lenny That's L O O M dot com slash Lenny. Something that I wanna help people understand and I don't even a hundred percent understand this is what's the simplest way to understand the job of a researcher? versus, say, a model designer and other folks involved. Like what's the simplest way to understand what researchers do at open air. So the project that I described that mostly like product oriented, like research is mostly like product research. Another part component of my team is actually more like longer term exploratory projects. And it's more about like Developing new methods. Understanding those methods. Under variety of circumstances. So Ли бисикли девалні мати You kinda like me to follow very similar kind of like recipe of like building e balls, but it's much more sophisticated evaluation, like you kinda want to have like outer distribution or like if you want to like measure generalization Um If I don't need to like capture thoughts. Uh, but it's basically more science y in a way where You know, it's it's the Talk about synthetic data. Like one of the hardest things about synthetic data is like how do you make it like more diverse? Diversity and subject data is like one of the most important questions. Uh right now. And so it's like exploring like ways to ench like diversity as a general method that will work for all is like a one of the like research explorations. other ones is like more like about making new capabilities. I feel like it's all just about like you know, like You You Work on this like new method. And you have like signs of life that it's working. I did just think of like how do you make it more general Or you think of like How do you make it very useful or like And this is how like longer term projects become more like medium like short term projects. That makes sense. Essentially working on developing ways to make the model smarter, oh four, or five, oh six. Yeah. Like O one was a big breakthrough, right? The way it um operates where it's not just here's your answer, it actually Things and has Right. Takes time to think through the process of coming up with an answer. Okay. Yeah. Very helpful. Speaking of that, of thinking about the future, where things are going. I wanna spend some time on Just this. I insight that basically you are building the cutting edge of AI. Like at the very bleeding edge of where AI is going and where it is. And so Uh I'm very curious to hear just your Take on How you think things are gonna change? In the world. And how people work. Based on where you see things are going and I know it's a broad question, but let's say like in the next three years. How do you see the world changing? How do you see people's way of working changing? It's a very humbling experience to be in both labs, I guess, like To me when they first came to and daring I was like Oh no, I really love from then engineering. And then like the reason why I switched to like research It's because I realized at that time is like, oh my god, like Cloth is getting better at like front ends. Like Cloth is getting better at like coding. I think Cloth can like develop new apps or something. And so like it can like There are new features for the thing that I'm working. It was kinda like this meta realization where it's like Oh my god, like The world is actually changing and they're like When the first like launched Hundred K context at that time. Obviously You know, I'm thinking about like From talk there, so it's like Yeah, like File uploads were like very natural, very familiar to people. But you could imagine we could just like make like Infinite chats in the called that AI app, right? Like as if like it's like a ha in a hundred key context. But because like file uploads It's like for and follow his function, it's like The form factor, the file uploads kinda enable people to just like literally upload anything, the books, like any reports financial, and like ask any task to the model. And then I remember it was like Yeah, enterprise customers like Um like financial customers are like really interested in that as like Oh wow, like actually the It's actually one of the very common task that people do. Uh In that setting it was like Kinda crazy to like see uh how some of the redundant tasks are getting like Oh, to me, basically. Buy this like smart models. And we're entering the the era where I actually don't know, for example, sometimes like if a one Gives me the correct answer. or not because I'm not an expert in that field. And it's like I don't even know how to verify the outputs. Um Other models is because like Oh, my experts know in life. They can like verify those. So Yes. So basically there are Trends that were going on. The first trend is the cost of reasoning. and intelligence. Is drastically going down. I had a blog post about this. Maybe I should update them like Latest benchmarks because at that time like MMO everybody was like Do I um Not one like one benchmark and then be like quickly saturated the benchmarks and like now be you need to like do the same flop but was with another like frontier evil. But the cost of intelligence is like going down because it's it becomes like much cheaper. S smart small models are becoming It was smarter than like large models. And that's because of like The Distillation research. This happened was like Claude High Cool. I was like working on like post annual clothes high cool. And I realized it was much smarter than like Claudio, which was like way, you know, bigger or something like that. Um, but like the power of like small models become very intelligent. And fast and cheap. We are moving towards the world. Dale has like multiple implications But The news that like People who have more access to AI, and that's really good. Like builders and developers. will have much better access Two AI But also it means like all the vertical that has been like bottlenecked by the intelligence. will be kind of like unblocked. So Anyone Like in I'm thinking about like healthcare, right? Like If I have Instead of going to the doctor, I can like Аск чачі пітії лаків чачі пітії A list of symptoms and ask me like Uh which Like would I have like a cold flu or like something else like I can literally got The access to like Uh doctor almost and there's like been some like research studies around that. Yeah, there was a New York Times story about that where they compared doctors Two doctors using chat GPT to just chat GPT and Chess Chat GPT was the best. Yeah. Like doctors made it worse. Yeah. Yeah, that's crazy, like right, like education, I think Uh I will have drands if Like I had the tool like ChatchPT and when I was like young and like would learn so much, but it's like People can now learn almost anything. From these models so they can learn New language. They can learn how to build new look ups like I don't know, anything that you want in like And so Oh, like it's humbling to like have like launch canvas and like bring that thing to the people, enable them to do something else that they couldn't have ever before and I think this is there's something like magical around this experience. Uh so education has will have massive implications like I guess like scientific research, right? Like I I think it's like the dream of like any AI research is like upmate AI research. Uh it's kinda scary, I'd say. Um Which makes me think that like people management Well stay, you know, it's like One of the hardest things to is like emotional intelligence. for the models or like creative we creativity in itself is like one of the hardest things. So Writers I I don't think like People should be worried as much. I think it's like I think it's L Eviate for a lot of like redundant tasks. Um for people. This is awesome. Okay, I wanna follow this story for sure. And it's funny that what you described is like you were an engineer at Anthropic and you're like Okay. Claude is gonna be very good at engineering. This isn't gonna be a Potentially career long term. So I'm gonna move into research. And The AI's gonna need me for a long time to build it, to make it smarter. I would say we still have like I think Canvas team has still have like a really cool Like from the engineers that I Really like Yeah, people who like really care about like Interaction design, like interaction spirits. Like I don't think like models are there yet, but like I think If y but if you can get the models to like this top one percent of like front end or something. Um Sure. So what I wanna move on to next along these lines is just uh and this is just speculation, but uh What skills do you think will be most Valuable Going forward. for product teams in particular. So folks are listening and they're like, Okay, this is Scary. What should I be building now to help me Stay ahead and not. Damn. in trouble down the road. What skills do you think are gonna be most more and more important to build? Yeah, I think like create a Thinking Like you kinda want to like Um come up like generate a bunch of ideas and like filter through them. And it's like build the best product experience. Listening You know, you want to like build something that like The most general model will not replace you. And oftentimes you You build something and you Make it really, really good. For Like Specific set of users. And I should emote. is now in like Your user feedback the mode is like More in life. Y whether you listen to them, like whether you you can like rapidly trade, like The mode is like in here. I I don't think like V yet to like There are so many ideas. I think there's an abundance of like Ideas that you shouldn't look recalls like. I wouldn't be worried. I feel like in fact I just think like people NAI fields are like I w I wish they were like a little more creative and like connecting the dots across like different like fields or something like that to like develop really cool new Like Generation. A new paradigms of interactions with this AI. Like I don't think we've cracked this problem at all. Um A couple of years ago I was like telling some people I was like, you know, you kinda want to like built for the future. So it's like It doesn't Necessarily Matter whether the model is good or not. But you can build Product ideas. such sounds like by the time the models will be really good, it will work really well. Um I think it just like happened naturally. Like for example like an anthropic light, right? Like Um The clawed artifacts. And I feel like early days of Canvas was like back in like twenty twenty two, like before Chai Chi P T, like writing I D was like on all Ch P but I feel like Claude one point three model itself. was like not there to like mate like really extreme good, like high quality edits, for example, like coding. Um And I feel like I I feel like Startup is like Cursor and it's like doing super well. Like I guess because you like Iterates so fast, they like invent like new ways or like training models. Day move really fast. They listen to like users like Massive distribution is like Yeah. That's really helpful actually. So what I'm hearing is that soft skills essentially are gonna be more and more important powerful. You s talked about management, leading people, being creative and coming up with innovative insights, listening. There's a post I wrote that I'll link to where I look I I try to analyze what AI how AI will impact product management. And we're actually very aligned. And my sense was the same thing that soft skills are gonna become more and more important. And the things that are gonna be replaced is the hard skills, which is interesting'cause usually people value the hard skills like Coding design. Writing really well. And it's interesting that AI is actually Really good at that'cause it's Taking a bunch of data, synthesizing it, and Writing, creating a thing versus All these fuzzy things around of What influences convinces people to do things and aligning and Listening, like you said, creativity. Anything along those l along those lines come up as I say that? I think it's actually really, really hard to teach the model how to be aesthetic or like uh do like visual really good like visual design or like how to be Extremely creative in the way they write. I think like I still think like Chad GP kinda sucks at like writing. And that's because it's like it's like bottled mouth by this like creative reasoning. I think like privatization is like one of the most important like I think like Um For Matters, I feel like I actually like AI research progress is unlocked by like Management like research manager is because you have like Has change set of computers? And you need to like allocate the computers to The research path that you feel the most convinced about it was like you need to like really you need to have like a really high conviction in the research paths. to put the compute and like it's more like return on investment, um, kind of situation. As like, okay, yeah, like I I'm thinking a lot about like okay, like How do I across all my projects which projects a higher priority is like prioritization and also like on the lower level like which experiments are really important to run right now and which are not and like cut through the line. So I think you're like prioritization, communication, like um management, um people's skills like Another thing. Like Understanding people like kinda like collaboration, like I think like Canvas wouldn't be Like an amazing launch. If the it wasn't like about like People. And I think it's a it's a wonderful glad group of people I like. I get a chance to look at Rugges, like people like Lee Byron, who's like a co creator like Ralph Kiel, and like some of the best like Apple designers. And it's like So cool. To like see And like how do you create this like collaboration between people is just like Something that's still humane, I think. Let me just follow throughout a little bit because I imagine people listening are like, Okay, but once we have AGI or SGI, it's like it'll do all this. It'll you know, it's like there's a world where like why isn't all this done? I think it's easy to just assume all that. I'm curious this idea of creativity and listening. Why you think AI m isn't good at it? Other than it's just very hard to Train it. to do this well. Is there anything there of just like why this is especially difficult for AI Now lamps to get good at. I think currently it's difficult For Many reasons I think it's still like a now I feel like research area and is something that like I think my team is like working on is like Okay, how like how do we teach the model to be like more creative in like the writing? And so like I'm thinking like Doesn't you paradigm always that the morals think more should actually lead to like better writing. In itself. But like when it comes down to like Idea generation Oh like Um discriminating of like what is a good Like visual design and odd. I feel like if Hasn't had learned Like Examples from like people To discriminate it by well. I do think it's because like You know, there are not that many people who are like actually like really like Oh, it's not like accessible to like models to learn from these people, I guess. Um So I definitely that's why it's talking. Yeah, that makes sense. Basically there's not enough of you yet. Uh researchers. T teaching it to do these things slash. People that have incredible taste and Creativity that can teach these things. You could argue this will come. But I'm not going to We don't need to keep going down that thread. Let me ask you a specific question. In this post I wrote, I I made this argument that a lot of people disagreed with that strategy. is something that AI tooling will become increasingly Great at and take over. There's the sense that That's the thing that People will continue to be much better at and you can't offload AI basically developing your strategy. Telling you what to do to win. My case is Isn't strategy just take all the inputs, all the data you have available. Understand the world around you and come up with a plan to win. Yeah, I would be Like an LM would be incredibly smart at this. What's your take? I think so too. I think like Again, like we g we you teach the model all sorts of like tools and like capabilities and like reasoning, right? And it's like when it comes down to like This is for Canvas right now, it would be very cool to like for the model to just like aggregate all the if you've bought from users like summarize me like the top five like most painful flow flows like user experiences. And then like The model itself is like very capable of like Like thinking of like knowing how it's Being made? Uh figure out like how to like treat a data sets for itself to like train on it. And then they look we have far away from that kinda like self improvement. Models becoming like self improved. buyer Wait. Then like the part the Wildman is basically kinda like self improving like s it's kinda like its own like organism or something. Um Yeah, like I like strategies like It's more like data analysis and like Um Um Coming up was like Like I think what models are really good at is like um Well connecting the dots, I think. It's like Okay, if you have Usual feedback from the Sores. But you also have like internal like dashboards with metrics. And then we have You know, like other kind of like feebal. Um Oh like Input. And then like it can co create like a a a plan for you, like recommendations, even. I think this is like one of the most common one use cases for Chaptwood is like Coming up with like this sorts of things. That makes sense. Like essentially A human can only comprehend so much information at once and look at so much data at once. to synthesize takeaways and As you said, these context windows are huge now. Here's all the information. What's the most important thing I should do? Yeah, same as like scientific research, is because like you like ideally the model would be able to like suggest like ideas, like new ideas or like iterate on the experimental like Given the empirical results of the previous experiments, like how do you Like Come up with like new ideas or like the methods. Yeah, no man. Uh okay, so just to close the loop on this conversation, this part of the threat is The skills you're suggesting people Focus on Building and leaning into soft skills. Like creativity. Managing influence collaboration. Looking for patterns. Is that generally Where your mind is at? Yeah, I'm thinking a lot about like how do we make organizations more effectively. And I think this is m most of like management, I guess. It's like How do you organise like research teams like generally teams like combine c compose teams such that they will be at their maximally succeed or like at the maximum like performance of What can possibly Like if you can like literally create like The next generation of computers it's just like the matter of conviction and like The way you manage through that. Organizations, all of scaling Product research does. Yeah, I think what like you're basically building this thing and not efficiently doing it is like limiting the potential of the human species right now is mismanagement within the research team and open and anthropic and some of these other models. Yeah. Yeah, kinda crazy to think about it. Okay, so speaking of anthropic and open AI, you've worked at both, very few people have worked at both companies and seen how they operate. I'm curious just what you've noticed about the differences between these two, how they operate, how they think, how they approach stuff. What can you share along those lines? It's more similar than different. Uh obviously there was a lot of like There are some like differences also comes to like nuances. I'll take culture. I really love Anthropic and they have a lot of friends there. And I also love Open AI and they still have a lot of friends, though. So it's like it's not about like enemies. I feel like there's like in the I was all like, yeah, the competitors and there's like enemies. This is actually like one big community and like of people like doing The same thing. What they would have learned from Antarctica is This like Real Karen Croft? Towards like moral behavior, model crossed, model teething. And I've been thinking a lot about like okay, like what makes cloth cloth and what makes chip chip hands like I suppose sometimes sounds like Operational processes. That kind of leads to the outputs. To to the model. uh is the odd quoted model. And it's like the reason why Claude has so much more personality and like Uh is more like a librarian, I don't know, like I don't know, I am like visualizing A claw being like a librarian. Like a um very like naughty or something. Um It's because I feel like It's a reflection of the creators who like making this novel and like A lot of like details around like the character and the personality and like whether the model should follow up on this question or like not like was the correct like ethical behaviour for the model to like in this scenario is like A lot of like craft, um And like trouble you did like the stats? And where I learned. That part of like Art I guess. Uh I don't know. I say on through is like much smaller. Like when it joined it was like what, like six seventy people? when I last the most comfortable people and like obviously the culture changed so much I really enjoyed the like early days startup like vibes. And like people knew each other as a family, but like the culture shifted. I would say like Underwick I learned from Undartic that like They're much better at like Focusing and like preaching like very, very hard like very hardcore priorities, I guess. And then you need to do it. Like but I think like opening eyes look much more um You know what else? and um much more like risk takers in terms of like Product, or like we start actually Yeah, we I don't know. your full time job can be just like teaching the model how to be like creator writers. And it's like there's s some luxury in this like research freedom. That comes to scale, maybe? I don't know. Um But it gives you It it's like you'll have I feel like I have much more creative like product freedom to do. Almost anything, I guess. within like opening eye like a Ross Chashi Pikine to like The Ugand. It's like more like Yeah. Yeah, that's how I was I was thinking about it. It feels like opening eyes more Mm-hmm. Bottoms up. Uh distributed people, bubble up ideas, try stuff. There's more Well and that in more leads to more th products launching, I imagine, more things just kinda being tried. Versus more of a Let's just make sure everything we do is awesome and create and craft and Right. That's really interesting. I've never heard it described this way. Uh Karina, we've covered so much ground. This is gonna help a lot of people with so many uh ways of thinking about where the future's going. Before we get to our very exciting lighting round, I'm curious if there's anything else that you think might be helpful to share or get into One of my regrets I guess when I was Early days I don't know it was That like I think there was like some luxury of the time. pre chai chippy to actually like come in with like a bunch of ideas and like prototype like almost every day. Um And I think like we did a lot of cool ideas, like Claude and Slock was actually one of the first like uh tool usy like product It's like Yeah, cloth could operate in like your workplace now. It's like But then you like Add Claude to summarise the thread. So maybe you have like a entire conversation with someone and then you want to like a summary of like what happened, like you can say like At the cloud summarizes. Also it was really fun to like even like iterate on the model itself. It's like when you just like talk to the model in like slog forever. Um It created like a some social element that's kinda cool, it's kinda like me joining me and like Um this discord. Like people learned So much about prompting and like how to work. Was like Claude? I should have the Features that was like Early tasks part of time was like, you know. Every Monday Claude would just like summarise the entire channel. Or like every Friday we just like summarise like Bunch of channels. Uh and give like the news about the organization or something. So it's it's kinda like really cool, like Phone factor? I think they're thinking about like phone factors like A really important like question like in AI. Especially if you have them like even thicker odds, like How do we create like an awesome like product experience was like oh serious models. It's like the paradigm between like synchronous real time. Give an answer. paradigm into like more asynchronous Perdonal like. Agents working on the background. But then now the question is like the ancients should both trust with you, right? And trust both over time, which is like with humans. And um You know, you you saw this collaboration, which is why like a collabor like This collaboration model was like you and the model is like so important. Because you both trust and the model learns from your preferences. So that it can become like more personalized? And it will start predicting the next like action that you want to take. on the computer or something. And it was like Kinda like more predictive, much more of a We we could we went from like personal computer to like personal model uh basically here. Yeah. That seems like such an obvious feature that every L M should have is a Slack bot version of them. Is that is that a thing I can have you install or is that not a thing right now? I know that Claude and Sug were sunseted in like twenty twenty three or something, but that's because like I think Uh I think it was like after Chipt it was mostly like the focus on like Consumer use cases or like enterprise use cases. Uh I think it's the one like I think the form five like clawed and slot'cause like Um Was kinda constrained a little bit. Uh when you want to call it new features. Never I want that. I know that J B had like Slack parts, I don't know, like maybe it will come back. All right. I would I would pay for that. Uh any other m memories from that time of early days?'Cause that's a really special place to have been is early days anthropic. Any other memories or stories from that time that might be interesting to share? I think the very first launch when we felt like When clips and use the game was like a hundred key context. Um, launch is like when the models could input The entire like book. um give you like some read a book or something. Um or the entire financial or like have like multi files. financial rewards and then like give you an answer. Um to the question. To very specific question. I think there was something in there that kinda like Oh my god, this is like a real cool New capability, not like model capability, but more like The capabilities came from the product form factor itself rather than like The model capability as much. Um I think like other prototypes that we m Right. Thinking about like Yeah, like Cl uh there was like one part of the clog work spaces and it's like kinda same like idea of like Claude and I would have the shared workspace and that shared workspace is like a document and we can like it to another document. And I feel like sometimes the ideas like private ideas log. And they'll last for like two years. Um, just like in this case. It's interesting there are these milestones that kind of uh open up our view of what is happening and where things are going. Chat GPT, I think was the first of just like, wow, this is Much better than I would have thought. You talked about One hundred K context windows where you could upload a book and ask you questions, have it summarized. I actually use that all the time when I have interview guests and they wrote a book. I sometimes don't have time to read the whole book, so I use it to help me s understand what the most interesting parts are and then I actually dive into the book just to be clear. Uh huh. Uh And then I don't know, maybe like voice Was another one where you could talk to say chat GPT. There any other moments there that you're like wow, this is much better than I thought it was gonna be. Yeah, I think like Uh The computer use agents like The model operating The desktop. And you can essentially think of like You know, new kind of like experience where The model can learn the way you browse. And from that preference it can just like brass as just like you and like simulation simulated Persona. And it's actually very similar to the idea of like Okay, like Maybe Some album doesn't have a lot of like Um time maybe I want to like talk to like his simulators like his simulation and ask like or like for example like yeah like I I I really appreciate some of the technical measures that like Ya'cok like but he doesn't have a lot of time so it's like I really want to like ask him this question like how do you respond Like simulated environments like those. Um That's a great place to plug Lenny Bot. I have one of those. It's trained on all of my podcasts and newsletters. And I it sits on many models, I don't know which one exactly. They use, but it's exactly that. It's Uh and it's not even me, it's All the guests that have been on the podcast on the newsletters I wrote, and you could just ask it, How do I grow my product? How do I develop a strategy and it's actually shockingly good. Do you feel like it reflects Yeah, like what is the again. The best part of it is you can talk to it. It's built there's an eleven labs voice. Version that's trained on my. voice on the films podcast and it's actually Very good. And people like have told me they sit there for hours talking to it. And somebody Uh. told it, interview me like I am on Lenny's podcast, ask me questions about my career, and he did the half hour podcast episode with Lenny Buck. That's so fun. It's incredible. Future is wild. Yeah, I think like Content transformation is like You know, like I would imagine sometime live, you know, um When you generate a sci fi story. In Canvas. Like you can like transform this into like audiobooks like very have like very natural like content transformation like one media to another medium. I think like One of my In earless inspiration. Um, is like One of the last episodes of like Westworld. Where uh I wanna slow, but where Dolores comes to her work. At the time when she comes to like There's like New workspace and she starts like Writing a story. And then if she writes a story, like a three D, like virtual reality stuff, like Creating on the fly. So I kinda want to Um Kind of cool. Wow. Speaking of medium uh uh I guess I I was I was wondering if I should go in the direction or not, but real quick. Uh Kevin Wyl slash Kevin Weil. I don't know exactly how to pronounce his last name. The CPO of uh of Open AI. Uh. Is a while or wheel? I think real. Real. Okay, okay. Let's just say that. Okay, we open uh He was uh he did a panel at Millennium Friends Summit last year and he made this really infascinating point that Chat is a really interesting interface. for these tools because They're just getting smarter and smarter and smarter and smarter and smarter, and chad continues to work. As a paradigm to just interact with them. Similar to a human. You could talk to Albert Einstein, you could talk to someone not very smart and it's all conversation still. And so it's a really flexible way to interact with increasingly Good intelligence. At some point it'll not be so great and You're talking about all these ways that You're adding additional ways to interact. But it's interesting chat proved to be a really powerful layer on top of all the stuff. Yeah, that's really cool. I feel like Chad also has like social element, which is like very Yeah, actually Yeah, you sometimes want to like get into a group shot and like Yeah, having conversations with AI is kinda like a group channel in itself, is a message thing. I actually think this this idea of like How do you build like features like this? Like I see tasks as like This like Um, general kinda like feature that will scale very nicely as the models will develop like new capabilities and so it's like Like the model will be able to like do better like searches and like Yeah, create new like Come up with like a more creative like writing on like render, you know, React apps and like HTML prev like apps and like you can have like Every day and you puzzle for you, like every day, like continue the story from the few days. It's like It it scales very nicely. You mentioned something as we were Getting into this extra section that we ended up going down is This idea of Uh your the agents using a computer. I know this is actually something you are gonna launch today, the day we're recording it, which will be out by the time this comes out. Call operator. Very cool feature that people will have access to. Yeah, so uh I unfortunately did not work enough, but I'm really really excited about like This launch. Um It's basically Imagine that can Complete. the task in its own like virtual computer. Like in its own virtual environment. You can do any literally task, like Order me. A book on Amazon. And then Ideally is a moral. Well, I don't like follow up with you like which books do you want Or like know you so well that it like start recommending like oh here's the five books that I might recommend to you To buy and then like you It's like, yeah, help me Help me by and then uh the model goes off uh into its own like virtual Little browser and like complete the task and buy the book in my Amazon and then If you give the model like credentials. credit cards. Obviously it comes with like a lot of trust and like safety. Um Then it will just complete The thing for youth. There's a virtual Assistance. It's interesting how this just sounds like obviously this should happen. Like why is this not a yet a thing, which is also mind blowing that Should exist. Like just some AI doing things for you on a computer. You just ask it to do. Like it's absurd. really hard. And I think like Um You're still cracking this, but I feel like I don't know if you're used like top all, it's like a pair programming No. But um I don't know if you love their product. Oh yeah, Shopify uses this. I remember came up on a podcast episode. Oh nice. Yeah, so it's it's a very cool product where you can just like call anyone. At any time. and then like share screen and the other person can like have access to the screen and like start like Literally operating your computer. And it's very like real time like the Legions is like very Um It's like h very high quality. Um And it's just like I kinda want the same, it's like I wanna like Pair Program with like my model. And like the model should You wanna talk to me or like draw like very specific like section in my code and VS code and like Tell me like I mean teach me and we can have like different modes. It's like right here. It does like A product right here for you. I don't know. Um People some people should both build up. It sounds like a startup just gotten birthed. Yes. From someone listening to this. You mentioned that it's very hard to do this Uh agent. controlling a computer as you and helping out what makes it so hard. For whatever however much you can explain briefly. Much of it is like Uh because right now the models Yeah. Operating on like pixels? Instead of like Language or whatnot, like Pixels is actually really, really hard for the models because like perception or visual perception. I think there's a still a like a lot of a lot of like multi model like research that's going on. Um But I think like language scale so much like easier compared to like multi models because of that. Another s like thing that I guess like my chimney store and I is like How do you derive human intent? Um very correctly. It's like Sometimes like Does the model know enough information to ask a follow up question or like To complete the task. You kinda don't want like an agent to like go off for like ten minutes and then come up with like An answer that you didn't even want that actually creates like much more verse user experience. And this is comes with like teaching the model like Like People skills. It's like You know, like What did people like? Like kinda like cre creating like the mental model of the user and like care about the user in order to ask certain questions like Actually That part is like hard. Two for the morals. That relates to what we talked about earlier, where the kind of the soft skill people skills pieces. Not where these models are strong yet. Okay. I'm gonna skip the lightning round. I wanna ask just one question from Lightning Round something fun. Uh Okay, so when AI replaces your job, Karina. I'm curious what you're getting and it gives you a stipend. Gives you a monthly stipend. Here's your here's your salary for the month. What w what would you want to do? What do you want to spend your time on? What will you be doing? In a future world. I've been thinking about this Oh, because I have I feel like I have a lot of Jobs Options I would love to be a rider, I think. I think that would be super cool. Um, just like a ride like short stories, like sci-fi stories. Um No I really like Arn T studio. So you know it's like um Contribution of this to like in the museums. We just like Try to preserve like art paintings, but just like painting through A little bit. I think that would be really cool. Um To do. Um Yeah. That sounds beautiful. I don't know. Uh what I'm hearing is you need to nerf these models to not get very good at writing. So that you can continue. Although at that point you don't need to do it from like you don't need people to buy. You're just doing it for fun. So it doesn't even matter if they're incredibly good at writing. Or art art conservation. Oh man, what an episode, what a conversation. What a wild time we're living in. Karina, thank you. So much for being here. Two final questions. Where can folks find you online if they wanna reach out and follow up on anything? And how can listeners be useful to you? You can find me on Twitter. Amen. Um you can also shoot me at email. on my website um And I'm my team is hiring and so like I'm looking for Research engineers, research scientists, as well as like machine learning engineers like people who come from like part engineers who want to like learn like model training. Um actually hiring for like My team, my team is called Crontier Product Research. And The tree models we develop new My first but for product oriented outcomes. What a place to work. Holy moly. Uh what's the best way for people to apply for these uh very lucrative roles. I think you can shoot me a DM on Twitter. Okay. Or um I'm yet to create a job description. Okay. This is the job description. Or you can apply under like post training. Team. Yeah. Okay, I was you're gonna get a flood of DMs. I hope you're prepared. Kurina, thank you so much for being here. This was incredible. Thank you so much, Lenny. Bye, everyone. Thank you so much for listening. If you found this valuable, you can subscribe to the show on Apple Podcasts, Spotify, or your favorite podcast app. Also, please consider giving us a rating or leaving a review, as that really helps other listeners find the podcast. You can find all past episodes or learn more about the show. at Lenny's podcast dot com. See you in the next episode.