Transcript
189: Agentic Loops
0:15 Programming Throwdown episode one hundred and eighty nine. Genic loops. Take it away, Jason. Hey everyone, how's it going? Just my microphone here a little bit.
0:26 Um So We should talk about About programming. So
0:32 You know, I have been For the first time. Since I was Probably like nine years old. I have not coded
0:42 Or uh a number of days. Um And I was thinking a lot about this. I mean You know, it's like
0:49 Ever since I was even before a teenager, I was constantly writing programs on the computer Um Yeah, we started I don't know if you started with basic, but I used to do a whole bunch of stuff in basic. I made like a choose your own adventure and basic and uh um a whole bunch of things on the on the commodore and then
1:09 Um, I went from basic to I used basic for a very long time. I think I went straight from basic to C plus plus, which was uh Terrifying. I definitely like uh Nothing worked for a really long time. Um
1:24 And I've just been coding ever since. And The weird thing is I've built so many more things. Like my my the things that I've built my my build rate or my efficiency or my progress is like Super, super high, but all of it is in English.
1:42 And um It's almost at the point where I don't even really look at the code. Um, except for very specific things that I'm trying to find and and even that is starting to dwindle. And it's just such a weird
1:55 Feeling. Um, I mean, even the the premise of the show is kind of, you know, learning every language. And uh and now it's like uh Everything is just code in English. Or or in plain language. You could even code it in
2:08 in whatever language you're most familiar with. Um by typing whatever that is into the Let's It's really starting to hit me.
2:17 That uh Uh That cod coding in plain language is like Really Probably here to stay. It's kinda wild.
2:25 Yeah, it is a bit of a mind bend. I mean I think this is Getting to to the point where I mean, especially with the framing of it's people say that a lot. Like it's the worst it's ever gonna be. Like this is the There's no reason to believe it would regress.
2:42 Um I guess like if we somehow ban all GPUs or AI, isn't that the Dune thing to have a big war because they anyways AI is banned? Um, I don't think that's gonna fail. I I doubt it's gonna happen. Uh so Yeah, you're right though. Like it Totally changes the way you approach things.
3:00 And I think There's various people have a lots of opinions on it. And I you know, I don't think we should Get into all of them. But one of the ones that you hear is like oh it can it can make bad code or it can make mistakes.
3:12 And I'm just like, have you worked with programmers? Like have you looked at my at your own code? Exactly. Have you looked at my like how many times have I found a stupid just like off by one error just like yeah, I did all the time. Like There's this weird thing that like, I don't know, somehow we were infallible or people were infallible. And I'm not saying like good engineers, bad like I'm just saying uh random people. Like everybody makes mistakes. And there's a lot of people I've read code from. Where it's clear like I'd have preferred to read the L L M code than what they wrote.
3:43 Um So yeah, it It's a bit of an interesting transition time and I think we've been talking about for a little while. But it definitely feels like there's been a uh knee in the curve or a change or
3:55 Some threshold has been reached where It's no longer like oh it's It's the yeah, that was uh chess and humans, right? The original idea is the Chase would help the humans. Then it's like now the humans are really holding the just engines back. And so now like modern ones there's
4:10 Like it is much better to use a chess engine directly than to try to do some Cooperation. Um and human computer interop is not not advantageous. And I feel That's not there yet. But it kind of feels like we're on that journey.
4:25 Uh there are definitely people who can work with the AI better and and giving instructions because it's you know We're not gonna get into like a a free will discussion, but like you know, having the AI pick something to do. It's still something that seems to be in the human domain, but Uh you know, certainly the tactical open the code editor, debug the code.
4:44 Like I Yeah. I'm not sure how much longer that's gonna be a thing and I wonder what like gets lost a little. Like you know, we can talk about it, but even one of the news articles I have this time.
4:55 There's definitely things that I've framed in a certain way out. Well say at least in today's AI that helps it. It's like I want you to take this approach, or I want you to use this technique, or I want you to do this thing. And I don't have to write the code for it, but I needed to know when and what I was asking for.
5:12 Uh again, maybe that's just because of limitations of today. But certainly, you know It has come up. mm in modern, you know, usage. So
5:22 We'll see where it goes, where the human computer interop boundary ends up being, but Yeah, it is crazy to do so much time. Like typing text. To get what you want? Yeah. I mean now here's the here's the flip side is I've been I the word I was looking for was productive. I've been
5:41 Order's probably at least an order of magnitude more productive. At least, if not more. So I'll give you like a number of examples. So I I run this open source thing called Mamehub. And it's basically a a network layer Built into this arcade.
5:58 Emulator. So if you and I wanted to play Pac Man together, Pac Man's a bad example, but night was that nineteen forty two, the game where you're in airplane. Yeah, like nineteen forty three. Yeah, I'm gonna go. So If we wanted to play one of these arcade games together. You you live in Florida, I live in Texas.
6:18 You can use Mame Hub, we'd fire it up. And it basically plays the emulator Um Yeah, it only works for emulators that are deterministic, right? But it plays the emulator.
6:28 And when you press up It actually cues that up command in the future for both of us. And how long how far ahead in the future it uses that command is based on you know, an estimate of our latency. So I try to fit
6:46 Um I try to fit our latency to a distribution. And then pick the upper confidence bound of that d distribution. It's like ninety nine point seven of the packets arrive on time. And Yeah.
6:58 Some s tiny percent of the packets don't and when that When it doesn't arrive on time we both have to wait. So if you if you um So if you say in the future My joystick looks like this and I don't get that future
7:11 that I have to wait for it and then you're also having to wait, et cetera. So that's the premise behind Maim Hub. It uses some statistics to figure out How much to schedule those joystick updates so that we both get them And we can play the game and it It runs.
7:27 You know, the same on both of our computers. So that's main hope. And I used to spend a ton of time just rebasing it. You know, because it's a fork of Mame. It's literally a fork of the MAM. Get project. And you know, they're making a zillion changes and you know, every few years or so
7:46 I would uh Bite the bullet. uh, you know, people would ask for it on on our Discord and I'd bite the bullet and I would rebase it and it'd always be a total disaster, right? 'cause I I didn't really have it. factored out maybe as good as I could have.
8:00 Um So now I just have a git. Job. I just have like a action and GitHub. That uses Copilot to rebase it.
8:10 And uh When there's conflicts. Whether there's conflicts or not, it like Even plays a game. And make sure that the game works. Um
8:18 And so that's something that, you know, I I told uh an LM to go build some GitHub action to go do that. And so Um Similarly, oh oh here's an even better example. So
8:31 This podcast What you all are listening to right now. Um, runs on a ton of software. Right, we have assembly AI for generating the transcripts. Um
8:42 We use uh Um Uh we have oh we have a transistor for actually hosting the site. All of that. Um and
8:52 With AI I just Basically brought all of that in house. I said, Hey
8:59 You know, right and by the way, assembly AI amazing. They've given us phenomenal service. If you need to transcribe something, And you don't want to ask AI to do it. um ask these people to do it. They're phenomenal. So this isn't a slight them or any of these people who have helped us so much for the years over the years, but Um I basically went to an AI and I said, Hey, here's a system that can do transcripts.
9:20 go on the internet, find some open source models. Generate some transcripts. And keep working until your transcripts look like their transcripts. And it took a while, but AI eventually did that. So so my point is Is uh you know
9:34 I feel like I'm building way more than ever. But I'm not coding. in in code. And so I think what we're doing is coding in plain language. I think that's really the the terminology. And so it seems like the future is gonna be how do we
9:49 code and plain language more effectively and how do we build these harnesses? Um that That Yeah. kind of uh make all of that as as as fluid as possible. So that's that's what we're gonna talk about this show. That might even be
10:04 Yeah, the most important thing for a for a while. Talking about important things. Maybe it's time to move to our uh our our news section and I think you got the first. I don't want to spoil it, but but uh you take it first. Yeah, I mean this is uh late breaking news. I'm sure nobody's heard of this But Ram is out of control. The ram prices are totally out of control. I think they're gonna keep
10:30 becoming out of control. I actually think they're gonna get a lot worse. Um And uh Um, and so this is just something to keep an eye on. I mean, this is a Tom's hardware article. I'm sure you can find plenty of others, but If you are
10:45 If you're saying yourself I'm not gonna buy a computer now because the RAM prices are really high. Um, you might have to just do that, just bite the bullet and and get the computer. Um
10:59 You know, I I just feel like the the uh The ram shortage is going to take years. And uh It's um it's not going to it's gonna get a lot worse before it gets better. So So if you if you have a computer you need to buy or an upgrade you make to your hardware or something.
11:16 Um You know, I really suggest kind of doing it now and locking in Yeah, even though it looks like a high price locking it in now. Yeah, it sounds like they're saying all the RAM is spoken for for like the next year over year uh, you know
11:31 even higher prices, so Yeah. Yeah, I didn't buy a Steam Deck as much as I wanted one because I always I actually this is this is kinda sad, but I went to the Steam Deck page. Probably quarter. And I got so close to buying it and I was like, Ah, you know, if I do that
11:50 I'll play even more games that I'm already playing. And uh So I didn't pull the trigger on it and now it's twice the price. So Uh I I decided that uh I'm fine with the switch. But uh but if I wanted uh I I don't expect the Steam Deck to come down in price, so
12:07 Yeah, I did get one, but now I feel like I should sell it. Oh man. Yeah, I had one from early on, but you're right, like I don't know how people are justifying it. Uh
12:18 Price it's you know what I mean? Like I I definitely yeah, even before it wasn't necessarily cheap, but I've I've enjoyed it. You know, it's been good. I definitely wouldn't have gotten the current price value out of it. It would have been too expensive. So you gotta really you know what you're gonna use the crap out of it, I guess. Well, Here's a question. I mean if the price of computers
12:41 Triples and stays that way. How is that going to affect you? Like are you still going to upgrade? Your computer, I mean, I feel like at some point you just kinda need to upgrade if it breaks or you know. I mean, I think it depends, right? So What does it mean like you need to upgrade your computer?
12:58 So for me, I mean I think I have a number of options. So For one, you could just like subscribe to the gaming streaming services. They have gotten a lot better if you have good internet. And some people will call that it, but Again.
13:11 If it's whatever, I don't know, twenty dollars a month, thirty dollars a month to to rent something, even one of the ones where you rent your own computer and install it, you know, on shared time or Like That's a form of arbitrage I've considered, you know, for running some of the local models. Like I think we'll talk about a little later. Like you could rent time. Um and time share. And I think that's That's one way.
13:31 You know, also we have lots of other like phones are more powerful. So some stuff can it's not ideal. It's just be done on like gaming on a phone. you know, uh consoles, if you're having a console, you know, sticking with older consoles and playing games you just wouldn't play before. I I think the people are gonna get hurt the most if people like want the max performance on the latest video games. Thankfully I just not in that race. So you know, for me short of like something straight up breaking.
13:56 You know, I I'm probably like you probably have like lots of electronics we could share between members of the family laptops or something if we had to. But you know. At some point, yeah, you would just break down and buy it if you needed it. But I think then you have like a There's a difference between it's broken and I have nothing and I need to acquire
14:15 uh a cost and I have something that is just less than I would want. Like you know what I mean like the uh the demand in both are are slightly different. So for me It becomes about that. But I've certainly Wanted to upgrade the RAM.
14:27 And I just waited too long and like That that window is closed. Um, you know. At this point upgrading the RAM the rest of the system is just basically free. Like that's true. No hard drives and uh specifically like SSDs and MVME have also gone nuts.
14:45 That's not just RAM. But if you try to buy you're talking about like the Steam Deck, I wanted to get a bigger SD card and I hadn't paid attention because everyone talks about RAM. But even S D cards have gone through the roof. For people, the shortage of NAND chips. So
14:59 Yeah it's pretty wild. I think uh Um Um yeah, I think it's probably a good time to replace your phone too,'cause again, phone has a has a lot of RAM in it. I I had an issue where my phone has started to degrade. It's about five years old. And um I was like, Yeah, I should it's it's time. Yeah. Plus I was worried about the prices going up. And so I got a phone actually arrived yesterday.
15:21 And about two or three days ago the GPS chip like completely failed. On on my older phone. So I'd I really got super lucky. I timed it. where a GPS ship fails next day the new phone is there. But there's there's about a day or two where my kids thought it was hilarious that we were we were in the car and the The blue dot was just like teleporting.
15:43 All over the city. And uh kids thought that was the funniest thing, but But uh Yeah, I mean even for a phone, I would say if you're on like a four or five year old phone You know, and you could and you can you can stomach it, you know, get the phone now before the prices go up.
16:00 My next article is not Per se about AI. But it's kind of about AI. So this is an article titled Mario meets
16:09 Perito. And this is about the Pareto frontier, the efficient frontier. But It's a it's in uh everyone's favorite competitive video game. Okay, that's a lie. It's the only one I play competitively with my family and not online because I'm not that good. Um, but my kids think I'm good, uh although not anymore because they've gotten a lot better. Okay, Blue Turtle is Pareto dominant, right?
16:31 Uh oh in terms of weapons. Uh but But it's basically going through which is a very actually, you know, interesting thing, even if you don't care too much about Mario Kart. But the setup is you have uh some number of characters, I don't know how many it's like sixteen characters you can choose from. Each of them have some number of uh you know vehicles they can be in.
16:49 you know, go karts or motorcycles, they have Some kind of tires you can put on and then some sort of like uh glider device in the uh more recent most recent, no, nearly most recent Mario Kart I think Mario Kart eight. Uh deluxe. Um, and each one of those changes various parameters. There's also some hidden parameters that that even aren't shown. Uh and so the question is like, how do you pick?
17:12 And the answer is there is no one right answer, you know. Growing up it was always, oh, you pick this this person or this car or you know, this is the team you pick in Madden, you know. The answer is like more complicated than that. Depends on like how you define, and this is something that comes up a lot in software engineering, which is why it's kind of an interesting article. Um and also in finance, but like what are you optimizing? Four.
17:35 Um, and so if you think about to to to spin it in terms of like finance Is it better to buy stocks or better to put your money in a savings account? And it's like well How much do you care about risk and how much do you care about, you know, growth? Because you know, there's no free lunch, right?
17:51 Uh okay, that's stealing another finance idiom, I guess. Um economics. But The idea is that There are Strictly worse.
18:00 positions you can be in. So there are things you could do with your money. That are more risky For less expected payoff. And you should not choose those things.
18:09 Just like in Mario Kart, although there is no singular best answer. There are certain combinations which depending on what metrics or even all the metrics, choosing them is strictly a worse choice. Unless you're just trolling. Um, I guess that's not a metric. But against the normal metrics, you know, there are certain combinations that are strictly worse. So if you care about
18:28 you know, the trade off between two variables like acceleration. And top speed, which are very common ones. then there are you know combinations which are not are strictly inferior to other ones. And so the idea is anything that lives For some combination not
18:46 in the interior, not inferior. We call those things Pareto efficient or on the Parrito frontier. Which means they're they're sort of like normally in the graph to the right and up, right? So there's this curve along all the choices. And this applies through so many things we do. Like I mentioned, trading off what you put in your investment portfolio.
19:05 How you choose characters in video games. You heard a lot coming up about Lodels. So their L M models have various Abilities. And various price per
19:14 Token. And that forms a efficient frontier. So you know, you may argue whether Open AI or Anthropic today has the the best model. They certainly don't have the cheapest. So you can go
19:27 Cheaper if you need less. And there's like a different position, but there are also places you could pay just as much. For crappier AI. Right. So there is a Parrito frontier. And when we talk about improvements, we're talking about moving that frontier
19:41 Four. moving to a new spot that no one has been before. And forcing others to play catch up. So A very interesting, well written, uh and lots of cool graphics. You know, everybody loves a good video game story if you've played Mario Kart, which lots of people have. You can definitely read this, but also an introduction to an important concept.
19:58 Yeah, this is super cool. I always pick to Bowser. And now I'm learning that that was Not Pareto dominated by Donkey Kong, apparently. So That's awesome.
20:12 Yeah, check it out. Yeah, this is a great uh visualization of a really important concept. I love this. Um All right, uh mine is Quinn three point eight versus Muse Glimmer. So
20:26 It's interesting, like the sizes of the models have gone through different, you know Like uh hype cycles over time. Um so You basically have the people who are trying to build the biggest models. And so currently, you know, these are you know the close
20:43 Closed labs. And then and then people like uh GLM Or like Kimmy K three, which has I think a two and a half trillion parameter model. And so um Uh, yeah, and so at least for Kimmy, if you wanted to serve Kimi K three, you would need
20:59 I think like two hundred thousand dollars worth of equipment minimum. So Um So you know definitely out of reach for for almost everybody. uh on an individual level.
21:12 Um And and not practical for most companies unless you think you can uh really serve that model, uh, you know, keep it busy twenty four seven. Um Yeah.
21:24 Then on the flip side you have models that are very small that are meant to really run on edge devices. So for example you have the GEMA four E two B. model And the E T B means effectively two billion.
21:40 So There's basically a router and without getting too much into the weeds here. There's kind of a router that decides which experts and which modules to activate. And um Um
21:52 And and it's it's I don't know where they get the word effective from, but you know, on average, you know, they're activating about two billion of the weights. Um So Um Uh so that can run on your phone.
22:05 Especially after you quantize it. Yeah, even with the K V cache and all of that, you know, it can run very comfortably on your phone. Um So what's emer what's emerged is kinda the sweet spot in the middle where you know, it's too big to run on your phone or even on an entry level.
22:21 uh machine. So imagine a machine with like sixteen gigs of RAM. Uh you know, these models are too big to run on those. But They're small enough to run on a single high end machine.
22:33 um uh cur consumer high end machine. So think like a MacBook Pro with forty eight gigs of RAM. Um you know, or or a Windows laptop at sixty four gigs of RAM, et cetera. Um And so Muse Glimmer and Quen three point eight. Are both uh
22:50 Around thirty billion. I think the Quinn one is twenty seven. Um And so, you know, as I say, if you have a high end laptop or desktop, you can run those pretty comfortably. And uh they're very, very powerful. Now I've tried it myself. I I haven't quite got the level of performance for the tasks that I'm doing.
23:10 Um that people are claiming in the benchmarks. Mm. So Um, so it does feel like maybe the Models not quite living up to the hype there, at least for Quinn. Um
23:23 Use Glimmer, I think is pretty good. Um the the big thing that I've noticed from these smaller models is um And this is kind of interesting is is they don't Necessarily call the right tool at the right time. And they also don't really know when to stop. Working.
23:40 Um Um so so they'll work for too little or they'll w they'll think for too long. And I found this really fascinating, like Because if you think about it, what is a bigger model Well, one of the the most obvious thing that it is is
23:54 Is is more parameters and so more facts. Right. So if you say What's the capital of France? And it says Paris. Well, there's some parameters or some set of parameters in the model That are responsible for knowing.
24:08 That fact. And so it stands to reason that you know if you have a giant model, it knows a lot more facts and it can bring a lot more facts to bear. on a particular problem, and that's true. But it just Seems like the critical thinking is better.
24:23 And I don't know if anyone's really The closest I can imagine people have come to explaining this is this J space. theory paper from Anthropic where Uh models like project problems into some high level space and then try and solve it in that space. And so
24:40 Maybe the bigger models have a bigger space to work with. Um But even things like browsing the web. Um And doing kind of pretty rote things on the web, like pretty easy things that that you know even a a child can do on the web.
24:54 Uh you know, like change the dates. on this uh set of assignments or something on this website. Um the smaller models will Will tend to struggle with that. Um, so it's just on the common sense level there is a correlation there, which I found interesting. But
25:09 But uh Um, but the models are getting better and better. These thirty B models are definitely a huge, huge improvement over the nine and twelve B models. that people are running locally. uh you know a year ago. So definitely a ton of improvement made. And um I'll just end it by saying
25:26 I think ninety percent of the questions people are asking can be handled by this 30 B model. You know, like what are the kind of questions people are asking ChatGBT? Yeah, how do I fry an egg? Uh I I yesterday I took a picture of my kids schedule. that they gave us on a on a sheet of paper.
25:44 I said, you know, create a calendar ICS file for this. You know, all of these things a thirty B model could do, I'm confident. Of it. And so So it's like on one hand the gap is is there. And uh it seems like it might not be
25:59 Easily closed. On the other hand You don't need the more expensive model for ninety percent of tasks. Yeah I I agree. I mean it it's definitely an interesting space. In fact, one of like the most interesting spaces because
26:18 I saw I I won't I don't know, I don't want to go too much into like personal beliefs, but I still see this just like Oh Hi, our you know, chat GPT, you know, Claude you Gemini app is happy to ingest your data. Would you like to share your calendar your health information, your whatever, and it's like
26:39 No You know, whatever. I I again I'm not gonna get into privacy as like a thing that's you know personal And different for a lot of people, but
26:51 I local models to me are a way where you can say, Look, I I can do this, there's not Like I'm fine doing what I'm doing here, you know, but it's staying on my computer. Yeah. And even just something like Hey, I have a can you go through my tax returns and just make a plot of like my effective tax rate over the last few years? I would love to know that. But I'm not
27:10 Upload it. One, and that's like a lot of data to upload. And two, like I don't really want to expose that level of detail For you know Who knows what reasons, what happens, Now all of a sudden you ask how much does Patrick Wheeler make or whatever? And then all of a sudden it's like, Oh, I know that's my dream.
27:27 Um Yeah, so it's just like I think local models like important thing, and if you've not tried it, it's definitely worth trying. And like and like you said, I think, Jason, there's a lot of like questions that Yeah, it it just
27:41 I I mean, I guess like to flip it around, Google search, like Things I just couldn't search for before or whatever, just get answered super fast by but I assume it's a relatively cheap. you know, Gemini model that Google is using. Um and so a lot of those things can be done.
27:57 And I do think the the mixture of experts, the effective ones like you're saying, are really interesting in that They give you opportunity for streaming from like an SSD into RAM if SSDs weren't so expensive as well. But like streaming the data off of disk and into RAM for like that stage of of inference. Um, and just in general, what they mean as well. But I I'm excited for that level of like
28:22 what you can do and just speed as well. Like Oftentimes you get a response pretty quick. Whereas if you're If you've ever used one of those other services extensively, sometimes I think they just get throttled. Like there's just too many people using them. And so the responses are just so slow.
28:37 Um here it's a single user thing, right? Like I I just get the response when I want it. Yeah. So I tried uh both of these. I tried Muse Glimmer and I tried Quinn three point eight pretty extensively. Um Mm.
28:53 I found that For and and this has been always a problem with these Quen models. Is is there there tend to really struggle with text and formatting text and and outputting text. correctly. So for example Um, I had both of them
29:09 um work on the same task which involved creating a lot of markdown. content. And um What I found is that the Quinn three point eight model, um, it wouldn't close the bold.
29:23 So you know how in Markdown you double asterisk to make something bold? So it'll start something bold. Like a word, but then it won't close the bold. And that they'll get real confused and you'll either end up with like a giant amount of bullet content or it'll just literally draw the two asterisk.
29:39 So it just wasn't able to respect the formatting of Markdown. Uh Muse Glimmer was Was able to respect it. Um The other thing I noticed
29:52 Personal anecdote is is um Olama, which is which is what I was using to serve these models, it has an MLX. mode for almost any model. Like for every model you can put dash MLX. And it will If you're on a Mac OS machine.
30:09 It it'll use like M L X, which I I have to admit, I have no idea what it is. But it's some kind of thing on the Mac that lets it do these vector operations faster. But what I came to the U Find out and I and later on I
30:24 Um uh confirm this is The way Olama does the MLX They Um, they have some way of automatically converting a model to MLX.
30:36 And I don't think that that is a hundred percent reliable. There there's something lost in that process because When I compared the MLX version to the regular version, the regular version was a lot better. So um So I guess all of this to say at a high level
30:55 It's kind of like uh Ender S one three D printing days. You know, or We're beyond like the Oh, you have to build your own three printer from scratch and it probably won't work.
31:07 But we're not quite at the like Oh, I buy a bamboo. And everything just works. So we're somewhere in the middle. where these thirty billion perimeter models are powerful and you can see the future right around the corner.
31:20 But It's not just like drop in yet. I'm here for that analogy. Yeah. I figured I figured you'd appreciate that.
31:32 All right. Uh my next one. Uh sort of Incongruous, I think, with a lot of the stuff we've been talking about so far, but that's okay. is everyone should know sim D. So we talked about SIM D uh
31:44 number of episodes ago. Someone's gonna be like, oh it's a hundred episodes ago. Um I don't think it was. Um SIMD is a single instruction multiple data. Uh and in line with I was saying before, this specific example is actually about um zig programming, which I've never done zig programming. There's some nuance there. But in general, just like I think we talked about in the other show. Just walking through like what
32:09 the pro what it is, um and why it's important. I think this is an example where Yeah, at least today. It can be difficult to balance Um, you know.
32:20 implementation I'll even say with the LMs. Of doing something one way and then like completely redoing it in another. So when you build a program sort of normally, the most common thing is to have an array of structures.
32:34 Right. So you have an array in your array, you put your data That's interleaved as like, you know, field A, field B, field C, then field A, field B, field C. That's very common and that's how you would for loop through it. It's you know how an L M would write stuff by default. But
32:50 understanding when and where sometimes you want to rotate To uh structure of arrays. where you say I have all my A channels together, all my Bs together, and then all my Cs together. It's a huge unlock in some cases.
33:06 for doing things like N Cymdia and and other optimizations. And a lot of times Um If you don't know that's possible, you don't know to dig into it. And there are lots of libraries. But again Compilers can do some SIMD.
33:19 But There's only so much uh range they can act in, right? So they won't refactor they won't recompile your whole program. Uh in order to get C optimization. And similarly, I think By analogy, even in the LMs, I think this is where, at least for today, I'll I'll say
33:36 I think understanding to ask for it because It's not something that is gonna know To reach for Immediately. And it's also
33:44 Trying to which we all want, you know. We don't want them spending extraneous tokens just doing constant refactoring and and trying stuff, except when we do. Um but like, you know, they're trying to limit it. So doing massive refactors like this and not knowing is so sometimes you gonna need to push it to say, look, I want to implement some D optimization. I want you to build it in this way. And when and how to ask for that and whether it even makes sense is something
34:09 you really gotta think about what hardware are you running on, what are you and these are things I'll say to the normally doesn't figure it out. Like you normally have to tell it. You know, like hey. You know, I want you you know, write a simple like I was doing something with uh just like a a cheesy ML something. Um, and it was like trying to run CUDA stuff and I don't have GPU that's NVIDIA like on my do it's like why are you like stop.
34:30 No, like you didn't even think to check it. So See this blog article, but just I think to talk about how useful SIMT is and a very efficient thing. Um and the instruction sets have gotten better over time.
34:46 From when I first use them to now they're they're really kind of awesome and libraries use them. But also as a means to back into talking. to I think this is the kind of understanding it still is important to have. Um, because I think it's something where we still need to help give direction. to uh what we want out of our programs.
35:04 Yeah, this is a good point, you know. I think that the Um the LLMs They make a lot of category errors. And by that I mean um just quick recap, five second on what a category error is.
35:17 The common example is Um You give you give a student you give a child a a tour of a university And you say, Here's like the math. Building.
35:26 Here's the computer science building, here's the literature building. And then the student says, Well, where's the university building? And and it's like well no, the university is like an abstraction. That's like you know, a collection of these physical buildings. It's not a building itself, right? So that's a category error. And
35:45 I've noticed the AI is uh a lot of these AIs really struggle when it comes to uh to to categorization. So for example Yeah, you'll tell on AI. Make this code faster. And
35:58 What you really want is for it to You know, find the Hot spots. And rewrite them in SIM D. But what it actually does
36:08 Is like Delete like the slow parts of your code and now it's not functional. Or or you know rewrites the whole thing in Rust or something like that. Like it rewrites the entire program, right? So it's like it doesn't know at what category at what level to operate.
36:25 Um And so this is I think You people people use the term taste. I think taste is something different. But I think independent of taste, I think
36:34 Knowing at what category Um you need to operate to solve the problem you're trying to solve is still kind of Firmly in the realm of of of human beings. And so Um, and so you have to know that
36:48 Hey, here's an option, you know, if you have this Python code. Um Uh, you know, one option you can do is Do you take the slow parts of the code and use Cython uh you know, to to bring them into C and then use CIMD after that.
37:03 Um, as opposed to like rewriting the entire thing. I think it's time for book of the show. I've got the first book. This is a book I have and am getting ready to start reading. So unfortunately I can't uh give the the full review yet. Um, but this is Waves in an impossible sea by Matt Strassler.
37:28 Uh, and this is a book uh that's I don't know how you call it. We've talked about some of these books before. have a name it's slipping in my mind now where it's like sort of a a casual read sort of like just a compelling uh story, but talking about facts. So I think Simon Singh does this for like for Maslast Theorem, um and the the code code book, cryptography book.
37:49 Where it's like stories, but through the stories he's explaining like the history and the operation of something, uh, in sort of like layman's terms or whatever. Uh anyway, so this is is that. about sort of subatomic uh physics. And I do not know why through my entire life I've just been sort of like fascinated. by elementary particles, subatomic physics, even though like it's not my background. I never ended up studying it. I know really nothing about it. I won't talk about it because I'll like horribly screw it up.
38:17 Um, but just like the way that people describe there's a couple of YouTube channels where People talk about like uh the sort of math and physic behind uh sort of like discoveries uh uh you know, why's uh spectrographic lines were, you know, uh befuddling people in the early twentieth century. Just like all these things and I don't know, it just always really uh excites me. So I'm I'm happy to dig into this book, uh, and just sort of like go on adventure because
38:43 For whatever reason, this domain has always uh is a nonfiction, I guess I should clarify. Has always been something that Uh is awesome. There's so much more going around at like a very, very small level that that we kinda just never think about. Um And it's it's kinda crazy, but it's it's cool at the same time. So is this a is this a fictional tale that explains a real phenomenon or is it a non fiction book?
39:06 It's a nonfiction book. It's not like an allegory or something. It's just uh a sort of Talk through Uh in in a in a engaging narrative way. Through the history and state state of these things. Oh, very cool.
39:21 Um Yeah, I mean on my I have just been diving into so many research papers that are Mm. Not barely interesting. To me and my friend Group.
39:33 Yeah. So I'm not gonna like uh uh make everyone suffer through the the list of research papers I've been reading. But one of them in particular I wanted to talk about So Okay. When When we started with natural language processing with neural nets.
39:51 Right, people started with these recurrent neural nets. And so the idea is Think about like a memory. A block of memory. But instead of it storing very explicit things, like you know, if you're to store the word dog in memory using Python or something.
40:08 Yeah, it'd be a bite for the letter D, a byte for letter O, a bite for letter G, and now you have dog, right? But instead of that, it's this really abstract, really high dimensional space. Where you're storing uh you know, a whole library full of concepts. And then you can sort of pop things off of that. So
40:26 You know, uh the early transformers did this where You had an encoder. That took What the user is asking for and the
40:36 you know, partial answer, which could be nothing at the beginning, right? Encoded it. Into uh you know this memory bank And then a decoder.
40:47 That popped off Um um, you know, tokens. off the memory bank and then also updated it. So So the idea is
40:56 Um, you know, user input comes in, typically not with a partial answer. So user input comes in, it gets encoded in this memory bank. And then there's just this really tight loop. It's like Based on the memory bank pop off You know, the first word of the answer. And then
41:12 mutate the memory bank. And then ask it again. And just keep popping words off the memory bank. Until you pop the end of answer word, which is like the special token, right? Um
41:26 And so the the problem with that is That you have to compress Your question into this memory bank, right?
41:37 And then as you're popping words off the memory bank, you're also kind of Having to store what you've popped off so far. So If you say what is the capital of France.
41:48 That has to get crushed into this vector. Right, that represents that question. And then if if you pop off, you know, the capital is Now that same memory bank has to know your question and know that you've already completed part of the answer. This is why early transformers would would do things like
42:06 the capital is is is is is is because it couldn't figure out how to Store. the the the fact that it's already said that word, right? Um And so now transformers are decoder only.
42:21 Which means They Um Uh what goes in is is the question And
42:29 You know, the part of the answer you have so far. And what comes out is a single word. So like what is the capital of France? What comes out is the And then this whole thing starts again. So like All that work went in just to say the word the
42:44 And then start the whole process over again. You know, what is a capital of France? The and then goes through this whole process and then outputs the word capital, right? And then so on and so forth. So Um now they've they've used K V cache and a whole bunch of tricks so that
43:02 This doesn't waste that much computation. But it's still pretty weird, right? I mean, as a human being, I don't think we really operate this way. I think yeah, we internalize the question and then we sort of roll out the answer. So we've kind of deviated. From Uh
43:17 the way a normal person or the way we expect a brain to to think. And and reason. It's pretty unnatural, right? So people are constantly trying to go back to this encoder decoder idea, and world models Have to be encoder, decoder.
43:33 Um Just because of their nature. We talked about world models on the last episode. And so the question is like why How can you encode these things better? So you don't have to do that.
43:47 Really expensive thing that we said earlier. Um And I think part of it is Yeah, the memory is flat. Right.
43:55 And so on your computer You know, when you do malloc or something like that and you get a flat block of memory You can then go and put an image in it. Um Because you know, you're storing some metadata and as a programmer, you always understand the code is an easy way to remember.
44:14 Oh, this huge chunk of This huge line Of memory. Is actually a two D image. Um
44:22 But I don't think that the LLM internal Is really good at that and so I think what we actually need Are You know, two dimensional, three dimensional, like n dimensional hypercubes.
44:36 of information that they can read and write to instead of just a line. of information. Um And so this paper kinda talks about that. So this is a
44:47 A paper where instead of a single line of data that you can write to You can now write to a grid or or a hypercube, and and the space kind of matters. So if you write something to the left side of the cube. And then you write something else to the right side of the cube. Yeah, those those those two locations are far apart.
45:07 And it actually matters. Um, it makes it harder for them to affect each other. So Um So so yeah, so so they basically are using in this case diffusion. Other people have used convolution, there's a bunch of different ways to do it.
45:23 But I think that Storing information this way is going to unlock Something really powerful. in uh in LLMs and a world model. So It feels like there's something cool here. We're just kinda on the cusp of it.
45:38 That's exciting. Yeah, I It feels like There's this play, but it it hasn't been true almost where like traditional I think there's on the output of LMs like this DSPY sort of like instructions and things like that for attempting to kinda get to it, but almost where like you're talking about like memory storage or whatever.
45:58 Like Two very limited tool invocation within the you know, actual inference itself, right? It's like Know that you can put stuff here in this way.
46:10 Uh rather than just like learning it completely. But You know, I that complication becomes uh incredibly difficult. So You know, I I it is definitely exciting though for them to figure it out themselves and potentially even a better way of doing it, I guess, than than maybe naively We would make tools to do it.
46:29 Yeah. Totally. Oh. All right, time for Tool of the Show. Patrick, what's your tool? Uh this is a very no, I'm just kidding. I was gonna s try to make it funny, but it's the same every time. It's a game. This game is available on many platforms. Uh PC, phone. I think it was popular for a little while. I'm probably like past it, but I think I picked it up on a
46:50 On a sale and it is Nubby's number factory. Which? Like the game is pretty good. Like I got pretty into it. I like it. It's a very casual Think it's called like a Plinko like Uh roguelike.
47:03 Okay. Is that where you drop a thing and it pings off the ball? Yeah, yeah, exactly. Yeah, it's pings off numbers. Uh you try to make number bigger and then there's various power ups and that are randomly chosen and you know, your your your exponential growth in the what you need to hit at each level makes it very hard. uh very careful. There's not a lot of skill to it, I would say. I mean like a little bit, but that's not the the kind of main point. Um But the aesthetics are are trip.
47:31 It's like a trip down nineties, you know, Web One point oh, I guess that's like mid nineties, like It's so good. Just like You know, just I don't know, just pull up a screenshot of it and if you you'll know instantly if it's like this is this is uh this is my jam or not. So for that reason alone, just like I don't know. Decent gameplay, but incredible art aesthetics.
47:53 Uh but I mean horrible, but like amazing. It sounds like Bellatro as far as the gameplay. Yeah, I mean same same kind of idea. Like a very basic game, like you're trying to play poker hands. But then like it's scaling. So you have like it's really all about the power ups. Yeah. In that way I guess it's kinda similar.
48:11 I I mean definitely it's not at the Bellatro level. I was watching people do like speedruns of Nan Imp, like Okay. That stuff's crazy. Oh god, yeah, the Belatro stuff. Yeah, I still haven't beaten Belatro. Like even in one deck, I haven't gotten through all the challenges and uh Then I watch these people and they're like Oh yeah, I have one card. I play this one card and I just win.
48:36 But I found out a lot of them are apparently using various mods like I and then you know saying they don't or Just to Titans. They really want to play a certain card, so even the like when I was watching, which yeah, interesting but I was like
48:52 Trying to get Nan E M flag very quickly. And so there's a certain card that they want. In order to be able to do it. And so they were like doing a seeded run where they had searched for seeds that would guarantee that card comes up. Within like the first or second shop or something. Um
49:09 And so they've never played that run before. But they know like the card they want is going to be there. Got it. Interesting. I I'm not saying everybody does that. I just like it's more common than I realized. So yeah, the game's actually not that easy. We're being diluted.
49:25 Well, I think there are people who just play it a ton. Uh, which is probably like it the same like Solitaire. Like not all games of Solitaire are winnable. Um yeah, it's probably like that. Like I'm not saying not all games of Black are winnable. But some you know just you're gonna have to play a certain amount of times in order to get hundred percent completion, like no matter how skilled you are.
49:44 Yeah. That makes sense. Makes sense. My tool to show is open code. And actually, Patrick, earlier you talked about doing your taxes with an LLM. Okay. I haven't done Noes. Reviewing my taxes. Reviewing your taxes. Reviewing your taxes with an LM
50:04 The IRS is on their way. Do Patrick, open up. Um So I Um
50:12 Okay, so our tax accountant, which is this really nice lady that we met Through church uh a long time ago. uh retired. And um And so
50:24 We had this debate in the household. I think that an LM can do my taxes. Oh, and Yeah, and and uh I'm the only person in the family who thinks this. Um
50:35 So I was like, okay, we'll do this. We'll we'll have uh somebody do our taxes. Um And I will also Do our taxes with an LM.
50:47 And then we'll see If they match up. Or mine's better even, right? And if they match up or if it's even close. Then we know that we're we're good. Um
50:57 But similarly I was thinking, you know, taxes, even like controlling the web browser. Where like all your cookies and your passwords are stored. Like installing the anthropic Chrome extension. Or having, you know
51:10 uh uh chat GBT do my taxes is where I draw the line. Yeah, like that's like yes. Yeah. Like and I feel like I'm I'm pretty You know, open when it comes to privacy. But but I feel like I draw the line at, you know, these kind of things. And so
51:27 Um And so I downloaded open code. uh, which is one of many harnesses. Uh I definitely don't claim uh that open code is better or worse than any of the other ones. Um and maybe we should do that could be a whole show, but uh we'll I could do a deep dive on that. Um but at the moment I'm trying open code.
51:46 Um with uh with these Muse and and Coin local models. And um It's really, really good. Um, as far as the the harness goes, you can spin up sub agents, you can do loops, which we'll talk about. Um, all the things that you can do with with claud code are pretty well supported.
52:03 I'm with open code. There is a little bit of jank. Um There was a situation where Um Muse like
52:12 When it returned the tool response, it wasn't quite formatted correctly. And instead of you know, dealing with that or retrying it. Uh open code just hung. Indefinitely, and so I had to kill it.
52:24 Um so definitely got some jank. You know, got some J. But But uh Uh yeah, the fact that it's free and open source and uh Works reasonably well.
52:35 Um Uh, I think it's pretty nice. I ended up writing a very simple like watchdog timer. to uh Um, to kill open code if it's processing and it hasn't finished in a certain amount of time. And restart it.
52:49 Um But again, I I think it's just like we talked about with the thirty B models where Yeah, we're we're maybe six months to a year away from having something that's really polished here. And and open code just seems to be very popular. It's got a Python SDK. If you want to Non interactively.
53:07 um, you know, ask a bunch of agentic questions. So Overall really powerful piece of software. The other one I I mean, we like you said, we may have an episode about it, but was the pie coding agent. Which tries to strip down the harness to like kind of bare minimum and then have you add to it what you want. So you ask Pi to like improve itself.
53:28 Which I think is really interesting. I haven't gotten into it yet, but it's it's on my short list and so I wonder if something like that Could be for you know Open code. by the name, it's probably built mostly for coding. So if you wanted to use it as like a a multi purpose harness, I wonder if something like that might, you know, also work.
53:46 Yeah, totally on the same page. Yeah, I want to desperately try Pi. I've heard a lot of good things. Uh I went to lunch with somebody uh about a week ago who's really into it. Um, so it's definitely on the list. So I can try it and we can do a show on it. Okay. All right. Sounds good. Uh but you left us hanging. Who was right?
54:05 Oh well so uh so so Janet, you know, our friend, she retired literally like a week or two ago. Um, and so this would be for next year's taxes. I could back test it. It's a good idea though. Maybe I should try to to my current year, you know, the ones that were already submitted. Try to do that without Without hindsight.
54:26 Um, that way I wouldn't have to wait until January or February to to try this experiment out. Oh, one thing I did do So um Uh kind of related. I downloaded CSV files uh for all my credit cards and
54:41 And checking accounts and all that. So I got basically uh all my transaction data. Over the past ninety days. And I asked uh through open code, I asked Muse, like, hey, what you know, give me a rundown of our expenses and what are we spending too much money on? Where could we cut down? Things like that. And it did a pretty good job. Actually a really good job. Because you that's one the LMs are really good at. If they see
55:04 Like a bill that just says chilies, it knows that chilies is a restaurant. Right, like that kind of stuff. Um So so uh that actually w turned out great. Um, so I'd highly recommend I mean, maybe there's even something there around
55:19 You know, like uh open source project or something where someone can just install some desktop app. And uh I don't know how it would connect to their bank. They'd probably have to do that part manually, but But some desktop app that just like goes through all their finances and You know, just for people who have a hard time setting up open code and all that.
55:39 Yeah, I'm more to say here, but well, let's let's go ahead and go to our topic. All right. No. We should do a show on uh yeah, automatic finance management.
55:51 Um Oh, we're burrito efficient. B Oh man, I I it's you know, I've had the same Twitter handle
56:04 For like I don't know, twenty years or something, but I think I need to change it to burrito fishing. Oh man. So good. Uh all right. Agentic loops. So Um a bit of a history lesson here, so
56:19 Um, modern history. Um Okay, when LMs first came out Uh when the technology first came out. Um
56:29 People companies were afraid to roll it out. Um, I don't know if you remember this era, but But there was definitely an era where You know, companies would Skip their toe in
56:42 The chat bought water And there would be so much backlash. People would find the worst possible thing that it could say, the dumbest thing it could say. And that would make the news headlines Um the one I'm remembering of specifically Um
56:57 was there is a Facebook I think it was called Was it called Liberatus? We'd have to look this up. But there is this Facebook project where It would create research papers. So you would say
57:09 Um, you know, here's a bunch of uh uh related work. I want you to write a research paper on uh, you know, spatial neural nets or something. Right. And it would go off and write a five page research paper that looked like you could submit it straight to ICML.
57:25 And so people wrote ridiculous research papers, like, you know, they would Like write a research paper on running a computer with hot dogs. And it would do it, right? It wouldn't push back. And then people would Submit it. Uh not submit it to ICML, but they'd submit to like the New York Times or something. It's like look how dumb the Facebook font is.
57:43 And Facebook pulled it down. Uh they were embarrassed they took it down. Um And so nobody wanted to release the chatbot. The chatbots kept getting better and better and better. Nobody wanted to press the green button until OpenAI pressed the green button. And when the initial chat GBT came out They got
58:03 And so much there's so much flack for it. Right. It's like oh this is evil. It's gonna teach people how to make bombs, it's also wrong. It's telling you to put glue on pizza and eat it, right? All these things. Um And to this day
58:17 You know, whenever you go to any of these websites, the first thing it will tell you is, you know, Gemini makes mistakes. Right, ChatGBT makes mistakes, right? I was driving behind a sema. And it's like sometimes You know, the autonomous vehicle makes mistakes. I'm just kidding about it.
58:34 That's terrifying. We just have come to terms with the fact that These models. Make mistakes. Right.
58:45 And Um and even if they don't make mistakes, again there's the category error. Where You know, it did technically what you wanted, but it didn't follow kind of the spirit of what you know, a common
58:58 person would would have expected. Right, for that question. Um And so In comes agentic engineering.
59:07 And the idea is Um We know we're not gonna get it right every time. So what we're going to do is feed in our answer. Back in the year.
59:16 and see if there's more work to do. And just keep doing this until we reach some kind of a conclusion. And and that's what uh, you know, Claude Code. kind of brought into the mainstream. Right.
59:28 Um So Um So that all works out well. And the but the question there is when do you stop? And again All of the same errors from before are still there. So
59:42 So Um if you might you might say something like Hey, uh Get to ninety percent unit test coverage. Uh and this is, you know, this is about a year ago. It's not really true now because of
59:54 Loops. You say get to ninety percent test coverage a year ago. And Claude Code would Write some unit tests. And then say, Hey, I got up to thirty four percent. Isn't that awesome?
1:00:05 I'm gonna stop here. Right. And and if you wanted ninety percent You had to Tell it to do it again.
1:00:13 And there's I literally about a year ago At a bash loop. that asked Claude to get to ninety percent test coverage in a loop. And just ran that for twenty four hours to get to ninety percent test coverage. Um
1:00:27 So that is extremely primitive. Agentic loop. Um That worked. It was successful. Um and nowadays if you say something like get to ninety percent test coverage
1:00:38 Claude code recognizes That you've set a goal And converts your question into a loop. And that's how uh it's able to run for much longer nowadays.
1:00:50 Um So Um sometimes the LM will do it for you. Um, oftentimes it won't. Um and and again, even if it does, it might
1:01:00 choose a a terminating condition that's not sort of in the spirit of what you are thinking of, right? If you say something like Um, I want the Um I want the accuracy of this model to go up. Well, like what does that mean? Do you want it to get to ninety percent? Do you want it to get to ninety nine point nine percent? You know, if you let the LM decide that, then
1:01:23 Uh, y you don't really know what you're gonna get. out of that. So the rest of the show we're gonna talk about how you can build Um how other folks have built these loops, how you can build these kind of loops and and best practices there. So I think T
1:01:40 Like a half step back. I think one of the earliest things people were figuring out and this was popular with things like Lang Lang chain and what is that other. Yeah, yeah. Where there were You wanted to and it is kind of a loop, like
1:01:56 you know, poll, watch an email queue. And any time there's like a customer service question coming in and you want to like Label it, tag it. dispensate it, you know, whatever, and move it through. Um, and I think for me this this just like turns into sort of like cron jobs, but Is actually a big unlock.
1:02:14 Which is You know, hey, I if if you have a computer system, like wake up and ask the LM every so often if like the price on this, you know, website has changed. You're like, Oh, you can write a scraper for that. Yeah, try to write a scraper and then just watch like it not work because They move the stuff and the tags change.
1:02:33 You know, versus giving it a screenshot, you know, and having it convert it to text and then do it. This is like a big deal. Um, and so just replying to recurring events or doing the same task, which again I think is a half step before you know, the introduction. But I this was like one of the earliest
1:02:49 Places where You know, we saw like the looping L M. Sort of Take over. But then to transition to to more
1:02:58 You know, what what Jason said. For me, the big one, the sleep there was Was it like the Ralph Wiggum loop? Oh yeah, that's right. Yeah, where they would like write the special skill right or like do the same thing over and over again. Like You know
1:03:13 pick a you know, bug and and you know, try to fix that bug. Pick the next bug and try to fix that bug. But the first one to really like I think go I'll say viral with it was uh Andre Carpathy, which feels like he has an act for going viral. Um or maybe we only a survivorship bias. I'm not sure. I actually don't know what says otherwise. Um but
1:03:32 He did this thing where he posted Something that he titled auto research. I'm not sure if it was from the beginning or he eventually that. But where he took a very small I think it was nano GPT. Um, there's like a small GPT on a set training set. And said, you know, basically can you make this run
1:03:50 Faster, train faster. And can you try various fixes to improve the the metric? the output performance. And to Jason's point, the terminat wasn't really present. But just this desire to go again and again and try something and then revert You know, if it didn't work or it made it worse.
1:04:08 Now of course like It went viral because it was hugely successful. But there are lots of Gachas along the way as well. Like there may be you need two, three, four things Combine
1:04:19 So you know, you take a step back to take two steps forward, you know. Like there are all of these places. So you have to be very aware. you know, of what the problem scope being and what the domain is and how much lease you're doing it. But I do think there's this and we talked about a little earlier, like the tension between especially a cloud LM provider Not wanting to be accused of just evaporating your tokens.
1:04:42 Uh so therefore they don't want Loops that aren't going anywhere. And you who may be saying Like I'm not paying the bill. I want all of it.
1:04:52 And so, you know, you can end up with these tensions and so specifying something almost external to even the harness. Or even a lot of these stuff is incredibly useful because Then you're really crafting, sculping, setting Trying it, encouraging it, whatever you want to do.
1:05:07 and manipulating, you know, that prompt. And normally you would need a human sitting there. But actually like you can ask an LM between runs Shit. Like should we revert that? Does that, you know, make sense? Is there something further along those lines? You don't need
1:05:22 The pre planned complete trajectory to to go down. And so I think this is really Caught. You know, people's attention.
1:05:30 Um and now and I think this has been popular before, but now it feels within and reach. Just to kinda like the final maximal extent of this Is when these other limbs are If they can reach recursive self improvement.
1:05:45 So the idea is if one of these you know, major AI providers or a new up and comer gets an AI that can improve itself. by you know, ten, fifteen percent per training cycle. then they can basically just keep doing that without human intervention and it's just bottlenecked by
1:06:03 Speed, compute, power. Whatever, uh, you know, and then they'll just be on a quote unquote escape velocity. I still think that makes a bunch of assumptions. But you'll hear it, you know, tossed around. You we've heard A GI It kinda went by the wayside a little.
1:06:18 um the artificial general intelligence. Um You know, but now you'll have this RSI, this recursive self improvement where These loops are what is being done to try to say Can the LLM figure out what the L M needs? Uh and of course that
1:06:34 The evil laughter one is You know, hey, we need to make this thing more efficient. And then the robot decides the humans are the inefficiency. And so that annihilates all the humans because they're just consuming needless resources. But yes, I well, you know, that's the the the the the terminator scenario, I guess. Yeah. Um yeah, it's it's uh Um
1:06:56 It definitely lends itself to going off the rails. I mean I think Um like Leading up to this was this idea of in context learning. Um, there's actually a paper that just came out not that long ago. Or somebody showed that
1:07:11 If you add Periods. To the end of a question. Or it's the end of a prompt. that it um it it actually get better answers.
1:07:21 So and the more periods you add, the better an answer you'll get. Um and the explanation was Well, you know, every single period requires the model to have to spend more compute. And so that's just more opportunity for the model to think. And so
1:07:38 You literally had this graph where it's like It's like, you know, statistically significantly smarter if you just add more periods at the end. And so So clearly like You know, if you were to add instead of periods, but add, you know
1:07:53 Useful you know, auxiliary content or a failed experiment or something like that, then clearly it's going to be you know, more intelligent. of an answer. If if periods do it right, then useful information's gonna do it even more.
1:08:06 Um And so And so that's kind of spawned this idea of of hey, like, you know, let's try something And as long as you document your success or failure. And that's going to result in in a better answer next time.
1:08:20 Um And uh Um yeah, and so Yeah, then you run into issues around the context limit.
1:08:28 And compaction. And all of that. And so Um yeah, that's a whole other issue is how do you how do how do you sort of give the right paper trail? So these LM that's
1:08:40 Maybe even another show, but Um Uh, but but the first part of this is is kind of setting up you know, appropriate
1:08:49 terminating conditions and an appropriate way to step forward. So if you If you say Um, you know, hey, build Facebook. All right. If you just
1:09:00 Go into Cloud Code and say build Facebook. Well, you know, that's a very nebulous question, and probably what it's going to do is it's going to You know, build a a front end That looks a lot like Facebook.
1:09:16 Um And And in in you know, maybe a back end for for doing some basic things. uh like posting messages and receiving them and all of that.
1:09:26 Um But Yeah, if you wanted, for example, live video It's probably not going to build that feature. Right, at least not in the first prompt.
1:09:36 Right. Now on the flip side, if you set up a loop. And you said You know. Build Facebook.
1:09:43 And then also like You know, search the internet for like the top One hundred most useful features in Facebook. And uh go through all hundred of them.
1:09:54 And uh Yeah, tell me which one of those has have been built and haven't been built and get out get get to a hundred out of a hundred. Right. Um Well that's very different,'cause now the model
1:10:06 has a easy way of knowing whether it's completed the task or not. Um, so it'll go to the web, it'll grab a hundred features based off, I don't know, Reddit or Whatever people people are talking about, their favorite Facebook features. Um
1:10:21 And uh And so the loop will Well well no Based on trying out your program that Only seventy of the features are done and it will continue
1:10:32 And uh it's at that point it's pretty determined. So Um Yeah there's There's like a
1:10:39 I've actually seen it where if it runs for too long, the harness will kill it. Um, and things like that. But generally speaking, if you give it a very easily verifiable target. Um loop indefinitely. to achieve that target. And there's things that you can do. to um
1:10:57 to to make sure that it's it's committed, which we'll which we'll talk about. Um so So there's kinda two categories of loops, or maybe three categories. One is sort of Firing on A trigger.
1:11:11 So you know, a GitHub issue has come in, a email has come in. Uh, et cetera. Um, there's sort of a cron, you know, firing every day or every five minutes or fifteen minutes. Um And then there's this third one where it's almost like a while loop.
1:11:25 Where it's you know. loop until a certain condition has been met. And so it's kind of like we're reinventing Basic uh all over again. But uh Uh, but we're doing it with like a much more capable instruction set.
1:11:43 Yeah, I w I wonder. I I feel chaining a lot of those things together is gonna be really interesting. And I think, you know, obviously there's lots of Opportunity, but it's tough to build right now because Um I I don't know. In that space, you just don't know what'll get scooped, what'll get you know
1:11:59 Implemented. uh by by the major major providers. But certainly I think that crafting very targeted directions. Um And also doing some form of which I don't know, you might have
1:12:13 have some insight to But also like sort of space exploration. No, not like outer space. But sort of like exploring design space is something Yeah, I've never seen a ton of progress towards like we talked about earlier Sort of like hey, you could
1:12:29 Please refactor this to use Sem D, right? It's like a very specific thing. But exploring that space in a like sensible way. uh with you know sort of branching and picking up that's how humans, or at least my brain ends up working. kind of like all of these potential things and you're sort of trying to balance the explore versus exploit uh in in the design space.
1:12:51 Uh and that's not something that's something I still need to provide strongly, but with loops. And this is something that I think in the future as Though we're Lo smaller models. Get faster and faster and faster.
1:13:06 There becomes a point where you can just, you know shotgun approach, try all the things in parallel, um and sort of like, you know, obviously explore very quickly. But then choosing which ones to pick up and continue or not, you know? G Yeah, I don't know. There there are some interesting ways of approaching it.
1:13:24 Yeah, I think actually. That's a really good point. Yeah, we should Uh Talk about sub agents. So
1:13:31 So a sub agent is basically Um, think of it as Yeah, it's basically recursion, right? So think of it as clawed code. spins up another clawed code and asks it to do something.
1:13:45 And so the nice thing about that is The sub agent has its own context. So here's an example. If you are Let's say you're coding up Facebook. So I wanna build Facebook.
1:13:57 And Um The UX the the the the UI doesn't look quite right. So you know it's The the the columns are too wide and uh
1:14:08 Yeah, the heading is too too thick and doesn't quite look like Facebook. Well, so you're gonna spend a ton of tokens. Going back and forth about that. Right, you're gonna say Well, you know, go to the Facebook dot com, look at their website, get their columns right.
1:14:23 And it's going to go to Facebook dot com and take a screenshot. Take a screenshot of your version. And each of those screenshots are gonna be hundreds and hundreds of tokens. Right. And
1:14:34 You're going to hit your context limit comparing all these screenshots, and then it's going to do something called compaction. And what compaction does Is it It basically summarizes All of the content up until now.
1:14:49 And You know, it's sort of a black box, right? You can't really count on it to do anything. uh uh you know, uh as you would expect. So So compaction might Just delete all those images, which is probably what you want.
1:15:03 Or it might keep them and like delete all the Interesting design work you did before you went into this rabbit hole, which is what you don't want, and you can't count on one or the other. So You could say, you know, spin up a sub agent. And have that sub agent
1:15:18 In a loop. Um Um uh have that sub agent compare visually your site and the original Facebook site in a loop.
1:15:28 And uh and and and iterate until it's complete. Um, and so while that sub agent is off doing that, you could even do other things. But concurrency aside, if nothing else, it manages the context. So So when that sub agent returns And it says okay, I've got it visually
1:15:44 Yeah, a match. Um when it returns all of that context is deleted. Which is in this case what you want, so that you can move on to the next thing without polluting the main The main context
1:15:56 And so Um if you if you use uh get work trees or other sort of technology where Multiple programs can edit the same code at the same time and
1:16:08 Basically have a whole git workflow. Uh locally. Then you could spin up you know. You ha come up with ten ideas
1:16:18 Website. Uh More flashy. And here's this benchmark, here's this like black box you can run that gives you a flashiness score. So come up with ten ideas.
1:16:29 Have ten sub agents do ten totally different things to make the website flashy. And then You know, tell me which of those ten increase the flashiness score. And uh um and then keep those and throw away the others. So so you can start to get to this
1:16:45 Like uh Simulated annealing. Kind of approach where You try a bunch of ideas and keep the ones that are better. Um, so I think we'll start to see a lot of the stuff.
1:16:56 come to bear more formally. But right now you can build it yourself. No, I'm just thinking about Nubby's number factories flashiness score. Yeah. Um
1:17:12 Yeah, I wonder if there's a way to quantify You know, like how Engaging something is. Yeah,'cause that's one thing an LLM I don't know if it can Really? Can it could it look at
1:17:24 That game. And the E T game on Atari. And know that one is better than the other. You know. Other than from just popular sense of it, right.
1:17:35 Uh nobody could definitely hire people on Fiverr and do a poll. That's true. Um Uh oh the another thing that's worth mentioning, we kinda wrap up here, but Um
1:17:48 Loops are early days. And I mean, we've this is now the third time that I've mentioned this, so it does seem to be kind of like a trope for this episode, but But kind of like thirty B models. Um loops are early days. And so one thing that I've found is if I tell it Let's say the ninety percent test coverage case.
1:18:07 Um Sometimes it will just end. Like sometimes it won't respect your terminating Sometimes the loop just ends and you don't know why and it's not It's not really clear.
1:18:19 Um The one loop that The the the one loop that seems to be reliable is the cron job. And maybe that makes sense. Right, because it's the least ambiguous. So
1:18:30 Now if you say Run something every five minutes. It will almost certainly run every five minutes, indefinitely. So
1:18:39 Um, so what I've learned to do is I have something that says, Hey Um Monitor This run. Like train a model Monitor the metrics of the model.
1:18:50 And um You know, if the model is better then Put it in this cat in this folder full of really great models. And if the model's a regression, then abandon it. Right. And sometimes that loop will work and sometimes it'll just stop. They'll say, Oh yeah, I trained
1:19:05 I trained my third model. I'm done. And it's actually not done, right? So I've set up another loop. Which runs in parallel.
1:19:13 Which basically says Yeah, wake up every fifteen minutes. And if if this other loop has stopped and just, you know, started again. And uh
1:19:24 And having both of those loops seems to be a a way to like keep that first one. You know, uh from dying. Was gonna crack a joke about us being sub agents in the loop, but uh My loop got stuck. Apparently we don't have the other agent that kicks us every two weeks to make an episode. We need a w
1:19:50 That's the agent that's we're missing. Um But Uh maybe just like a you know kind of a Call to action here. I mean
1:19:58 I've done so much with loops in the past month. As I said, I've kind of Gotten the podcast to uh transcribe and do all of that stuff locally with local LLMs for free. Um So many other projects, the Maim Hub thing where it just goes off and does its thing now.
1:20:17 Um, it's a super, super, super powerful technology. Highly recommend folks learn it and leverage it. Um Uh, but it's also it's just janky. It's just early days. Uh early days are actually kinda the best days. In hindsight, you know, when all of this stuff is
1:20:34 solved and kind of frozen and We're all just using it, it becomes a little bit less interesting. I mean at this point you have a chance to actually shape um the way that these things end up. So So definitely if you're not Using loops, you should
1:20:47 learn it. Um some people have gone further and done graphs and Um Basically this whole like, you know, communication protocol between many different agents and I mean I think that the this is becoming like a maybe more efficient ways of doing this.
1:21:05 Um, but you're not really typing that much anyways. So if you have like three loops Uh, it's probably fine at this point in time. Um, but yeah, definitely uh you know, something that you should keep your eye on. This space is moving so fast. I don't know. I feel like uh There are topics for us to discuss in instead of like the backlog like, Oh, someone suggested this a few years ago.
1:21:29 I feel like we're in this stage of like This development happened. We we should talk about it. Yeah. Yeah, it's wild. I mean, I definitely think we should cover pie uh I actually have it installed, but I haven't done a whole lot with it.
1:21:44 I wrote it on my sticky. I'm gonna go do it. I I gotta get a face and and do something productive. Yeah, yeah. So uh folks out there, if you are uh coming across tech uh, you know, whether it's harnesses or Or really anything that you feel like uh we should bring to the attention of the audience, just shoot us an email, hit us up on Discord. Um
1:22:04 And uh as always, thank you so much for all of your support on Patreon and and the other platforms. And uh we will catch you all next time. Music by Eric Farnballer. Programming Throwdown is distributed under a Creative Commons Attribution Share Alike 2.0 license. You're free to share, copy, distribute, transmit the work, to remix, adapt the work, but you must provide uh attribution uh to uh Patrick and I. And uh
1:22:51 Share alike in kind.
What you see above is a preview of the first minutes. One unlock costs 10 credits and covers this episode forever: full segment and word-level timestamps on this page, plus .txt, .srt, .vtt and word-level JSON downloads, as many times as you like.