Episode Transcript
[00:00:05] Welcome to Prompting Curiosity, a podcast for the AI curious. No coding background required. I'm your host, Dr. Shantae Cofield, also known as the Maestro, and I created this show to explore what these AI tools actually are. Really, though, are the files in the computer, how to use them, and what they might mean for how we think, work, create, and move through life. Whether you're skeptical, intrigued, or already experimenting, you're in the right place. All that I ask is that you stay curious. All right, let's get into it.
[00:00:38] Hello, hello, hello, my curious people, and welcome Back to episode 52. By welcome back, I mean welcome to episode 52 of prompting curiosity. I am your grateful host, the Maestro, and today we are talking about a term that is everywhere in the AI world and that is agents. So I'm realizing that I said welcome back because I. This, for me, it is welcome back. I did a bunch of those episodes the past, like, three episodes. I batched them because I went to Switzerland, and I wanted to get them all done beforehand because I wasn't trying to be recording podcast episodes in a different country. So I am m back on the mic, and I'm excited about this episode. Uh, this episode, based on the outline that I'm looking at that I made, this episode may be a little longer. Um, but I'm excited about this episode because it answered a question that I have had for quite some time. And as tends, or seems to be the case with AI terms, similar to the last episode, which was about mcp, um, you start to see these. These terms, these acronyms or words or phrases, and especially if you're, like, in the AI world and you're kind of reading stuff, learning stuff, and sometimes it's, like, hard to understand what it is, but it doesn't really matter because you're like, I'm not really using it. I'm not really doing anything with it. I don't really care.
[00:01:50] Uh, I'm. By no means that person is like, I have to know everything. I know the things that are relevant to me.
[00:01:54] And this word, this term, agent, it is not new. I have done previous episodes where we spoke about it. Um, episodes where I talked about Claude Code and Claude Cowork and. And explaining that those are agents, and, you know, speaking to what agentic AI, AI is and what agents are, but still never necessarily having the fullest, most complete grasp of what it meant, especially, like, just from a level of. Of certainty that I would like. And, uh, it didn't really matter, right? It didn't. Honestly didn't matter. But we will, as we'll see and as I'm gonna go through, um, some of the model releases that happened in the past two weeks, like, this is what AI is pushing towards, right? It is pushing towards agents. And so, uh, I was like, let me do an episode about this. So before we jump into agents, I want to run through all of the new models that were recently released. Um, both to just keep you in the loop because we're curious, but also because these new models, they are being developed in order to improve agent. Agentic capabilities. In order to improve agents. Right? So let's first start off by talking about the literally 11 billion releases that went down, like in the last two weeks or so. And then we'll get into stuff about agents. So the releases, Anthropic, OpenAI and XAI, that company, uh, they all released new models within literally like days of each other.
[00:03:16] Which is a heads up to all of us that these companies continue to be competing in the self created race to AGI. Right? Artificial General Intelligence. So anthropic, um, Fable 5 and mythos 5, it came back, um, it was suspended on June 12th. I spoke about some previous episodes. I will link that in. I will link the actual episode that I spoke about what Fable and Mythos were. Um, but a quick recap of what happened with the suspension and then the reinstating of both models. So Fable 5 launched June 9th. Three days later. Amazon basically said, hey, we found a, uh, you know, a jailbreak. This is. Which is basically a way to trick Fable into ignoring its safety rules, right? And that they could get the model to, you know, identify software vulnerabilities and produce code, um, that, that it shouldn't be able to show, um, and that code could be used to break into systems. So it was, it was like, hey, we have found a way to hack this thing. The US Commerce Department then ordered Anthropic to cut off access to that model for any foreign national anywhere in the world. Obviously, Anthropic cannot do this because it doesn't. It can't verify a user's nationality in real time. So it just shut the models down completely. Right. So for the next two weeks, Anthropic's working with the government, going back and forth, um, and working with Amazon to build a new safety classifier that like, targeted that jailbreak. Um, and basically they got the sign off. Eventually, the restriction was lifted on June 30 and Fable was returned to usage to everyone on July 1. Anthropic said the whole time they were like, hey, the models are safe, but we're going to fucking do this because we want people to be able to use it. Um, and then it gave affected users, um, 50 extra weekly usage through July 12th to make up for that outage.
[00:05:07] They are saying that they're going to pull it from the regular Claude subscription, that this would be like another subscription on top of it. Another. It's like pay per usage on top of it. But I would not be surprised if they backtrack on that because OpenAI is doing their own ship. We're gonna get to that in a second. So, um, just lastly with, with Anthropic, they also released Claude Sonnet 5. And that's now the default model for anyone that's using Claude, whether the Free tier or the Pro plan or whatever. Um, so, yeah, that is all the stuff that has been released by Anthropic. What I was saying before about OpenAI and why Fable may stay part of the. The monthly subscription is because OpenAI, they released GPT 5.6 which comes in three versions, Soul, Terra and Luna.
[00:05:59] All right, that's Soul being the most powerful and fastest. Um, most powerful, excuse me. And then Luna, that's gonna be the fastest and the cheapest. So there's a bunch of benchmark testing that gets done. And GPT 5.6, Soul, the highest model, scored 59 on that intelligence index, which was only one point behind Fable 5's 60. Right. So Anthropic is the owner of Fable 5. And that was just like, it's just been touted as like the most, most capable model on the market. And then here comes OpenAI with a basically equally capable model that get. This costs roughly a third as much per task. That's it, a third.
[00:06:38] So, uh, this makes it, you know, arguably the best pick specifically for coding work at a meaningfully, you know, significantly lower price, which is what I'm seeing on Threads, people talking about that.
[00:06:48] So expect to see more back and forth battling between OpenAI with their soul model and Anthropics Fable.
[00:06:54] Um, in addition to the models that OpenAI released, um, they also launched a separate product called Chat GPT Work, which is where the agent conversation is actually starts. Right. It's actually more, that's, that's more of an Agentic thing. It can do things. Um, but I just wanted to cover that like it has released these things, not specifically going into what the things are. Okay. Next Xai, they released Grok 4.5. Uh, this is their first model that's built specifically for coding and agent style work. Um, it was trained partially on real session data from Cursor. I spoke about Cursor before and I spoke about in the last episode maybe. Um, Cursor is an AI coding tool and SpaceX bought Coast Cursor for like $60 billion last month. And I maintain that this acquisition is actually hugely problematic. It's very bad but no one is talking about it. So I'm just gonna like sit here and wait for someone with like way more knowledge about this to say something.
[00:07:55] Um, but Musk, uh, of course being the worst, he said that the model is Opus class and roughly comparable to Opus 4.7. Um, but that's Musk just being Musk, so feel free to ignore that.
[00:08:11] Uh, but independent benchmarking, uh, done by artificial analysis, they ranked Grok 4.5 fourth overall on its Intelligent index which is behind fable and behind GPT, uh 5.6 and 5.5. Um, and Opus, but it is still ahead of every open weight model and all of Google's Gemini lineup. So worth noting, um, where it did lead the pack was Agentic tool use. Specifically, um, specifically Agent Agentic tool use.
[00:08:42] And it's again a fraction of the cost of Opus. So just worth throwing that out there. Fuck that company. But I want to give you the updates on everything.
[00:08:53] Does this mean anything for all of you? Like I almost hesitated to put this in the episode but I was like listen, curious people, they could also skip through this. But uh, I think this is worth knowing these things that all these companies are coming out with these new models. Um, but will you notice a difference? Likely no. If you're in the 20amonth, you know, um, tier like me and you're using it for our, the stuff that we use it for, you will probably not notice a difference. Uh, the benchmark that these companies use, they measure things like you know, bug fixes and um, how they hold up, how the models uh, hold up in like these very long multi step tasks.
[00:09:31] Um, and if you're a developer, a coder that then you're running it at scale. Yeah, you'll notice these things. But if you're not doing that kind of stuff, you're using it for emails or thinking through decisions or kind of how I guess most of us are probably using it, you likely won't notice a difference. And you know, the models that came before were perfectly capable and these will be also be perfectly capable and they're all very similar to each other. So you know, pick the one you like the most. Um, as per, you know I've said A million times. The gap really shows up in the specialized technical work, not in the regular use cases. Right. All right, so we're 10 minutes in. Let's talk about agents, right? Every single one of those releases that I just spoke about, they were models, right? They are, yes, more agentic than the model that they replaced. Meaning that the model itself is better at planning multiple steps and using tools with less hand holding from the human. But none of those models that were released are an agent, right? An agent is a computer program. It is an actual piece of software that is built around one of these models, right? So it's built around Sonnet 5 GPT, 5.6, Soul Fable 5, etc. And it gives the model three things. Number one, it gives it tools that it can call, right? So it can have search function, code interpreter, it can browse, you know, it can click around in a browser, has access to your files and other apps, things like that. A loop. This is a fun one to understand.
[00:10:59] A loop means it does that. The model doesn't just take one action and stop and then wait for you to tell it to do the next thing, right? It keeps cycling and going on its own. It does a step, it looks at what happened, it decides what to do next. It does that step, it repeats that, right? All without you having to send a new message each time just to keep it moving and nudging it along.
[00:11:17] Third thing is that it gives it a way to check its own work, right?
[00:11:22] Uh, an agent gives the model a way to check its own work. It looks at the result of each step before deciding what to do next, right? The model.
[00:11:33] So Sonnet 5 GPT, 516, whatever that is the part that makes the decisions, the agent, it's software. It is a program that's wrapped around it and it turns those decisions into actions.
[00:11:46] Okay, this is fun. I'm sitting here and like, I'm like. As I'm saying it, I'm like, is that correct? I went over this thing a million times before I this outline, before I'm recording it. But I'm like, is it correct? Yes, that is, that is what's happening.
[00:12:02] The, uh, agent will continue until the model decides that the dot, the job is done, or it hits a set point, a stopping point that you have set, right? This differs from a regular chat. Hopefully you're seeing this. But just to articulate it, this differs from regular chat. And that, uh, when you're just having a regular chat with ChatGPT or Claude, right? Uh, you ask it something, it answers, you Go do the thing yourself. You go execute the thing yourself. With an agent, it plans, it acts, it checks its own work and it keeps moving without you approving each individual step.
[00:12:36] It doesn't matter if you're sitting there with it or you walk away or right. It is that autonomy that makes it an agent. Okay, um, so the next part here is I kind of want to go through some terms that you may or may not have heard. If you are in the AI space, you're reading any kind of threads, things like that, you may have seen some of these terms. I know that I saw them and I was like, what exactly does this mean? So I did some diving and that's why they are in the episode. If you haven't heard these terms, well, now you will. Okay, so the next phrase that you may hear is building an agent. So when people say, I built an agent, it can mean wildly different things and wildly different amounts of work depending on the context. There's low effort and there's high effort. Low effort means basically just configuring something that already exists. So we spoke about this with things like building a custom GPT or a cloud project, right? So you're giving something like custom.
[00:13:24] Wow, I almost had a stroke there. You're giving something like a custom GPT or a cloud project. You give it instructions, reference files, you can turn on some tools.
[00:13:33] Um, I went through this process and explained some of this in the custom GPT episode. I'll link that. Um, but there's no code involved. Like you're just telling existing software what to do, what it can touch, what it has access to. That is one way to build an agent and specific to the tasks that you, that you want executed.
[00:13:48] High effort version of this is when you're actually writing the program, right? You're using something like Claude, agent SDK. Don't worry about that means don't worry about it.
[00:13:56] And this is where a developer would decide exactly, again, which tools it gets, how many steps it can take. But it's building the pro. The the program out. Okay? So both of these things get called building an agent. They are not the same job. But I just wanted to expose you to this terminology, okay, Spinning up or running multiple agents. So this is something that you'll see people bragging about. The brochachos on threads love the tech brochachos love to talk about this. And they're like, I'm spun up a bunch of agents and they're running on my Mac mini. I'm like, what the fuck? Why Are you doing that? What is happening here? Go make some friends. But to understand what this means, very simply, um, this typically means someone started multiple instances of an agent program and handed each agent a task, right? So imagine, uh, five copies of the same program. Each is working on a different piece of a problem at the same time, right? Sometimes there is one lead agent that's, like, in charge of the other ones, right? It splits up the work and it pulls things back together at the end. But it's not five different AI models that are having a meeting. It's literally this same model running several times in parallel. The same way you might have, like, five different browser tabs open instead of one. Okay? So running multiple agents. That's what that means. Computer program. Same computer programs running multiple times, working on different things at the same time. Um, something. The next part here, something that I was like, that I went back and forth with, talking to Cloud about. I was like, wait, is regular Cloud an agent? Because, like, it's not. I know it's not. But, like, cloud code is. But, like, is regular Claude an agent? And the short answer is not quite. And the reason this came up is because for this episode and for many episodes, I have Claude help me out with the episode. I have it help me to research, um, and, like, compile things and summarize things for me, right? So in this case, I had Claude, specifically Sonnet 5, do some research for the episode, because I was like, hey, what was all the releases? Can you aggregate all of those? And can we go back and forth, um, with what an agent actually is? Right? So I had it searched the Internet multiple times. It read the results, it decided what to search for next based on what it found. You can see it doing these steps, right? Um, it wrote summaries and it checked its own work.
[00:16:06] Claude, in this case, was able to plan, act, check, and decide the next steps without telling me which, you know, individual search to run. So that is agentic behavior. Absolutely. This is largely why I did this episode, because these models are moving so much in that direction where before it was like, you had to, like, explicitly. It was like Google, right? That's largely what, um, this. This felt like there was no decisions made. You just type in a thing, and it just, like, gave you the thing back. Now I type in something and it's making a decision about, like, which sources to go and check and how to aggregate them and which ones to pull and. And what. What to write out to show me that is agentic behavior. And so I wanted to do this Episode, because that is the direction that all these models are headed, right? But does this make it an agent? Not quite. And the reason is that the autonomy is capped at a single response, right? All of this searching and deciding, that happened, it happened inside of one single reply. Once it answered, it stopped. It didn't just go on and do other things, right? It didn't keep working after that. It didn't work across time, didn't own the task across time. It's not deciding what to do next until I tell it what to do, I send another message.
[00:17:16] Cloud code, on the other hand, or cloud cowork, they are built to keep running across a longer stretch, right? And if you're using Fable as the model underneath, that's like built to run the longest. Um, and you often with those, with those, um, with cloud code or cloud cowork, you don't need to check in after every single step. It will just do the things, right? So the real line, in my opinion of agent versus non agent isn't simply can it do things, it is the scope of autonomy.
[00:17:49] Does this system only work within one exchange, even if it's a complicated one? Or is it built to. Is. Is. Is this thing built to own a task across time with less need for me, the user, to show up at every checkpoint? Right. Regular Claude, it gives you a taste, nice little taste of agentic behavior. And it's all inside of that chat wrapper, which makes it magical.
[00:18:12] A dedicated agent, right? A dedicated agent has that same capability, but it's inside of a slightly different wrapper, right? One that is built for longer, less supervised work instead of that single back and forth, okay? Set independence there. Um, if any of you have used Claude code, and maybe you are sitting there and being like, wait, Claude does ask for permission, though. It does stop. It doesn't, like, just keep doing all these things. And you're right, it is still agentic, right?
[00:18:42] That is different because Claude is not asking you what to do. It is asking you if it is okay to do what it's already determined to be the next step, right? It's asking, is it? Okay, I've already decided. This is the plan. Can I do the next step? Can I do the next step? Right? Very, very different. What should I do next? That is the human forming the plan. That's what we'll see. You know, that's what you see when you use regular cloud. It's like, what would. I could do this next? But it's asking you, what should I do?
[00:19:10] Whereas when you are in cloud Code, it's saying, can I run this command? The agent has already decided what it wants to do. It's formed a plan, it's chosen the next step. There is literally a section inside of CLAUDE code that says planning mode. And it plans it out for you, and then it goes and executes it with your permission.
[00:19:30] Uh, you are not the planner. You're just a gate. The agent is still driving.
[00:19:35] So in between these permission checks, that agent can read files. Maybe it, uh, weighs out some approaches, it decides on a plan, changes the plan, all without you involved.
[00:19:46] That does not happen in a regular chat where you are re engaged literally every single step, uh, because there's no ongoing plan. So we see very immediately, ideally, you see the difference here between it asking for permission versus it asking for what to do. Right, Permission versus plan.
[00:20:05] So, uh, this another term here, human in the loop. So the pattern where the agent is pausing to ask for approval with, you know, typically more risky things that has a name, it's called Human in the Loop. And this is actually the more, the most dominant way that agents are currently being used. Um, and it's more common than fully autonomous because you run fully autonomous, where this thing just does all the things it wants on its own and you can get some bad mistakes. Right. This approach, this human in the loop approach, it is also adjustable. So, you know, they call it a harness. I believe that's what they call a harness, but don't quote me on that. Um, but basically you can loosen the reins here. Um, so like for CLAUDE codes, permissions, you can have it where, like, they're all accepted, or you can have it auto approve certain actions or has to ask you for certain actions. Um, the fully autonomous version or setting where CLAUDE never checks in with you at all, just does all the things that is the command for that is dangerously skip permissions. It's a legit thing, real thing. Um, but that name should tell you about how, you know, anthropic feels about making that the default setting. It's a very much a use at your own risk, your whole hard drive may get deleted. It's not on us type of setting. Right. So, um, again, just reiterating that permission asking isn't, you know, it's not proof that something is less of an agent. Right. And it is actually the opposite. And it's just a safety setting that's usually sitting atop a, a very capable agent. So getting to the end here, um, you are probably already using an agent if you're using something like Claude code, cloud, cowork new, you know, ChatGPT work that's new Claude and Chrome, um, Grok build and Cursor. Those are all agents. There's a good chance if you've used any of those, you listen to this, this podcast. I know I've been real heavy on, on cloud code. I did that whole, um, workshop on, on vibe coding. That is an agent. Um, if you're not using it, that's totally fine, right? Because who is actually using this? You know, I think that if you're, if you're really in the AI world, you'll think, you'd, you'd think that like everyone, the whole world is using it. And while I do think that uptake and use of AI in general is definitely increasing, and I'm just saying this based on like, who I'm hearing say, like, oh, I asked ch amputee as mentioned as people want to on that, that statement, to me it's very telling and very eye opening of like, oh, more people are saying that it means more people are using it. I have a, an older, an older man that I know named Chuck. He, uh, he's retired. I don't know how old he is, but, um, he's older than me and he plays, he plays volleyball. I've never played with him. He was on the court next to me and he's very nice and I hadn't seen him for a little bit and we also, part of that was we were gone, but I came back and I was like, chuck, where you been? And he was like, oh, I hurt my foot. And then, you know, he goes on to talk about it and then he's like. I asked chatgpt, right? Uh, he asked chatgpt about his foot and it's just like to me, that really stuck out and I'm like, people are using this thing, right? So as it relates to this agentic side of things. And they're using agents. Yes, people are using it. Um, it's not at all this discussion of like, nobody's using it. It's, you know, the reason I brought this up and have this episode is because I wanted to give you the name and an understanding for something that is absolutely already being used, um, and is only going to continue to be used more and more. So I think what gets thrown around a lot and put to the forefront when we're looking at uptake and usage are enterprise numbers. So enterprise just refers to like large companies and big corporations, um, and you know, them talking about usage and just in general, AI usage in general, AI adoption And now in this case, um, agent usage and what the studies that, that Claude sent to me found, um, where that most companies are still just experimenting with agents rather than running them at real scale. Um, and even among companies claiming that they've adopted agents, only a small fraction have rolled them out broadly. Right? So adoption is absolutely happening right now and um, happening outside of the tech world. It's just in its early days. But I do think that just given how capable these are, adoption, uh, and usage will continue to grow.
[00:24:04] Right. And part of this I think is almost the word that comes to mind is like subversive, but that's like not actually the right word. I don't think. It's almost like secretive like these, as we look at what these models are able to do, it's almost like, well, they're just being, they're just pushing into agentic use cases and agentic usage and becoming agents just because of that's like ask you, do we want them to do this? You want me to do this? And as people start saying yes and you know, getting more familiar with it and getting a better understanding of things, it feels like it's just a natural progression. Um, so do you need to be using an agent? You know, I think in general AI continues to be very much, here's a solution, go find a problem. Uh, and agents are no exception to this. So depending on the use case, they can be genuinely helpful. Um, but they're not like the most necessary. But I did put together like a little mini cheat sheet of when might be a better when might be a helpful case versus not. Um, and basically a good fit for an agent is for repetitive, well defined tasks that you already know how to do yourself. Right? And this is one of, we talked about this a bunch where it's like let the robots do the work. And this is also, um, for tasks where you're totally fine, uh, reviewing the output before anything goes live. Right?
[00:25:19] That's how I do a lot of my stuff as well.
[00:25:25] Um, cases where you know, proceed with caution regarding an agent. Anything that you can't easily verify anything that's high stakes, right?
[00:25:32] Spending money, publishing anything, you know, without a review step in between. Like, I would be very hesitant to use an agent for that. Right.
[00:25:42] I think in general most of us do not need to become agent power users tomorrow. But again, like I've been harping on all episode, I do think it's worth understanding what they are so we can identify when, you know, it might be a good use case. So last things last I'm going to wrap it up how I used AI this week.
[00:26:03] So each episode I share a quick example of how I or someone I know used AI that week. This time I'm sharing about Tom.
[00:26:11] Uh, he is the husband to one of my friends from my, My good friends from volleyball, Diana. Um, we all went to Switzerland. Like I said in the beginning episode. We, uh, went there for a volleyball tournament and Tom came along. He's one of the six people that, you know. Six of us. Uh, Tom is way more of a hiker than he is a beach volleyball spectator. Um, so he used chat GPT to plan out an entire day trip to a location that was like three hours away from what the tournament. The tournament was. Um, he put in what he wanted to do. He wanted to like, see.
[00:26:41] I think he went and saw a castle type thing, but he wanted to go and like be able to hike, he wanted to be able to be able to eat, he wanted to be able to go in a lake.
[00:26:49] Um, and I think he probably put in his budget and kind of just like the distance and it created an entire itinerary for him. I told him the train that he needed to take, um, planned out the time frame for things and Diana actually wound up meeting up with him. And I think it was actually a two hour train ride right now, not three hours. Um, but that's not too shabby, right? I had done something similar when we went to Hawaii. Um, and that's how I figured found out about, uh, the Jurassic park stop, uh, on, you know, that was on the way to where we were staying. So I think it's a very cool use case for AI and um, yeah, just share that one with you so that my friends looking at the times, a little bit of a longer episode, but I told you it would be, uh, that my friends, is all for today. Hopefully you found this episode helpful. If you did consider leaving low rating or review. I haven't checked. I'm gonna do that after. I haven't checked if you had any new ones. Um, but for those of you that have left anything, y' all are the best. It really does mean so much. I read the things I look at the stars. Like, it really does mean a lot. Okay, so thank you for that. Don't, uh, forget I have a companion newsletter and blog called the Curious Companion that drops every Thursday that is basically. And by basically I mean exactly the podcast episode in text format. So if you prefer to read or you just want a written record, you can join the newsletter. You can check out the blog, head to prompting curiosity.com forward/newsletter or forward slashblog.
[00:28:07] Or you can just check out the links in the show notes. Much easier, right? All right. As always, my friends, endlessly, endlessly. And one more time, endlessly appreciative for every single one of you. Until we chat again next Thursday, stay curious.