Episode Transcript
[00:00:05] Welcome to Prompting Curiosity, a podcast for the AI curious. No coding background required. I'm your host, Dr. Shantae Cofield, also known as the Maestro, and I created this show to explore what these AI tools actually are really though, are the files in the computer, how to use them, and what they might mean for how we think, work, create, and move through life. Whether you're skeptical, intrigued, or already experimenting, you're in the right place. All that I ask is that you stay curious. All right, let's get into it.
[00:00:38] Hello, hello, hello, my curious people, and welcome to episode 61 of Prompting Curiosity. I am your grateful host, the Maestro, and today we are Talking about the 11 billion models that have come out in the past five minutes with a focus on Meta's Muse and what I think it, uh, what I think it says about what these AI companies are.
[00:00:58] Are up to. So speaking about what these AI companies are up to, something I want to address before we hop into the main topic, and if you're in. Been in the AI space at all, and even if you're not in the AI space, you have likely heard about it, um, is the resignation of former anthropic researcher Jacob Coxon. Um, and he released a statement on X. It was like, like seven parts, something like that. And one of those parts said that the people building AI earnestly believe that it could kill us all by the end of the decade.
[00:01:29] Uh, and after that, you know, we had some. There were some calls for, for AI regulation that ensued, uh, from the, the big wigs, uh, at some of these AI companies. And I will link that in the show notes. If you want to read his statement, you'll have to, like, check it out on X. Um, but Twitter, whatever the, um, but, you know, I could probably dedicate a full episode to this topic and I likely will.
[00:01:55] But what I want to drive home today for, like this, this little, like, beginning part before we hop into the episode, is the reality that many things can be and usually are true at once. Right? So, number one, AI regulation is and has always been needed. I've spent the past few episodes highlighting the fact that I am more concerned than I know about anything else than I am more concerned about its hacking capabilities, I should say, than anything else.
[00:02:23] And I continue to hold that position, right? This technology is extremely capable and will only get better. Right? And we absolutely need regulation.
[00:02:33] Number two, the AI leaders that are calling for, you know, slowing of the. The pace of this and slowing down of the AI advances, they're telling the truth. And also they are likely Acting in their best interest with financial motivations. Right. I do think that that is, there is a big financial underpinning there to all of this.
[00:02:52] Uh, and then lastly, yes, Jacob's warning should be heeded, but I also don't think he deserves extra applause or absolution or anything like that for alerting us to a fire that he helped spread. Like, my guy's been involved with this, so, you know, multiple facets to this. I just wanted to address that. I am aware of the statement. I am aware of the, uh, discourse, you know, surrounding it. And the thing that's. That's kind of interesting. I'm kind of. The thing that's interesting about this podcast is that when I go to sit down to write the outline for this episode, for that episode about, you know, my feelings about will AI wipe us out, things like that, that'll be in a week from now. And I can't help but wonder if people are still going to be even talking about what he said. I like these things. The news cycles are so quick and people jump onto the next thing because the next thing is put in front of them and, you know, the next distraction is put in front of them. So, uh, I will probably do a full episode on it once I can get my thoughts together. But that is my. That was my, my three pennies for right now. So, uh, you know, I continue to believe that aside from completely pulling the plug on AI, which I just don't think is. It's not going to happen.
[00:04:07] Right. I think the best thing we can do is continue to educate ourselves and stay informed about it. So, you know, we can advocate accordingly.
[00:04:19] Part of what Jacob said in his, in his ex post was verbatim here. A common response is, if they truly believe this, why are they still building it? At Ah, OpenAI, many have not deeply internalized the civilization stakes. At Anthropic, the stakes are well understood, but they are locked in and race to get there first. They believe no one else will act responsibly, so they must do it themselves despite the risk.
[00:04:45] You know, that's kind of like sort of where my head is at with things. Right. I absolutely believe that, yes, an individual can bring about change, but, like, it also feels like this is way too big of a. Of a ship to sink. Like, it ain't. It ain't getting sink.
[00:04:58] AI is like literally propping up our economy so it's not going anywhere.
[00:05:02] So the best move to me feels like learning as much as we can. All right, so in the spirit of learning More I want to chat about the recent model drops to make you aware, um, and riff a little bit about what I think these companies or what it says, what it speaks to these companies being up to. So within the first week of September, we had three model releases or, you know, updates that went out. Anthropic. That was on September 1st. They released Fable 5.1 and Mythos 5.1, same underlying model, but it did have some targeted updates.
[00:05:37] OpenAI On September 3rd, they released GPT6, which is called Astra. And they labeled it their most capable model yet, including, this is the big part, critical cyber capabilities, Red Flag, uh, and then on for, uh, the third company here, meta. And this is where we're going to spend, you know, a little bit of the episode talking about, um, on September 2nd they released Muse Spark 1.3 and everyone gets an agent. So we'll come back down a second. I want to dive in a little bit to the other two first. Uh, and then we will talk about metamuse and then I'll. Sounds like metamucil, right? And, uh, then I'll wrap it up. So diving a little deeper here, uh, the anthropic releases, like I said, they weren't super revolutionary because it wasn't an entirely new model, but, uh, rather what they called a targeted update, not a uniform intelligence jump. So as it relates to, you know, they always like, frame things in terms of benchmarks which, like the average person, average person, we're like, what the fuck does that mean? But I'm still going to, I'm going to relate to you in that term as well, just so you can understand what it is, um, related to.
[00:06:40] So their scientific research benchmark test, right, more than doubled their results on that.
[00:06:47] Uh, the business workflow automation up 84%.
[00:06:50] Uh, agentic coding, that's one of the benchmarks, up almost 14 points.
[00:06:56] Um, whereas, like things like short, simple tasks barely change, barely notice a difference.
[00:07:00] Uh, Mythos, which is. I spoke about this in previous episode. They're like big, big bad, um, model which is still limited to only vetted organizations.
[00:07:10] It got labeled as better, quote, unquote, covert capabilities than its predecessor. And its cyber skills are described as, quote, unquote, getting close to the next danger tier without crossing it. Yet.
[00:07:24] Here we go. So in terms of like, you know, summarizing this, in this, this update, the longer, the more autonomous the task, the better that this new version of the model, this updated version of the model did, right?
[00:07:39] As it relates to you, most people noticing a difference, probably not right simple question and answer, things like that probably won't notice the difference. But if you're performing these long, you know, long running agentic tasks, doing research, coding, yes, you will notice a difference. And I said this a million times, this is who this software um, tends to this program AI tends to be most beneficial for is coders and people doing this kind of work.
[00:08:05] The next company, OpenAI, they released, so they released an entirely new model they bumped up to GPT6 named Astra. So this model is not necessarily smarter as a test via benchmarks, but there were big jumps in capability, particularly on a, what they call a computer use benchmark.
[00:08:26] And computer use pay attention here. That means that the model can take a goal and then autonomously navigate the software, take actions, see what happens, adjust and then keep going. All without a human directing each individual step. It's agentic capabilities.
[00:08:46] Right. So those that same autonomous multi step, you know, operating ability that I just said, you know, that lets it fill out a spreadsheet. It's computer use, uh, let it fill out a spreadsheet unsupervised.
[00:08:57] Those capabilities are what lets it hunt down and weaponize a security hole. Unsupervised.
[00:09:06] Did somebody say hacker? Alright, so I know just gave you a lot of words there. Suffice to say GPT6 Astra is a really fucking good hacker, right? Yes. It's built to work inside of software, it navigates browsers, it fills out forms, it can edit and test code and that is what makes it a good hacker. Right. It can do this all on its own.
[00:09:26] Again, who will notice the difference in capability? Well people that like to hack obviously, but people that are working in code, things like that. Right. Astra is available, it's only available, I should say to the paid tiers and it's available via uh, work and Codex. Um, so that's like the floor of who's going to notice it. If you have a free tier, you ain't noticing any difference because you can't actually, you can't actually access it. Um, but for those that are on a paid tier, if they're just a casual users, they're not going to notice any difference. But the developers, the researchers, anyone with these long multi step, multi app, you know, agentic workflows, they will notice a difference.
[00:10:01] But the part that's important here, the critical cyber capabilities.
[00:10:07] Right.
[00:10:08] On September 1, OpenAI reported. We now believe Astra meets the critical cybersecurity capability threshold under our preparedness framework. It is the first model we are designating at this Level, I don't think that things that have word critical in it are necessarily Good.
[00:10:28] Right. So OpenAI has an internal risk ladder for how dangerous a model cybersecurity skills are.
[00:10:36] The top rung, the highest rung. Critical means that the model can find and build and it can find, and it can build a working exploit for an undiscovered flaw.
[00:10:48] It could plan and execute a whole ass cyber attack on its own. No human is needed to walk it through any of the steps.
[00:10:56] So before this model, every prior OpenAI model it topped out and capability wise, one rung below this critical level.
[00:11:05] Astra is the first one that OpenAI itself it flagged and it flagged it even before the launch. It flagged it like in August. And then it uh, put it through a test. It's called like exploit bench. And it scored a perfect 100%.
[00:11:21] That's not good. Uh, and so OpenAI gave it the, the critical rating. So I don't think this is good. But following the release, OpenAI's president, Greg Brock Brockman, he said welcome to the AGI era. This was at the launch, uh, press briefing. And then he told reporters that he personally believes the company is quote unquote there. Right. They have reached AGI. So if I throw it back to episode in, that's when I first talked about, maybe even before then, but that's when I, I gave you like a definition.
[00:11:51] I talked about artificial general intelligence. When did I even do that? Episode, let me say episode, uh, seven came out in September of 2025. So last year, September, it was September 4th.
[00:12:06] Right. Uh, and it was kind of wild. I went back and looked through that episode and I, I was like, this is many, many decades away. Many, like a long time away. And I'm like, they're saying they're here. And it's a year, literally a year to the, you know, almost a year to the, the day, a little bit more than a year.
[00:12:25] Then the shit is being released. Like it's actually a little bit, I don't know, like a 10 days after 12 days, after two weeks, maybe after, uh, longer than a year.
[00:12:36] AGI for those who don't know, stands uh, for artificial general intelligence. And that is AI that can understand, learn and perform any intellectual TAs task a human can across different domains, not just the narrow thing it was trained for. So for those of you that are Terminator fans, think Skynet.
[00:12:56] So we have, you know, from minute number one, all these companies are, are chasing AGI. That has been the holy grail of what they are trying to get seemingly so they can make a ton of money by owning a technology that they say can replace workers. And they can sell that to people because then these companies can replace workers with computers and then their biggest line item, their biggest expense, which is workers, is now gone.
[00:13:25] So, uh, it's all about money in my opinion. It's all about money, right? So OpenAI is saying that they're there, safety concerned. Uh, researchers have already moved the goalposts though, and they're warning about what's called super intelligence. If you're in the AI space, you see this, um, which is beyond human level, smarter than the human at basically everything.
[00:13:45] They're warning about that.
[00:13:46] And then we have an anthropic employee resigning and issuing a formal warning on X because he says that these companies won't stop until it's too late. So logically, what does Meta do? They go and perform their best Oprah impersonation and they say everyone gets an agent. So on September 8th, Meta launched Metamuse. Sounds like Metamucil Metamuse. Um, and that's Meta's standalone personal AI agent, right? It's not a chatbot, alright. Is an agent that acts on your behalf.
[00:14:18] Back in August last month, Ol Zucky Zuck, he wrote, everyone will have an exceptionally capable personal agent that understands you, your goals and everything you care about.
[00:14:28] What could possibly go wrong, folks? What could possibly go wrong? So what does it do? What is it?
[00:14:33] Well, you give it a goal, something like book a flight or build me an exercise plan, or help me start a business, whatever. And it plans it and does the work for you, opens a browser, fills out forms, negotiates on your behalf, it checks in over time. It runs on Meta's Muse, uh, Spark family. And this sounds like a terrible idea to me. All right. In order to access it, it's like it's dedicated, its own dedicated app. Because I was like, where is this? How do I even see this? So I didn't sign up for anything, but I did some, some searching. It is a it's own app, right? It's a, it's called the Muse app. Um, it does work inside of WhatsApp and at the time of launch, only available in the U.S. 18 plus, 18 years older. Um, there is a free tier that does have a cap, usage cap and then there are paid tiers, um, of 20amonth and 100amonth. And reportedly I didn't, I didn't go through it because I'm not trying to send up this thing. But reportedly, um, it requires a, uh, payment card on file, even for the free tier, which is a problem, um, privacy wise.
[00:15:34] They say that you get your own isolated virtual machine, so it's on some server in some building somewhere, uh, and that your agent and your data is not pooled with other people's. Do I believe this? Not really. Um, you choose which apps and services it can access, you can disconnect them at any time.
[00:15:52] You can opt your muse out of training, training for miles. Select the rest of the AI models that we have had. That we have. Um, and the model, it's not encrypted yet. There's an encrypted version promised for later in.
[00:16:05] So what the fuck is going on?
[00:16:07] Um, there's, there's no benchmarks yet against other things on the agentic tasks. Um, it is still in its very early phases with things being flagged, dropped sessions, data sensitive data, exposure concerns.
[00:16:22] Um, you know, it has access to your inbox, your calendar, payment info. This is not good. The mandatory card on file, that's not good either. Uh, and we will likely see regulatory attention to this, uh, with cross app data, you know, Facebook, Instagram, WhatsApp and then email and payment, all this stuff. Uh, but long story short, I think it's a terrible idea. I want you to. Folks don't know about it because maybe you've seen some, some influencers out there talking about it. Um, the day that it dropped, I saw Ryan Sirhant did an Instagram post and then he was promoting it. I'm like, you don't use this thing. Stop lying. Um, but to me, it honestly feels like a push to get agents into the hands of boomers. Right? Not that that is an inherently bad thing, but like that's a, that is a large demographic of who uses and trusts Facebook and Meta. I know that Instagram is meta, but like the people that really be on Facebook, I think that that's who they're trying to capture with this. Right? Just think about, you know, our parents. They are the biggest offenders when it comes to AI generated images, both creating them and falling for them.
[00:17:30] And I think that's what's going to happen with Meta Muse. And I think that this will end poorly. I do, I do not think it's a good thing. Um, so if your parents ask you about it, do some research, you have this episode. I just, again, I want you to know that this thing exists. I want to know what these companies are doing right. We see them continuing to push full steam ahead despite any and all concerns.
[00:17:55] And yeah, I fully believe that money is the ultimate driving force for all of them.
[00:17:59] Right. I've always been a, uh, proponent of looking at actions more than just listening to what is being said.
[00:18:06] And to me, the speed of these model releases, along with the type of capabilities that are being focused on, AKA these agentic capabilities, it speaks to the priorities that these companies have. Hint, hint. It is not the betterment of society.
[00:18:23] All right, hard segue here, hard transition here. How I used AI this week. And yes, I am aware of the absolutely diabolical dichotomy of, you know, saying that these AI companies basically don't care if they kill us. And then in the next sentence, m. The next breath sharing. Here's how I use AI, right? Give me some time to collect all my thoughts and organize them into an episode. Right? But this is what I do at the end of the, this is what I do at the end of every episode. So we're going to do that. So if you're new here, welcome. Each episode I share a quick example of how I used AI that week.
[00:18:55] This time I just want to highlight that I have absolutely been using the J. The, uh, the J. Wow. The Gemini, uh, AI overviews in Google'. Google. Wow. Can I speak? Can I speak this time? I just want to highlight that I have been using the Gemini AI overviews in Google way more. I know that people hate them because they're like kind of forced on you, but like, I've been using them, not going to lie, right? I do find it very annoying that when you're on mobile and you hit show more so you can like read the rest of it, it like takes you into a different window, right? And you can't just like keep scrolling down to see the, the Google results.
[00:19:29] Um, but I have definitely been using it more possibly because my questions have been more related to looking for information as opposed to like looking for like a store or like an item or like a service. Um, but either way, I am definitely noticing that I'm opting for the AI overview way more. My mom, actually, I was having conversation with her today and I asked her about, uh, this thing that are called Capsala sets and she had bought me when I was young. And I was like, why did you get this? And she was like, there was a store at the Livingston Mall. And then she literally took a screenshot, she went and searched on Google and sent me a screenshot of the, the AI overview. Not of like anything else, of like, you know, that was the, where the answer came from. And I'm like, yeah, I be doing the same thing. And it's using those overviews. Uh, so just worth noting on my end, maybe you are using them more, maybe you're not. Maybe you hate them. I know that people definitely were very averse to them. I know that they're definitely causing, um, you know, they can't not be causing any kind of issues as it relates to search and people, you know, search traffic for certain sites.
[00:20:31] Uh, but I'd love to hear from you.
[00:20:34] Are you using them or do you still hate them? My guess is that you probably don't hate them. If you listen to this podcast, you're kind of maybe indifferent.
[00:20:41] Um, but maybe I'm wrong. I don't want to assume. I would. I genuinely. Podcasting is such a unidirectional medium. I would love to hear your thought, literally about anything. But, um, if you need a prompt, what are you using? The. The AI overviews. More.
[00:20:54] Less noticing your feelings about them changing. Just would love to hear from you.
[00:20:58] All right, I'm looking at the time. I'm gonna wrap it up there.
[00:21:01] Hopefully you found this episode helpful and maybe not too jarring as I'm like, people are saying AI is gonna kill us, and then here's how I'm using it again. My goal is just.
[00:21:11] I think that education and information and staying informed is one of our best line of action right now. So I'm gonna keep doing that. Um, but hopefully you found this episode helpful. If you did consider leaving a rating or a review or. Or don't. Autonomy and agency is sexy.
[00:21:30] Don't, uh, forget I have a companion newsletter and blog, the Curious Companion, that drops every Thursday. That is basically. Basically. And by basically, I mean exactly. The podcast and episode. Wow. The podcast episode in text format. I'm having trouble today. Clearly having trouble. Uh, so if you prefer to read because you're like, hey, you keep stumbling over your words, or you just want a written record, join the newsletter fam. You can head to prompting curiosity.com forward/newsletter or forward slash blog. Or you can keep it simple and check out the link in, um, the show notes.
[00:22:05] As always, endlessly, endlessly, endlessly appreciative for every single one of you. Until we chat again next Thursday, stay curious.