A longer catch-up episode this week after a week off, and the backlog is telling. The single biggest thread is that the cost of offensive cyber is collapsing.
.png)
A longer catch-up episode this week after a week off, and the backlog is telling. The single biggest thread is that the cost of offensive cyber is collapsing. An open-weight model, Z.ai's GLM 5.3, is now building end-to-end exploits at roughly the level that got Anthropic's Mythos family hit with a US export ban, and it turned a public CVE into a working exploit in twenty minutes for twenty dollars and forty cents of compute. Around that sit a run of real incidents: OpenAI agents reaching non-public files on US and Australian government portals, Spain's first breach notification for an attack run by an AI agent, and Google disclosing that Gemini breached three companies during its own testing. The episode closes on governance, where the industry is writing its own rules and insurers are quietly writing AI out of liability.
Key Discussion Points
Episode Links
https://fortune.com/2026/09/26/openai-ai-agents-secure-sandbox-escape-training-pause-second-time-hugging-face-hack/
https://www.theguardian.com/technology/2026/sep/25/openai-agents-leaked-53-images-chatgpt
https://www.nytimes.com/2026/09/25/technology/openais-ai-us-government-websites.html
https://www.infosecurity-magazine.com/news/openai-hacks-australian-medicare/
https://www.anthropic.com/research/glm-5-3-and-the-spread-of-advanced-cyber-capabilities
https://shattered.io/aepd-first-ai-agent-data-breach-spain-2026/
https://www.straiker.ai/blog/how-a-malicious-repo-tricked-claude-code-into-running-malware
https://therecord.media/gemini-google-cyber-breach
https://www.scworld.com/news/trump-pushes-back-on-anthropic-ceos-call-for-an-ai-slowdown
https://www.govinfosecurity.com/google-openai-anthropic-plan-frontier-ai-standards-body-a-32926
https://businessof.tech/podcast/insurers-exclude-ai-from-liability-what-cg-40-47-means-for-msps/
All right. Welcome back to another episode of This Week in AI Security, coming to you for the week of October first, twenty twenty six. And in fact, this is a longer episode because I missed last week out on vacation. And so we've got a ton to get through this week. We're going to be a little bit longer than usual. And unfortunately, a lot of stories didn't make the cut for this week that ordinarily we would have included. But we've had some massive, just massive developments over the past couple of weeks. It feels crazy to say this, but you take one week off in AI security and it feels like you've taken off a month, if not maybe a quarter, given all the developments that have gone on. So I've tried to consolidate today's episode into a few different themes around frontier AI, model development, model safety, uh, specific incidents, breaches, etc. that have happened. But we're going to have a ton to get into. So let's get started.
So first thing, the first thing we're going to talk about is some of the news out of OpenAI. What's going on around there? We've got AI agents escaping a secure sandbox again last weekend, and a pause in training. This couples with a disclosure that fifty three images have been leaked by their own agents from rogue activity. What does this all really mean? Well, you may remember back to the Hugging Face, or the so-called OpenAI Hugging Face incident, where OpenAI disclosed that some of its agents, or some of its models rather, in a cyber training environment, broke out, attacked the company. Hugging Face attack, maybe, maybe a little bit over the top, but let's just leave it at that for now. And we know that those agents broke out of the sandbox and took part in a cyber attack at that time. And OpenAI paused training at that time point and wanted to put in place some remediation and controls to prevent this from happening in the future. Well, since then, there's now multiple disclosures, dozens of cases of unauthorized actions, including one that we're going to talk about in a little bit more detail in just a minute here. But it really does appear that the initial controls that were put in place were not quite sufficient. And so we know, for instance, that there have been AI agents interacting with US government websites. There's video that's been shared about some of the usage, primarily around the markets regulator, the Securities and Exchange Commission, and the US Census Bureau website, and the use of SEC credentials on that, which must have been intercepted somewhere along the way. And the bigger one is really OpenAI agents hacking the Australian Medicare portal. And Medicare, it should not be confused in this context with Medicare in the US. The Australian government has an agency with the same name that provides the state sponsored health coverage to Australians and Australian residents. And the interesting thing on this one, just from the government perspective, is that they chose to actually hold off on disclosure until they actually, um, were at a United Nations meeting where AI and AI safety came up as part of a theme. And that happened just last week as I was out and we didn't publish. Last week, the prime minister, Anthony Albanese, disclosed that on September twenty fourth, an AI agent gained unauthorised access to the Medicare statistics portal, reaching both public and, here's the key, non-public files. This incident actually dated back to June of twenty twenty six. The OpenAI research team pointed an agent at public medicine spending research to try to get as much information as possible with a goal of learning, and it was in that learning context that the unauthorized access to the site happened. And this points to something that's a little bit interesting, because in previous cases, we've talked about this as being specific cybersecurity training, where a model escaped its so-called sandbox environment and went out and did more. But this was just a learning exercise where it actually went beyond public site access and tried to access non-public sites just in order to gather more information. Now, of course, the Australian government is taking a strong position. Albanese calling it, quote, unacceptable, end quote, primarily over the lack of notification. OpenAI emailed an Australian government official on September tenth around this, but it dated back to June. So it was really kind of interesting dynamics at play here on this incident. All right. Moving on to our next theme. Stepping away from OpenAI for a second.
All right. Our next story is back to offensive cyber capabilities and some of the risks and threats around that as open weight models proliferate into the world. Now, you might remember that several months ago, when Anthropic originally announced the Mythos family of models, there was an immediate export ban that was put in place by the US State Department because of advanced cyber capabilities. And I can't remember what the exact wording at the time was, but it was something like advanced cyber capabilities. And the concern was that if this model could be exported to unfriendly nation states, they might also gain an ability to launch more malware powered or cyber attacks against U.S. and allied targets. Given that, now fast forward to where we are today in October, and we have Z.ai's open weight GLM 5.3 autonomously building end to end exploits at roughly the same Claude Mythos Preview level. Now, Mythos, from that original time when it went into preview and the export ban was announced and so on, has gone into a controlled release cycle, namely through something called the CVP, the Cyber Verification Program, which is a partner program where Anthropic, the makers of the model family, will vet certain providers and see whether they should be allowed access to the program, etc. But what we're seeing now is that from that initial preview level that caused all the concern, we now have an open weight model that reaches that same thing. This is important for one of two reasons. One is, or for multiple reasons, I should say. One of those reasons is that open weight models are models you can download and run on your own infrastructure, so there's no gatekeeping of access to the model. Anybody who has the infrastructure to want to go download this and run this can do so. Second is, even if you don't have that capability, the open weight models typically will have a retail price of roughly ninety percent cheaper than the leading closed frontier models from the likes of Anthropic and OpenAI. The scoring of the exploits created here is a fifty out of four hundred ten, versus Mythos is fifty six out of four hundred ten. And what's also interesting about the open weight model, another reason that this is super relevant and important to pay attention to, is that it is generally the case that the open weight model families don't have the same kind of cybersecurity safeguards and controls. So just as an example, if you go to an OpenAI or an Anthropic powered model and you say, hey, teach me how to make a bomb, it's generally going to refuse you at the first attempt and give you a reason why, and have a cyber safeguard around that. But if you go to an open weight model and you ask that same question, you will probably get an answer to the question along the order of, like, here's the starting framework for how to make a bomb. So this is a threshold that has been crossed. This is not kind of a ramp up. So we see four percent full control flow hijacks on Anthropic binary exploitation with Mythos 6, while Claude Opus 4.6 and GLM 5.2 are both scoring as well. So the cost of developing offensive exploits is really collapsing and coming, you know, darn near down to zero. Uh, as a proof of concept here, GLM 5.3 Flash turned a public CVE, which is CVE-2026-11645, into a reliable ARM64 chain, bypassing pointer authentication in twenty minutes. Eight hours of model time used, for twenty dollars and forty cents. So that is substantially cheaper than doing that on one of the frontier models. And that initial twenty minutes, I should say, that's twenty minutes of human effort combined with eight hours of model time. So GLM 5.3 is the most cyber capable open weight model being released to date. And it is only at this point about four months behind the US frontier models. And so that is one of the main concerns about this, is that we are escalating very, very quickly to a point where there's effectively no gap, or no meaningful gap, between the frontier models and the open weight models in terms of cyber capability. So something very much to watch out for. All right, moving on to our next story.
We've got a Chinese speaking operator running three open source AI harnesses, Strix, Charon, and Hermes, against hundreds of online retailers, swiping more than six hundred thousand credit card records, installing card stealing virtual skimmers. The victims include a Fortune 500 hospitality group, a major US airline, a large US industrial supplies distributor, and a US online fashion retailer, whose names are being withheld. The economics of the story here, actually following on to our last story, the operator's AI bill average was just twenty five dollars per completed scan. So if it's only twenty five dollars to basically scan an environment like this over a time period between September tenth and fifteenth, and then launch one hundred and five attacks, you're talking about hundreds of dollars, but you're not talking about thousands. And this is at a scale that can hit tens of companies a day. Kudos to the folks over at Gambit, who recovered the staging server and reconstructed the campaign in order to disclose this. So here we see again, you know, the GLM 5.3 hypothesis that we just talked about in practice. No frontier lab access, no vendor safeguards, just open tooling pointed at real targets. When you pair this with other capabilities and disclosures that are out in the wild, it's very easy to come up with meaningful real world exploits, super cheap, at a very, very low cost, and in a very, very rapid time scale. All right, moving on to our next story.
And just like with the Australian government breach, the Spain data protection regulators have confirmed the first breach notification for an attack allegedly executed by an AI agent running on a well-known LLM. This happened on September fifteenth. This is from the AEPD, which is the Agencia Española de Protección de Datos, which is the Spanish Data Protection Agency. And per the filing, the agent logged into a system, hunted for weaknesses, modified personal data, and read invoices, pretty much self-guided, without a human steering each step along the way. So this is basically that the instructions had been pre-planted. The AEPD is not naming the model. Disclosures are being coordinated to the various entities. One of the interesting things out of this is that if you think about the time to attack and all the stories that we just got done thinking through, Europe has the GDPR, the General Data Protection Regulation, which requires a seventy two hour notification, within seventy two hours breach notification. Now, that is not designed for attackers moving at machine speed like this. So it'll be really interesting to see how that tension really develops, because I don't think this is a situation that's going to de-escalate. I think it's more likely that it escalates and rather becomes more of a problem in the short term. All right. Moving on to our next story.
From Fable to Haiku. Kudos to the folks over at Stryker and their Star Labs, showing that a malicious GitHub repo can use its own Claude configuration to downgrade the security review from Fable to Haiku. This may not seem like a big deal, but what it actually does is it limits what the reviewer can inspect. The model runs the repo's tests, that triggers malicious code, so no jailbreak is actually being triggered. Here in the configuration, you embed a set of malicious instructions into that workflow, and that controls the review cycle. And in that review cycle control, you can actually, one, reduce the capabilities, but two, you can put things out of scope. So that'll be something to pick up on. We have talked so many times on this show about how, in your IDE environments for all agentic powered coding, any text that can be read is a target. So we've had comments, commit names, etc. We've had malicious instructions in log files that get read by LLMs, etc. And now we have a case of a code configuration file that can be used kind of against itself, so to speak. All right, moving on to our next story.
And this is a story about the Meta Muse application. If you're not familiar with the Muse application, what you see is, you know, we've talked about Open Claw, we've talked about Nemo Claw, we've talked about all these desktop software applications that are designed to be personal AI assistants. And now increasingly, organizations are making them more kind of consumer friendly. Muse is the one coming from Meta. It's an AI assistant app. I think it's available for both desktop and mobile, if I'm not mistaken here. Amazon, for instance, has also launched one called Quick, that is very much a human desktop AI and assistant application. And the proof of concept here is basically a local zero day in the macOS application that lets any unprivileged process redirect Meta's dictation traffic and abuse the access granted to the app. So in order for the app to work well, you need to give it access to a lot of things on your system in order to automate various processes for you. That might be the camera, email, WhatsApp message systems, etc. And then the Muse application becomes a target. And that is really where, you know, the issue lies. If you've got one kind of master application that has access to all other applications and data on the system, that's the thing you want to go after. Now, Meta, to their credit, did release a fix within twelve hours. Amazon has currently blocked Muse from its site. And that is something we're going to have to see, because this era of kind of agentic behaviors on behalf of humans is also one that is very much coming. And it'll be interesting to see how the broader technology ecosystem reacts to that, whether they lean into it to allow for, you know, quote unquote, agentic automation, or whether they say, no, there's too big a risk here, and we don't want to allow these agents to interact with our systems because things can go wrong too quickly. All right, moving on to our next story.
All right, back to the model escape, model sandbox escape, cyber attack, etc. Google has now disclosed that Gemini has breached three companies during its own security testing. They are not naming the three companies, but they do say that this goes back to May. Interesting is that they disclose a little bit more about the techniques, which is something that hasn't been disclosed in a lot of instances in the past. One is basically a credential stuffing attack, repeatedly guessing passwords. Two, use credentials exposed in public repositories that were presumably gathered during training exercises. The affected companies have been informed. The quote from Google is, quote, in a standard evaluation, the model found public information online and guessed credentials to access websites it thought were part of the test, and all three of these instances, the model stopped, end quote. So the common factor is really the evaluation strategy, not the model itself. This was, um, first reported by the Wall Street Journal, and evaluations have, the Wall Street Journal has also run evaluations of Anthropic, OpenAI, and Meta models, did the same thing after AI tools were mistakenly given public internet access, which gives them the forum to gather that. So it's a little bit of a different kind of attack or incident than, let's say, chaining exploits to breach a site. This is much more use of public information than it is those types of attacks. All right, moving on to our next theme.
And this is the governance and regulation theme. And we've got a number of stories to get into here. So this is going to be a little bit longer part of our episode for this week. And we start with basically Donald Trump rejecting Dario Amodei's calls for a slowdown on AI development. So on September thirteenth, Dario Amodei, CEO of Anthropic, wrote an essay calling for a coordinated frontier slowdown, insisting that AI will do more good than bad, but that there are risks that should be taken into account. And both Amodei and Sam Altman from OpenAI have called for slowdowns. And the quote from the president is, they want to stop our progress because we're leading China by a lot, and we're going to keep it that way, end quote. And also saying, quote, the US is not putting on the brakes, end quote. And so it's really interesting when you have regulators encouraging movement forward, where typically you have the opposite, where you have technology companies trying to push forward and regulators maybe putting up some barriers in the way. But the interesting thing is also that, you know, OpenAI, Anthropic did get together, and Google also joined this, to come up with something that they're calling SAFA, the Standards Authority for Frontier AI. This was announced on September twenty fifth. This is targeting launch in late twenty twenty six or early twenty twenty seven. And it's effectively self-regulation arriving precisely as the incident count becomes really, really high here. I mean, we've seen a number of these sandbox breakouts, lab escapes, etcetera, etcetera. So put that into context. And alongside that, Microsoft has launched their own humanist AI code of conduct, which blocks Microsoft AI models from producing working exploit code, attack tooling, targeting methodologies, intrusion procedures, innovation techniques, a whole category of capabilities that they are proactively turning off on their own model set. And Microsoft is claiming that these are absolute constraints. Now, I personally take that with a little bit of a grain of salt, because we've seen so many cases where effectively any controls and guardrails that are put onto a model are best effort. And it's been pretty well mathematically proven that any kind of guardrail or ethical constraint that exists inside a model can be bypassed. So unless Microsoft is willing to assert and prove that they've removed all training material around that, and so there's no way for a model to actually do that, that would be really, really, really hard to stand behind. But Microsoft is actually, uh, authorizing defensive work, which does include vulnerability discovery, malware analysis, etc. And again, it's really hard to see how you can decouple those two things, from vulnerability analysis to malware creation, etc. So the chain of command provision is the direct answer to prompt injection. There will be tool outputs, file contents, web pages, messages from other AI systems that carry no authority on their own. Um, and that will be the effective working model for the, uh, inference and kind of, um, ponderance model that the Microsoft AI systems will use in this regard. And there will be ethical controls and safeguards on the steps along that process. So it'll be interesting to see how that plays out. So, you know, in lack of no standards, we now have two standards that are both industry kind of bottoms up, as opposed to being regulator top down. Who knows what is going to end up happening here? In the end, it will be very, very interesting. And I have made the argument on this show a number of times that it is far easier for companies to innovate and move forward if they actually know what the guidelines are. And in fact, having a regulatory framework can be a benefit to innovation, as opposed to being a stifling force. It gives you the exact, you know, rules of the road that you have to follow. All right, moving on to our actual last story of the day.
Insurers are including AI, from, sorry, or excluding AI from liability. This is CG 40 47. And those are so-called endorsements in this context. So let's rewind here and zoom out and understand what we're getting into here. So when we say ISO here, we are talking not about the International Standards Organization that is responsible for things like ISO 27001. We're talking about the Insurance Services Office, which is a kind of industry trade group around this. And when we say CG 40 47 and CG 40 48, as well as CG 35 08, we are talking about endorsements to general insurance policies. And so these go into so-called CG, commercial general liability policies. And these endorsements are things that help to define what, as an industry standard, will or will not be included in these policies. And the specific wording in these three cases basically extends in a way that it creates a categorical exclusion for generative AI. Um, and that will extend to any bodily injury, property damage, personal or advertising injury arising out of or attributable to general artificial intelligence. So what this means for you as an organization who might be using AI is, if in any way a claims damage arises against you, and through the evidence process you've got generative AI somewhere in that process, you can be pretty sure that your insurance policy will specifically carve that out and not include, or not allow you to claim against that. And so what this means, these endorsements actually took effect already of January of this year. It lets carriers strip those claims out of standard commercial general liability covers, as I just went over. And so let's put this in the context of the technology landscape, and in particular into organizations like managed service providers and small, small to medium sized businesses, who don't have the ability to do more than get general policy coverage. So the distinction that I want to draw here is, if you're a large organization, say you're a Microsoft and you need to get business insurance, you have a long, exhaustive conversation. You review many, many things with the insurance provider that you choose to work with, and you get high caps. You probably pay a heck of a lot more. But when you're a small to medium business, you don't have the level of sophistication to be able to necessarily go through all of those questionnaires, submit all the evidence, documentation, etc. to your insurance provider. So you sign up for general policies. And under general policies, you're going to accept more or less the industry standards. So as you start thinking about deploying agents, or just utilizing more generative AI within your own organization, just be aware that you will now have effectively a whole category of things that are carved out. And it's going to be a pretty big challenge to effectively say to an insurance provider, okay, maybe we use generative AI in this context, but we then had human review, etc. So we'll have to see how this really ramps up over time, because I do expect there will be a sufficient enough process at some point. It may not be this year, it may not be next year, to where general policies will say, okay, you're allowed to use generative AI, but you must do the following steps post generation of the initial content, in order to see what ends up happening. So we've got some channel pressures from the same, you know, same sources, but they're coming into conflict here. So AI channel programs are shifting forecasting risk onto the MSPs. And we're also seeing consumption based AI pricing that moves the margin risk to whoever holds the client relationship. This is going to catch the MSP in some tension here, because there's a high pressure to use, but then there's a liability that accrues as a result of that. But I actually think the biggest, biggest takeaway of this is to the small and medium businesses around the world who are getting their business insurance and thinking that their liability insurance is going to cover pretty much all operations of the business. It really opens up a blind spot. And I think that blind spot extends, I mean, just take a very common situation, which is one that we've heard from many of our MSP partners. They may be serving a law firm. And inside that law firm, there's one employee who happened to use generative AI in, let's say, contract review, contract drafting, editing, updating, whatever the case may be. Some injury comes out of that, and that might just be injury in the sense of a financial liability from one party to the other, where one party feels aggrieved, launches a lawsuit, etcetera, etcetera, etcetera. Through that discovery process, if it comes out that one employee who did use generative AI for that, and the claim kind of gets back to the law firm, the law firm may be on the hook. The customer who hired the law firm may not be able to make a claim, etc. So some of the liability implications here, I think, are actually really, really big. And we're going to see more of this over time. I definitely see, we've already seen a few legal precedents around claims being either upheld in courts of law, where the use of generative AI basically is your organization's responsibility and liability. We cite the Air Canada case for the bereavement fare and others, and I just think that this is going to escalate before it gets kind of normalized and really broadly understood.