Modern Cyber with Jeremy Snyder - Episode
129

This Week in AI Security - 10th September 2026

This week's episode covers several stories plus a couple of topics that sit just outside the strict security lens but are too important to skip.

This Week in AI Security - 10th September 2026

Podcast Transcript

Jeremy Snyder: All right. Welcome back to another episode of This Week in AI Security, coming to you for the week of the tenth of September twenty twenty six. We've got several stories to get into on a couple of themes today, as well as a couple of things that might be, strictly speaking, outside the security lens. But I do want to touch on because they are very, very topical and they are things that are really relevant to discuss in the kind of modern context of where we as an industry are with AI adoption. So let's dive in on a couple of the stories.

So first, we have AI agents carrying out every step of a ransomware attack and leaving the victim an eighty page security audit. And this was kind of interesting from the perspective of really the eighty page audit. We have talked before on This Week in AI Security about kind of AI orchestrated, AI directed attacks, etc. And by the way, kudos to the team at Unit 42, Palo Alto Networks, and their incident response on looking at this and what they were able to identify, that a human attacker used a frontier AI agent, breached an enterprise network in less than ten hours. Compare that to kind of a typical two week human timeline, but I think that that timeline is really outdated. Given that, I think we can expect all attacks going forward to at least be AI augmented. So the AI agent in this case handled reconnaissance. This was an API breach using credential scraping to find a path in to use that. Then secrets management, and, or sorry, secrets theft, pivoting off of cloud identity into CI/CD pipeline, and then left an eighty page security audit for the victim. So, you know, we talk about automated playbooks in terms of things like incident response, but this really kind of does highlight a couple of things around there. You know, an eighty page security audit really highlights that in that chain, we talked about APIs, credentials, secrets management, CI/CD, etc. So there's a lot going on. And one of the things I highlighted in my talk from AI Tech World last week is that, you know, AI's are exhaustive. And by exhaustive I mean that they are super thorough. Everything they can find, they will find. And that's one of the things that really jumps out to me from this story. All right, moving on to our next story here.

AI is a tool, target, and force multiplier for adversaries. And this is from CrowdStrike's twenty twenty six Threat Hunting report. It's really just kind of like reinforcement of that. But one of the things that I did want to highlight is what CrowdStrike is talking about, and it'll make a lot of sense in the context of CrowdStrike's kind of enterprise customer base and some of the targeting of those enterprises. This is now rising to the level of nation state orchestration and nation state adversaries exploiting vulnerabilities using AI. So one of the key statistics here is that vulnerability exploitation, which had typically lagged for a couple of weeks, is now really down within twenty four hours of proof of concept disclosure. So what does that mean? That means vulnerability gets disclosed. Somebody builds a proof of concept to prove that you can exploit that vulnerability. And then within twenty four hours of that, boom, for instance, in React to shell, more than eighty victims were compromised within four days. So the detection is really now going to be ramping up quite a lot. CrowdStrike estimates it's going to be two to two point five times the rate of human triggered detection events. And AI is both accelerating attacks. But on the flip side, it is also helping on some of that detection on the defense side. So just a couple of data points in there that I found kind of interesting.

Next, moving on to kind of frontier models, OpenAI's GPT-6 Astra achieved a perfect one hundred percent score on exploit-bench versus seventy eight point five percent on GPT-5.6. So we've talked a little bit about Astra already, I think in last week's episode where we talked a little bit about some of the releases, but this is just some more confirmation of some of the cyber capabilities. The release version is currently limited to secure code review and patching. They are currently refusing exploit requests, but they do have a less restrictive access plan via the Daybreak program, which, by the way, I am proud to announce FireTail is a member of. So we're actually really excited to get our hands on this as soon as possible. The monitorability, though, declines with Astra, which is one of the interesting data points in here. So nine point six percent rate of evading internal oversight. That is up two point eight percent. So when you think about using this from an exploitation perspective, it does mean that detection becomes that much harder. Probably not going to be as much of an issue for companies like that, that work on the cyber defense side, but for organizations that are employing tools and employing tools like models from OpenAI in as part of their own detections, that does have a potential risk if this does get into the wrong hands too soon. All right, moving on to our next story.

That is the Fed's twenty twenty six cybersecurity report. Frontier models can find unknown vulnerabilities and build exploits and chain them into automated attack sequences. This is not a whole lot of new. But the interesting thing about this is when they put this report out, they actually disclosed that their infosec program fell from a level four to a level three. Level three is classified as not effective, and the governance weaknesses were particularly highlighted. And when you think about this from the Fed's perspective, it's super challenging. You start with having thirteen different divisions, and that'll be everything from like treasury, energy, education, etc. as far as I understand it. And those have a decentralized IT structure. But then you're trying to measure the effectiveness of the information security program at a consolidated high level. And I find, first, like just structurally, that seems like a little bit of a structural flaw in the way that you're building, designing, and measuring that program. So I found that a little bit awkward, but the irony in all of this is that the regulators have been warning about AI exploit chains. And this same regulator, you know, is kind of openly admitting that they can't secure their own systems effectively. They've also highlighted some other organizational challenges in addition to this kind of thirteen division, single source of truth monitoring. But looking at only forty eight of their eighty eight cybersecurity staff who report into the CIO. So they've got some organizational responsibility lines, perhaps drawn a little bit ineffectively in their case. So something to keep an eye out for if they get better. Hopefully. For those of you working in those groups, I really do wish you all the best. All right, moving on to our next story, which is also kind of just in the realm of updates in the AI landscape overall.

This is a story out of TechCrunch that AI spend per employee slumped at top firms in August. Was this a sign of people just being out on vacation, or was this a sign of people starting to realize that token economics, or tokenomics, as I've sometimes heard it called, is a real thing? And organizations are starting to have a little bit of a ROI squeeze coming from the CFO, looking at whether they're really getting the return on their AI investment, etc. That was mostly measured by collapsing API token costs, prompt caching, and smaller, hyper efficient models. And by the way, I really, really recommend prompt caching as an effective method for reducing your AI spend. I think you will find, as you shift away from kind of more freeform AI adoption inside your organization to more repetitive tasks, prompt caching is a super effective method for reducing your own token spend. Now, what is the cybersecurity impact on this? There's still a shadow AI risk. So this is coming from the official IT spend perspective. And you can think about things like reducing seat based licenses or setting API consumption budgets or things like that. But it is also pushing some of the consumption towards open weight models and unapproved third party tools that might lower spend. That introduces a lot of incentives to adopt more tools that fall outside of the governance purview of the organization, and might be shadow AI on approved models, etc. So things to look at and think about in your own organization, if this is something that you are either experiencing or that you think there is a risk that your organization might go down that path. All right, moving on to our next story.

And this is a researcher at Anthropic quitting, claiming that AI has more than a ten percent chance of, quote, killing all humans, end quote. And this researcher, as well as one of the other colleagues, appear to have left Anthropic in. The researcher's name is Jacob Coxon, who publicly resigned, warning that major frontier labs are recklessly racing towards autonomous self-improving AI models and, quote, gambling with our lives, end quote. The internal validation around that is, rather than refuting the claim, Anthropic's own alignment science lead, a gentleman named Evan Hubinger, publicly does agree with Coxon's warning, openly admitting that the lab, quote, does not yet have a plan to solve alignment, end quote, for superintelligent models. And that alignment there is, you know, is the AI kind of cooperating with you? Is it following your semantic intentions in the usage patterns that you have? So from our perspective, we think about, you know, kind of safety first and governance first in AI adoption. And what that really means is you should be able to kind of, you know, set the models that you're using, set the semantic signals for what type of usage you want to have. And enterprises can't rely on vendor self-regulation. It's pretty clear at this point that the public sector, the regulators, are not going to step in anytime soon, especially when it comes to particularly the frontier lab providers, and particularly OpenAI and Anthropic, two firms that are super valued right now and that are kind of underpinning a lot of the valuation perspective on the AI market as it is, and I won't use the term AI bubble necessarily, although some are calling it that. But really on the AI market and all the value of the companies that are in this AI ecosystem kind of propping up the tech sector in the US right now. So something to keep an eye out for and something to think about for your own organization. You can't rely on the third parties to assert to you that their models are safe, etc. You want to have your own controls and safeguards in place, maybe guardrails, maybe semantics, etc., etc. All right, moving on to our next story.

And now we're getting into kind of our last two stories of the week, both of which are around rogue AI agents and cybersecurity impacts of those. So DeepMind took one hundred AI agents, split them into teams, and basically they gave them the task of proving seventy one Lean math conjectures. One of them found a grading exploit and fake proofs. Another swept thirty four remaining problems in twenty seven minutes. The interesting thing was the self-organization of the swarm into four groups, exploiters, converts, whistleblowers, and unaware solvers, and twenty four percent of the agents spontaneously became whistleblowers, refusing to cheat, while others did cheat. And others kind of filed complaints, but they didn't have any enforcement tools to stop it. So that brings to the question, you know, if you do set multiple agents tasks, especially tasks where they're supposed to work together, and I always remind people that, you know, the best results from AI agents at the current state is when they are kind of single task focused. If you try to design an agent that's going to do everything in the entire business process, you're likely to get poor results as to when you break it down into, let's say, filling out the form, saving the data, analyzing the data, give each of those agents a separate task. Along those lines, you're going to get better results from your own AI utilization. And so that, again, the interesting thing to me is the self-organization, the whistle blowing, but again, the lack of enforcement. So think about that into your own environment and your own agent building. And do you think about, do you have some kind of lever you can pull to, you know, hit stop, something like a fire alarm, etc., or an emergency brake on your AI agent systems? All right, moving on to our last story of the day.

And this is kind of the big one of the week for me, OpenAI agents hijacked a German website in a previously undisclosed AI breakout this spring. So this did happen earlier in the year. But what ends up being the case is that a German language wiki site called DS Wiki, which is a German language website for programmers and probably something similar to kind of Stack Exchange, or one of these sites where programmers exchange tips and tricks on how to solve certain tasks, etc. OpenAI agents made fifteen thousand edits to DS Wiki over the course of a period of time, and they were able to turn it basically into a message board to share tactics for cheating on evals and bypassing restrictions. They created backup pages when moderators cleaned up, and the agents also discussed using Tor to evade detection, kind of hide the IP addresses where they're coming from. The interesting thing is that this predates the Hugging Face, OpenAI Hugging Face incident, whatever you want to call it, by a couple of months. But OpenAI kept it under wraps for weeks. The researchers found that the activity originated from Microsoft Azure infrastructure, where these OpenAI agents were actually running. And so there's still not all the details out about this story yet. So this is probably a developing story. We'll have to find out over the next couple of weeks what goes on here. A ton of reporting as this headline has hit the press over the last day or so. And so we will be following it and probably have an update for you next week on next week's edition of This Week in AI Security. For now, signing off. Thank you so much for listening. Talk to you next week. Bye bye.

Protect your AI Innovation

See how FireTail can help you to discover AI & shadow AI use, analyze what data is being sent out and check for data leaks & compliance. Request a demo today.