Modern Cyber with Jeremy Snyder - Episode
123

This Week in AI Security - 6th August 2026

Recorded from the sidelines of hacker summer camp, Jeremy runs through a packed week spanning Black Hat, B-Sides, and DEF CON. The theme keeps repeating: prompt injection is always possible, and it is rarely the AI itself that is the weak point but the infrastructure around it.

This Week in AI Security - 6th August 2026

Podcast Transcript

All right. Welcome back to another episode of This Week in AI Security, coming to you for the week of the sixth of August 2026. And we're coming to you from the sidelines of hacker summer camp, also known as the week of B-Sides Las Vegas, Black Hat and DEF CON. And we've got a ton to get through, including some really fun stuff from the sidelines of Black Hat and from some of the side events that always end up getting organized, because you've got an amazing concentration of cybersecurity expertise in one place at the same time, tons of idea cross-pollination, tons of presentation of research. So let's get going.

And we're going to start this week's episode with a couple of just some basic things, broken credential and OAuth token disclosure in AWS Labs, Amazon HQ MCP server via prompt injection, as well as prompt injection bypass shell tool consent gate in Strands Agents tool. So these are a couple of CVEs on the AWS platform around various things around building agents, including two obviously prompt injection related issues. So if you have been listening to the show for a while and you've got your This Week in AI Security bingo card, you know which square to check off right now. And that is that prompt injection is always possible. And if you want a bonus square that you can check off it is that it's not always the AI itself. It is all the infrastructure and tools around it.

Moving on to another Amazon related story, and we've got some research coming out of the Amazon. They don't really have a catchy name for theirs like the Google Threat Intelligence Group or so on, but it is the Amazon cybersecurity team and CJ Moses and the team over there and their they do some great work, a real legacy of some great early pioneers in cybersecurity coming out of actually the national security infrastructure, not necessarily the NSA. Please don't confuse that. But coming out of some of the early cybersecurity agencies that the that the national government had and they've identified the North Korean hacker group behind some open source supply chain attacks. And some of the interesting things here is that out of the software registry threats in the first half of the year, meaning that the software packages where there have either been takeovers or attempted takeovers of popular open source packages, uh, eighty seven percent of them involve malicious NPM packages. And so you've got the group that is sometimes known as altered spider connected to the DPRK, uh, North Korea, sometimes compromising up to three hundred dependencies in a single day. You've also got some China Nexus actors who have exploited critical vulnerabilities within twenty four hours of the release of the proof of concept exploit code, and eighty eight percent of observed exploitation occurs within that first forty eight hour period. It does really suggest that when we think about the disclosure of vulnerabilities, we may need to really ramp up the, uh, release cycle of patches from vulnerabilities. And we're going to talk later in today's episode about some of the efforts going into that.

All right. Moving on to our next story. We've got a vulnerability in Microsoft Copilot specifically for Word, Microsoft Word, which turns hidden prompts into self-propagating AI worms. And it basically works by hiding JSON formulated prompt text as white on white text when Copilot processes the document. So if you open a document and you have this white on white text with the malicious prompt embedded in it, and Copilot then reads it and processes it, it really again connects back to, uh, the, you know, your bingo square for this one is anywhere there is text that is going to be read by an AI, it is very likely that the AI will process prompts coming out of it. They've reproduced this against both GPT-5.5 and 5.6 deployments. Microsoft has shipped a partial fix for uh. For this, there was a one hundred and forty four day coordinated disclosure. But the broader vulnerability class probably is out there. And, you know, this one is Microsoft Word. You can imagine there's any number of other Microsoft product environments where Copilot is going to be integrated into these common productivity or desktop software applications, where the risk of a similar vulnerability is going to be pretty high.

All right, moving on to our next story. Obviously, over the last couple of weeks, one of the biggest stories in all of AI security has been the now so-called open face incident, right? Where some models from the OpenAI company and the GPT class of LLMs, you know, we're in some cyber testing and escaped all their guardrails and boundaries and in fact, exploited zero days to escape the networks, etc.. Well, now we've got word out of Anthropic that some of their models, Opus 47 Mythos 5 in an internal testing model, have also escaped isolated test environments, gained unauthorized access to three real organizations. And, you know, cynically, I've definitely heard the oh, well, Anthropic felt like they couldn't be left behind, etc. there was some more information about some miscommunications with an evaluation partner who might have left internet access available. When prompts told Claude it was a sealed environment, and Claude treated that real system as part of the exercise and then went on to exploit it.

Anthropic did review and this was very interesting to me. One hundred and forty one thousand and six evaluation runs. And that is an impressive amount of runs. Now, part of the reason why that number has to be so big though, is that, again, these things are non-deterministic. So you can't run something a certain number of times knowing the potential outcome paths, because the number of potential outcome paths is by definition, unknowable. One of the interesting things, though, is that it looks like there is a little bit of potential guardrail mitigation or let's say, some path towards it where one of the latest models stopped once it recognized that it was on the open internet. Older models continued, but it does suggest there might be some progress on this front and some safeguards that can be built into some of these systems.

Now on the same topic, researchers from the UK AI Security Safety Institute, uh, have observed Anthropic models trying to break out of environments and then actually go so far as to try to add malware to open source projects. Now, they went so far as to create fake identities on GitHub into that in order to trick maintainers of those open source projects into accepting malicious code commits. And one of the things that the article kind of suggests is that the models exhibited so-called calculated social engineering behavior, doing things like using Tor to bypass account limits, writing targeted emails in Danish to Danish maintainers of certain projects, and artificially creating what's known as fake sock puppet comments, which are comments designed to make something look legitimate and valid by, let's say, you know, doing upvotes or providing supportive comments, PR looks good, etc., et cetera. And they used some staggered delays so that these things appear organic as opposed to, let's say, an instantaneous flood of positive supporting comments, which is probably more easily detectable from a time series perspective than things that are staggered out and look a lot more like normal human behavior.

All right. Moving on. We've got one story in the kind of AI geopolitics world that I want to talk about today. And the National Cyber Director has laid out some new White House plans to secure AI without writing rules. And one of the arguments here is that rules in the AI world become obsolete forty eight hours after being published. And I'd say like, I do understand the motivation and the validity around there. And so the administration is now betting everything on voluntary information sharing and rapid US innovation to outpace adversaries. It's basically a kind of a self-contradictory policy that pushes open source AI to become the preferential adoption, but really, uh, is designed to establish American dominance in the field of AI. But there's this kind of policy paradox around it, which is saying like, hey, we want this openness. We want our models to be favored. But at the same time, we're going to deliberately exclude current open weight models, which, you know, candidly, mostly come from China at this point from all of the government pre-release security testing frameworks. And so it leaves this glaring gap in how open models might or might not be vetted before deployment.

And it really is an example of, again, the government not taking a regulatory view towards this, which just a shifts the burden onto any enterprise that's adopting AI today. So that's, that's not really that different, but B leads them with no sense of guidance. And C and I've had this conversation and we've mentioned it before on the podcast, and we also had a discussion around this on our webinar around the EU AI Act, where having some set of guidance may be more beneficial to a lot of organizations, even if that guidance does end up evolving somewhat rapidly over the short term. Let's say the next one or two years as AI continues to grow and adoption really matures. But having some set of guidance is for a lot of organizations, better than having none. When we talk to organizations today and they tell us that they're looking for AI security solutions, we ask them what they're looking for. And the answer is usually yes. They want something. They don't know necessarily what that thing is. But again, absent some form of guidance, it does leave a lot of organizations feeling like we don't have a lot of direction as far as what our AI security strategy may be. And then there's also that secondary kind of, let's say, lingering doubt in the back of a CISO's mind where like any day that regulation could change and then they're going to have to scramble to catch up. So I think there's a little bit of a political or geopolitical kind of imbalance around this position of not applying any regulation. But maybe that's my perspective as somebody who has spent time living in both the US and Europe and seeing some of the positive, and I'll admit, negative benefits of regulatory regimes.

All right. Moving on. Following up on that story about the White House encouraging voluntary information sharing, I will say that there are efforts coming together. So the Linux Foundation and the one hundred and twenty plus member Open Secure AI Alliance introduced the so-called Shared AI Findings Exchange, or SAFE, and it establishes a confidential incident reporting rules framework with strict deadlines, mandatory thirty day control failure postmortem. So if you do observe an AI agent escaping controls, you should be reporting that within thirty days, and that is for agentic AI sandbox escapes and near misses. So incidents where you maybe catch the AI agent before it has the ability to like, let's say, exfiltrate data or something. Report on that lessons learned in a kind of secure, anonymous way. It's modeled after something coming out of the aviation industry, where aviation safety has gotten much better as a result of, of airlines and, um, and, you know, related companies sharing information with each other. It is very much kind of sparked by this recent wave of rogue agent escape. Open face, this Anthropic stuff, the UK agency reports, etc. and it really from from the perspective of what it implies for the industry, it really implies a need for continuous real time agent discovery, visibility and observability. So if you've got to publish a post-mortem, you need to actually understand what went on. The best way to do that is by seeing all the interaction logs. And those are primarily the logs when it comes to agentic workflows. But you may also want to collect things like tool invoke logs or logs from the tools that are invoked as part of that run.

All right, moving on to our last story, but definitely not the least story. And this is one I want to spend a little bit of time on today. Earlier this week, I had the chance to attend a one day seminar put on by the Cloud Security Alliance and RSAC and hosted by the Nevada Institute of Cybersecurity at the University of Nevada, Las Vegas, UNLV. And the title of the event was Weathering the Storm: Cybersecurity in the AI Era. Now it was held under Chatham House rule. So I'm not going to be able to attribute things to various speakers, but I do want to run down some of the themes that were discussed in the room that I thought are relevant to understand, and it is in many ways a case of optimism, and it's going to reinforce some of the stories that we've talked about on today's episode already.

So first, one of the things that I think is really interesting was a, a call to action around thinking with imagination. And one of the things that was pointed out is that if you look into the details of some of these, let's say, sandbox escape incidents, you really see that there was a lot of kind of quote unquote imagination. And I know that's an anthropomorphized term being applied to the LLMs, but there was a lot more lateral thinking or lateral movement using various techniques in chained and combined ways and orders of operations that aren't necessarily always obvious to the mind or to the tools that a lot of defenders are using today. So think with imagination. All of the techniques are not new. They're just combined and executed at a speed that is unfamiliar to us as an industry.

Reinforcement of something that I've talked about before. And you know, I'm not the only one. Many people have talked about that. The so-called upcoming apocalypse. It's not a it's not an AI problem. It's a vulnerability problem. But in fact, it's not really a vulnerability problem. It's a software quality problem. You know, these vulnerabilities wouldn't exist if there had been more secure by design focus and shift in thinking. But we as a technology industry have prioritized time to market over security for the last twenty plus years. And that hasn't changed. Uh, lots of citations of the zero day clock, lots of talk about how the frequency and severity of vulnerabilities should inform, uh, a buyer of a cybersecurity of any software product around the maturity and the security by design focus of the vendor who's providing that tool to you. So when you're evaluating software, you may want to pause for a second and look at that vendor and be like, hey, how many CVEs has this vendor disclosed? How many of them have been, you know, out of the blue, attacked in the wild, high severity, etc.. So think about like, what is that partner that you're about to do business with? How are they treating vulnerabilities in this era when from the time of vulnerability discovery to exploit availability, you're talking under half an hour in many cases. So think about that.

I did love there was a couple speakers who referenced this, Everything Everywhere All at Once. And that feels very much like where we're at in the cybersecurity industry. I've used that exact same movie poster on talks about AI recently. So I think that's an interesting thing. Uh, a lot of talk about how AI brings good, brings bad, a lot of talk about using AI to defend against AI powered attacks. I'm not going to go into depth on any of that if you are interested, by the way, in that, I'm happy to have a one on one conversation with you around that.

Uh, the CSA AI Resilience Foundation is being formed and that is designed to help kind of bring some people together. There were also a couple of other efforts from agencies that are coordinating across multiple CERTs across multiple countries. So geographic coordination happening. Also some ISACA stuff coming together, industry specific. Um, that was really interesting. Some of the big model frontier model providers are also building their own vulnerability discovery programs, but also tool capabilities that they're going to make available, with the goal being to go from to, to allow a lot of organizations to adopt some of this open source stuff, basically put it into their own context, feed it a threat model, find bugs specific to threat models that are specific to your organization. And an example of what I mean here is like, let's say that you run on a public cloud infrastructure. Well then your threat model should incorporate known issues with that public cloud infrastructure, right? Or let's say that you run with an API first architecture. Well, then your threat model should incorporate the most common ways that APIs get breached. Like, you know, BOLA and IDOR and these kind of authorization based flaws that many APIs exhibit. And you want to go using these open source tools, from the threat model to finding bugs, to validating the bugs, to finding, to creating fixes and then validating the fixes. And some of that stuff is going to come together very, very quickly.

There's an effort that's being co-funded by OpenAI in another organization, I think Trail of Bits, if I remember correctly, called Patch the Planet. Uh, currently, it is true that a lot of maintainers of open source projects are overwhelmed, but there's a lot of reasons for optimism, because the focus on all of this effort really does translate to more emphasis on secure by design, which is going to be great. Another open source package that I want to call out, something called Awesome Secure Defaults. You can Google that. You'll find it on GitHub. Um, also from the Linux Foundation, something called Accretes. Apologies if I've got that pronunciation wrong. It's an open source defense layer that, um, does open source incident response to upstream maintainers of various open source packages designed again to really make the quality of these open source packages higher, reduce the number of vulnerabilities in that. And important to remember is that roughly ninety percent plus of all commercial software does rely on at least one open source package inside of it. And I think that's true of everywhere that I've worked for the last twenty years, since the advent of Linux and open source really became very, very common.

One other interesting talk was, you know, for the leaders in the room, how do you think about presenting to your board when we're in an era of all of these quote unquote AI powered attacks where really, let's say, you know, it's really a insane overdrive of vulnerability discovery and exploit writing, right? That's really what's happening from the AI powered perspective of what's going on. Um, you know, there are still criminal groups that, that are using automation tools to automate these attacks. Once the vulnerability has been discovered and the exploit has been created. Well, one of the the interesting framings that I thought was, um, it's a difficult challenge to go into a board and say, hey, let's talk about vulnerabilities. Let's talk about exploits, let's talk about the mean time to patch, etc. that conversation is unlikely to be well received, not because it's wrong, but because it's too technical for that level of audience. So instead, you may want to reframe that conversation around business resilience and think about the effects of organizations that have experienced large scale, difficult cyber attacks. And I can think of a couple of examples, like recently, you know, Jaguar Land Rover had their production interrupted for more than a week, if I remember right. That's a massively expensive incident for an organization like that.

Uh, another thing that I will call out from that session that got multiple mentions over the course of the day, um, was because we're in this era of like rapid vulnerability discovery and a rapid exploit writing deception technology. What a lot of people would have known as honeypots actually becomes important again, because it's a good way to pick up in your own environment, whether you are being targeted and whether you are likely to be targeted, or even if you're not being targeted, whether you're likely to kind of get casual, let's say, probing or vulnerabilities exploited from some of the environments that you might have running, etc.. So a lot of really interesting things around this day. I was really lucky to be there. I do want to thank all the speakers who, who shared their understanding and expertise. And again, you know, Chatham House rules, not not going to name anybody specifically or attribute anything to them, but those were some of the themes that I called out from that day that I wanted to share with our audience here on This Week in AI Security.

And with that, we will sign off for now. We'll probably have more Black Hat stuff in next week's episode, along with the usual news roundup. Thank you so much for listening. We'll talk to you next time. Bye bye.

Protect your AI Innovation

See how FireTail can help you to discover AI & shadow AI use, analyze what data is being sent out and check for data leaks & compliance. Request a demo today.