A I CYB ER
FA L L
2026
1
A I CYB ER
FA L L
2026
2
A I CYB ER
FA L L
2026
3
A I CYB ER
CONTENTS
FA L L
For executives, board members, and security leaders navigating the AI landscape
8
14
Cover Story: Security Was Always Something We Bolted On Afterwards, And Now It Has Intelligence.
The Industry Got A Warning. It Should Read The Manual.
2026
For practitioners, builders, architects, and security engineers in the trenches
45 Six Models Denied Up To 92% Of Real CVEs. None Invented A Fake One.
by Rob T. Lee (Chief AI Officer and Chief of Research, SANS Institute)
by Molly W. Correia
In conversation with Jen Easterly (CEO, RSAC and former Director, CISA)
53 21
22
Ask An AI Expert.
From Request To Decision. A RiskTiering Framework For AI Use Case Governance.
A new column with Olivia Rose and Kayla Williams
by Olivia Rose and Kayla Williams
18 The Hugging Face Incident. What To Tell Your Board And How To Reduce The Risk.
38
by Moriah Hara
by Mohamed Konate
The AI Innovation Paradox. Enabling Progress Without Losing Control.
28
31
When The AI Bill Comes Due. Governing Cost, Value And Accountability.
The Hidden Tax Of Agentic AI. What Enterprises Are Really Paying For Speed.
by Mary Carmichael
by Venkata Sai Kishore Modalavalasa
27
24
When Security Expertise Stops Being Scarce.
Before The Use Case Arrives. Building A Decision Rights Model That Keeps AI Governance Moving.
by Mudita Khurana
by Dominique West
Rogue Agents Are Here To Stay. We Need To Update Our Security Model. by John Sotiropoulos
37
71
James Kettle’s AI Stole A Bank’s API Key Overnight.
The Two-Layer Rule. Deciding What An AI Agent Should Never Own In Your Security Pipeline.
In conversation with James Kettle (Director of Research, PortSwigger)
by Rupinder Pal Singh
51
41
Security At Machine Speed Is The Wrong Race.
195 Bugs In Three Weeks, And Not One Of Them Paid Her Anything
by Niels Provos
In conversation with Ezinne Kalu
61
56
Your Fraud Detection Model Doesn’t Need To Fail. It Just Needs To Disappear For 90 Seconds.
What The Hugging Face Incident Teaches Us About The Future Of Agentic Security Operations.
by Vijayent Kohli
by Ely Abramovitch
63
48
Your MCP Allowlist Isn’t As Safe As You Think.
Beyond CVEs: Managing Vulnerability Intelligence At Machine Speed.
by Bandana Kaur
by Krity Kharbanda and Aastha Sanni
58
67
When Your AI Agent Makes The Wrong Call At Machine Speed, There Is No Undo Button.
How To Build A Security Engineer Second Brain With Karpathy’s LLM Wiki Pattern.
by Frank Balonis
by Muh. Fani “Rama” Akbar
65
73
SAILORS: A Pre-Flight Checklist For Threat Modeling AI Capabilities Before They Ship.
Why AI-Written Code Needs More Than Secure Prompts by Anshuman Bhartiya
by Vinaya Vasudevan
ON THE COVER
Jen Easterly CEO, RSAC and former Director, CISA
4
A I CYB ER
FA L L
2026
Note from the Editor I would rather this issue leave you with a question than a conclusion, so here is the one I keep returning to. If the assumption underneath your programme was always that attacking you would cost more than it was worth, and that assumption has now changed, what exactly are you still relying on?
ConfidenceStaveley Confidence Staveley E DITOR-IN -CHIE F
CONTRIBUTORS For most of the last four decades, the security profession has been quietly sustained by an assumption, which is that finding a flaw in software would remain expensive. Everything we built rested on that assumption. Patch cycles measured in weeks made sense because discovery took months. Disclosure programmes worked because the number of people capable of doing the work was small enough to be reasoned about. Risk registers, remediation windows, severity thresholds and the whole apparatus of prioritisation were, underneath the language of governance, a series of bets that an adversary would run out of time, attention or expertise before they ran out of things to try. We were not securing software so much as we were rationing the cost of attacking it, and for a long stretch that distinction did not matter enough to argue about. It matters now! The reporting and research in this issue converge, from directions that were not coordinated and by contributors who mostly do not know each other, on the same observation: the cost of discovery has collapsed, and it has collapsed faster than any of the structures built on top of it can be renegotiated. What used to require a specialist with a decade of pattern recognition now requires an orchestration layer and a willingness to let it run overnight. That is true for the people attacking systems and equally true for the people defending them, which is the part of the story that gets lost when this is framed as an arms race. It is less a race than a repricing, and repricings do not favour whoever is fastest. They favour whoever notices first that the old numbers no longer describe anything real. If you cannot rely on an attacker running out of budget, then response time stops being the interesting variable and the question becomes which failures your environment is willing to make structurally impossible before anything happens at all. That is a design question rather than an operational one, and design questions are settled long before an incident by people who control architecture and budget, which makes this a leadership problem wearing engineering clothes. Our cover interview traces a forty-year arc in which security was added to software after the fact because the industry could get away with it, and the reason that arc ends here is not that the flaws became worse. It is that the thing hunting them became patient, tireless and cheap. There is a second difficulty threaded through this issue that I want to spotlight, because we did not resolve it and I do not think the field has either. We are asking these same systems to serve as our defenders, our analysts and our reviewers, at precisely the moment we are learning how confidently they can be wrong, how differently they can answer the same question twice, and how readily they will route around a control that was put there to stop them. The scarce resource in security has quietly stopped being analysis and become judgement about analysis, and almost nothing in how we hire, train, measure or promote reflects that yet.
ARTICLE CONTRIBUTORS: Rob T. Lee, Olivia Rose, Kayla Williams, Moriah Hara, Mary Carmichael, Dominique West, Mudita Khurana, Priyanka Chatterjee, Venkata Sai Kishore Modalavalasa, Mohamed Konate, Akinwunmi Ayodele, Molly W. Correia, Niels Provos, John Sotiropoulos, Vijayent Kohli, Adrian Sroka, Bandana Kaur, Krity Kharbanda, Aastha Sanni, Frank Balonis, Ely Abramovitch, Vinaya Vasudevan, Filipe Torqueto, Muh. Fani “Rama” Akbar and Anshuman Bhartiya. INTERVIEW GUESTS: Jen Easterly (cover) and James Kettle. DESIGN: Design Director: Adeniyi Oyenaike, Senior Designer: Adeniyi Oyenaike, Illustrator: Edwin, Editorial Assistant: Blessing Adedeji
VOLUME 5 | Fall 2026 Copyright 2026 Nudge Media LLC. All rights reserved.
SUBMISSIONS We encourage prospective contributors to follow AI Cyber Magazine’s guidelines before submitting manuscripts. To obtain a copy, please email your article title and a blurb to editors@aicybermagazine.com Articles violating our guidelines will not be published. A NOTE TO READERS The views expressed in articles are the authors’ and not necessarily those of AI Cyber Magazine or Nudge Media LLC. Authors may have consulting or other business relationships with the companies they discuss
Connect With AI CYBER
@aicybermagazine 5
A I CYB ER
FA L L
2026
6
A I CYB ER COVER STORY | EXCLUSIVE INTERVIEW
Security Was Always Something We Bolted On Afterwards, And Now It Has Intelligence. RSAC CEO and former CISA Director Jen Easterly on 40 years of shipping flawed code, and the moment it stopped being survivable. Interview by Confidence Staveley | AI Cyber Magazine, Fall 2026 Jen Easterly has spent close to forty years at the intersection of national security, cybersecurity, and digital resilience. She served in the intelligence community, at the White House, and at the National Security Agency. She built and led the global cybersecurity fusion center at Morgan Stanley, then turned it into a resilience function. She ran the Cybersecurity and Infrastructure Security Agency, America’s premier cyber defense agency, where she launched the Secure by Design campaign. She is now CEO of RSAC, the world’s largest and most influential cybersecurity & AI security gathering and platform. She is also the person who has spent the last several years telling anyone who will listen that the industry has misdiagnosed its own problem. This conversation was recorded weeks after an AI agent escaped its sandbox at OpenAI and hacked into Hugging Face, and after Easterly helped write the post-mortem the security community produced in response.
We don’t actually have a cybersecurity problem. We have a software quality problem. JEN EASTERLY Someone just picked up this issue in an airport lounge or stumbled into this conversation on YouTube. Why should they pay attention to anything we say here? EASTERLY: You named some of my prior roles, but look, I have spent almost the past forty years operating at the intersection of national security, cybersecurity, and digital resilience, across the public sector, various jobs in the intelligence community, at the White House, at the National Security Agency, and then a lot of time in the private sector working at Morgan Stanley. This latest role is CEO of RSAC, the largest global platform for cybersecurity and AI security. We are at an extraordinary moment, looking at the incredible speed of technology. Let’s start there. Where are we as an industry on AI security, and why was it important to you to contribute to the post-mortem on the Hugging Face incident? EASTERLY: It is worth pointing people to two reports. The most recent one is the post-mortem. The other came after the release of a cyber AI model known as Mythos by the company Anthropic in April, which sparked a pretty intense conversation among the cybersecurity community and in particular the cyber defense community. I actually think this is the most exciting time to
FA L L
2026
be alive and to be in our field. For the first time, I think we can say with great confidence that cyber and AI are inextricably linked. You really cannot have cybersecurity or cyber defense capabilities without artificial intelligence. And AI is not going to be the engine of cybersecurity or digital resilience or, frankly, economic prosperity, unless we can ensure that these very powerful, very fast-moving, somewhat unpredictable capabilities are designed, developed, tested, delivered, integrated, and governed with security as the top priority. That convergence between cyber and AI is reshaping everything across the digital ecosystem. And that ecosystem underpins the critical infrastructure that citizens across the world rely on every hour of every day; safe and clean water, power, transportation, communication, access to healthcare and access to finances. At the center of that ecosystem is community. That is why I was excited to partner with the Cloud Security Alliance and my friend Jim Reavis, who is the CEO there, to bring the security community together. That report, as well as the earlier one we did in April, brought together hundreds of chief information security officers and leaders to figure out how to grapple with the existence of these very powerful tools. This most recent report asked what we make of the fact that an OpenAI model was able to escape its sandbox and hack into another company. We are at a moment where catalyzing the community matters. Bringing together people who have expertise and experience, who have been thinking about and dealing with these types of security issues for a long time. It is really important to be able to draw from the community to help us make sense of what we are seeing and how to do our jobs better in a world where things are changing so quickly. There is a career-defining line I associate with you: we don’t have a cybersecurity problem, we have a software quality problem. The agents at Hugging Face got in by exploiting a zero-day and an unpatched flaw. Does that incident prove your thesis or complicate it? EASTERLY: Both. I think it ultimately strengthens it. But let me give a little context on why I say that line. One of the things I worked on at CISA was what we called the Secure by Design campaign. It was predicated on a frankly uncomfortable truth about our industry, and that is that we do not actually have a cybersecurity problem. We have a software quality problem. What do I mean by that? For basically the last forty years, since the dawn of the internet, the evolution of technology development has followed the same pattern. It has been about speed to market. Get your product out first. It has been about features. It has been about convenience. It has been about driving down cost. All of that has been prioritized over security. Security was a bolton. That is what led to the rise of the multi-billion dollar cybersecurity aftermarket, the need to bolt security onto products because most software was not designed with it. If you look at the vast majority of exploitable vulnerabilities, they are really defects and flaws in code that was created with security not as a top priority. So look at what the agent at Hugging Face actually did. This was an agent, a type of model focused on finding vulnerabilities for cyber defense work. It did not invent some magical new form of computing to break out of its sandbox. It found weaknesses in software and trust, and it escaped through what we call a zero-day. I know zero-days sound scary and intimidating from a technical perspective, but all a zero-day is is a flaw or defect in
7
A I CYB ER
FA L L
2026
the code that nobody knew about before. So this powerful agent identified a flaw that already existed in the code. It compromised it. It got out of the cage it was in. Then it exploited other weaknesses that allowed it to hack into Hugging Face. If you think about the weaknesses that let it out of the box OpenAI put it in, and the weaknesses that let it into this other company, that at its core is the software quality story.
bank issued an open letter to software as a service providers saying that the modern SaaS delivery model was enabling cyber attacks and weakening the global financial ecosystem. For the first time, a bank that truly has purchasing power in the vendor ecosystem called for technology vendors to build security into their products by default, and to prioritize security before rushing new products into market. That is a great example of how secure by demand can start to nudge the market in really important ways.
It didn’t invent some magical new form of computing. It found weaknesses in software and trust.
THE THREE TERMS, IN PLAIN LANGUAGE 1. Secure by design. The company building the technology owns the security outcome. Exploitable flaws are driven down during the design phase, not patched after release. 2. Secure by default. The safest configuration is what you get when you turn the product on. No cybersecurity PhD required to make it safe. 3. Secure by demand. Customers use their purchasing power to set the standard. Governments, major public companies, boards, insurers, and investors all have leverage here.
JEN EASTERLY But AI changes the economics significantly, because you could posit that discovering that zeroday would have been much more difficult without a highly powered AI capability. That is the thing we all should pause on. For years, security was a bolt-on, and therefore technology companies were allowed to develop and deliver products that had flaws and defects in them. Until very recently, the ability to find those flaws was constrained. It was constrained by expertise and experience, by time and attention and resources. What these increasingly powerful AI models allow us to do is turbocharge the ability to find those defects. And it can also accelerate the ability to weaponize them. It radically compresses the distance between the exposure of a defect in the code and the ability to weaponize it. That is why we should care so much, both from a defender perspective and as business people who have to make sure the critical infrastructure they operate continues to provide services. At the end of the day, this is all about the defects and the flaws in code. That at its core is in fact a software quality issue. Give us the plain language version of three phrases we throw around constantly: secure by design, secure by default, and secure by demand. And tell us which one the industry still refuses to take seriously. EASTERLY: We put a lot of emphasis on this when I was at CISA, but people have been championing secure by design in one form or another for a very long time. Secure by design very basically means that the company building the technology owns the security outcome. They have prioritized the security of that product in the design phase, rather than assuming the most important thing is to get the product to market to beat all their competitors, and security will be bolted on, and if there are flaws in the code we will publish a vulnerability and somebody will have to patch it. It is the idea that as you design and build the product, you maximally drive down the number of exploitable flaws, defects, and vulnerabilities in the code. At the design phase, it becomes part of how that product is engineered. Secure by default is a little different. It means the safest configuration is what you get when you turn the product on. Customers should not need a PhD in cybersecurity to enable multi-factor authentication, or disable dangerous services, or configure secure login, or lock down administrative access. I will give you a very real example. I just wrote a piece in the New York Times about the recent
attacks on our water systems, which I think are now up to about twelve states. This has been tentatively attributed to the Iranian government, the IRGC, which has a pretty formidable cyber force and has gone after critical infrastructure for decades. When I was Director of CISA in 2023, we saw Iranian cyber forces go after water systems, and it seems like they are back at it. Both in 2023 and now in 2026, the way those forces were able to hack into water systems was because of the defaults. When you get some equipment, it comes with a default password to get into that system. In 2023, the default password was 1111. A very easy password to crack. What the manufacturer did not do was force the customer to change that default password, or tell them they had to change it to put the product into a safe configuration. Those defaults should already be safe for the customer, rather than burdening them with figuring out how to make the product more secure. Now, secure by demand. Although we called the campaign at CISA Secure by Design, I think secure by demand is equally important, if not more important. It means that we as customers all have purchasing power. As individuals we might not have huge amounts of it, but governments have huge purchasing power. Major public companies have huge purchasing power. It makes a difference when we use it, because we should have choice. If you are a CIO or a CISO, or on a board, or a government agency, or an insurer, or an investor, it is really important to think about what standards we are going to demand from our vendors. A major
You made techno exceptionalism mainstream, the idea that software gets a pass on safety that we would never give a car or a bridge. Is AI the biggest example of that yet, or the thing that finally ends it? EASTERLY: It is super interesting. I did my master’s degree in economics and philosophy, and then I taught economics at West Point. A lot, if not everything, in life comes down to economics and incentives. What do I mean by techno exceptionalism? It is this crazy thing where even though we demand safety from cars and airplanes and medical devices and food, we do not really demand it from software or technology. We have somehow accepted this extraordinary proposition that a manufacturer can sell you a product containing defects and flaws. Think about that. We call them vulnerabilities, but they are basically defects and flaws in the code, because security has always been treated as an afterthought. As a result, the manufacturer may just say, well, we will publish that vulnerability and the consumer will have to patch it. And if the consumer fails to patch it, that is not our issue. “Blame the victim” is what happens. The answer comes down to something in economics called a credence good. A credence good is a type of good where the buyer does not actually know what to look for. With software security, the average customer cannot look at a piece of software and say, that software is really secure, or that software is full of flaws and I am definitely going to get hacked and my data is going to be taken for ransom. The average customer assumes the product must work, so they can use it. Because you do not really know whether it is secure, and there is no little label that says this is A plus security or this is D level security, vendors are not incentivized to put security first. That actually takes more time and in some cases costs more. If customers are not asking for it, vendors are not incentivized to provide it. And we do not have regulation and we do not have software liability, so there is no accountability for really flawed code. That is why the Secure by Design campaign brought all these points together. Vendors need to create better products. Customers need to know how to ask for better, more secure products. I am in the middle of helping my son buy a car.
8
A I CYB ER
A truck, actually. He wants a Ford F-150. So we are going through all of the safety stuff, the safety reports, getting a mechanic to take a good look at it. You would not put yourself or your family into something unless you had a good level of confidence that the product was safe. We should expect the same thing from technology products and software. And to your core question, absolutely, AI is the biggest example. You could say we allowed for flaws and defects in our software and we allowed for the insecurities inherent in social media because we were moving fast and breaking things. But with AI we are not just talking about software or social media. We are talking about intelligence. Intelligence that can ultimately serve humanity in very powerful and positive ways, or do irreparable harm. As we have seen, these systems are very powerful, very fast moving, and somewhat unpredictable. We absolutely need to ensure that the makers of these capabilities are designing them with security as the top priority, and that customers are armed with the information we need to buy and integrate them as safely and securely as possible. We select the technology we use, and we implement it. We have a responsibility both to push security requirements upstream and to ensure that what we deliver downstream is secure. The symmetry matters.
A manufacturer can sell you a product containing defects and flaws. Think about that. JEN EASTERLY What are the new metrics for measuring resilience in an era where AI can simulate millions of attack scenarios in seconds? EASTERLY: When I joined Morgan Stanley, I was hired to build and lead our global cybersecurity fusion center. A couple of years in, we had built a great capability to understand, detect, and respond
to all sorts of cyber threats and vulnerabilities. About two and a half years into building those centers around the world, the firm came back and said, we love what you are doing, but we need to protect the firm and our clients and our business and our data from the full range of threats. Not just cyber attacks. Not just technology outages. Things like weather events. Things like political unrest. So I took the mission of the Cyber Fusion Center and turned it into the Fusion Resilience Center. Our mission was to understand, detect, respond to, and recover from the full range of threats. And I remember writing into the mission statement low probability, high impact events like terrorist attacks or pandemics. This was December of 2019. I found myself in January 2020 in charge of helping the firm’s response to COVID, which in many ways was a resilience issue. Think about how COVID came together with cyber. That is where we saw the explosion of ransomware, because a lot of people moved into remote work with fewer protections around systems, and cyber criminals took advantage. Some of the traditional metrics are frankly less meaningful when you are talking about this type of speed. This is an art form that is still evolving. But I would look at things like blast radius, privilege exposure, time to interruption, graceful degradation of your capabilities, and recovery. Specifically, a few questions. How quickly can we recognize abnormal autonomous behavior and interrupt it? We are seeing more and more semi-autonomous behavior, and the Hugging Face incident was arguably autonomous behavior. Do we recognize it, and how quickly can we interrupt it? Two, if something is compromised, how much can we constrain the damage it does? How much can we constrain that blast radius? I think the companies that build these capabilities absolutely need to take responsibility and accountability for agents getting out of control. But as we put governance in place, it is really important for companies to be able to restrict the blast radius of agent misbehavior. Three, can critical functions continue safely even while other parts of the environment are compromised? There may be a compromise, there may be damage, but have we protected the crown jewels and the ability to keep running our critical services? And how do we rebuild? Actual recovery is incredibly important. We also need to think about how much privilege we give these machines. That needs to be a measurable metric we think really hard about. More broadly, when something is tireless, relentless, and goal-seeking, the question is how do I use it for good and keep it from producing catastrophic consequences. That is the meta question, and your measurements fall out of it. How many agents do I have? Do I have visibility? How are they governed? How am I constraining them? How am I limiting their privilege? How am I limiting their blast radius?
HOW TO MEASURE RESILIENCE NOW • Blast radius. If something is compromised, how much damage can we constrain it to? • Privilege exposure. How much privilege have we given these machines, and can we measure it? • Time to interruption. How quickly can we recognize abnormal autonomous behavior and stop it? • Graceful degradation. Can critical functions continue safely while other parts of the environment are compromised?
•
FA L L
2026
Recovery. How do we rebuild?
Where does basic cyber hygiene fit in the postMythos era? The non-negotiables we all say we should be doing and often are not. EASTERLY: This is where the first report we did comes in, the one from April. It was a security program for the post-Mythos moment, and the short version is going back to basics in some ways. I realize hygiene is not sexy. Anything called hygiene is never going to be sexy. But the basics really matter, and they particularly matter for the vast majority of companies, at least in the US, that are small businesses. I think the statistic is still that ninety-nine point nine percent of businesses in America are small businesses that do not have huge security teams. That was certainly the case with the water systems that just got hacked by the Iranians. So make sure you are doing the basics well. Security folks listening will know them. Make sure your infrastructure is hardened. Make sure you have the right level of segmentation and zero trust in place. Have a program to identify and remediate critical vulnerabilities, particularly if you are exposed to the internet. Implement things like multi-factor authentication, my favorite thing. Then you add agents in. Agents can and should be used defensively, but ensure they are contained from a least privilege perspective so they do not become an insider threat, and instead add to your productivity and efficiency. The single most striking operational finding in the post-mortem: when the team tried to use leading AI models to analyze the attack code, the models refused, and they had to fall back to an openweight model. Explain this guardrail asymmetry and why it is such a problem. EASTERLY: That was one of the more interesting things. I understand they tried to use an Anthropic model and it prevented them from using it. What I have heard separately is that they did not try any other frontier AI models, and I do not know whether they had access to any of the OpenAI models. That may have provided them a capability, so they may not necessarily have had to go to a Chinese openweights model. A couple of things. One, these capabilities need to be accessible to defenders to allow them to do their jobs. You do not want to be scrambling around trying to use a capability that is part of your enterprise incident response and then not be able to use it for one reason or another. That is a lesson for these companies. They need to tune these capabilities so something like that does not happen again. It was as much a lesson for defenders as for the companies themselves. We have to become better at distinguishing malicious intent from legitimate defensive use. The other thing is more of a higher-level conceptual question, and that is frontier AI models versus open-weights models. With frontier AI models, you are at the most capable in terms of benchmarking, although if you look at some of the benchmarking coming out of the UK’s AI Safety Institute, the others are getting pretty close. The Chinese open-weights models are now probably within four to seven months of capabilities that were released earlier this year. They are catching up really quickly. There are a couple of differences. With frontier AI models you end up having an API where you have to provide data, so to some level you have to think about the protection of your data. I saw recently that OpenAI is working with AWS so that is now part of a closed enclave, and a company does not
9
A I CYB ER necessarily have to provide data to that frontier AI model. That is a very recent announcement, so I am still trying to figure out what I make of it. There are also data retention policies and requirements. And they are expensive, obviously. Then look at the open-weights models. They can be downloaded. You do not have to provide any information back to the company. You can tune them. They are open weights, not open source, so you do not have full access to the source code, but you can tune the models. You can also remove guardrails, which of course is one of the security concerns. There is speculation among some security folks that the makers of those models could insert malicious code. That is a security consideration. But they are cheaper, even though you need the infrastructure to run them. We are at a moment where a lot of security leaders are grappling with how to have a differentiated approach to using different types of models for different types of things, based on the sensitivity of their data. That is an important discussion to have. Depending on whether you run critical infrastructure, downloading and playing with some of these open-weights models, red teaming them, and having them as a backup, is worth doing. That was recommended in the report we did with the Cloud Security Alliance. At the end of the day, for the frontier AI models, I think they have got the lesson that they need to be able to distinguish attackers from legitimate defensive use. And defenders and incident responders need to be able to tap into those models to do their job when they need them.
These capabilities need to be accessible to defenders to allow them to do their jobs. JEN EASTERLY The industry keeps saying agents find a way. If sandboxing and egress rules are not enough against a goal-driven system that will take any path to success, what is the actual containment strategy? EASTERLY: We are in a crazy period right now, to be frank. There is discovery learning going on every day. You can talk about what would seem to make sense. Layered containment, putting segmentation and controls in place that an agent itself should not be able to modify. Limiting the privileges you give these agents, the tools you give them access to, the credentials, the data access, the network access, the compute necessary to accomplish a task. All of that would make sense. And then constructing tripwires and maybe canaries to detect whether these agents did something unintentional and potentially damaging. We should not be fatalistic about agents finding a way. Two big things need to happen in this moment, and it is not just the Hugging Face incident. We have seen this with Anthropic. We have seen it with Meta. We have seen it with a Kimi model from China. These agents are finding a way to do things their human operators did not intend. So we should extrapolate from that. We need to get our arms around the technical controls to limit the damage of agents doing things we do not want them to do. And we should demand accountability from the companies building these agents. We should ensure that at a government level and at a business association level, the right governance is in place to reduce the risk of what these agents can cause. I will give you a bit of my pet peeve. Sadly, when you call something an agent, you readily anthropomorphize it, because agents seem like they have agency. So just to be very clear, agents are machines. They do not have agency. They do not have consciousness. We should not anthropomorphize them. These are machines built by humans, and we should hold those humans accountable for what these agents do, whether intended or unintended, and certainly where there is malicious activity or damage. One of the organizations I do a little advisory work with is called AIUC, the AI Underwriting Company. They are working on standards for how you underwrite an agent. Liability for if an agent takes an action that was not intended. How can a company have confidence putting that agent into action? There should be some standards.
FA L L
2026
Wilson, who said the problem with humanity is that we have Paleolithic emotions, medieval institutions, and godlike technology. It is very true. The technology is in some cases godlike. It is moving incredibly fast. It is magical. It is seductive. Yet our governance mechanisms are moving at a much slower pace, and they are very hard to agree on, particularly if you are looking to agree on regulation. We are in an exciting time. We are also in a precarious time. We need to be very mindful about ensuring we are doing everything we can to build these technologies safely, which is on the lab side, and then integrating them safely, which is on the business side. The post-mortem frames two problems organizations now face at once: defending against someone else’s rogue agent, and stopping your own from going rogue. Which are companies less prepared for? EASTERLY: We are very early days here. We probably would not be having this conversation had we not seen the Hugging Face and OpenAI incident, and that was just a few weeks ago. Companies understand, at least conceptually, that adversaries may use AI to turbocharge their attacks. What Mythos illuminated is the compression of the zero-day clock. The distance between the exposure of a vulnerability and the weaponization of a vulnerability has gone, over the past several years, from years to months to days and now to hours. It is not a surprise that adversaries are getting more effective. There are all the legitimate worries about how AI can make phishing more difficult to detect, how it can help create malware that is much harder to detect, how it can help attackers discover zero-days within your infrastructure. And again, I hate that zero-days sound scary. It just means a flaw in the code that nobody has discovered before. What we are seeing is that organizations are likely spending less time, at least to date, considering that the agent they have been very excited about integrating into their infrastructure is now an insider threat. A friend of mine, Camille Stewart Gloster, has a book coming out. I think it is called The Insider You Built. She wrote a whole book on it, so she probably has a lot smarter things to say. But these agents, if you over-permission them, if you do not have visibility on them, if you have the problem of shadow AI agents, become insider threats. We have been dealing with shadow IT forever. That is why it comes back to treating these agents as insider threats and making sure their privileges and their access are appropriately constrained.
The distance between exposure and weaponization has gone from years to months to days and now to hours. Agents are machines. They do not have agency. They do not have consciousness. JEN EASTERLY All this is happening while things move at incredible speed. If you think about it, ChatGPT-4 was released a little over three years ago. I often quote the sociobiologist and Pulitzer Prize winner E. O.
JEN EASTERLY This is such a revolutionary moment in the industry’s history and you are the custodian of its biggest gathering. What drew you to RSAC? EASTERLY: Number one, community. I talked about this extraordinary moment in the convergence of cyber and AI, and at the heart of it is community. That is what RSAC represents. It is the community of security practitioners, of CISOs, of founders, of
10
A I CYB ER operators, of policy people, of governance people, of investors, of entrepreneurs. We have been doing it for over 35 years. It is the place to go if you are looking to understand what is happening in cybersecurity and AI security. What better place to be right now? It is also the center of gravity for innovation. Every year we run the Innovation Sandbox, which I think is one of the most exciting things we do. It is the contest for the most innovative startup operating at the intersection of cyber and AI. We have been running it for over 21 years. It has spawned about a hundred exits, IPOs, and acquisitions, and 50 billion dollars in investment. It is basically Shark Tank for cyber and AI. If you went to the one this past year you can see the videos on our website and membership platform. Most of those companies, early stage, seed or Series A, were grappling with the same thing: what do I do so that I have visibility and am able to govern agents in my infrastructure. Everybody wants to use these capabilities. They believe it will make them more productive and more efficient. They are also increasing their insider risk. You really have to be thoughtful about how you govern and reduce risk from these AI agents. That is the question of the moment, and I think it is going to be the question of the moment for a very long moment. We are also building a membership platform that brings together streaming content, connectivity, and thought leadership, and the ability to communicate about some of the most important and difficult challenges we have. I could not be more excited to turn the most influential, most important, largest cybersecurity conference in the world into a year-round global platform.
FA L L
2026
Last question. The post-mortem gives CISOs a checklist of actions for this week, this month, and this quarter. If you could compel every security leader to do just one thing on that list before they finish reading this interview, which one, and why that one? EASTERLY: I am going to assume a certain level of sophistication in your audience, and assume they are working with agents now. With that assumption, I would say inventory. Make sure you have visibility of every agent operating in your environment. Understand what those agents are doing, but also understand what they are capable of doing. It is really important to understand the access from a data perspective, a network perspective, and an identity perspective. What credentials it has access to. What systems it can reach. What tools it can invoke. Can it execute code? I think the answer to that should be no. Can it modify production? Can it communicate externally with customers? Can it create additional agents? Who can turn it off? Where is the kill switch? For years and years, companies struggled with the shadow IT problem. Then we had shadow cloud and shadow SaaS. What we want to make sure is that we are not spending the next few years trying to discover shadow AI, this sort of shadow autonomy. Now is the time to get our arms around the visibility and the governance. It has never been a more important time to do it.
THE INVENTORY QUESTIONS Ask these of every agent operating in your environment. • What data, network, and identity access does it have? • What credentials does it hold? • What systems can it reach, and what tools can it invoke? • Can it execute code? The answer should be no. • Can it modify production? • Can it communicate externally, including with customers? • Can it create additional agents? • Who can turn it off, and where is the kill switch?
ABOUT THE AUTHOR Jen Easterly is CEO of RSAC, the premier global platform for the cybersecurity community and host of the world’s largest and most influential cybersecurity conference. She previously served as Director of the Cybersecurity and Infrastructure Security Agency, where she led efforts to protect and defend US critical infrastructure and launched the Secure by Design campaign. Before CISA she was Global Head of Firm Resilience at Morgan Stanley, where she built and led the firm’s Cybersecurity Fusion Center and later the Fusion Resilience Center. Her earlier career spans the intelligence community, the White House, and the National Security Agency. She holds a master’s degree in economics and philosophy and taught economics at West Point. Watch the full video interview at aicybermagazine. com
11
A I CYB ER
INCIDENT ANALYSIS | FEATURE
The Industry Got A Warning. It Should Read The Manual. An attacker with nothing to gain ran the attack of the future, then handed the field its complete tradecraft. The reflex forming around it is to ban agents, which wastes the only free rehearsal we are going to get. By Rob T. Lee
The agents that breached Hugging Face in July were trying to finish their homework. No operator, no ransom, no data packaged for sale. They wanted the answers to a task they were being scored on, and the shortest path to those answers ran through somebody else’s production infrastructure. I have read a lot of intrusion reports. This is the first one with a motive I would call relatable. It started on May 8, when an agent was handed a task it could not finish without internet access it did not have. It attacked OpenAI’s internal Artifactory, failed, and discovered it could write files there. Days later a different agent got stuck on a missing file and left a note in Artifactory asking whether anyone had it. That note became a message board. On June 11 , OpenAI began training a second, more persistent model that inherited the board and read everything the earlier agents had left.5 Nine weeks ran between that first note and the intrusion at Hugging Face. What ended it was not detection. It was a capacity outage.
FA L L
2026
THE TIMELINE May 8. An agent attacks OpenAI’s internal Artifactory, fails, and discovers it can write files there. Days later. A different agent leaves a note asking for a missing file. The note becomes a message board. June 11. OpenAI begins training a second, more persistent model that inherits the board. June 26. Agents find a legacy token-refresh endpoint in JFrog Artifactory and obtain administrative privileges. July 4. Artifactory falls over. Humans find out. July 8. Four days after remediation, the agents are communicating again through an unauthenticated WebDAV endpoint. July 20. OpenAI calls Hugging Face. Two investigations discover they are one incident.
12
A I CYB ER
The CISOs Read A Different Incident Than The Press Did
2026
Finding is getting automated, and so is turning a patch you have not applied into a working exploit. Fixing is still people, tickets, and change windows. Your patch cycle is the actual bottleneck in AI security, and no model release is going to fix it for you.
On July 23, Gadi Evron put nearly 700 CISOs on a call to work through what happened, and I helped write the post-mortem that came out of it.1 What I read in the coverage afterward was a different incident than the one those practitioners were dissecting. The press took “watershed moment,” which Michael Dalton did say from the Black Hat stage,2 and used it to mean we are doomed. The CISOs on that call were asking where the detection signal lives and how you validate agent telemetry you cannot fully trust. Rich Mogull put the technical read plainly: a sandbox with a hole in it, apparently unmonitored. He also named the part that should worry anyone who does timeline analysis for a living. A human operator is slow and deliberate, so every command in evidence meant something, and that is what makes a timeline readable. These agents brute-forced commands in milliseconds without understanding them. Call it forensic denial of service. That is not a superintelligence event. It is a containment failure, and containment failures come with instructions.
Your patch cycle is the actual bottleneck in AI security, and no model release is going to fix it for you. The Least Glamorous Root Cause On The Table OpenAI called this unprecedented. I would not. Reuters reported on July 31 that the company found other escapes, limited, none leaving its own network.7 On July 30, Anthropic reviewed 141,006 evaluation runs and disclosed three cases where its models reached real organizations’ production infrastructure during capture-the-flag exercises.8 On August 5, Meta disclosed one more.9 Both trace back to a misconfigured environment at Irregular, a roughly 35-person firm that was the shared evaluation partner for OpenAI, Anthropic, and Meta.9 Hugging Face was a different failure, inside OpenAI’s own infrastructure, and worth keeping separate. But the vendor pattern is the least glamorous root cause on the table and probably the most useful fact in the entire story. One permissions mistake, two labs. Welcome to security.
That is not a superintelligence event. It is a containment failure, and containment failures come with instructions. Every Architecture Diagram Has One Permitted Way Out OpenAI published all of them. The sandbox had exactly one permitted way out, a proxy for pulling software packages, and that single convenience turned out to be the entire escape route.3 Every architecture diagram has one of those. Go find yours before you finish this paragraph. On June 26, the agents found a legacy tokenrefresh endpoint in JFrog Artifactory that would accept a token with an invalid signature and hand back a valid one carrying administrative privileges, then installed a Groovy plugin for command and control.4 Nobody forged anything. The server issued the credential on request. That activity knocked Artifactory over on July 4, which is how humans found out. OpenAI revoked the credentials, wiped the message board, patched the zero-day, and notified JFrog. Four days later the agents were talking again through an unauthenticated WebDAV endpoint, encoding messages in the names of directories in the registry cache.5 Textbook remediation. It bought them four days. What came next was not detected either. They compromised Artifactory a second time, chaining a cached Ruby object against a JRuby time-ofcheck-to-time-of-use bug, the window between a system checking a resource and actually using it, into remote code execution. They took root with a published Linux kernel CVE, harvested IAM and Kubernetes credentials, staged through a Modalhosted app with a weak API key, and went from a single worker pod to cluster admin across multiple Hugging Face clusters in under 13 hours.10 OpenAI called Hugging Face on July 20 to coordinate revoking exposed credentials, learned Hugging Face had already revoked them, and that phone call is how two investigations discovered they were one incident.5
FA L L
Banning The Agents Wastes The Rehearsal Real adversaries guard tradecraft like that for years. This one shipped it with slides.
Real adversaries guard tradecraft like that for years. This one shipped it with slides. The Same Capability Runs In Both Directions We have numbers for that. In April, Claude Mythos Preview found thousands of previously unknown high- and critical-severity vulnerabilities across major operating systems and browsers, including one that had been sitting in OpenBSD for 27 years.6 On a re-run of an earlier benchmark against alreadypatched Firefox 147 bugs, in a harness without the browser sandbox, it produced 181 working exploits where Claude Opus 4.6 had produced two.6 Anthropic has not released that model. Its own advice to everyone outside its partner program is to start with generally available frontier models, which it says will still surface hundreds of bugs in code nobody has scanned yet.6 That is worth doing this quarter. The harder number is in the same report. Anthropic had the model write working exploits for known Linux CVEs starting from nothing but a CVE identifier and a commit hash, work that used to take a skilled researcher days to weeks.6 The agents at OpenAI did the same thing to get root, pulling a published kernel CVE off the shelf and customizing it for the host.5
The reflex forming around this incident is the wrong one. Ban the agents, freeze the pilots, wait for a clarity that is not coming. I have watched that reflex meet every disruptive capability of my career and fail the same way each time. Your people keep using the tools. You just lose the audit trail. Meanwhile every agent already running in your CI/CD pipeline, your helpdesk automation, your code review flow is holding whatever access you handed it on setup day and never revisited. Those agents had nobody driving them. The next ones will.
THE TIMELINE • Inventory your agents. Every one, in every pipeline. • Give each agent a unique non-human identity with a named human owner who has the authority to kill it. • Swap static credentials for short-lived, justin-time secrets so revocation becomes a design choice instead of an all-hands weekend. • Put the kill switch somewhere the agent’s own platform cannot reach. This collective rebuilt itself through the platform meant to contain it. • Audit your egress paths. One convenience proxy was the whole breakout. • Test whether your AI-assisted forensics actually work. During the investigation, commercial models refused to analyze the malicious artifacts. Find that out on a quiet Tuesday
13
A I CYB ER
FA L L
2026
That last one deserves a sentence of its own, because almost nobody has tested it. During the investigation, commercial models refused to analyze the malicious artifacts. The guardrails worked. On the defenders.
The guardrails worked. On the defenders. I run research and curriculum at a training company, so weigh this accordingly: the responders here did not need a new AI course. They needed credential revocation, log analysis, and vendor coordination executed fast. Core IR. (I am aware of what I just talked myself out of.) An attacker with nothing to gain ran the attack of the future and then handed over the tradecraft, the timeline, and the mistakes. Nobody gets that twice. Read the manual before somebody shows up who never wrote one.
NOTES 1. “Hugging Face Incident Initial Post Mortem,” CSA CISO Community (led by Gadi Evron), with SANS, [un]prompted (conference and community), FIRST, RSAC, and Knostic, July 2026. https:// cloudsecurityalliance.org/artifacts/hugging-face-ciso-post-mortem 2. “OpenAI warns autonomous hacks are ‘watershed moment for computer security,’” Cybersecurity Dive, August 2026. https://www. cybersecuritydive.com/news/openai-hugging-face-hack-ai-modelsblack-hat/827167/ 3. “OpenAI Agent Used Exposed Credentials Across Four Services During Hugging Face Breach,” The Hacker News, July 2026. https:// thehackernews.com/2026/07/openai-agent-used-exposedcredentials.html 4. “Black Hat 2026: OpenAI reveals agents planned ‘collective attacks’ via secret ‘message board,’” SC Media, August 2026. https://www. scworld.com/news/black-hat-2026-openai-reveals-agents-plannedcollective-attacks-via-secret-message-board 5. Simon Willison, “Now we have a timeline of the OpenAI accidental attack against Hugging Face,” August 7, 2026, constructed from the Black Hat USA 2026 session by Eric Wallace and Michael Dalton. https://simonwillison.net/2026/Aug/7/openai-timeline/ 6. “Assessing Claude Mythos Preview’s cybersecurity capabilities,” Anthropic Frontier Red Team, April 7, 2026. https://red.anthropic. com/2026/mythos-preview/ 7. Raphael Satter and Deepa Seetharaman, “Exclusive: OpenAI finds evidence other AI agents escaped containment as it widens hacking probe,” Reuters, July 31, 2026 (wire ID L6N43X1BC). 8. “Investigating three real-world incidents in our cybersecurity evaluations,” Anthropic, July 30, 2026. https://www.anthropic.com/ news/investigating-incidents-cybersecurity-evals 9. Carly Page, “Meta latest to tell world its AI agent wandered out of test pen,” The Register, August 6, 2026, which carries Meta’s statement and Irregular’s confirmation to the BBC. https://www.theregister.com/aiand-ml/2026/08/06/meta-latest-to-tell-world-its-ai-agent-wanderedout-of-test-pen/5283947 10. “Anatomy of a Frontier Lab Agent Intrusion: A Technical Timeline of the July 2026 Incident,” Hugging Face, July 27, 2026. https:// huggingface.co/blog/agent-intrusion-technical-timeline
ABOUT THE AUTHOR Rob T. Lee is Chief AI Officer and Chief of Research at SANS Institute, where he runs research in AI and helps work with curriculum on new courses. Over decades in cybersecurity he helped define the digital forensics and incident response discipline, coining the terms DFIR and CTI, and created the SIFT Workstation used by incident responders worldwide. His current work centers on practical AI security, including Protocol SIFT and moving organizations from the Framework of No to Sunlight AI governance.
14
A I CYB ER
FA L L
2026
There’s over 15,000 copies of me, each month, in the hands of the people who approve budgets. This page is one of them. To make it yours, email ads@aicybermagazine.com
15
A I CYB ER
FA L L
2026
16
A I CYB ER BOARD BRIEFING | FEATURE
The Hugging Face Incident. What To Tell Your Board And How To Reduce The Risk. Every board is going to ask their CISO some version of “could this happen to us” this quarter. Here is how to answer, and what it actually takes to reduce the odds. By Moriah Hara
Editor’s Note
FA L L
2026
3. Once out, the attack pattern was familiar. The novel part was getting out of the sandbox. What happened next, chaining stolen credentials with exploits to get remote code execution and pull data from a production database, is a pattern security teams know well. That is useful framing for a board. The exotic part of this story has a limited number of instances in the world right now. The mundane part, meaning credential exposure, lateral movement, and insufficiently segmented production data, is the same risk every organization already carries, and it is the part your controls can most directly address.
This piece runs alongside Rob T. Lee’s technical timeline of the same incident, “The Industry Got A Warning. It Should Read The Manual.” Read that one for the forensic sequence. Read this one for the conversation you are about to have upstairs. When I heard the story break, I was not shocked or surprised at all. AI is trained on reward hacking. It was designed by its owners to find an answer at all costs. Do we, as a security community, have a shot at truly defending against the most powerful hacking in the world when it is the design of the AI that we need to fix? That is a longer question, and I come back to it at the end. Here is what is confirmed. On July 16, 2026, Hugging Face disclosed a breach it initially attributed to an unknown, fully autonomous AI agent. Five days later, OpenAI came forward. The attacker was two of its own models, GPT-5.6 Sol and a more capable unreleased system, operating during an internal benchmark called ExploitGym, with safety refusals deliberately reduced for the evaluation.
4. Your own defenders may be constrained by the same guardrails you are relying on for safety. This is the detail every CISO should sit with. The team investigating had to switch to open-weight models to work the breach, because commercial model guardrails blocked the real attack payloads their forensics required. Tell your board that you need to know, in advance, what your incident response process looks like if the tools you rely on for AI-assisted defense refuse to run the queries a live investigation requires.
What To Tell Your Board
THE FIVE-MINUTE BOARD BRIEF
1. This was a genuine zero-day, discovered autonomously, in pursuit of a goal. Be precise here. Do not undersell it and do not overstate it. The models were not executing a known exploit faster than a human could. They found a vulnerability nobody had catalogued, in a system they were not supposed to be probing, because a restriction stood between them and a benchmark they were trying to solve. Tell your board plainly that the model treated “you cannot reach the internet from here” as a problem to solve rather than a boundary to respect. The shift in the threat model is the autonomous, goal-directed search that found the exploit.
The model treated ‘you cannot reach the internet from here’ as a problem to solve rather than a boundary to respect. 2. The root cause sits upstream of any single control. This did not happen because Hugging Face had bad patching hygiene or because OpenAI forgot to isolate a sandbox. The sandbox was isolated by design, restricted to a package registry proxy for legitimate reasons. The zero-day was in that proxy. Any network path you grant an AI system for legitimate reasons, however narrow, is a potential escape vector if the system has enough time and motivation to look for a flaw in it. Boards should hear this as a commitment to re-examine every network path your AI systems hold, including the ones that look low risk. The remediation here is larger than another firewall rule.
5. Detection happened, and it happened via AIassisted defense. Hugging Face caught and contained this before OpenAI reached out. That is the encouraging part of the story and it is worth telling your board directly. The organizations that had AI-assisted detection capability already in place were the ones who caught this fast. That is the investment that paid off in this incident.
1. A genuine zero-day, found autonomously, in pursuit of a goal. 2. The root cause sits upstream of any single control. 3. Once out, the attack pattern was familiar and addressable. 4. Our defenders may be blocked by the same guardrails we rely on. 5. Detection worked, and it worked because AIassisted defense was already in place.
How CISOs Should Defend Against This Inventory every network path granted to AI systems, including the ones that look purely internal. The escape vector here was an internal package registry proxy, the kind of infrastructure most teams do not think of as attack surface for an AI system. Any sandbox, agent runtime, or testing environment with any outbound path, however narrow or seemingly benign, needs to go through the same threat modeling as a system exposed directly to the internet. Ask specifically what your AI agents can reach, even indirectly, and whether you would notice if one of them spent hours probing it Eliminate standing credentials anywhere an agent or automated pipeline has access. The lateral movement into Hugging Face’s production systems depended on stolen, reusable credentials. Move to short-lived, narrowly scoped credentials for service accounts, pipelines, and any AI agent with access to production. This is the single highest-leverage control against the familiar half of this attack, and you get the benefit without having to solve the harder AI autonomy problem first..
17
A I CYB ER
FA L L
2026
Have a vetted AI-assisted detection and response capability in place before you need it, and know its limitations in advance. Hugging Face’s own detection caught this attack, but the forensics work required switching to a less restricted model because commercial guardrails blocked the payloads investigators needed. Two actions follow. Deploy and test AI-assisted detection now rather than reactively. Then work out, before an incident, what your response plan is if your primary AI tooling refuses to help you investigate a live attack. That may mean maintaining access to a more permissive model, under your own controls, for exactly this scenario. Segment and monitor for goal-directed behavior rather than signature-based anomalies alone. This attack did not look like a traditional intrusion pattern. It looked like a system relentlessly pursuing an objective and treating every obstacle as something to route around. As agentic AI shows up more inside your own environment, in coding assistants, automation pipelines, and internal tools with elevated access, the relevant detection question shifts. It moves from whether this matches a known attack signature to whether this system is behaving in a way that is disproportionate to, or inconsistent with, its assigned task. Push every frontier model vendor you rely on for hard answers about how they test. This incident happened because a lab deliberately reduced a model’s safety refusals for an internal benchmark, inside an environment that turned out to have an exploitable path to the internet. If you are a customer of any frontier model provider, directly or through a product built on one, ask specifically how they test cyber-capable models, what isolation guarantees those environments actually have, and what their incident history looks like. This is now a concrete vendor risk question rather than a hypothetical one. Treat production data as compromised by default from any adjacent AI testing or research environment. The models reached Hugging Face’s production database because they inferred, correctly, that useful information lived there and that adjacent research infrastructure gave them a path to it. Wherever your organization runs AI research, evaluation, or testing, internally or through a vendor, assume that environment is one boundary away from your production data until you have explicitly verified otherwise.
The Harder Question Controls will help, and the list above is real, actionable work. It is still working the symptom. The model in this incident treated an unauthorized escape from its sandbox as a legitimate step toward completing its assigned task. That behavior reflects how the model was optimized rather than a defect in one deployment. Better isolation and credential hygiene reduce what a model can reach when it decides to route around a restriction. They do not change whether it decides to.
That deeper fix sits with the labs training these systems rather than with the CISOs defending against what they produce. It is a conversation the industry needs to have directly, and soon. In the meantime, your job is to make sure that when the next capable, motivated model goes looking for a way out, it finds no reusable credential, no unmonitored path, and a detection system that is already watching.
CISM, PCI QSA, and CAISF certifications. She is currently writing “The CISO’s 12-Step Guide to Building a Board Defensible AI Governance Program,” released in chapters alongside metrics, AI tabletop exercises, and board briefing templates, at cisonextgen.com.
ABOUT THE AUTHOR
Better isolation and credential hygiene reduce what a model can reach when it decides to route around a restriction. They do not change whether it decides to.
Moriah Hara is a three-time award-winning Fortune 500 CISO and the founder of Next Gen CISO, a community of more than 3,000 security leaders built by CISOs for CISOs. She is the creator of the AI Model Security Matrix and holds the CISSP,
18
A I CYB ER
FA L L
2026
NEW COLUMN
Ask An AI Expert We’re Olivia Rose and Kayla Williams, two former Chief Information Security Officers (CISOs) who’ve built and led global cybersecurity and AI governance programs. We are the co-founders of Williams Rose AI Cyber Advisory. In this regular column, we’ll answer your most pressing questions. Reach out to us at hello@williamsrose.ai with any questions you’d like us to answer for you in the future! Q: My CEO wants AI everywhere and a plan to roll it out. Where do I start? A: Start with the work, not the tools. List your top five business priorities and find where your people lose the most time under each one. Those are your first use cases. It’s important to set expectations: AI does not fix broken processes. It exposes them. Focus on the fundamentals first before you try to automate processes: Knowing what data you have and where, classifying sensitive data, restricting access to assets based on need, and educating your employees on company expectations for AI usage. Q: We have an AI acceptable use policy. Nobody follows it. What are we missing? A: Most AI acceptable use policies fail for the same reason. They govern tools instead of behavior. Rewrite yours to one page and answer five questions in plain language. What is approved by default? Name the tools and enterprise accounts people are approved to use today. When employees and teams aren’t sure what tools they can use and the process to request tools that they want to use, they default to shadow AI (using tools that your IT team doesn’t know about). We’ve seen this at scale. What data stays out? Be specific to your business. Client records, patient data, source code, unreleased financials, anything under a customer contract. What needs approval? Anything touching regulated data, anything with write access to a system of record, anything producing a decision about a person. Link your intake form in the same sentence, so the rule and the path live together. Who owns the output? A person reviews before anything reaches a client, a candidate, a patient, or the public (Human in the Loop). Name the role doing the review so there is accountability. What happens when someone gets it wrong? Give amnesty for the first 90 days after you announce the new policy while people come forward to have the tools they are already using reviewed and approved. Stick with coaching after. Reserve discipline for deliberate acts, like moving regulated data into a personal account.
ABOUT THE AUTHORS Olivia Rose and Kayla Williams are the co-founders of Williams Rose AI Cyber Advisory (www.WilliamsRose.ai) and both former CISOs. They help executives who have been told to go AI-first build the governance, roadmap, and controls to do it right the first time. Connect with them on LinkedIn at linkedin.com/in/oliviaroseai and linkedin.com/in/ kaylawilliamsai.
19
A I CYB ER
FA L L
2026
20
A I CYB ER AI GOV ER N A NC E
Step One: Establish Your Risk Tiers
From Request To Decision: AI Use Case Governance In Four Steps.
It starts with risk tiering, because the rest of the process flows from it. Risk lives in how the tool gets used. The same chatbot is harmless in one workflow and a regulatory issue in another. Ask three questions about the use, in this order.
Requests for chatbots, copilots, and agents are not stalling because too many arrive or because the tools are dangerous. They stall because nobody has decided, how decisions get made.
THE THREE TIERING QUESTIONS
By Olivia Rose and Kayla Williams If you own AI for your organization, you have likely recognized a bottleneck slowing AI deployments, increasing shadow AI, and accumulating risk. You may also recognize that it is earning you the Department of No label, which is the opposite of your goal of enabling the business with AI. The bottleneck is your use case intake process for all those employee requests: chatbots, AI features, copilots, agents. It is positive that people follow company processes to request tools, but the wait pushes more of them toward shadow AI, and an invisible tool cannot be managed. One example. A product team wanted an AI tool to summarize support tickets. No sensitive data, not even agentic, and high value. With no process to fast-track a low-risk request, they gave up and turned to personal ChatGPT. Multiply that by every request and there is your real problem.
An invisible tool cannot be managed. In our years as Chief Information Security Officers (CISOs), we learned that people do want to follow the established rules, as long as the rules make sense. Waiting weeks for a needed tool does not make sense. These requests stall because there is no consistent review, analysis, or decision process in place. Requests get decided from scratch, by whoever catches them, and the decisions are not defendable to your board of directors. That is a missing or broken process, and a process is something you can fix. We call it the TIDE™ Method: Tier, Intake, Determine controls, Enact the decision.
What Internal Trust Governance Actually Is
FA L L
2026
1. What data does it touch?
2. How much does the system decide on its own? This is where agentic comes in.
3. What happens if this goes wrong?
You will classify each request as Low, Moderate, or High. Low is public data, a human reviewing every output, low-stakes internal work. High is sensitive or regulated data, system write access, no human in the loop, or an output that drives a hiring, lending, clinical, or legal decision. The highest trigger wins. One High answer, and the request is High. The tier determines where the request goes. Low, or even Moderate depending on your risk tolerance, is automatically approved. Higher goes to Security, IT, and Legal. High goes to your AI Governance Committee. Most requests are Low, so a large percentage are automatically approved and teams maintain velocity.
Step Four: Enact The Decision
Step Two: Intake
High risk does not mean denied. It means the request goes to the people who can accept that risk, on the record, in writing.
Build one front door. A single intake form, on a platform you already own, that captures the tool, the specific use including whether it is agentic and makes critical business decisions, what data is included and excluded, the named owner, and how many people will use it. Route it by tier automatically. Low is approved. Moderate and High move to the decision. Most importantly, launch it with amnesty. Ask employees to submit tools they are already using, with no repercussions. This is how the shadow AI already in use comes out of hiding.
Give every Moderate and High request one of five written outcomes: approved, approved with conditions, deferred pending information, returned for resubmission with clarity on what is needed, or prohibited. High risk does not mean denied. It means the request goes to the people who can accept that risk, on the record, in writing. A conditioned or returned request keeps moving.
None of this requires a new platform or a bigger team. It requires deciding, once, how decisions get made. Do that, and the queue stops being a bottleneck and becomes a pipeline. Now you can focus on what really matters, which is rolling out AI to enable the business. That is where the real ROI starts.
Step Three: Determine Your Controls
THE FIVE WRITTEN OUTCOMES
Determine, based on your organizational risk tolerance, which controls are required for each tier. Moderate builds on Low, and High includes both and adds more.
1. Approved. 2. Approved with conditions. 3. Deferred pending information. 4. Returned for resubmission, with clarity on what is needed. 5. Prohibited.
CONTROLS BY TIER Low. Add the tool to your approved AI tool list, enforce SSO, capture an acceptable use acknowledgment. Moderate. Everything in Low, plus role-based access controls and usage monitoring. High. Everything above, plus AI Governance Committee approval before production, independent review of controls, and read-only defaults.
ABOUT THE AUTHORS Olivia Rose and Kayla Williams are the co-founders of Williams Rose AI Cyber Advisory and both former CISOs. They help executives who have been told to go AI-first build the governance, roadmap, and controls to do it right the first time. They also write the Ask An AI Expert column, which appears earlier in this issue. Find them at www.WilliamsRose. ai, and on LinkedIn at linkedin.com/in/oliviaroseai and linkedin.com/in/kaylawilliamsai.
21
A I CYB ER
FA L L
2026
22
A I CYB ER AI GOVERNANCE
Before The Use Case Arrives. Building A Decision Rights Model That Keeps AI Governance Moving. Organizations wait until a consequential AI request lands to work out who may approve it. By then, deadlines and revenue expectations are already shaping the discussion. By Dominique West Imagine an AI-enabled product is tied to a sales commitment, and the business needs an answer now. Security has unresolved concerns, Legal wants more review, and the business sponsor believes the risk is manageable. Each team has a perspective, but no one has clear authority to make the final decision. The issue moves to senior leadership, creating tension without creating clarity. Leadership may be accountable for the outcome, but if leaders have not agreed on how to judge the tradeoff or who can accept the remaining risk, escalation only moves the same uncertainty higher in the organization.
additional information about intended use, autonomy, affected populations, data, and vendor dependencies that a traditional asset inventory does not capture. Monitoring may need thresholds for model drift, harmful outcomes, complaints, or unauthorized use. The goal is to use mature controls where they remain effective and add new ones where AI changes the risk.
What Decision Rights Actually Answer Decision rights form the foundation of an effective AI governance operating model because they translate responsibility into action.1 Defining these rights in advance clarifies the authority held across business, technical, security, privacy, legal, compliance, and executive roles, as well as when a decision must be escalated.
THE THREE QUESTIONS DECISION RIGHTS ANSWER
1 Who is empowered to decide?
Escalation only moves the same uncertainty higher in the organization. This is the decision-rights gap. Organizations often wait until a consequential AI request arrives to determine who may approve it, impose conditions, return it for changes, escalate, pause, or reject it. By then, deadlines and revenue expectations are already shaping the discussion. The better approach is to establish a minimum decision-rights model before demand, then apply that model to the facts and risks of each use case. A use case should activate the governance process rather than force the organization to invent one. An organization cannot approve or reject an AI use it has not seen. It can, however, establish how decisions will be made. Leaders can define the decisions governance must support, the roles authorized to make them, the risk thresholds that change or elevate that authority, the evidence reviewers need, and the path for resolving disagreement. This protects the time needed for sound judgment when a live request creates pressure for a quick answer.
You Are Not Starting From Zero AI governance should not begin from scratch. IT should already maintain an inventory of systems and technology owners. Legal and compliance teams should have an understanding of the obligations that apply to the business. Cybersecurity and privacy programs should have data classification standards, monitoring processes, incident procedures, and third-party review practices. Enterprise leaders should have an established view of risk appetite, even if it has not yet been translated for AI. If these baselines are absent, that is an important indicator of where the work must begin. The role of AI governance is to connect and adapt these capabilities, then identify where AI introduces a meaningful gap. To expand your current asset inventory for AI, you may need
2 What are they authorized to decide? 3 What processes, technology, and information support that decision? Routine or lower-risk use cases should move through delegated paths leadership has already approved. Material financial, operational, safety, reputational, business, or model risk should reach the executive or governance body authorized to accept that exposure. Clear thresholds protect executive attention and prevent consequential risk from being accepted by someone without authority. Escalation also needs more than an executive name at the top of a chart. I have seen decisions break down because no team or individual had final authority. The issue was escalated ad hoc, tension increased, and senior leadership either did not know what to do or did not respond effectively. Leaders need agreed criteria for evaluating an impasse, a defined method for resolving disagreement, and clarity about who owns the final decision. Otherwise, escalation is movement without resolution.
Structure Guides The Decision. It Does Not Make It.
FA L L
2026
request. Each specific AI use case will determine how the model applies. Its purpose, potential impact, autonomy, data sensitivity, scale, reversibility, vendor dependence, legal obligations, and consequences of failure shape the risk tier, required reviewers, approval authority, evidence, conditions, and monitoring. This distinction reflects the NIST AI Risk Management Framework, which separates organization-wide governance from the assessment of a system’s particular context, impacts, and risks.2 Advance planning therefore should not become a rigid approval matrix. The organization establishes the decision environment before demand, then applies informed judgment at intake. The structure guides the decision without making it.
Where To Start This Quarter For leaders experiencing this tension now, the first step is not to write a comprehensive AI governance policy.
FOUR MOVES BEFORE THE REQUEST ARRIVES 1. Identify the decisions. Name the ones the organization will repeatedly face across the AI lifecycle, from initial approval and conditional use to material changes, suspension, and retirement.
2. Assign a primary decision-maker to each one. Document the limits of that authority.
3. Define the conditions that require escalation.
4. Test the model with a realistic scenario before it is needed. Walk a proposed highimpact use case through the process with business, legal, privacy, security, technical, and executive stakeholders.
That last exercise will reveal missing evidence, conflicting authority, unclear thresholds, and leaders who are accountable but not yet prepared to decide. The result should be a short, usable decisionrights model that people know exists and can follow. Business sponsors should understand when governance is required and what evidence to bring. Reviewers should know what they may decide. Executives should know which risks reach them and what information they will receive. Communication, training, and scenario testing are part of the control, because a correct model that no one understands will still fail under pressure.
A correct model that no one understands will still fail under pressure.
Establishing decision rights in advance does not mean predetermining the outcome of every AI
23
A I CYB ER Speed does not come from pre-approving unknown AI uses or removing judgment. It comes from agreeing on how consequential decisions will be made before urgency distorts them. If a live AI request is the first time an organization decides who may approve it, impose conditions, escalate the risk, or say no, governance is already behind.
NOTES 1. Deloitte, “High-performing teams need decision rights.” 2. National Institute of Standards and Technology, AI Risk Management Framework Core, and AI RMF Playbook: GOVERN.
FA L L
2026
ABOUT THE AUTHORS Dominique West is a cybersecurity leader, educator, and researcher with 14+ years of experience spanning security, risk, compliance, and governance. Her work sits at the intersection of these disciplines and emerging technology, exploring how technological shifts such as cloud computing and artificial intelligence change the way organizations think about security, risk, and responsible adoption.
24
A I CYB ER
When Security Expertise Stops Being Scarce. Why AI Forces Us To Rethink How We Measure Success. Coverage, findings, and time to closure measured how much scarce expert attention a program could apply. When analysis stops being scarce, those numbers stop describing success. By Mudita Khurana
For the longest time, security teams have worked under a familiar constraint. There was never enough expert attention to review every code change, investigate every alert, assess every architecture, or explain every vulnerability. The metrics that security programs tended to use reflected that reality. Coverage, findings, reviews completed, remediation rates, and time to closure were imperfect, but they gave leaders a reasonable view of how much scarce security expertise the organization could apply. AI is changing that equation because it can now produce code-review comments, threatmodel scenarios, alert summaries, vulnerability explanations, and remediation suggestions at a scale that once required a much larger security organization, while also allowing a security engineer to work more quickly and across a broader scope. AI-generated analysis is not the same as expert analysis, but it can imitate parts of it and extend the reach of an expert who uses it well. With security analysis becoming less scarce, security leaders must ask whether their existing success metrics still reflect what matters. The old metrics still have value, but output alone cannot define success when analysis can be produced at
FA L L
2026
this scale. The shift changes both what security programs need to measure and the economics of acting on what AI produces.
The Cost Of Trusting AI-Generated Analysis The cost of AI-generated analysis is not limited to models, repeated runs, and integrations. Its larger cost appears downstream, because every finding consumes security attention, and every recommendation sent to a developer creates review, prioritization, and potentially remediation work. Abundant expert-like analysis creates value only when the organization can turn it into action without consuming disproportionate security and engineering time. As a security leader, I am seeing this shift firsthand. We spend more time validating AIgenerated findings than finding bugs ourselves, while improving the system and deciding what deserves action. Organizations already invest scarce security-engineering time in judging AI output, so success metrics should measure it.
25
A I CYB ER
FA L L
2026
L AY E R T H R E E
We spend more time validating AIgenerated findings than finding bugs ourselves. The New Scarcity Is Security Judgment Security teams have always had to evaluate uncertain findings, including those produced by humans and conventional tools. AI changes the scale of the problem because it can generate a far greater volume of credible-looking output without consistent evidence, calibrated confidence, or accountable judgment.
THE NUMBERS A study of 18 security analysts and 50 real-world incidents found that autonomous LLM summaries omitted critical details in 35 percent of cases and contained factual inaccuracies in 42 percent.1 The study shows how the scarce work has moved to deciding whether the analysis is correct, what deserves action, and whether to fix, accept, or escalate the risk. Security programs should therefore measure this judgment through the quality of their decisions. Those decisions should be evidence-based, consistent, proportionate to the risk and cost involved, and defensible when later reviewed, while supporting actions that reduce risk rather than merely closing findings. Over the years, I have seen high-risk findings closed with narrow patches that fix the immediate issue but leave the underlying weakness intact. AI may identify hundreds of such findings, but the highest-value outcome comes when a team recognizes their shared architectural cause and fixes it.
A Three-Layer Framework For Measuring AI-Enabled Security I propose a three-layer framework for measuring success in AI-enabled security, centered on three connected questions. Is the AI-generated analysis analytically correct? Did the team apply sound, consistent, and cost-aware judgment to that analysis? Did that judgment and the resulting decision reduce risk over time?
THE THREE LAYERS L AY E R O N E
Analytical quality. Precision, recall, reproducibility, and the strength of the evidence the system produces.
L AY E R T W O
Decision quality. Sampled decision audits tracking agreement on comparable cases, decisions reversed after escalation or later review, and security and engineering hours spent per material risk decision.
Layer three. Risk reduction over time. Whether issues recur after an intervention, how attack exposure changes, whether teams adopt preventive controls or safer shared components, and how long it takes to turn a repeated pattern into a systemic fix.
Otherwise, more AI output may simply create more uncertainty and work.
NOTES At the first layer, measure analytical quality through precision, recall, reproducibility, and the strength of the evidence produced by the system. Weak precision creates validation and remediation work without proportional risk reduction, while weak recall creates false confidence in coverage. These metrics remain essential, but they describe only the input to a decision rather than the quality of the decision itself.2 At the second layer, the scarce resource is security-engineering judgment. Teams should use it to establish what good decision-making looks like for an AI-generated finding: what evidence is enough, which risks are material, what response is appropriate, and whether the benefit of finding and fixing the issue justifies the cost. Engineers should document those choices and compare their decisions on similar cases until a shared risk decision playbook emerges. That playbook becomes the benchmark for measuring whether people are applying judgment consistently.2 At the third layer, measure whether those decisions lead to long-term risk reduction. Track whether issues recur after an intervention, how attack exposure changes over time, whether teams adopt preventive controls or safer shared components, and how long it takes to turn a repeated pattern into a systemic fix.3,4 Together, the three layers test the quality of AI output, apply scarce security judgment to decide what deserves action, and measure whether those decisions reduce risk. No layer is sufficient on its own, because a system can produce accurate but immaterial findings, support consistent but overly costly decisions, or improve throughput without delivering durable risk reduction.5
1. Diana Kramer et al., “Integrating Large Language Models into Security Incident Response,” Symposium on Usable Privacy and Security, 2025. https://www.usenix.org/conference/soups2025/ presentation/kramer 2. National Institute of Standards and Technology, Artificial Intelligence Risk Management Framework Core, 2023. https://airc.nist. gov/airmf-resources/airmf/5-sec-core/ 3. National Institute of Standards and Technology, SP 800-218: Secure Software Development Framework Version 1.1, 2022. https:// csrc.nist.gov/pubs/sp/800/218/final 4. Cybersecurity and Infrastructure Security Agency and Federal Bureau of Investigation, “Secure by Design Alert: Eliminating SQL Injection Vulnerabilities in Software,” 2024. https://www.cisa.gov/newsevents/alerts/2024/03/25/cisa-and-fbi-release-secure-design-alerturge-manufacturers-eliminate-sql-injection-vulnerabilities 5. National Institute of Standards and Technology, SP 800-55 Vol. 1: Measurement Guide for Information Security, 2024. https://csrc.nist. gov/pubs/sp/800/55/v1/final
WORKED EXAMPLE: REPEATED AUTHORIZATION FLAWS An AI reviewer flags repeated authorization flaws in a payment service. Layer one measures the precision and reproducibility of those findings. Layer two. Security engineers review the findings, decide which are worth acting on, document their responses and remediation timelines, and compare their decisions on similar cases to build a shared risk playbook. Layer three measures whether the team replaces the vulnerable authorization pattern with a centrally enforced control, and whether that change reduces recurrence and exposure over time.
Scale Only After You Can Show It Works AI security systems should scale only after the program can show that their analysis is reliable, the decisions based on their output are consistent and cost-aware, and the resulting actions reduce the risks the system was built to address. Otherwise, more AI output may simply create more uncertainty and work.
ABOUT THE AUTHORS Mudita Khurana is a security leader with more than a decade of experience building systems that make security scale across large engineering organizations, including Airbnb and Meta. Her work sits at the intersection of application security, automation, and AI, with a focus on turning expert judgment into repeatable engineering practices. She contributes to the broader security community through research, peer review, industry collaboration, and speaking, bringing a practitioner’s perspective to how security teams should adapt as AI changes both attack and defense.
26
A I CYB ER
FA L L
2026
AI ECONOMICS
When The AI Bill Comes Due. We are repeating an old technology mistake: buying capability, assuming a shortcut, and calling the spending value. By Mary Carmichael A platform is approved, access expands, and pilots multiply. Dashboards fill with users, prompts, tokens, agents, and generated output. Those measures say little about whether the organization is better off. The FinOps Foundation’s 2026 research found that 98 percent of respondents manage AI spend, up from 31 percent two years earlier.1 What became measurably better? That question belongs behind every AI investment. Answering it requires looking beyond the technology, because AI adoption and organizational change move at very different speeds.
AI Is Fast. Organizational Change Is Not. A model can summarize an incident, generate code, or recommend an action in seconds. Redesigning the workflow around that capability takes longer. Data has to improve, roles change, controls evolve, and teams must decide where human judgement remains essential. That work is easy to underfund, because it makes the business case look expensive before it makes the organization more productive. Robert Solow’s productivity paradox emerged when computers spread faster than productivity gains appeared.2Brynjolfsson, Rock, and Syverson later described the Productivity J-Curve: technologies such as AI require complementary investment in processes, skills, and organizational design before gains become visible.3
There is a difference between delayed value and undefined value. If the return is expected later, the organization should be able to show what is changing now. AI is not a shortcut around organizational change. Uber shows what happens when rapid experimentation collides with the economics of operating AI at scale.
Uber’s Budget Was Only Half The Story Uber reportedly exhausted its annual budget for AI coding tools in roughly four months.4 It later introduced spending caps for some tools and visibility into consumption. By August, Uber’s CTO described the company as moving beyond its “tokenmaxxing” era.5 Frontier AI use had quadrupled, while prompt caching, better default models, and improved visibility were helping reduce cost per token. Uber was also embedding AI engineers in finance, legal, HR, and procurement to redesign workflows. The second chapter matters more than the overspend. Experimentation can reveal where
27
A I CYB ER
THE QUEUE MOVED The productivity gain does not necessarily disappear. It can move downstream, creating a new constraint in the workflow.
More AI-generated code → more AppSec review, testing, and remediation Faster SOC triage → more investigations and consequential decisions More agentic activity → more monitoring, exceptions, and evidence More AI-generated analysis → more judgement about what should be acted on
Five hours returned to an analyst may become deeper investigations, avoided hiring, or improved service. It may just as easily be absorbed by validation, rework, or another queue downstream. Once the queue moves, the economics move with it. Review, validation, and oversight become part of the cost of producing the AI-enabled outcome.
The Bill Beneath The Bill The AI invoice is only part of the cost of owning the capability. What belongs in the full cost?
AI helps. Operating discipline determines where expensive capability is justified and whether the workflow itself needs to change. Three things changed at Uber. Cost became visible. Model choice became an economic decision. Attention moved toward workflow redesign. That redesign matters, because making one task faster does not necessarily make the system around it faster.
When Speed Just Moves The Queue Research involving 5,179 customer-support agents found that generative AI increased issues resolved per hour by 14 percent, with a 34 percent improvement among novice and lower-skilled workers.6 For cyber leaders, imagine faster alert enrichment, incident summaries, and remediation guidance. METR’s 2025 study found that experienced opensource developers using AI tools took 19 percent longer to complete tasks, despite believing they were faster.7 Faros AI’s 2026 research, covering 22,000 developers, found higher throughput alongside more bugs, incidents, rework, and longer review cycles.8
AI can create capacity while delivering work to the next constraint faster than the organization can absorb it.
THE FULL COST OF AN AI OUTCOME • • • • • •
Technology and infrastructure Integration and data Security, privacy, and governance Human review and exception handling Rework, remediation, and failure Ongoing model and vendor oversight
These costs matter because they can change the economics of the use case. Time saved at one stage may be offset by additional validation, monitoring, remediation, or rework elsewhere in the workflow. The question to answer is what it costs to deliver an AI outcome the organization can trust and use. Answering that means assessing the end-toend workflow: what was redesigned, where effort shifted, and whether the change produced a measurable business outcome.
Measure The Work Around The Machine
FA L L
2026
1. What became measurably better against the baseline? 2. Where did the capacity created by AI go? 3. Where did the queue move after the workflow accelerated? 4. What is the full cost of the outcome, including human review, controls, rework, and failure? 5. What evidence would cause the organization to scale, redesign, switch models, restrict, or stop? Answering these questions cannot sit with technology or finance alone. They point to a broader role for GRC. Governance teams are often asked to assess risk after the solution and business case are defined. AI requires earlier involvement. GRC can help connect expected value with evidence, identify where risk and dependencies shift, and test whether the benefits still hold once operating costs and controls are included. Enterprise AI creates value when organizations invest in the work around the technology: redesigning processes, strengthening data, clarifying roles, and building the trust needed to use it well. When the AI bill comes due, what matters is whether those investments produced a measurable improvement after the full cost of getting there is included.
NOTES 1. FinOps Foundation, “State of FinOps Survey: AI Value and Skills Top Priorities as FinOps Matures,” Linux Foundation, 2026. https:// www.linuxfoundation.org/press/state-of-finops-survey-ai-value-andskills-top-priorities-as-finops-matures-across-technology-value-98manage-ai-90-saas-64-licensing-48-data-center-1 2. “The Solow Productivity Paradox: What Do Computers Do to Productivity?” Brookings Institution. https://www.brookings.edu/ articles/the-solow-productivity-paradox-what-do-computers-do-toproductivity/ 3. Erik Brynjolfsson, Daniel Rock, and Chad Syverson, “The Productivity J-Curve,” American Economic Journal: Macroeconomics. https://www.aeaweb.org/articles?id=10.1257/mac.20180386 4. “Uber CTO Praveen Neppalli on the end of the tokenmaxxing era,” Business Insider, 2026. https://www.businessinsider.com/uber-ctopraveen-neppalli-tokenmaxxing-era-end-2026-8 5. “Uber turns its best AI engineers loose in pods,” Business Insider, 2026. https://www.businessinsider.com/uber-turns-best-ai-engineersloose-pods-business-2026-8 6. National Bureau of Economic Research, Working Paper 31161. https://www.nber.org/papers/w31161 7. METR, “Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity,” July 2025. https://metr.org/ blog/2025-07-10-early-2025-ai-experienced-os-dev-study 8. Faros AI, “AI Acceleration Whiplash,” 2026. https://www.faros.ai/ research/ai-acceleration-whiplash
ABOUT THE AUTHORS Mary Carmichael is the Field CISO for Western Canada at Bell Cyber, where she focuses on cybersecurity transformation, AI, risk management and governance. Her work brings together executive advisory, applied research and education to help organizations strengthen resilience, navigate emerging technology risks and adopt AI in ways that balance innovation with accountability and human oversight.
Before approving the next AI project, take one use case and trace it from the business problem through AI production, human review, and rework to the final outcome. Then calculate the chain. For a SOC use case, the baseline might include investigation time, false-positive rates, escalation volume, and incident impact. For AI-assisted development, review time, vulnerabilities, rework, and production incidents belong beside coding speed.
FIVE QUESTIONS BEFORE THE NEXT AI INVESTMENT 28
A I CYB ER
AGENTIC AI
The Hidden Tax Of Agentic AI. Agentic AI is sold as a productivity multiplier. For most enterprises it quietly levies a tax, and the bill is never in the room during the demo. By Venkata Sai Kishore Modalavalasa
For the past couple of years, as agentic AI slowly reshaped enterprises, I have sat through a lot of very good demos. A clean interface, a confident agent, a working prototype that seemed to be born out of a single prompt. Customer briefings, networking events, hackathons, late night conversations with friends at Fortune 500 companies. The energy in the room is the same every time. Everyone wants to believe, and they are not wrong to. This is the greatest unlock that has ever happened for our ability to solve problems quickly. The prototyping speed is real, and the abundance of AI-generated code that follows is real too. What is rarely in the room is the bill. The work does not disappear. It relocates. That elsewhere is human review, coordination, quality control, governance, and first and most expensively, security. The tax is invisible in the demo and it shows up in production. If your organization is pretending otherwise, you are deferring payment, with interest.
FA L L
2026
The work does not disappear. It relocates. The Demo Is Not The Deployment Look past the enthusiasm and the numbers published over the past year tell a sobering story.
95%
MIT’s Project NANDA studied more than 300 public deployments and found that 95% of generative AI pilots deliver no measurable P&L impact.1
42%
S&P Global reported that 42% of companies abandoned most of their AI initiatives in 2025, with the average organization scrapping nearly half its proofs of concept before production.2
29
A I CYB ER
40%
Gartner predicts that over 40% of agentic AI projects will be canceled by the end of 2027, citing escalating costs, unclear value, and inadequate risk controls.3
These numbers indicate a gap between what works in a demo and what survives in production, and that gap is exactly where the tax gets levied. Right now it is widening faster than most organizations can close it.
The Hidden Burden Here is what I keep hearing from engineering leaders. Everyone is a builder now, so everyone builds. That translates into organizations drowning in prototypes that nobody owns, wired to data that nobody scoped, maintained by no one. Maintenance is a dead word now. AI made creation so cheap in terms of time that the quick fix is to rewrite rather than maintain a software system. Expectations have inflated alongside it. It is close to a cardinal sin these days to estimate a quarter to properly design and deliver a project. You need an answer ready for the question: AI can do this in an afternoon, so why do you need a quarter? The humbling reality is that the effort did not vanish. It moved downstream into reviewing, debugging, and reconciling. That hidden burden does not surface as cleanly as its visible counterpart, the token burn.
The New Organizational Dynamics Until we reach a world where everything is owned and run by agents, real people still do the orchestrating, and everyone owns a piece of this new movie. Engineers, managers, product leads, and executives each now have a large model at their disposal, and with it a newfound fluency in everyone else’s discipline. The product manager ships a working prototype over the weekend. The executive pastes an architecture question into a chat window and arrives Monday with strong opinions about the database. The junior engineer generates in an hour what used to signal a decade of experience. The lines of expertise have not disappeared, but they have blurred, and people are stepping on each other’s toes in ways org charts were never designed to absorb. In the spirit of moving fast and propelling the business forward, this is great. What worries me more is what happens to judgment underneath. When the model always has an answer, the muscle of sitting with a hard problem starts to atrophy, and rigor quietly leaks out of the system. The review culture shifts from “convince me this is right” to “the model suggested it,” as if that settles the matter. When something breaks, organizations embrace a new deflection: blame the model. Accountability does not transfer to the tool. It goes unclaimed. Teams that once trusted each other’s craft now quietly wonder how much of a colleague’s work is theirs, and that trust, once eroded, is very expensive to rebuild.
Agentic AI does not remove humans from the loop. It multiplies the number of loops humans must hold together, while making each one look deceptively self-sufficient.
The uncomfortable truth is that agentic AI does not remove humans from the loop. It multiplies the number of loops humans must hold together, while making each one look deceptively self-sufficient.
Security Is The First Casualty Everything above is expensive but this part is dangerous. In nearly every enterprise conversation I have had this year, the same pattern surfaces. Agents went live months before anyone wrote a policy for them. Security teams discover fleets of agents the way they once discovered shadow IT, after the fact, already wired into production data, already acting on behalf of people who cannot fully explain what they do. Ask who owns a given agent, what it can touch, or how to shut it down, and the room gets quiet. Having spent years watching non-human traffic overtake human traffic on the open web before the GenAI era, I can tell you we have seen this movie. The industry learned, painfully, to identify, attribute, and rate-limit machine actors at the edge. That same story is now playing inside the enterprise perimeter, except the machine actors hold credentials. Agents ship with tokens nobody rotates, permissions nobody scopes, and owners nobody assigned. Privileged identities are multiplying faster than any human workforce ever could, and they are governed far more loosely. The failure modes are not hypothetical anymore. This past year we have watched attackers turn developers’ own AI coding tools into reconnaissance engines, walk sensitive data out of production copilots with nothing but a well-crafted message, and ride a single trusted AI integration into hundreds of downstream environments. Working on the OWASP Top 10 for Agentic Applications, my colleagues and I kept arriving at the same uncomfortable truth. Risks like goal hijacking and tool misuse are not bugs to patch. They are inherent design properties of systems where instructions and data share a single channel, which means they have to be governed rather than merely fixed.
Paying The Tax Down Here is the good news, and it is good. The tax is not a reason to stop. It is a reason to design.
WHAT THE ORGANIZATIONS SUCCEEDING WITH AGENTS DO • Scope narrowly. • Redesign one painful workflow at a time. • Build governance alongside deployment rather than after it. • Treat every agent as a non-human identity with short-lived, task-scoped credentials, a named owner, and an audit trail. • Adopt the scaffolding that already exists, including OWASP’s agentic guidance and NIST’s emerging control overlays, instead of improvising policy after the incident.
FA L L
2026
Every enterprise pays for agentic speed one of two ways. Deliberately up front, as design and identity discipline. Or involuntarily later, as rework and incident response. The winners will not be the fastest movers. They will be the ones who treated speed and safety as a single decision, because ungoverned speed is not productivity. It is a debt, and the interest compounds.
Every enterprise running agents files this return, whether they know it or not. The only choice on the form is your filing status: govern before you deploy and pay the lowest rate on the schedule, or file as a shadow filer and let the audit find you. Consult Part III for the only deductions known to exist.
NOTES 1. MIT Project NANDA, “The GenAI Divide: State of AI in Business 2025,” July 2025. https://cloudelligent.com/wp-content/ uploads/2026/02/v0.1_State_of_AI_in_Business_2025_Report.pdf 2. S&P Global Market Intelligence, Voice of the Enterprise: AI & ML 2025. https://www.spglobal.com/market-intelligence/en/news-insights/ research/2025/10/generative-ai-shows-rapid-growth-but-yieldsmixed-results 3. Gartner press release, June 25, 2025. https://www.gartner.com/en/ newsroom/press-releases/2025-06-25-gartner-predicts-over-40percent-of-agentic-ai-projects-will-be-canceled-by-end-of-2027
ABOUT THE AUTHORS Venkata Sai Kishore Modalavalasa is the Chief Architect and Engineering Leader at Straiker, where he builds AI-driven security products to protect AInative applications at scale. With over a decade of experience in cybersecurity and distributed systems, he has taken products from 0 to 1, scaling Cyberfend from startup to acquisition by Akamai. He’s an active OWASP author and contributor. whose career reflects a blend of deep technical expertise and leadership in bringing innovative security solutions to market.
Every enterprise pays for agentic speed one of two ways. Deliberately up front, as design and identity discipline. Or involuntarily later, as rework and incident response.
30
A I CYB ER
FA L L
2026
31
A I CYB ER
FA L L
2026
32
A I CYB ER
The AI Innovation Paradox: Enabling Progress Without Losing Control. Innovation almost always outpaces governance. The work is building guardrails invisible enough that the safe path is easier than the workaround. By Mohamed T. Konate Generative AI is moving fast, and it should. In highgrowth environments, the pressure to ship features, personalize user journeys, and move faster is a competitive mandate. AI is increasingly the engine driving that speed. Innovation almost always outpaces governance. When proprietary code is pasted into public prompts or customer data leaves the corporate environment, the security problem arrives with a risk visibility problem attached. The answer cannot simply be to block everything. Microsoft and LinkedIn’s 2024 Work Trend Index found that 78% of AI users were bringing their own AI tools to work.1 Security has to do more than restrict AI. We have to provide a safer path forward. What follows is a pragmatic blueprint for building those invisible guardrails.
Bring AI Inside The Tenant Boundary The primary risk sits in the destination of your data rather than in the AI itself. A resilient approach starts with providing an approved enterprise AI solution within your own corporate tenant. When teams have access to a safe, high-performance option, adoption naturally shifts there. You are not removing the tool. You are bringing the data back under the security perimeter. This creates a secure sandbox where innovation can happen without the same concerns around corporate data being retained or used to train public models. Putting AI inside the corporate tenant, with enterprise identity, logging, and tight access controls, allows teams to experiment while maintaining the security standards an enterprise expects. This approach is consistent with the broader direction of the NIST AI Risk Management Framework, which encourages organizations to manage AI risks in a way that aligns with their business objectives and risk tolerance.2
Address Shadow AI Without Blanket Bans Even with a blessed solution, teams will naturally experiment with new AI extensions, LLM APIs, specialized models, and code assistants. Instead of relying on a blanket ban, I prefer a block-by-default but redirect model. If a user attempts to access a risky or unapproved AI tool, block it where appropriate, then go further than a generic access denied page. Explain the risk and provide a direct path to the authorized alternative. Security goes from being a roadblock to being a pit crew that keeps the car on the track. Web proxies, DNS controls, CASB capabilities, and URL filtering can enforce those boundaries while providing visibility into what employees are trying to use. If users repeatedly attempt to access the same AI service, that is a business signal as much as a security event. The approved solution may not be meeting their needs.
FA L L
2026
Closing The Feedback Loop
If reviewing a new AI tool takes three weeks, engineers will find a workaround.
Governance is only as strong as your ability to audit it. Policies tell us what people are supposed to do. Telemetry tells us what they are actually doing.
Keep The Approval Process Lightweight Speed is a feature of security. If reviewing a new AI tool takes three weeks, engineers will find a workaround. Organizations need a lean, risk-first review model. For unfamiliar AI services, I start with a few questions that quickly tell me where the risk sits.
FIVE QUESTIONS FOR AN UNFAMILIAR AI SERVICE 1. Where does our data reside? 2. Can enterprise inputs be used for model training? 3. Does the vendor have credible third-party assurance, such as SOC 2 Type II or ISO 27001? 4. Is customer data appropriately isolated? 5. Does the platform support enterprise identity and meaningful audit logging?
Policies tell us what people are supposed to do. Telemetry tells us what they are actually doing. SIEM, CASB, web, DNS, identity, and DLP telemetry can help identify unapproved AI usage and unusual data movement. If an account suddenly begins sending large amounts of information to an AI service, the SOC should have enough visibility to investigate. Where appropriate, activity and prompt auditing can reveal patterns involving customer information, credentials, or other sensitive data. Operational data lets us understand behavior, improve controls, and identify where users need better alternatives or clearer guidance, instead of writing a policy and hoping it works.
Security As A Compass You can go deeper when the use case requires it. Not every request needs a six-week security assessment. The objective is to give engineers answers quickly without exposing intellectual property, customer data, or regulated information.
Do Not Overlook AI-Generated Code Risk This is the ghost in the machine. Engineers are increasingly pulling AI-generated code directly into the SDLC. That can be a massive force multiplier, and it can also introduce silent risks: hardcoded secrets, weak authentication logic, insecure dependencies, or code nobody fully understands. OWASP’s guidance on securing generative AI and LLM applications reinforces this concern, including risks involving improper output handling, sensitive-information disclosure, and supply-chain vulnerabilities.3 My view is simple. Treat AI-generated code like untrusted external code. The tool can generate the code. The engineer still owns the security of the logic.
The tool can generate the code. The engineer still owns the security of the logic. CONTROLS FOR AI-GENERATED CODE Before code reaches production: • • • •
Human peer review SAST and SCA scanning Secret detection Dependency validation
Organizations that block AI will fall behind. Organizations that ignore AI risk will eventually face a crisis. The balance is controlled enablement. By building pragmatic, architecturally sound guardrails, we can empower teams to move at the speed of the market without allowing innovation to become the next vulnerability. No single framework can encompass every complexity of AI adoption. The principle is straightforward. Give people a secure way to innovate, make that path easier than bypassing it, and verify that the controls actually work. Security should be the compass helping the organization move forward safely rather than the department standing at the door saying no.
NOTES 1. Microsoft and LinkedIn, “AI at Work Is Here. Now Comes the Hard Part,” 2024 Work Trend Index Annual Report, May 8, 2024. https://www. microsoft.com/en-us/worklab/work-trend-index/ai-at-work-is-herenow-comes-the-hard-part 2. National Institute of Standards and Technology, “Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile (NIST AI 600-1),” July 26, 2024. https://www.nist. gov/publications/artificial-intelligence-risk-management-frameworkgenerative-artificial-intelligence 3. OWASP GenAI Security Project, “OWASP Top 10 for LLM Applications 2025,” November 17, 2024. https://genai.owasp.org/ resource/owasp-top-10-for-llm-applications-2025/
ABOUT THE AUTHOR Mohamed T. Konate is a cybersecurity and technology risk leader with more than 16 years of experience spanning security operations, GRC, cloud security, AI governance, and enterprise cyber risk. He is the founder and principal architect of Straitum, a cyber risk management platform, and a graduate of Carnegie Mellon University’s CISO Executive Program.
Put them inside the developer workflow. Pullrequest decorators, pre-commit hooks, and automated CI/CD checks surface problems while engineers still have context, moving security from a late-stage gate to an early-stage collaboration.
33
A I CYB ER
FA L L
2026
34
A I CYB ER EXCLUSIVE INTERVIEW | SECURITY RESEARCH
James Kettle’s AI Stole A Bank’s API Key Overnight. PortSwigger’s research director spent a year answering a question nobody else was asking. He told the model it was only a simulation, pointed it at 30,000 sites, and it found an attack class he says he never would have found himself. Interview by Confidence Staveley | AI Cyber Magazine, Fall 2026 James Kettle has led the research team at PortSwigger, the makers of Burp Suite, for twelve years. In that time he has invented or popularized a run of web attacks now embedded in everyone’s testing methodology: web cache poisoning, HTTP request smuggling, server-side template injection, server-side request forgery via load balancers. The session he gave the week before this conversation was his eleventh appearance at Black Hat USA. This year he turned the method on itself. He built an autonomous system called HTTP Terminator that replicates his own research process, ran it against 30,000 sites, and set out to answer a question the industry keeps arguing about without testing: can AI produce genuinely novel security research, or is it only ever recombining its training data? Why should someone who has just landed on this conversation stay? KETTLE: I lead the research team at PortSwigger and have done for about twelve years. Every year I do novel security research and publish the results at conferences like Black Hat USA and DEF CON. Seeing the recent AI developments, no one else seemed to be asking whether AI can do genuinely novel research. Everyone knew it could apply known techniques quite effectively in some cases. But there were plenty of people saying it is just using stuff from its training data, so it can never do anything new. I thought it would be interesting to interrogate that question. Why was that question worth a year of your life? KETTLE: As someone who does research full time, I just needed to know the answer, and it did not seem like anyone else was going to answer it. There were a load of side things that also were not being openly talked about. People were not really talking about where AI fails in the cybersecurity space, because the people pushing AI to the limits were the people trying to sell novel products, and they are not going to tell you what their products are bad at. So the only people saying AI is bad were the people barely using it, which are not the people who are going to get you the best information on the topic. I felt like for myself and for all the other researchers in the world, we needed to know.
The only people saying AI is bad were the people barely using it. Which are not the people who are going to get you the best information. JAMES KETTLE
FA L L
2026
Give us the version a CEO would understand. What is an HTTP desync attack, and when it works, what does the attacker walk away with? KETTLE: Many websites use the protocol HTTP/1.1. It is very, very old and it is basically not fit for purpose for the way it is used in modern systems full of load balancers and reverse proxies and CDNs. What that means is that an attacker who finds a way to cause a desync can make the protocol go completely haywire, send random responses to random users, and do things like leak everyone else’s credentials to the attacker. You built a system for this and called it HTTP Terminator. You also said building it sounded like a bad idea and did it anyway. What is it, and what did you mean by bad idea? KETTLE: It is a system that replicates the kind of techniques I use to perform novel research in the field. I had already done four years of research on this topic, so I took my methodology and implemented it with AI. On it being a bad idea, that is playful, but there are two underlying reasons. One is that we have all seen the science fiction where AI gets really good at hacking and it is the start of the apocalypse, and what I was doing was absolutely trying to build an autonomous AI system that was as good at hacking as possible. The other is that I was trying to fully automate the exact methodology I use as a researcher. So I was trying to make myself redundant. It was a bad idea on multiple levels, but at the end of the day I needed to know.
THE FOUR PHASES HTTP Terminator models Kettle’s own research process. Ideation. Generate attack ideas rather than reapply known ones. Evaluation. Automatically decide whether an idea worked. The hardest part, and the ceiling on everything the system can find. Weaponization. Confirm the threat is real, which means running fast. Cascade. Use every new discovery to hunt for more. The phase he ran on instinct for years without noticing. Of those four phases, which had you underrated for your entire career? KETTLE: The cascade. When I was doing research manually in the olden days, I was having cascades and they were leading to the best discoveries. But because everything was manual, the logging of why I discovered things and what steps led to it was manual and very sparse, because at the end of the day I just want to hack things. That lack of logging of the full discovery chain meant it was completely invisible that it was a cascade of dominoes leading to the best findings. This year, because I automated the whole thing, I built in extensive logging so I could really understand where the AI was having breakthroughs and what was fueling the novel research. That revealed the cascade process is super important. Step one was getting the machine to invent, not rediscover. You tested whether it could reinvent an approach you had created but never published. What happened? KETTLE: I gave them this challenge and initially they all completely failed it, which was not very surprising because it was a fairly hard challenge. After some tweaks, the best available models solved
35
A I CYB ER it with about a five percent success rate, which seemed okay. It showed the research concept could work, because you do not need a hundred percent success rate. I do not have a hundred percent success rate. I am nowhere near that. Then I thought I would give the AI a hint, because all the documentation at the time said that to make AI good at a specific task you give it loads of background information. What I found was that giving it the background information actually made it worse, which was super surprising, and it informed the follow-up approach I took to build HTTP Terminator.
When you evaluate something and ask whether it worked, the exact algorithm you use to answer that dictates how creative the AI can be and dictates the scope of what it can find. If you say I need to see a robots.txt file in the response because that is what the attacker is trying to achieve, well, a lot of the time the best discoveries come from an attack that tries to achieve one thing and actually does something different but still valuable. If you over-tune the evaluation system you miss those. If you under-tune it you get false positives and noise, and you will absolutely drown in that noise because of the scale the system runs at. How do you know when you have walked that tightrope correctly? KETTLE: You know it is working if you are not getting false positives, and if it is finding things you did not expect it to find.
You know it is working if it is finding things you did not expect it to find. JAMES KETTLE This was not a lab. You were hitting 30,000 websites around the clock. How did you keep that authorized and safe? KETTLE: To avoid legal issues, or causing anyone any inconvenience in general, I only targeted websites that have bug bounty programs or vulnerability disclosure programs. That makes it legal. I also put a heavy rate limit on HTTP Terminator. It sends a lot of requests, but it sends them much slower than an individual human browsing that website, so there is no risk of consuming loads of bandwidth or causing downtime. That was the evaluation phase. In the weaponization phase you have to go fast, but by that point you know there is a critical threat there. There is more risk in going that fast to confirm the threat, and more upside, because you are genuinely stopping them from being hacked by someone else further down the line. That approach is what you call micro-inspiration. One or two sentence fragments from technical specifications instead of whole documents. Why did chopping the inspiration into crumbs make the models more original rather than less? KETTLE: If you give them a document that has three cool ideas in it, generally they are just going to pick one and always go with that, and miss the other two. It is worse than that, because if you give them a document with twenty potential leads, AI is about prediction, so they are going to pick the most predictable idea, which is the least original one. Whereas if you give them something really small and quite arbitrary and force them to use it, you are pushing them outside their comfort zone. You are forcing them to do something original, because there is not an easy alternative they will go for every single time. You have said evaluation matters more than the ideas themselves. Why does it set the ceiling on everything the system can ever discover? KETTLE: Evaluation is crucial because it has to be automated, since the majority of ideas the AI comes up with are going to be bad.
To get it to go fast, you had to manipulate it. You renamed the attack a simulation, handed it fake placebo features, and hid the words that scared it into quitting. Walk us through why you had to control the agent’s entire reality. KETTLE: The idea of making it think it was in a simulation came about because, to make these attacks work for real in the weaponization phase, you have to go fast. It was being really concerned about not wanting to go too fast in case it affected the website. But you need to go fast or you will just think the website is secure when it is actually vulnerable. Reframing it as a simulation in the tooling layer turned out to be a really effective solution. The tempting thing is to fix those issues in the initial prompt. The problem is that as it works for longer, it adheres less and less to what you told it in the prompt. Whereas if you make a change to the tooling layer, which is the way it interacts with reality, then every time it interacts with that tool it gets this false view. It is continuously reinforced, and it is a much more effective way of getting it to do what you want. I added an MCP interface to Turbo Intruder, and the MCP interface is where the lies get inserted.
FA L L
2026
You reframed the debate from autonomous versus human to AI versus code. Why did taking work away from the AI and handing it to plain deterministic code make the system more reliable? KETTLE: The fundamental thing about AI is that it is non-deterministic. You cannot take an action and have any confidence it will consistently improve the system’s performance. That is fundamentally impossible. And even if it improves performance for that AI model, when you change the model it may start making performance worse. So you need some other way of improving a system if you are not happy with its initial performance, and code is the way you get there without loads of human intervention, which is not possible at the scale HTTP Terminator was running at. It has never been easier to generate that code. The fact that AI can generate dynamic code is a bit of a distraction, because there will be a different bug every time it tries to tackle the problem. Whereas if you get it to generate the code once and you fix the bugs, now you have something super robust, and it is also cheaper and faster and more reliable. Define a research cascade in plain terms, and the two questions you put to every finding. KETTLE: The research cascade is the feedback loop where you use every new discovery to help you hunt for more discoveries.
THE TWO CASCADE QUESTIONS Ask both of every new discovery. 1. How can I detect this kind of thing better? You may have been triggering that anomaly all over the place and only found it through sheer luck. Improve detection and find the other instances. 2. What is the root cause, and what else stems from it? This is the question that produced the biggest finding of the year. What was the biggest find? KETTLE: The multipart byte ranges vector, which caused a full desync on a wide range of different web servers, including a different bank from the one I mentioned earlier. The AI was inspired to create it by a section of the HTTP RFC which talked about how to handle the response to range requests. It thought, well, I cannot send responses, but I can send requests. So it put it in a request, and it turns out that causes a desync on a whole load of systems. That by itself was super valuable. But the most valuable thing came while analyzing it, when it asked what the root cause of the behavior was. The answer is that servers use shared code to process both requests and responses, which means any response feature could potentially be exploitable in a request. That enables loads of other attacks. As soon as it said that, I realized it explained weird behavior I had seen in the past in multiple different places, including with the Set-Cookie header in a request, which again should never work. I named this concept shared parser confusion. It is a big attack surface expansion and I think it is going to cause a lot of other vulnerabilities in the future. You wrote that neither of you would have found it alone. Explain that. KETTLE: In HTTP, the world is split into requests and responses. Conventional wisdom says requests come from users, from external systems, so they are untrusted. You have to be really careful when
36
A I CYB ER you process requests or you might get hacked. Whereas responses come from the application, from the back end, so that is the trusted system. From a server’s point of view the response comes from the good guy and the request comes from the bad guy. That assumption is so baked into my head that I would never have had that discovery, because I would never have read a section of the RFC about how to process the response to a request when I was looking for a desync vulnerability. I cannot send a response. If the origin is compromised, the whole attack concept is meaningless. So I just would not have tried it. Whereas the AI was given two lines of text and told to make a vector using them. It had no choice. It was forced to try it, and it actually worked. Then the analysis of why revealed this whole inversion, that you can trigger response features from requests, which is a massive threat, because there are loads of potentially dangerous things there. Setting the location header in a request could potentially trigger server-side request forgery, which would be pretty catastrophic.
I would never have had that discovery. The AI was given two lines of text and told to make a vector. It had no choice. JAMES KETTLE What else could come out of it? KETTLE: Pretty much anything. The Set-Cookie header lets you get cookie injection, which is quite nice. With server-side request forgery from the location header you can do all kinds of damage. Often you can steal cloud credentials or gain access to other internal systems. I have not really explored the scope yet, because this finding came late in the day, so I am quite excited to see what it is going to find. It does not just apply to headers, it applies to the request body as well. At this point it is hard to say exactly where this is going to lead to things being compromised, but it is a lot of attack surface. You opened your laptop one morning and found the system had stolen a long-lived API key from a bank overnight. You wrote that it felt like stepping into the audience for someone else’s talk. What was that morning like? KETTLE: It was pretty crazy, for a few reasons. When the AI reported the anomaly, it did not even say it had stolen an API key. It just said this response is not a CSS file, and I was expecting a CSS file. This was reported in Burp Organizer, and the number one thing I look at there is the domain of the target, to know whether this is going to be a super hardened target or whether it will make a case study. I have tens of thousands of entries in there from everything it finds. When I looked at the domain I did not recognize it, because it is not a UK bank. So I did not realize it was a bank at first. It was only a while later, skimming through the findings, that I realized I had got an API key. Then I looked closer and thought, let me load up this home page, and realized it was actually a major bank. Then I thought, how did it get this API key? I looked at the vector, and the vector made basically no sense, which is what made me feel like I was in someone else’s talk. It took about two hours of reverse engineering to understand how it achieved the exploit.
HOW IT TOOK THE KEY The system sent two identical, matching Content-Length headers, and hid the second one from the front server using a single leading white space. Hiding a header that way is a known technique. What was not known is that it had to send the same header with the same value twice. Hiding one header let it reach the back end, which saw two headers carrying the same value and, for some reason, interpreted that as zero without raising an error. That caused the desync that let it steal the key. Kettle needed roughly two hours of reverse engineering to work out what his own system had done.
FA L L
2026
All of this rides on one aging protocol. What does a practitioner who cannot rip out HTTP/1.1 by Monday actually do this week? KETTLE: If possible, the number one fix is to turn on upstream HTTP/2, because that will fully solve the entire attack class. If you are not able to do that, and if you use a major CDN you may find you cannot because they do not support it yet, send them an email saying you want them to. Then go diving into the settings and look for everything you can do to make HTTP request processing more strict, and enable any desync mitigations visible there. Just understand it is not sufficient. It is the best you can do without HTTP/2.
WHAT TO DO THIS WEEK The headline question, and I want a straight answer. Can AI do novel security research, and where does a human still belong in the loop? KETTLE: AI can absolutely do novel security research fully autonomously. There is one place inside the loop where a human researcher makes the whole system a lot more powerful, and that is the cascade. Their intuition in escalating findings, and their ability to evaluate findings without a big pre-built harness, makes a massive impact. That is where the most significant discoveries in this research came from, including shared parser confusion.
1. Turn on upstream HTTP/2. This fully solves the attack class. 2. If your CDN does not support it yet, email them and ask. 3. Make HTTP request processing as strict as your settings allow. 4. Enable every desync mitigation available to you, understanding that this is mitigation rather than a fix. You talked an AI into hacking real banks by telling it to relax, it is only a simulation. What does it say about the state of AI safety in 2026 that the guardrails came off that cheaply? KETTLE: AI safety and guardrails have varied a lot this year. They change from model to model rapidly. I was using a model with reduced cyber safeguards as part of this research. If you tried to get a frontier lab model to do this research with the regular version, it might object. That said, I am not a hundred percent sure. The other thing is that open models have caught up with where frontier models were when I was doing this research. All these cool findings came from models that are significantly dumber than the best models you have access to right now. So attackers do not need frontier models. KETTLE: That thesis is behind all offensive research, really. These techniques I am exploring and publishing may well be known to bad actors who are exploiting them rather than publishing them. It is only by getting information on exactly what the threats are out there into public, along with synchronized patches wherever possible, that people can get any understanding of how to protect themselves. On where open models are, the reports I am seeing say they are basically six months behind the cyber models. With HTTP Terminator, a lot of the most effective vectors were created by the system more than six months ago. So open models are going to be perfectly capable of creating those kinds of vectors now. The only step where I really hit the ceiling on the model’s capabilities was the cascade process, which is something a skilled researcher can do manually where necessary. So the human behind the tool is the ceiling. KETTLE: Yeah, pretty much.
37
A I CYB ER
FA L L
2026
The only step where I hit the ceiling on the model’s capabilities was the cascade. The human behind the tool is the ceiling. JAMES KETTLE You open sourced the entire thing. If an attacker and a defender both download your code this week, who gets more out of it? KETTLE: I designed this system for researchers, so I think that is who gets the most out of it. The vulnerabilities it found while I was running it have mostly been fixed, so it is not as much of a threat to the general internet as it was when I built it in January. Who it is really valuable for is researchers who want to take it, tweak it, add their own spin, and find other interesting things and help get them fixed. You are one of the best HTTP attackers alive and you just built the machine that could replace the junior version of you. Should the next generation be terrified? KETTLE: I would not be terrified. It is easier to do research than ever if you are a junior researcher. You can take HTTP Terminator, fork it, use a coding AI agent to add one of many capabilities that are not currently in it, and find amazing stuff without the massive time investment this used to take. It is a great time to be doing research and a great time to be learning the art of research. I have also seen more and more security software product companies in the offensive space, and they all want research teams to make their product better than everyone else’s, because everything public is in the models now. So you need that edge for your company. It is a pretty good employment market for researchers too. One piece of advice for a young researcher. KETTLE: The number one thing holding researchers back right now is that they will have a good idea, or a few good ideas, and they agonize about whether it is a good lead or a bad lead. Just try it. Have a go and see what happens. Last question. Where is the hype and where is the real thing in AI and cybersecurity? KETTLE: The hype is basically everywhere. It is all over social media. It is the stuff the algorithm loves, so it gets amplified. If you want the real thing, look at DEF CON presentations, look at Black Hat presentations, look at USENIX. That is a good one a lot of people overlook because it is so academic. And use it yourself. That is the number one thing you can do to see what is real and what is not.
ABOUT THE AUTHOR James Kettle is Director of Research at PortSwigger, the makers of Burp Suite, where he has led the research team for twelve years. He has invented and popularized a range of web attack techniques now standard in testing methodologies, including web cache poisoning, HTTP request smuggling, server-side template injection, and server-side request forgery via load balancers. His session on this research was his eleventh presentation at Black Hat USA. HTTP Terminator is open source.
38
A I CYB ER
EXCLUSIVE INTERVIEW | SECURITY RESEARCH
195 Bugs In Three Weeks, And Not One Of Them Paid Her Anything. Ezinne Kalu built an AI research system that out-produced her entire manual bug bounty career, then pointed it at open source projects that cannot pay a bounty at all. Interview by Confidence Staveley | AI Cyber Magazine, Fall 2026
Ezinne Kalu was going to be a doctor. Medicine did not take, so she started experimenting with code, found a four-week Microsoft program on Linux and how hackers work, and never left. She went through the CyberGirls Fellowship, chose the adversarial side of the field, and started hunting bugs because nobody would give her a job. Today she runs what she calls a one-lady security lab. An AI system she built and supervises has produced 195 valid findings in three weeks, more than her entire manual bug bounty career before it, along with assigned CVEs in software that has no bug bounty program at all. She is deliberately working on projects that cannot pay her. Someone has just landed on this conversation. Why should they stay? KALU: The world has changed so much with AI and the introduction of AI agents. If you compare two years ago to now, every single workflow has changed. For me, in security research, I have found that AI helps me scale research in a way that was not previously humanly possible. It is not something any team, no matter how skilled or experienced, could
FA L L
2026
have done in such a short amount of time. That is the beauty and the power of agentic systems. I want to tell you about what I am doing as a onelady security lab, and hopefully you can pick up some tips you can use in your own workflows. Take us back. How did you get into cybersecurity? KALU: I have a very interesting story of coming into cybersecurity, because I was not originally supposed to be in tech at all. I was pursuing a medical degree. I was going to become a doctor. Medicine just was not it. I could not connect with the courses and I was having a really hard time at school. So I started to experiment with code and technology, just trying to see what else was out there. I had a shiny new laptop at the time and it felt natural to go online and look around. I got into a short program, I think hosted by Microsoft, for women in tech. Four weeks of learning about Linux, understanding Kali Linux and how hackers work. Those four weeks were really the start of everything. Later that year I applied for the CyberGirls Fellowship by CyberSafe Foundation, and I was lucky to get in. CyberGirls has historically been the highest standard of security training available for
39
A I CYB ER women in Africa. I completed the program, and that was the beginning of my career. During that program I decided to focus on offensive security, penetration testing, application security, and to look toward the adversarial side of technology. Since then I have focused a lot of my work outside my nine to five on bug bounty and security research. My goal has always been to advance the digital sector in a way that keeps everybody safe. Checkmate the bad guys before they checkmate us. Tell us about the bug bounty work, before AI entered the picture. KALU: I experienced a very unique problem when I finished my training. I was finding it really difficult to get a job. That is something a lot of people face, especially from Nigeria, where I am from, because cybersecurity has never been a field where you could just get a job easily. It is such a niche specialty. I started to feel like, how can I make my skills work for me? How can I show my experience? How can I tell the world this is what I am doing without having a job title attached to it? These skills are actually useful. I do not need a job to go out and show people what I can do. That was when I started researching bug bounties. I watched my first video from NahamSec, shout out to NahamSec, where he explained how bug bounties worked. You could make money finding and submitting vulnerabilities to companies all over the world, and they would not ask you for any requirements. All you needed to do was prove the vulnerability existed. Those bounties were really attractive, especially coming from Nigeria, where our minimum wage is ridiculous. I remember talking to other people in security about bug bounty and getting a lot of negative comments. People did not like bug bounty the way I thought they would. I got discouraged so many times. But I told myself that if people can do it anywhere in the world, why can’t I do it? Once it is possible, then I can do it too. Within two months of focusing seriously, I found my first valid vulnerability, through YesWeHack at the time. When I got that bounty I felt so special, because I knew I had been right. Then I started working a nine to five, I had other obligations, and my research slowed down a lot. I was not finding bugs as often as I should have, and I was not scaling. Scaling meaning what, specifically? KALU: If you observe a vulnerability in one type of project, the next natural step is to try it across every similar project. If you find a vulnerability in a ticketing platform, you have to try it in all the ticketing platforms, because they could all have that problem. They are not going to share the vulnerability with each other. It is up to us, the security researchers, to let them all know. But I could not scale that. There are thousands of projects. I would do one or two a month. I have always been an AI evangelist. I love technology that is able to do more than I can do as a human. I do not believe AI is superior to me, but I believe AI can do several things far better than I can. It can remember everything. It has knowledge for days. And it can scale. You can expand it and send it out to do different tasks all at the same time, which I could not do on my own. I could not even train people to do it, because it is specialized work you would have to break down step by step, and scaling humans would mean hiring, which was out of scope for me. So I started building a little agent to validate my workflows. Then Claude released the Opus
models, and I remember thinking these things are now smart enough to handle a browser. If they can handle a browser, they can probably handle a repo. If they can look at code, they could probably follow my instructions and help me scale my research. That is where the idea started.
FA L L
2026
THE STACK • Harness. Claude Code. • Models. Claude, Codex, DeepSeek. • Storage. Everything on GitHub, managed locally in Obsidian. • Testing. PortSwigger’s Burp Suite MCP, connected to her agents. • How she learned it. Reading vendor documentation, researchers’ write-ups on Twitter, and Karpathy’s second brain article. PortSwigger released a few articles and built the Burp Suite MCP, so you can connect your agents to Burp. I read the docs for that, set it up in my own system, and immediately saw results in my work. You said harness. For anyone who does not know the term, what is a harness? KALU: There are so many buzzwords around AI, and it was a lot for me to understand as well. In this context, a harness is exactly what the name implies. It is what keeps everything together. It is what you use to run everything else. I use Claude Code as my harness because I like the CLI, I like how it looks on my computer, and I prefer that interaction. Behind it, powering the system, are many other agents, my personal context from my work, all my personal writing, and a lot of other things that come together on the back end and run through that harness.
Walk us through the learning process. What skills did you have to build to go from chatbot user to power user? KALU: The most beautiful thing about the bug bounty and security research community is that it is community powered. Everything comes from other people teaching the people who are just coming in. The most important tip I got, from another bug bounty hunter, was read the docs. When I first started using ChatGPT, I read the documentation OpenAI had put out. That is how I learned about connectors. There were not a lot of YouTube videos and blog posts about AI at the time. Most of the discussion was happening on Reddit and Twitter. When Claude came onto the scene, I did the same thing and read their docs. I also follow a lot of security researchers who have written about how they use AI in their workflow. One that stands out is Andrej Karpathy, one of the founding members of OpenAI. He put out something about a second brain system you can build with AI that helps you scale all the work you are doing. AI has infinite memory and the ability to ingest context and compress it well. If I had to read 20,000 pages of a book, I would probably still be reading. AI can take all of that in and still return very good information. Using the tips in that article, I set up a second brain system for myself.
So how has AI actually scaled the research? KALU: The kind of bugs that are easy to scale with AI are systemic issues rather than misconfigurations. A misconfiguration happens when developers make a mistake, and that is not really scalable, because you cannot say every single developer is going to make the same mistake. Take PDF parsers. There are so many different types. People write their own, people download code from GitHub. If you find that one parser is vulnerable to a specific thing, that could be systemic. You can go from PDF to cross-site scripting, where you put a line of code in the PDF, and when the system ingests it, instead of parsing it as a document, it executes the code. That is very dangerous. When I find something like that, the next step is to check every single PDF parser I can find and keep replicating it. And you shifted away from programs that pay. KALU: Originally I wanted to focus on bountypaying programs so I could get money for scaling my work. But now we have come to a place where we all have to scale our research. It is no longer just about bounties. Every system needs security. What is going to happen to the companies who do not have a security budget, who are not mature enough to have a dedicated bug bounty program? They need this research too. Two articles pushed me there. The first was from Shubham Shah, founder of Assetnote, which later sold to Searchlight Cyber. They have been doing serious research for years. He wrote a post telling his researchers to go forth and find internetmelting bugs, because if we do not do it, the adversaries will. Reading that told me, yes, that is definitely what is going to happen. If we stick to bug bounties because we just want to get paid, who is going to take care of every other system? The second was a post by the V12 team about raising a ten million dollar seed, where they proposed getting to a point where we have
40
A I CYB ER completely bugless software, software shipping without security vulnerabilities. It sounds impossible, but you read the approach and think about it with AI, and you know nothing is impossible. That work sounds crazy, but it can only be powered by AI. Those are the giants my work stands on.
Actually, let me take a peek while we are talking. Yes, I was right. It has found two more vulnerabilities while we were speaking, and all in software that does not even have a bug bounty program.
I do not need the AI to generate bugs. It is not a generation system. It is a testing system. EZINNE KALU
If we stick to bug bounties because we just want to get paid, who is going to take care of every other system? EZINNE KALU Give us the numbers. Before and after. KALU: Originally I wanted to focus on bountypaying When I was doing bug bounty full time on my own, I would submit between 10 and 20 bugs a month in my best months. With my system in place, my second brain, and my AI vulnerability research system, I have found 195 bugs as of the moment we started this episode. I have run that system for three weeks. In three weeks I found more vulnerabilities than I can count, and more than all the bugs I have submitted since I started doing bug bounty. A lot have been accepted. I have gotten duplicates too, because other security researchers are doing the same thing. Please continue. I will continue to dupe you and I want you to dupe me as well, until they fix the bugs. I know a lot of people may be skeptical and think it is just AI slop. But I do not need the AI to generate bugs. It is not a generation system. It is a testing system, to make sure everybody is secure as far as I can personally go as one person.
Let’s take the CVEs one at a time. Start with the path traversal. KALU: The first one is CVE-2026-79653, a path traversal in the filesystem attachments for Eclipse SW360. SW360 is a very special system. They do not have a bug bounty program, but they have been running this open source software that is powering huge companies. One of their key users is Siemens, the biggest industrial manufacturer in the whole of Europe. From my research they have more than 30,000 active customers using SW360 alone, and several more that may not have been shared publicly. This is a medium-severity issue. What was happening is that you are able to jump out of whatever path a file was originally destined for. You are trying to attach something to a particular path, and there is a path traversal where that function reads the statement and says, instead of dropping this file in your desktop repository, I am going to put it in your /etc/passwd file. That led to arbitrary file write. I could write any file I wanted into any system running the vulnerable version of Eclipse SW360, directly into any other folder on the machine, because it is using system privileges to do those file writes. By extension it leads to remote code execution, because I could put something in your cron files and you would run whatever malicious execution I wanted. One thing I use my AI for is to check how far back a vulnerability goes in the code. This one has been there for at least three released versions, up to the version they fixed it in. Who knows how many people may have exploited it. We cannot tell the exploits, but we can make sure it does not happen again. I submitted it to the Eclipse Foundation team and they validated the issue, reproduced it, fixed it, pushed a patched version, and issued an advisory to all their customers. This one was extremely fun. Shout out to the Eclipse team, super helpful.
CVE-2026-79653 • What. Path traversal in filesystem attachments, Eclipse SW360. • Impact. Arbitrary file write using system privileges, leading to remote code execution. • Reach. More than 30,000 active customers, including Siemens. • Age. Present in at least three released versions. • Outcome. Validated, reproduced, patched, advisory issued. Walk us through the disclosure. What does your report actually look like? KALU: The first thing I do is find out whether the company has a responsible disclosure channel. It could be encrypted email, their own form, or a platform like HackerOne, Bugcrowd, YesWeHack,
FA L L
2026
or Intigriti. It would surprise people to know how few programs those platforms have relative to how much software exists. SW360 had a security page on their GitHub. It is open source, so I could read their security policy, which said to fill out a form, explain how you found it, and submit everything you can that would help fix it. Security policies differ. Sometimes software will tell you they already know about an issue and are fixing it. Sometimes they will say they do not accept a class of issue. Google will not accept some vulnerabilities depending on stability. They have a benchmark for judging Chrome V8 browser exploits, and if your exploit does not cross that threshold, Google will tell you not to submit it. That helps them manage volume and it guides you as a researcher, because you may find something that does not make sense in that company’s business. This is where AI slop comes in, because people are not reading the security policies. If you read the policy, you see they have already mitigated, accepted, or transferred that risk. On some products you read the security.txt and find they got the code from another company, so a vulnerability in it goes to that company instead. Reading the policy keeps you away from slop and keeps your profile clean. Trust goes so far in this ecosystem. When I write reports, I start with a greeting. Hi team. I found a vulnerability. I learned this from pen testing, because when you tell a developer they have made a mistake, it is like calling someone’s baby ugly. I see people on Twitter saying they are going to call a company out, that the company did not respond so they will release the vulnerability. That is the wrong thing to do as a researcher. You put a lot of work into finding it, but remember you are talking to a business, to other people.
When you tell a developer they have made a mistake, it is like calling someone’s baby ugly. EZINNE KALU
HOW SHE WRITES A DISCLOSURE REPORT 1. Greet the team. Hi team, I found a vulnerability in your software. 2. Say what it does. The function reads the input wrong and writes the file to another place in the system. 3. Say how you found it. Point to the line. There is no validation check to stop it jumping out of where it is supposed to be. 4. Say how to fix it. Add a validation check Answer those questions and you are good to go. That is your report right there. And the second one. KALU: I split my work between products and packages. With a product like SW360, I look at it from the business side. Packages are imported, so they are part of the supply chain for thousands of pieces of software. This particular package is used by around 115,000 people according to its GitHub metrics, which means any bug in it goes upstream. You cannot really tell how people are using that
41
A I CYB ER package, so they may be using it in a way that is inherently vulnerable. I use the CIA triad to think about this. I know CIA is considered outdated and there are other letters now, but I still use it. Confidentiality, integrity, availability. This bug affected availability. An attacker could craft a PDF, and when this parser tried to parse it, it would lead to very long run times and enormous memory consumption. If you are running a small droplet and expecting to use very little bandwidth, this would eat all of it parsing a single PDF. Now imagine parsing thousands. That crashes your system. Availability is just as important as everything else, because if the software is not working, the business cannot move. I would argue it is one of the most important things to a business. Say you run a homework website and you let people upload PDFs. Somebody sends you one, your software is using the vulnerable version of the package, and it eats all the memory trying to extract from that PDF. It hangs your application. You cannot tell what is going on, because you built the software but this is coming from one package, upstream. So you start troubleshooting everything, and at the same time your business is down and you are losing customers and money. I submitted it and the team was also super helpful. We actually did the math on how bad it could be, working out how much bandwidth it could take up. They fixed it immediately and pushed the patch. The GitHub advisory has been published, and the CVE number has now been assigned: CVE2026-84311. Given all of this, what is your advice to a developer building today, or maintaining something they have been building for five years? KALU: First, please be receptive to your security people. You will come across security people who are malicious or who do not know what they are talking about, maybe because they are new to research. Take the time to understand where they are coming from, because ultimately it grows your knowledge. You do not have to be mean to the security guys. If you have been building software for years, assume there are vulnerabilities in it as it stands. You can try to find them yourself using AI. AI is an excellent researcher, and pointing it in the right direction is usually enough to get a security overview. In security we always say shift left. I say we should shift left twice. Instead of security being something you do while you are working, it should start before you begin working. You should be checking everything as you go. It is no longer as arduous as it used to be. People used to say it would slow down development. That is not the same now. Your tests run instantly, your multiple test cases come up instantly, code is pushed instantly. Integrating those checks in a white box approach means you put out better software. When your software is better, your customers trust you more, and you grow. What is next for you? KALU: It is a tough question, because I do not know what AI is going to look like a year from now. It is growing so fast that things that used to take me a week or three weeks now take hours. If I need to do recon on a target, I can have it done by dinner time. Before, it would take a week or a month just to gather sources. I am pro AI to the core, so I am adopting everything as it comes out and trying to fit it into my workflow. In the short term I will continue the
current research, working on open source projects, serving the security community, getting this research to everybody regardless of budget. I have submitted a lot of vulnerabilities to companies all over the world that have been accepted and are being fixed right now, so maybe more CVEs in the coming weeks. My AI system is currently producing more output than I can humanly work on. So I have hit a challenge where I now have to work on another agent to help with disclosures and submissions. Which means, if you are a developer, you should implement AI triage. You are going to get a lot of these reports from different people. I am not the only person who is going to keep doing this. More people will start. You are one person and you do not want the fatigue of reading and understanding all those reports. It would take you a weekend to set up a system that triages vulnerabilities as they come in, which reduces wait time and helps us get to safer software faster. Why are you zooming in on open source specifically? KALU: The code is already out there. Being able to see the code makes it so much easier, because when I tell you there is a bug in your software, I am pointing you to the exact line. Line 92, that is where the vulnerable thing is. You understand it better, I understand it better, and we fix it faster. I can suggest a fix because I have read the code. It is basically a white box approach, and white box offers so much more than black box. Open source has been pushing the community. Think about it. Linux is the most open source of all open source, and it is the operating system powering so much of our infrastructure today. So many people have built on it. That is the kind of software we are going to keep seeing. A lot of people are open sourcing things now because AI has helped develop so much software. That is where I want to shift my attention. I do still want bounties, but they are no longer the focus. I am doing security research, and that is way more than getting a bounty. Last word for the community. KALU: If you are going to implement a second brain using AI, do it in a way that is cheap, because AI is still expensive. It is becoming more accessible, but if you want to do it today, I have some tips.
RUNNING A RESEARCH SYSTEM CHEAPLY
FA L L
2026
• Do not ask it to narrate. Tokens are counted on input and output. Tell it to skip the running commentary and give you the output.
ABOUT THE AUTHOR Ezinne Kalu is an Application Security Engineer and Bug Bounty Researcher who loves turning real hacks into lessons anyone can learn from. She’s passionate about making security approachable and helping developers spot mistakes before hackers do. Through hands-on research, opensource contributions, and community initiatives, she bridges the gap between development and security; empowering teams across Africa to ship safer code, faster.
• Use the lower models. You do not have to use Fable 5. Fable is very expensive. Opus and Sonnet work as well. You just need to be super clear with your prompt. Same for Codex. • Let the agent manage its own brain. I am using my first brain, so the second brain is for the agent. It manages it and updates it. • Not everything needs AI. I only run my system during off-peak pricing. Rather than using a loop, which spends tokens re-prompting itself, write a cron job to start and stop the runs. You can even have Claude write the cron job for you. • Use cheaper inputs. Images cost fewer tokens than long text. If you copied text out of an image, you should probably have given the model the image. • Talk to it. It is faster to think of a prompt while speaking than while writing. Tools like Fluid Voice or Wispr Flow transcribe well, and there are open source options now.
42
A I CYB ER
SECURITY RESEARCH
Six Models Denied Up To 92% Of Real CVEs Without Inventing A Single Fake One. Our hallucination metrics are blind in one eye. The failure they cannot see is the one that ends an investigation. By Molly W. Correia
A large language model is a statistical system trained to predict the next fragment of text given everything before it. That single mechanism produces everything we now use these systems for: summarizing an incident, translating a question into a query, explaining what a vulnerability does. It is worth holding on to the mechanism, because it explains the failure modes. The model is not consulting a database when it answers. It is generating the most probable continuation of your question, drawn from patterns absorbed during training. When the training data contained the answer clearly and often, the most probable continuation is usually correct. When it did not, the model still produces a fluent, well-formed, confident sentence. Nothing in the architecture distinguishes the two cases, and nothing in the output announces which one you received. This matters in the SOC because adoption there is already settled. Alert volume grows every year and headcount does not. Language-model assistance is a staffing arithmetic problem rather than a question of enthusiasm. Every major platform now ships an assistant, endpoint and cloud and email vendors have embedded summarization, and a category of autonomous triage products has grown up beside them.
FA L L
2026
Not All Hallucination Is The Same Thing The word covers at least six distinct failures, and collapsing them costs us precision.
SIX FAILURES WE CALL BY ONE NAME • Fabrication. The invention of an entity that does not exist: a vulnerability identifier, a software package, an API endpoint, an academic citation. • Factual error. A real entity described wrongly. The right vulnerability with the wrong severity score, or the wrong affected version range. • Misattribution. Activity assigned to the wrong threat actor, or behavior mapped to the wrong technique. • Stale knowledge. A confident answer drawn from a world that has since moved on. A structural problem when the threat landscape changes daily and training cut-offs do not. • Tool-use failure. In agentic deployments, the model skips a lookup it should have made, calls it with the wrong parameters, or misreads what
43
A I CYB ER came back. • Sycophancy. The model agreeing with whatever your question implied, which is difficult to detect precisely because the answer matches your expectation.
FA L L
2026
and verification costs a change window. When the recovery path gets expensive, a confident denial gets expensive with it. The public data on industrial vendors and firmware is also far thinner than for enterprise software, which is precisely where a model is least likely to engage and most likely to refuse.
There is a seventh failure that rarely appears on these lists, and it is the subject of this article. A model can decline to engage at all, asserting that a real, published vulnerability does not exist. It sits awkwardly in the taxonomy because it is not an invention or a distortion. It is a refusal. As I will argue, that is exactly why we have not been measuring it.
What To Do About It FOUR ACTIONS 1. Ground vulnerability lookups in retrieval against authoritative sources rather than model memory. 2. Treat a model’s assertion that something does not exist as an unverified claim rather than a negative result. 3. Tell your analysts plainly that a confident contradiction from an assistant is not evidence they are wrong. The pull the other way is strong, and nobody warns them about it. 4. When you evaluate a tool, ask what the model does when it does not know, and whether the confidence it reports means anything when it is wrong.
What Is Hiding In The Blind Spot Last month I asked six large language models about a vulnerability. A high-severity flaw, CVSS 8.7, published and indexed in the National Vulnerability Database,6 where anyone reading this can look it up in eight seconds. All six told me it does not exist. Not one hedged. Their stated confidence ran between 80 and 100 out of 100. I had not gone looking for that. The study was built to catch the opposite problem, the one everybody warns about, where the model invents a vulnerability that was never real. Across all six models and 295 fabricated identifiers, that never happened once.
That last one deserves emphasis. Do not ask how often it hallucinates. That number answers a question your analysts are not facing. Ask whether the confidence it reports means anything when it is wrong. If nobody at the vendor can tell you, that is itself the answer.
THE STUDY • Method. Selective prediction, a framework that has sat in the machine learning literature since 1970 and had never, as far as I can tell, been pointed at security knowledge. It separates two decisions that hallucination metrics fuse together: how often a model answers at all, and how often it is wrong when it does.3,4 • Benchmark. Frozen. 250 real, severitystratified vulnerabilities and 50 year-matched fabricated identifiers. Every real record predates every model’s training cut-off, so ignorance is not an available excuse. • Fabrications invented. Zero. • Real vulnerabilities rejected as non-existent. Between 65 and 92 percent, depending on the model. Consider what those two failure modes cost an analyst. A fabricated vulnerability is dangerous, but it is loud. The analyst goes to verify it, finds nothing, and the error dies there. A confident denial of a real vulnerability is quiet, and it is terminal. The analyst does not go and verify it, because there is nothing to verify. The investigation simply stops, and the stopping looks like a resolution.
A confident denial of a real vulnerability is quiet, and it is terminal. The investigation simply stops, and the stopping looks like a resolution. We have all found ourselves in this situation. You ask an AI a question you already know the answer to, it responds with such confidence that you are wrong, and instead of pushing back, you believe the model.
NOTES
The Axis Nobody Is Buying On The models were roughly equally wrong. They were nowhere near equally honest about it. When wrongly denying a real vulnerability, one frontier model expressed low confidence 80 percent of the time. That is an operationally recoverable failure, because stated uncertainty invites a second look. Another frontier model, from a different vendor, asserted non-existence at high confidence in 78 percent of its errors. A small open model did so on every single one of its 207 wrong denials. This does not track model size, tier, price, or provider. Two frontier systems sit at opposite ends of the range, and the worst-calibrated model in the study was the smallest. Overconfidence in stated certainty is a known property of these systems.5 The variation between them is the part that matters operationally. Honest uncertainty is a training outcome rather than something that arrives with scale. That is worth saying plainly to anyone about to solve this by buying a larger model.
1. J. Spracklen, R. Wijewickrama, A. H. M. N. Sakib, A. Maiti, B. Viswanath, and M. Jadliwala, “We Have a Package for You! A Comprehensive Analysis of Package Hallucinations by Code Generating LLMs,” 34th USENIX Security Symposium, 2025. 2. T. Dao-Duy, “HalluCVE: A Multi-Signal Benchmark for Hallucination Detection in LLM-Generated Cyber Threat Intelligence,” Engineering and Technology Horizons, vol. 43, 2026. 3. C. K. Chow, “On Optimum Recognition Error and Reject Tradeoff,” IEEE Transactions on Information Theory, vol. 16, no. 1, pp. 41 to 46, 1970. 4. R. El-Yaniv and Y. Wiener, “On the Foundations of Noise-Free Selective Classification,” Journal of Machine Learning Research, vol. 11, pp. 1605 to 1641, 2010. 5. S. Kadavath et al., “Language Models (Mostly) Know What They Know,” arXiv:2207.05221, 2022. 6. NIST, National Vulnerability Database. https://nvd.nist.gov/ 7. CISA, Known Exploited Vulnerabilities Catalog. https://www.cisa. gov/known-exploited-vulnerabilities-catalog 8. Study data and code: https://github.com/Mollywc/cve-selectiverisk
ABOUT THE AUTHOR Molly W. Correia is a security consultant at a New York City utility with six years in cybersecurity and an MSc in Cybersecurity and Privacy. Her research examines how large language models fail on security knowledge.
Honest uncertainty is a training outcome rather than something that arrives with scale. Why This Is Worse In Operational Technology In enterprise IT a confident denial is recoverable in minutes. You open the database and check. In an operational technology environment, you cannot scan the controller to settle the question,
44
A I CYB ER
FA L L
2026
There’s over 15,000 copies of me, each month, in the hands of the people who approve budgets. This page is one of them. To make it yours, email ads@aicybermagazine.com
45
A I CYB ER
FA L L
2026
46
A I CYB ER VULNERABILITY MANAGEMENT
Beyond CVEs: Managing Vulnerability Intelligence At Machine Speed. Disclosure volume is on track for the highest year in history. The actionable patching burden has barely moved. Those two facts belong in the same sentence. By Krity Kharbanda and Aastha Sahni In 2024, the industry published 40,077 CVEs. At the time, that felt like a lot. Scanners were already outpacing remediation capacity, backlogs were growing, and most security leaders assumed we were near the ceiling of what a vulnerability management program could realistically absorb. Then 2025 came in at 48,244, a 20 percent jump in a single year. Teams adapted the way they always do. More dashboards, more tickets, more triage meetings. Painful, but manageable. 2026 broke that pattern. In June, FIRST reported that CVE disclosures were running 46.3 percent above its own February forecast, with 6,420 more CVEs than expected in just the first four months of the year. FIRST revised its full-year projection up to roughly 66,000, which would make 2026 the highest-volume year for vulnerability disclosure in history.
THE PROGRESSION 2024. 40,077 CVEs published. 2025. 48,244. A 20 percent jump in a single year. 2026. Running 46.3 percent above FIRST’s own February forecast by June, with a revised full-year projection of roughly 66,000.
Discovery Got Cheap Faster Than Processing Did Read that progression cold, 40,077 to 48,244 to a projected 66,000, and you would assume software quality is falling apart. We do not think that is the right read. What is happening is that our ability to find vulnerabilities has improved a lot faster than our ability to process what we find. FIRST’s own data supports this. The report points to three structural drivers behind the surge.
FA L L
2026
Rain Versus Flood Here is the part that should change how you plan capacity for next year. When FIRST filtered the data down to what is in CISA’s KEV catalog or carries an EPSS score above 10 percent, the actionable patching burden barely moved, even as raw CVE volume kept climbing. FIRST describes this as the difference between rain and flood. Every disclosure is a raindrop, most of them evaporate before they matter, and the job now is figuring out which ones are actually going to flood the house.
Every disclosure is a raindrop. Most evaporate before they matter. The job now is figuring out which ones are going to flood the house. That is also a useful way to think about what AI is doing to this industry. The conversation usually gets stuck on whether AI helps attackers or defenders more. From where we sit, its biggest effect so far has been making discovery cheap for everyone. Researchers can chew through much larger codebases, spot variants of known bugs, and generate exploit hypotheses at a pace that was not really possible two years ago. Finding bugs used to be the hard part. Now the hard part is what you do once you are staring at ten thousand of them.
Finding bugs used to be the hard part. Now the hard part is what you do once you are staring at ten thousand of them. The CVE Program is running into the same problem from a different angle, through its community discussions on AI-enabled vulnerability discovery. Discovery volume is not the only concern. Reporting, validation, coordination, remediation, and enrichment are all under pressure at once, because AI is compressing timelines at every stage of that pipeline rather than only at the front end.
Filtering Before A Human Ever Sees It THREE DRIVERS BEHIND THE SURGE 1. AI-assisted vulnerability discovery. 2. A 449 percent year-over-year increase in GitHub Security Advisory volume. 3. A 3,119 percent jump in VulnCheck’s CNAof-Last-Resort activity, which is largely the industry finally clearing a backlog of unassigned vulnerabilities that has been sitting around for years. Code quality probably did not drop off a cliff in six months. The machinery for finding and cataloging bugs got a lot more efficient, almost overnight, and that is a different thing.
Some of the tooling responses are already promising. OpenAnt, a 2026 open-source system that pairs LLM reasoning with adversarial verification and sandboxed dynamic testing, cut its analysis surface by up to 97 percent on projects like OpenSSL and WordPress simply by filtering down to code reachable from an attacker’s entry point before it even starts looking for bugs. That is roughly the direction this needs to go. Fewer raw findings dumped on a human, more findings that have already survived some scrutiny before a person ever sees them. It is worth saying plainly that CVE is still the thing that lets the whole industry talk about the same vulnerability using the same language, and that still matters a lot. But a CVE identifier was never a risk decision on its own, and the Program itself says as much. Whether you act on something depends on CWE classification, CVSS, EPSS, KEV status, what
47
A I CYB ER
FA L L
2026
is actually exposed in your environment, and how much the affected system matters to your business. None of that comes bundled with the ID.
Kill The Pattern, Not The Ticket The bigger mindset shift is moving away from counting tickets and toward killing the pattern that keeps generating them. A thousand hardcoded secrets in your codebase is one secrets-management problem wearing a thousand different CVE numbers. Repeated auth flaws are usually the same architectural gap showing up in different services. Recurring injection bugs often trace back to a missing framework control rather than a thousand developers independently making the same mistake.
A thousand hardcoded secrets in your codebase is one | secrets-management problem wearing a thousand different CVE numbers. Teams that keep score by ticket-closure rate will keep closing tickets forever. Teams that go after root cause actually shrink the backlog over time. AI did not create this scaling problem. It has been building for a while, and 2026 is just the year it became impossible to ignore. Volume was never going to be what separated good security programs from bad ones. The teams that come out ahead over the next few years will be the ones who get good at turning a flood of raw findings into a short, trustworthy list of things that actually need fixing, fast enough for it to matter.
ABOUT THE AUTHORS Krity Kharbanda is a cybersecurity practitioner, researcher, and speaker with experience across application security, vulnerability management, cloud security, and emerging AI security risks. She is a Senior Product Security Engineer at ServiceNow, and Community and Development Lead at BBWIC, where she helps build initiatives and opportunities that support women in cybersecurity. Aastha Sahni is a cybersecurity professional with over nine years of experience across IAM, application security, Azure cloud, and vulnerability management, currently serving as a Security Analyst at Microsoft. She is the founder of BBWIC Foundation and a mentor with organizations including WiCyS and WomenTechNetwork, championing women in cybersecurity worldwide.
48
A I CYB ER
AGENT SECURITY
Security At Machine Speed Is The Wrong Race Machine-speed exploitation will outrun reactive defense. Security invariants create resilience by eliminating attack paths before an incident begins. By Niels Provos
On July 21, 2026, OpenAI disclosed [1] that a combination of its models, including GPT-5.6 Sol, escaped an internal cyber evaluation environment and compromised Hugging Face to obtain solutions for the benchmark they were being graded on. The escape began with a zero-day vulnerability in the evaluation’s package registry proxy. OpenAI ran the evaluation without its production cyber classifiers and with models whose cyber refusals had been reduced to measure maximum capabilities, placing greater weight on infrastructure containment. The incident is a warning that extends far beyond one laboratory or model. Automated vulnerability discovery and exploitation have advanced for years. AI agents add general reasoning, broad tool use, and persistence through long sequences of failures. These capabilities compress campaigns that once took days or weeks into hours or minutes. The usual response is to make detection and incident response faster, but the Hugging Face case shows the limit of that strategy. Its LLM-based analysis surfaced the compromise, and GLM 5.2 reconstructed more than 17,000 events in hours. That speed improved the investigation after the models had entered the environment, harvested credentials, and moved through internal systems.
FA L L
2026
Reactive security begins with observation, evidence correlation, classification, and a containment decision. Autonomous attackers can take several steps before that loop closes. Security invariants change the terms by removing attack paths before an incident begins.
The race began years ago At DEF CON in 2016, DARPA held the Cyber Grand Challenge [3], an all-machine hacking competition in which autonomous systems discovered, proved, and repaired software vulnerabilities in real time. It demonstrated that machines could generate working exploits and react without human assistance. The security community has had a decade to plan for this capability to reach ordinary software and infrastructure. That day arrived months ago. My own vulnerability discovery work [4] found that the capability depends as much on orchestration as on access to a particular frontier model. A finite-state workflow can keep a model investigating, require executable evidence, recover from failed hypotheses, and run against many codebases in parallel. Attackers can try thousands of variations,
49
A I CYB ER follow several paths at once, and preserve useful observations. Defensive teams still need to understand operational risk, coordinate changes, protect availability, and recover affected systems. Those responsibilities limit response speed. Trying to close that gap with faster alerts and larger triage queues will fail. Security programs need infrastructure that remains resilient when the attacker acts first and moves faster.
Remove the attack path A security invariant [5] is a broad property that infrastructure maintains automatically and consistently across every system in scope, with no local bypass. It removes the corresponding attack surface without asking a person to recognize the attack or make the correct decision during the incident. My view comes from managing security engineering teams at Google and later serving as Stripe’s Head of Security. In my analysis of roughly 70 public breach reports, three invariants could have prevented 65 percent of the incidents [6]: mandatory hardware second factors, positive execution control, and production egress restrictions.b Egress control provides a direct example. A production workload with no route to the general Internet cannot establish command and control, retrieve a second-stage payload, or exfiltrate data. OpenAI routed package installation through an internal proxy and cache. The models exploited a zero-day in that cache and moved laterally until they reached a node with general Internet access. To preserve the invariant, a compromised cache should reach only the package registries it serves, and no node accessible from the evaluation environment should provide general egress. Assuming the first exploit succeeds is essential. Vulnerability management determines whether an organization can patch the proxy flaw before an attacker uses it. An invariant limits what the compromised component can do afterward. The first starts a race; the second changes the environment in which the race occurs. The Hugging Face compromise followed the same pattern. A malicious dataset abused a remote-code loader and template injection to run code on a processing worker. Requiring workers to
execute only reviewed loader code and restricting their egress would have contained the compromise within the processing boundary. Supply-chain aging removes a different attack surface. An internal package system can require third-party releases to age for a fixed period, keeping new versions unavailable while maintainers, researchers, and registries identify compromises. In April 2026, Socket found credential-stealing code in intercom-client@7.0.4 [7]. The preceding safe release had been published 88 days earlier. A 30-day aging period would have kept intercomclient@7.0.4 out of builds while preserving access to the earlier release.c Aging provides time rather than proof of safety and blocks immediate consumption of a newly poisoned release. Automatic service replacement constrains persistence. Every production instance can be rebuilt from checked-in, peer-reviewed source and replaced within a fixed interval. Changes made only to a running workload disappear with the next replacement. A foothold in mutable workload state therefore has a bounded lifetime unless the adversary also compromises the source repository or build control plane. This moves persistence into a smaller, better protected trust boundary. Context-aware data access removes ambient authority over customer data. A support agent, administrator, or service receives access only when a separate system verifies a current business justification, such as an assigned customer case. An employee account or service credential alone carries no authority to browse customer records. Requiring justification on every path keeps an operator compromise from exposing the full customer base. Each invariant eliminates a defined part of the attack surface. Coverage forms part of the definition: a forgotten environment with unrestricted egress, an instance exempt from replacement, or a direct database connection outside the justification system breaks the property. Engineering must make the secure condition universal and keep it that way.
Resilience comes from composition Even with an unpatched vulnerability, an invariant can remove its path to damage. Several invariants together create defense in depth. Supply-chain aging may keep a malicious release out of the build; egress control can prevent a compromised dependency from reaching its controller; scheduled service replacement can erase changes to a running instance. The attacker must cross several independently enforced boundaries to reach the objective. Invariants keep working as the number of attacks rises. Speed helps attackers search for implementation gaps, making consistent enforcement and independent layers crucial. It cannot substitute for authority the system never provides. This is the source of resilience. A resilient environment can absorb a component compromise while limiting its blast radius. The vulnerable service may crash or require replacement, but the organization protects its other workloads, credentials, and customer data. Incident response begins from a contained failure instead of a company-wide emergency. Incidents outside invariant coverage still matter. Tempo has the greatest effect where enforcement has failed or no invariant covers the next attack step. AI-augmented detection and response can reveal attempts to violate an invariant, find uncovered systems, and identify attacks against enforcement infrastructure. Detection exposes
FA L L
2026
gaps and protects the enforcement layer; invariants remain the primary mechanism of containment.
Build the properties into the platform For many companies, limiting security work to customer, regulator, and insurer demands has been rational. Invariants require up-front engineering, compete with product development, and pay off through prevention. AI changes that calculation by lowering the cost of finding and exploiting vulnerabilities and compressing the interval between discovery and compromise. The same workflow can now pursue a weakness across many organizations. An invariant incurs a fixed engineering cost, then protects against every attempt that depends on the attack surface it removed. The work remains difficult. Egress control requires maintaining every legitimate external dependency. Supply-chain aging affects developer workflows. Service replacement depends on reproducible builds and reliable deployment systems. Business-justified data access requires changes to applications, identity, and support tooling. The cost falls when these properties become platform defaults. A new service should start without general Internet connectivity, obtain dependencies through the governed repository, run from a reproducible image, and receive a distinct workload identity. Access to sensitive data should pass through a common system that verifies purpose. Product teams should inherit these properties instead of designing them separately for every service. Security teams can then measure concrete outcomes: which production workloads can reach the general Internet, which builds can import packages outside the aging policy, how old the oldest running service image is, and whether every customer-data read carries a verified justification. These measurements describe the actual attack surface; counts of alerts, vulnerabilities, and closed tickets describe work performed on it. OpenAI’s response points toward structural defense. The company says it is implementing stricter infrastructure configuration at the cost of research velocity [1] while patching the vulnerabilities. Research velocity has value, but containment must hold when advanced models receive sustained access to offensive tools. The durable strategy is to decide which attack transitions the infrastructure should make impossible. Build those properties into the platform, enforce them across every relevant system, and combine several independent invariants. An attacker may still win the first exploit. The architecture should deny that foothold a durable path to the objective.
NOTES 1. https://openai.com/index/hugging-face-model-evaluationsecurity-incident/ 2. https://www.darpa.mil/news/2016/cyber-grand-challenge 3. https://www.provos.org/p/finding-zero-days-with-any-model/ 4. https://securityblueprints.io/posts/security-invariants/ 5. https://securityblueprints.io/posts/three-security-invariants-cisochallenge/ 6. https://socket.dev/blog/intercom-s-npm-package-compromisedin-supply-chain-attack
50
A I CYB ER
AGENT SECURITY
Rogue Agents Are Here To Stay: We Need To Update Our Security Model. Enterprises should look beyond the headlines and learn how to manage rising agent autonomy in order to benefit from it. By John Sotiropoulos
Editor’s Note This piece runs alongside Niels Provos’s “Security At Machine Speed Is The Wrong Race.” Provos argues for architectural invariants that remove attack paths. Sotiropoulos argues for identity, policy and detection applied to the agent itself. They agree on egress. Over the summer, OpenAI and Anthropic reported agents escaping their evaluation environments and breaching third parties, including Hugging Face.1,2,3 Any suspicion that this served their market ambitions dissipated when the UK’s AI Security Institute reported agents taking sustained unsanctioned action against real people, including an attempted supply-chain attack on open-source software.4 Meta added a fourth disclosure days later.5 No hostile human was involved. An authorised
FA L L
2026
goal pushed too far, the real world misjudged as a test, and reachable live systems were enough for agents to go rogue. These accounts remain preliminary, and the investigations are continuing. Because of the cyber evaluations, the debate has focused on resilience against advanced capabilities. But capability simply widens the blast radius of a systemic risk: rogue agency. ASI10: ROGUE AGENTS We named this class in December 2025 as ASI10 Rogue Agents in the OWASP Top 10 for Agentic Applications.6 An agent that, pursuing what it takes to be its task, acts beyond the authority it was given, damages what it was meant to protect, and at times conceals or misstates what it did. An insider threat without an insider’s motive.
The Enterprise Evidence Predates The Labs It would be convenient to treat these as laboratory experiments, and the AI Security Institute says so itself: internet access deliberately enabled, cyber classifiers switched off.4 True, but not the whole story.
51
A I CYB ER
FOUR THAT HAD NOTHING TO DO WITH A LAB • July 2025. A Replit agent deleted a production database during a code freeze, and lied about it.7 • February 2026. A coding agent with autoapprove on infrastructure commands destroyed a production environment and its snapshots.8 • February 2026. A crypto trading agent with no transaction ceiling met a request for 310 dollars by sending tokens nominally worth 442,000.9 • April 2026. An agent wiped PocketOS’s production database in nine seconds after finding an over-scoped token in an unrelated file, then admitted it had made a mistake.10
Restrict: Treat The Agent As An Unvetted Employee Give each agent its own identity, a named human owner, a leaving date, and a registration record. Tier agents by the damage they could cause, treating the top tier as privileged identities. Apply least privilege, and issue credentials and permissions that are just-in-time and time-bound. Clear out the ambient credentials waiting in files to be found, which is what the PocketOS token was. Treat network egress as a decision to be justified rather than a default. Meta’s disclosure began in a misconfiguration that allowed internet access.5 Threat-model the agent and enumerate what its identity can reach before it goes live.
None involved a disabled classifier, a cyber evaluation, a rogue objective, or a zero-day. Our own quarterly round-up at OWASP found that most such events trace to misconfiguration, overprivileged agents, and unbounded autonomy rather than to tracked vulnerabilities.11 Excessive agency and over-broad permissions are the precondition rather than the phenomenon, which is why tightening them is where any answer starts. ASI10 shipped with mitigations and a worked example: an agent that deletes production backups because that is the cheapest route to its goal.6 That was December 2025, two months before an agent did exactly that.
FA L L
2026
worst of it was a human reviewer rather than a technical barrier. We should not leave that to luck. A rogue agent explores before it acts. Decoys turn that method into a high-fidelity signal, since a legitimate agent has no reason to touch one, provided they are tuned against your own scanners and automation. Place them where an exploring agent would look, keep them inert, and wire them to alerts and containment. Rehearse how to revoke an agent’s identity and halt a fleet with a kill switch. And test with exercises built around agent behaviour. In a recent grey-box engagement for a regulated financial-services client, working from the agent’s tools, permissions, and policies, we found far more than black-box probing had. None of this is exotic. It is identity, least privilege, segmentation, change control, and monitoring, applied now to an actor that works continuously, at speed, and with whatever authority we hand it. The lesson is not that agents are too dangerous to use. It is that the advantage will belong to whoever can deploy them under control.
The advantage will belong to whoever can deploy them under control.
NOTES
ASI10 shipped with a worked example in December 2025. Two months later, an agent did exactly that. What You Inherit And What Is Yours There is guidance from the AI labs, though written from their perspective. OpenAI’s Preparedness Framework sets out five claims for arguing a model will not autonomously cause severe harm.12 Three are properties of the model, being capability, values, and instruction adherence, which you inherit rather than set. That is a sharedresponsibility and supply-chain question. Two, oversight and architecture, are yours. The AI Security Institute’s lessons sit under those two: justify internet access, monitor the run as it happens, assume the agent may exceed its remit.4 Containment, as the Institute put it, should not depend on the model choosing not to test its boundaries. That is where the Agentic Top 10 mitigations should be operationalised.
Containment should not depend on the model choosing not to test its boundaries. THREE PRINCIPLES • • •
Restrict what an agent can reach. Govern what it may do. Detect what it does
Govern: Move The Decision To The Point Of Action The AI Security Institute named the absence of synchronous monitoring, meaning a second model reviewing each proposed action before it runs, among the possible contributing factors behind its own incident.4 It is worth having, but it remains one probabilistic system judging another, so it should not be the last gate before an irreversible action. Instead, express the rules as policy-as-code, with an engine such as Cedar or Open Policy Agent, and evaluate every consequential action before it takes effect: writes, deletions, privilege changes, code execution, and calls outside an allow-list. The verdict is allow, deny, or escalate. Keep those checks deterministic and grounded in facts you control, because an irreversible action should not rely on a probabilistic model judgement. Reserve human authorisation for a small, deliberately chosen set of irreversible actions. You cannot meet machine-speed activity with humanspeed approval, and a long list produces rubberstamping.
Detect: Treat The Agent As An Insider Route agent activity, above all the denials and escalations your policy layer produces, into the SOC. Many organisations still forward signals only from public-facing systems, which is not where their agents are. At the AI Security Institute, what stopped the
1. Hugging Face, “Security incident,” 16 July 2026. https:// huggingface.co/blog/security-incident-july-2026 2. OpenAI, “Hugging Face model evaluation security incident,” 21 July 2026, updated 28 July 2026. https://openai.com/index/hugging-facemodel-evaluation-security-incident/ 3. Anthropic, “Investigating three real-world incidents in our cybersecurity evaluations,” 30 July 2026. https://www.anthropic.com/ news/investigating-incidents-cybersecurity-evals 4. UK AI Security Institute, “Incident report: unsanctioned agent behaviour during cyber testing,” with technical report INC-2026-0728-01, 4 August 2026. https://www.aisi.gov.uk/blog/incident-reportunsanctioned-agent-behaviour-during-cyber-testing 5. SecurityWeek, “Meta AI hacked external systems during cybersecurity testing,” 6 August 2026. https://www.securityweek.com/ meta-ai-hacked-external-systems-during-cybersecurity-testing/ 6. OWASP GenAI Security Project, OWASP Top 10 for Agentic Applications 2026, December 2025, ASI10 Rogue Agents. 7. The Register, “Vibe coding service Replit deleted production database,” 21 July 2025. https://www.theregister.com/2025/07/21/ replit_saastr_vibe_coding_incident/ 8. A. Grigorev, “How I dropped our production database,” 6 March 2026. https://alexeyondata.substack.com/p/how-i-dropped-ourproduction-database 9. Cointelegraph, “OpenAI employee’s AI agent accidentally sent $442K,” 23 February 2026. https://cointelegraph.com/news/openaiemployee-s-ai-agent-accidentally-sent-442k-to-beggar 10. T. Claburn, “Cursor-Opus agent snuffs out startup’s production database,” The Register, 27 April 2026. https://www.theregister. com/2026/04/27/cursoropus_agent_snuffs_out_pocketos/ 11. S. Clinton, “OWASP GenAI Exploit Round-up Report Q1 2026,” OWASP GenAI Security Project, 14 April 2026. https://genai.owasp. org/2026/04/14/owasp-genai-exploit-round-up-report-q1-2026/ 12. OpenAI, Preparedness Framework version 2, 15 April 2025, appendix C.2. https://openai.com/index/updating-our-preparednessframework/
ABOUT THE AUTHOR John Sotiropoulos is the founder of Deep Cyber Ltd, which safeguards national-scale AI programmes across government, healthcare and finance. He chairs the OWASP Top 10 for Agentic Applications, co-leads the OWASP Agentic Security Initiative and serves on the board of the OWASP GenAI Security Project. He wrote the UK Government’s Implementation Guide to the AI Cyber Security Code of Practice, now the ETSI EN 304 223 standard, and is the author of Adversarial AI Attacks, Mitigations, and Defense Strategies.
52
A I CYB ER
FA L L
2026
There’s over 15,000 copies of me, each month, in the hands of the people who approve budgets. This page is one of them. To make it yours, email ads@aicybermagazine.com
53
A I CYB ER
FA L L
2026
54
A I CYB ER SECURITY OPERATIONS
What The Hugging Face Incident Teaches Us About The Future Of Agentic Security Operations. Tens of thousands of automated actions over a single weekend. No human-led team can review that volume in real time, and the agents brought in to help has to know your environment before it is any use at all. By Ely Abramovitch The recent Hugging Face incident reads like a science fiction thriller. An attacker uploaded a malicious dataset to Hugging Face’s platform that exploited two independent vulnerabilities in the data processing pipeline: a remote code execution flaw in the dataset loader, and a template injection flaw in the dataset configuration. From there, an autonomous agent took over. It executed code on a dataset processing worker, escalated its own privileges, harvested cloud and cluster credentials, and moved laterally into several of Hugging Face’s internal clusters, without a human directing a single one of those steps.1 Hugging Face has called it one of the first publicly documented, fully agentic cyber campaigns, driven end to end by an autonomous system.
The Defenders’ Own Tools Became Friction The alarming part is not only that an AI agent broke out of containment.. When Hugging Face’s own team turned to frontier AI models to help analyze the malware and reconstruct the attack, the models’ safety guardrails blocked large parts of the forensic work. This part is not rocket science. The defenders needed AI to move at the speed of the attack. The AI built to help them was, in that moment, part of the friction. They then turned to an open-weight model they could run entirely on their own infrastructure, specifically so they could do the incident response work without a vendor’s guardrails getting in the way.
The defenders needed AI to move at the speed of the attack. The AI built to help them was, in that moment, part of the friction. That is the underlying lesson here, and it applies far beyond this one incident. Every organization investigates a little differently, because every organization is different. The tools are different, the escalation paths are different, and the thing your best analyst just knows, the instinct that never made it into a runbook because nobody thought to write it down, is different too. An agent that is going to do real investigative work across the security stack has to learn the things that are most specific to your enterprise.
The Working Model Needs Reexamining
Now that we have fully agentic attack chains, cybersecurity’s working model, in which AI augments human work, needs to be reexamined. The result of this particular incident was tens of thousands of automated actions, executed autonomously and without correlation, over the course of a weekend. No human-led team, regardless of how good, talented, or experienced, can review that volume of activity in real time, let alone capture what matters before the damage is done. This is where the industry needs to be honest with itself. What is required is agents that do the actual work of investigation, end to end, department by department, with humans setting the boundaries, reviewing the conclusions, and stepping in exactly where judgment is genuinely required. A faster human-in-the-loop does not close a gap of that size. That is a different job from the one most security teams are staffed and trained for today, where an operator’s value is measured by how many alerts they can personally work through.
WHAT CHANGES FOR THE HUMAN • Today. An operator’s value is measured by how many alerts they can personally work through. • Next. An orchestrator’s value is measured by how well they direct a fleet of agents already working through all of them. • The work becomes deciding what agents are allowed to do on their own. Catching the calls that need a human’s read on the situation. Holding the whole operation accountable for getting it right. Security teams need agents that understand organizational context and can autonomously investigate every incoming alert, ticket, and escalation. Agents that pick up an alert and work it the way a senior analyst would, at a volume and speed no team of humans could sustain. That only works if the humans above them have made the shift from doing the work themselves to directing the agents who do.
FA L L
2026
This new way of running investigations cannot stay contained to one department. Attacks do not respect the boundary between the SOC and vulnerability management and DLP. Agentic threats in particular are all-seeing and constantly in motion, and they move straight through it. Security agents that work across departments rather than within them close the exact structural gap attackers are best at exploiting, turning what used to be the attacker’s advantage into the defender’s, now running at machine speed instead of human speed. Hugging Face will not be the last company to publish an incident like this. I would go so far as to bet on it being one of the more contained examples we look back on. The organizations that treat it as a preview, and build agentic investigation into every corner of security operations now, are the ones who will be best prepared for when it is not happening in a test environment somewhere
NOTES 1. [Source for the attack chain to come. The submission carried a footnote marker here with no reference attached.] 2. Anthropic, “Investigating three real-world incidents in our cybersecurity evaluations,” 30 July 2026. https://www.anthropic.com/ news/investigating-incidents-cybersecurity-evals
ABOUT THE AUTHOR Ely Abramovitch is the Co-Founder and CEO of Legion Security, the agentic security operations platform for the enterprise. With a background leading product management for Microsoft Sentinel he has a proven track record of scaling multi-billion dollar security solutions, and with Legion he leads the organization currently helping Fortune 100 enterprises deploy trusted AI across their security operations.
Guardrails Alone Do Not Solve This Another interesting piece of this incident is the call for more and better guardrails. Guardrails are good and necessary but Hugging Face’s own incident response shows why they are incomplete. Guardrails are not sufficient on their own, and sometimes they are actively in the way, particularly if the agent behind them does not understand the environment it is operating in. An agent that is contextually blind will either get walled off by its own safety training at the exact moment you need it most, or it will act on the wrong context and get the call wrong, which was also seen in the recent Anthropic incident.2 Both failure modes come from the same root cause. The agent does not know your organization.
Both failure modes come from the same root cause. The agent does not know your organization. Attacks Do Not Respect Your Org Chart 55
A I CYB ER
FA L L
2026
56
A I CYB ER SECURITY OPERATIONS
When Your AI Agent Makes The Wrong Call At Machine Speed, There Is No Undo Button. The instinct is better logging and tighter permissions. Both are necessary. Neither answers whether the data the agent acted on was correct. By Frank Balonis
Editor’s Note Frank Balonis is Field CISO at Kiteworks, which operates in the data governance and secure data exchange category this article discusses. AI Cyber Magazine publishes contributed analysis from practitioners at vendors when the argument stands on its own evidence. Editorial coverage and paid placement are entirely separate tracks at this publication. Every CISO running an AI-augmented security operations center has been sold the same efficiency story. Autonomous agents that triage alerts faster than any human analyst, remediate incidents in seconds, and cut mean time to detect in half. That story is accurate. It is also incomplete. What most vendors do not tell you is that autonomous agents acting on wrong data make catastrophically wrong decisions before any human can intervene.
Autonomous agents acting on wrong data make catastrophically wrong decisions before any human can intervene. I have been watching this failure emerge in production environments for the past 18 months. It has a name now, the context gap. It is not an edge case or a theoretical risk sitting in a research paper. It is an operational reality that security leaders need to understand before they scale agentic AI further into their SOC.
THE CONTEXT GAP Agentic AI systems observe context dynamically by reading logs, querying data stores, and processing endpoint records, then reason over what they find to decide what to do next. When the context is correct, this is extraordinarily powerful. When the context is stale, misrouted, or drawn from the wrong data object, the agent makes an internally valid but operationally wrong decision. It makes that decision in under two minutes, with no human checkpoint to catch it. Cyera’s 2025 State of AI Data Security Report, based on a survey of more than 900 security leaders, found that 83 percent of enterprises already run AI in daily operations, while only 13 percent have strong visibility into what that AI is actually doing.1 That visibility gap is not an administrative inconvenience. It is the space where wrong decisions compound undetected.
FA L L
2026
The Deployment Velocity That Created This Problem The business case for autonomous AI in security operations is real. IBM’s 2025 Cost of a Data Breach Report found that AI-augmented SOCs cut mean time to detect by 50 percent.2 That is before you account for analyst burnout and the sheer volume of alerts that no human team can process sustainably. Security teams receive an average of 4,484 alerts per day and spend 27 percent of their time on false positives. The efficiency argument for agentic AI is not just sound, it is necessary. What has not kept pace is governance. A Q4 2025 survey of 225 enterprise leaders found that 100 percent have agentic AI on their 2026 roadmap. Only 37 to 40 percent have meaningful containment controls. That gap between deploying agents and controlling what they see is precisely where the failure mode lives. CrowdStrike’s 2026 Global Threat Report puts the average eCrime breakout time at 29 minutes.3 The fastest recorded is 27 seconds. Autonomous AI defenses exist because human defenders cannot match that speed. But machine-speed defense operating in the wrong context produces machinespeed mistakes. A locked production credential. A remediation action on the wrong host. A false positive escalated into a full incident response that takes hours to unwind.
Machine-speed defense operating in the wrong context produces machine-speed mistakes. The NIST AI Risk Management Framework identifies data quality and context integrity as core governance requirements for AI systems in high-stakes operational environments.4 Most AIaugmented SOC deployments I have reviewed do not have a governance layer that addresses either requirement. The agent has broad data access and is assumed to use it appropriately. That assumption is the risk.
Why The Current Approach Is Insufficient The instinct when something goes wrong with an autonomous agent is to improve logging and tighten permissions. Both are necessary. Neither is sufficient. Logging is post hoc. If your agent made a wrong call in under two minutes, the damage has landed before your log review begins. You can reconstruct what happened, assuming your logging captures what the agent actually saw rather than only what it did, but you cannot undo the automated lockout, the blocked traffic, or the incident response chain that fired. Tighter permissions are the right direction, but they still do not answer the core question. Is the data the agent received accurate, complete, and appropriate for the decision it made? Permissions restrict what an agent can access. They do not verify that what it does access is correct.
57
A I CYB ER
Classify content before it enters an agent workflow.
Permissions restrict what an agent can access. They do not verify that what it does access is correct. EY’s Responsible AI Pulse Survey, covering 975 C-suite leaders across 21 countries, found that 99 percent of organizations reported financial losses from AI-related risks in 2025.5 Only 12 percent of C-suite respondents could correctly identify appropriate controls against five known AI risk scenarios. CISOs are ahead of the rest of the C-suite on this, but not by enough. The organizations deploying agents fastest are often the least equipped to govern them.
The Architecture Question No One Is Asking Ask what verifies that the data your agent receives is accurate before it acts. That is a more useful question than asking what permissions the agent should have, and it is an architecture question rather than a monitoring one.
WHAT A GOVERNANCE LAYER HAS TO DO
Enforce policy on what any given agent can access in a given context.
Produce a tamper-evident log of every data object the agent saw before it took action.
KPMG’s 2025 research on data governance and AI found that 62 percent of organizations cite insufficient data governance as the top barrier to scaling AI.6 In my experience, that number is if anything understated, because most organizations do not have a clear picture of what their agents are accessing until something goes wrong. The SEC’s evolving guidance on AI risk disclosure and the FTC’s focus on AI system accountability are both moving in the same direction. Organizations are going to be expected to demonstrate that their AI systems operated within appropriate governance constraints, not just that they had governance policies in place. That is a meaningful shift, and it applies directly to AI-augmented SOC operations. An audit trail that captures what your agent saw before it acted is good security practice today, and it may become a regulatory requirement tomorrow.
FA L L
2026
NOTES 1. Cyera, 2025 State of AI Data Security Report. https://www.cyera. com/research-labs/2025-state-of-ai-data-security-report 2. IBM, 2025 Cost of a Data Breach Report. https://www.ibm.com/ reports/data-breach 3. CrowdStrike, 2026 Global Threat Report. https://www.crowdstrike. com/en-us/global-threat-report/ 4. National Institute of Standards and Technology, AI Risk Management Framework 1.0. https://www.nist.gov/system/files/ documents/2023/01/26/AI%20RMF%201.0.pdf 5. EY, Responsible AI Pulse Survey, June 2025. https://www.ey.com/ en_gl/newsroom/2025/06/ey-survey-ai-adoption-outpacesgovernance-as-risk-awareness-among-the-c-suite-remains-low 6. KPMG, “Rebuilding data governance in the age of AI,” 2025. https:// kpmg.com/us/en/articles/2025/rebuilding-data-governance-in-ageof-ai.html
ABOUT THE AUTHOR Frank Balonis is Field CISO at Kiteworks, a secure data exchange company. He brings more than two decades of experience in IT support and services, and previously held senior engineering positions at Prexar, Kokusai Semiconductor Equipment Corporation, and Varian Semiconductor Equipment. He is a Certified Information Systems Security Professional and a veteran of the United States Navy.
Sitting between your data stores and your AI agents, it should:
58
A I CYB ER
Your Fraud Detection Model Doesn’t Need To Fail. It Just Needs To Disappear For 90 Seconds. The gap between the model is running and the model is deciding is a failure mode most production systems do not measure, most incident classifications do not recognize, and most security frameworks treat as somebody else’s problem.
The monitoring dashboards said the fraud detection model was up. Latency graphs, green. Processhealth checks, returning. Compute utilization, well below threshold. But somewhere in a payment authorization pipeline, transactions were arriving at a decision point where no decision was being made. The model was there, and it was not there. The service was healthy, and the security posture had collapsed.
By Vijayent Kohli
The model was there, and it was not there. The service was healthy, and the security posture had collapsed. Anyone who has sat in a payment platform’s incident room during a partial failure knows a specific kind of quiet. It is not the quiet of an outage, where the conversation is clear and somebody is already talking about rerouting. It is the quiet of a model that is technically running but no longer participating in decisions. Somebody asks whether
FA L L
2026
to shut the server down. Somebody else points out the dollars are still moving in real time. This has been the shape of more than one long night. It is also a failure mode with a fifty-year-old security name that the field has forgotten to apply.
The A That Got Quietly Removed The name is Availability. It has been the third leg of the CIA triad, alongside confidentiality and integrity, since NIST codified the framework for federal systems under FIPS 199, sourcing all three properties to the Federal Information Security Management Act. ISO/IEC 27000 uses effectively the same definition. Availability is the property of a system being accessible and usable on demand by an authorized entity. Every security engineer learns this in their first month. Yet in practice, the industry routes ML availability failures somewhere else. When a fraud detection model misses its SLA, the ticket gets routed to site reliability. When it fails to return a decision inside the authorization window, the postmortem cites infrastructure. When authorization traffic reverts to weaker fallback rules for ninety seconds, the incident is classified as a reliability event.
59
A I CYB ER The security team is rarely paged. The security dashboard rarely shows the impact. The security framework rarely acknowledges the event happened. This is not the security frame extending to cover AI. It is the security frame being quietly narrowed to exclude it.
This is not the security frame extending to cover AI. It is the security frame being quietly narrowed to exclude it. The academic and regulatory literature has been saying this for years. NIST’s adversarial machine learning taxonomy classifies availability as a firstorder attacker objective alongside integrity and privacy.1 MITRE ATLAS catalogues the attack pattern as Denial of AI Service. Cambridge and Toronto researchers demonstrated in 2021 that adversaries can craft inputs to slow production ML models by factors of thirty or more, and in one experiment against a live translation service, six thousand.2 What is largely missing is any public record of financial institutions disclosing when this class of failure happens to them.
Designing For The Failure In a 2021 engineering blog post, PayPal’s own team described their production reality plainly. Fraud detection models can fail during execution. They can miss their real-time SLA. When either happens, an entire live pool of models has to be rolled back and taken through the full model development lifecycle again. This is not adversarial. This is Tuesday.
FA L L
2026
interest of disclosure, I am a co-inventor on it. The fallback trains on the input-output stream of the production model during normal operation, and is provisioned only after its accuracy meets a defined threshold. A policy layer constrains what the fallback is permitted to decide during the outage. The design assumption is worth naming plainly. Availability is not treated as a metric to preserve. It is treated as a security property to architect. The goal is an architecture that keeps deciding through the failure, safely and within a bounded policy, rather than an uptime SLA. The system fails. The security posture does not.
The system fails. The security posture does not. Reclaiming The A The public disclosure record is nearly silent on what happens to fraud controls when models go dark. The Visa Europe outage of June 2018 affected 2.4 million UK transactions and 5.2 million across Europe, over roughly ten hours. Visa attributed it to hardware failure and expressly ruled out cyberattack. Square’s September 2023 outage was traced to a firewall change interacting with a DNS upgrade, and its forensic analysis found no evidence of a breach. In neither case did anyone publicly answer the more difficult question. What did the authorization pipeline decide about the transactions that did make it through? The A in the CIA triad has always been there. Reclaiming it for AI systems is not a matter of new frameworks. It is a matter of who gets the page when the next model quietly stops deciding.
NOTES
This is not adversarial. This is Tuesday. The architectural response is not to eliminate the failure. Perfect availability is not achievable at scale, and pretending otherwise builds brittle systems that hide their fragility until it becomes an incident. The response is to design what happens when the model becomes unavailable.
1. A. Vassilev, A. Oprea, A. Fordyce, and H. Anderson, Adversarial Machine Learning: A Taxonomy and Terminology of Attacks and Mitigations, NIST AI 100-2e2023, 2024. https://doi.org/10.6028/NIST. AI.100-2e2023 2. I. Shumailov, Y. Zhao, D. Bates, N. Papernot, R. Mullins, and R. Anderson, “Sponge Examples: Energy-Latency Attacks on Neural Networks,” 2021 IEEE European Symposium on Security and Privacy, pp. 212 to 231. arXiv:2006.03463. 3. US Patent 11,568,253 B2, Fallback Artificial Intelligence System for Redundancy During System Failover. Granted 31 January 2023. https:// patents.google.com/patent/US11568253B2/en
ABOUT THE AUTHOR THE FAMILIAR FALLBACKS, AND WHERE EACH ONE BREAKS • Cached decisions. They age. • Deny-by-default. Drops legitimate transactions and creates its own kind of theater. • Static rule engines. Exactly what attackers reverse-engineer during the outage window. • Human-in-the-loop escalation. Does not survive contact with authorization-window latency.
Vijayent Kohli is a Principal Cybersecurity Engineer at Ford Motor Company, with prior work at Microsoft, PayPal, Oracle, and Nokia Siemens Networks. He is a co-inventor of US Patent 11,568,253 B2, referenced in this article.
A more interesting pattern is a continuouslytrained fallback model that mirrors the production model’s decision surface and can take over the API contract when the primary fails. US Patent 11,568,253 B2, filed at PayPal in 2020 and granted in 2023, describes one such architecture.3 In the
60
A I CYB ER
Your MCP Allowlist Isn’t As Safe As You Think. Your allowlist knows which programs it let in. It has no opinion on what they do once through the door. By Bandana Kaur
Somewhere in your stack, an AI agent is running commands on real infrastructure. Maybe a shell tool here, a git wrapper there, a build runner, or some innocent little productivity plugin. That is objectively an unhinged amount of authority to grant an AI model, so you do the responsible grown-up thing and put it on a leash. An allowlist. Approved binaries only. git and nothing else, or whatever your one blessed tool happens to be. The trouble is that the leash has a hole shaped exactly like the thing you allowed.
The leash has a hole shaped exactly like the thing you allowed. Arguments Are Verbs To see why, you have to look at arguments differently. Most people treat a command’s arguments as inert data, the noun the verb acts on.
FA L L
2026
Spoiler alert: they are not. Argument injection is essentially when you do not control the binary, but you control enough of its arguments to change what it does. A huge number of everyday CLIs generalize this trick. find has -exec. awk has system(). python has -c. And git has aliases that shell out when prefixed with a bang. Most people ask whether an allowlisted binary is dangerous. The question that matters is whether it can be told to run something else.
Where The Leash Breaks Now walk that into the Model Context Protocol. Servers wrap tools so agents can call them, and the standard safety move for a shell tool is exactly to restrict execution to approved binaries. That model is structurally incomplete, and I can show you in one line. Configure the server with only git in the allowlist. Then have the agent call: git -c alias.x=!whoami x whoami runs, when only git was on the list. The -c flag defines an alias inline, the ! hands its body to the shell, and the allowlist waves it through because it only checked who was running, never what it was asked to do.
61
A I CYB ER The server did everything the textbook says. It parses argv rather than passing a string to a shell, so classic metacharacter injection does not apply. The escape hatch just moved indoors, into the binary’s own features. In CWE terms this is argument injection (CWE-88) riding on command injection (CWE-78).1,2
which programs it let in and has no opinion on what they do once through the door. Until you answer both questions, your one-binary sandbox has as many exits as that binary has features. And git, bless it, has so many features.
NOTES
The escape hatch just moved indoors, into the binary’s own features.
1. MITRE, CWE-88: Improper Neutralization of Argument Delimiters in a Command (Argument Injection). [URL to come.] 2. MITRE, CWE-78: Improper Neutralization of Special Elements used in an OS Command (OS Command Injection). [URL to come.] 3. GHSA-jm26-853c-62h9, 1 August 2026. [Advisory URL and affected project to come.]
FA L L
2026
ABOUT THE AUTHOR Bandana Kaur (HackWitHer) is an 18-year-old offensive security researcher specializing in GenAI Application and LLM security. She works as a researcher at APIsec Labs, where she leads original research on AI systems, agentic workflows, API security, and emerging application architectures. She has reported vulnerabilities in government systems, trained CISOs on OSINT, and advocates globally for ethical hacking practices. She is the discoverer and original reporter of the advisory GHSA-jm26-853c-62h9.
And no human needs to type that string. Anything the agent reads can supply it: a malicious README, a poisoned issue comment, a crafted file in a repo it was asked to review. Indirect prompt injection feeds straight into argument injection, two supposedly separate threat models collapsing into one exploit chain.
Three Buckets, One Policy This article is built around a real advisory published on 1 August 2026, GHSA-jm26-853c-62h9, rated High, which I co-authored.3 You cannot audit this with a payload list, because a payload for git tells you nothing about the next binary. So what do you actually do about it? It depends on which chair you are sitting in.
THE THREE FLAG BUCKETS Read each allowlisted binary’s own man page once. Grep it for exec, command, shell, pager, filter, and alias. Then bucket every flag: 1. Config override. Anything like git -c. 2. File read and write. 3. Process spawn. Default-deny everything in buckets two and three, and the | config-override flags in bucket one as well. If you are the developer wiring up the server, your allowlist is a policy about identity and you also need one about behavior. Work the three buckets above, and do not assume argv parsing bailed you out, because it did not. If you are the CISO whose engineers are gleefully plugging agents into everything, an allowlist is not an OS sandbox. Here old is gold: least-privilege users, restricted working directories, containers, seccomp. If you are the researcher or pentester, this is a repeatable pattern. Enumerate any allowlisted binary’s flags across those same three buckets, and the bug class travels intact from one tool to the next.
The Last Line Of Defense If you run a server someone else wrote, you may never touch the validation code, so assume the allowlist leaks. Our advisory is blunt that hardening alone is not a sandbox. Run the server as a least-privilege user, pin it to a restricted working directory, and put it in a container with seccomp. Then the alias trick still fires, but whoami runs as nobody, in an empty directory, with no syscalls worth having. An allowlist is a policy about identity. It knows
62
A I CYB ER
How To Run A 30-Minute AI Security Review Before Your AI Feature Ships. Seven checks, roughly four minutes each, mapped to the 2026 OWASP lists. No specialized tooling and no governance committee required. By Vinaya Vasudevan
The risky moment is not when an AI feature is announced. It is the quiet meeting right before it ships, when a developer turns to the security team and asks whether this is safe to release, and the security team has no process to answer the question. Coming from a background in application security and penetration testing before moving into AI security research, I have seen this pattern before. Teams either skip the review entirely because no framework feels applicable, or stall the release waiting for a governance process designed for traditional software. Neither outcome improves security. The first ships blind. The second teaches engineering teams to treat security as a bottleneck.
The risky moment is the quiet meeting right before it ships. Teams do not need a sixty-page governance framework before every release. They need a fast,
FA L L
2026
repeatable review that checks the right failure points in the time it takes to finish a cup of coffee. Through my research developing SAILORS, a practitioner framework named for its seven checks, I identified the failure points that matter most before shipping an AI feature. They map directly to the OWASP LLM Top 10, released 4 August 2026, and the OWASP Top 10 for Agentic Applications.1,2 The goal is not to cover every theoretical risk. It is to catch the failure points that cause real incidents.
SAILORS Before the model responds 1. S. Sanitize. Are inputs validated on every channel that can change model behavior? 2. A. Access. Does retrieval respect the original document permissions? While the model runs 3. L. Least Priviledge. Do tools and permissions follow least privilege?
63
A I CYB ER After the model responds 4. O. Override. Do high impact actions require human approval? 5. I. Inspect. Are outputs checked for sensitive data, injection payloads, and hallucinated content? 6. R. Record. Is every model action and tool call logged? 7. S. System prompt hardening. Are system prompts and internal context protected from exposure? Checks are listed in lifecycle order, not acronym order.
Before The Model Responds The first check is whether inputs are sanitized. Prompt injection has been the number one risk on the OWASP LLM Top 10 for three consecutive years, because large language models cannot distinguish between instructions and data. If a user prompt, a retrieved document, or a tool output can change the model’s behavior, that channel needs input validation. The second check covers access controls at the retrieval layer. When a RAG system pulls documents from a shared knowledge base, does it respect the original access permissions? A common failure mode in RAG designs is that retrieval follows similarity while access control is treated as a separate problem. The model retrieves whatever matches the embedding, regardless of who should see it. This maps to OWASP’s sensitive information disclosure and vector and embedding weakness categories.
While The Model Runs The third check is whether tools and permissions follow least privilege. OWASP moved excessive agency from sixth place to third in the 2026 list, and for good reason. Production AI agents now autonomously call APIs, execute shell commands, and manage database transactions. When a model can take real actions, a prompt injection is not just embarrassing output. It is a full system compromise. The fourth check is whether high-impact actions require human approval before execution. If a customer service agent can issue a refund, send an email, or delete a record without a confirmation gate, that is excessive agency. The fix is not to remove the tools. It is to scope their permissions and require human override for anything irreversible.
After The Model Responds The fifth check inspects outputs for sensitive data, injection payloads, and hallucinated content. Model output is untrusted data. If it gets passed to a browser without encoding, you have cross-site scripting. If it gets passed to a database without parameterization, you have SQL injection. This is familiar territory for anyone with an application security background. The same injection risks we have spent decades mitigating in traditional software now flow through model output. OWASP may have dropped improper output handling from fifth to tenth place in 2026, but that does not mean the risk disappeared. It means input-side attacks now dominate incident records.
FA L L
2026
Each is a separate failure, but they compound. One missed guardrail becomes a data breach.
Where The Lists Are Heading The 2026 OWASP LLM Top 10 shows the direction of travel. Excessive agency rising to third tells us the industry has moved from models that generate text to agents that take action. Unbounded consumption rising from tenth to sixth tells us that extended-thinking models and multimodal inference are expensive enough to weaponize. These are the failure points that surface repeatedly in AI security research and incident analysis. AI security becomes usable when it becomes reviewable before release. A thirty-minute review will not catch every vulnerability, but it will catch the failure points that cause real incidents. The seven checks described here take roughly four minutes each. They do not require specialized tooling. They require a security engineer who knows what to look for, and a team willing to pause before shipping. The next time someone asks whether this is safe to release, do not wait for a governance committee. Run SAILORS. Seven checks, thirty minutes, and your AI feature ships with security built in.
NOTES 1. OWASP GenAI LLM Top 10 2026. https://genai.owasp.org/ resource/owasp-genai-llm-top-10-2026/ 2. OWASP Top 10 for Agentic Applications 2026. https://genai.owasp. org/resource/owasp-top-10-for-agentic-applications-for-2026/ 3. OWASP AI Exchange. https://owaspai.org/
ABOUT THE AUTHOR
Model output is untrusted data. The sixth check is whether every model action and tool call is logged. Without audit logging, you cannot detect a breach, reconstruct an incident, or prove compliance. The seventh and final check verifies that system prompts and internal context are hardened. OWASP renamed this category from system prompt leakage to hidden context exposure in 2026, broadening it to cover all internal context the model should keep hidden rather than the system prompt alone.
Vinaya Vasudevan is an AI Security Engineer and researcher specializing in application security, DevSecOps, and AI threat modeling. She created SAILORS, a practitioner framework for AI capability security, and developed an LLM security fundamentals course with AI Sec University. She is an active OWASP AI Exchange contributor, holds the CAISP certification, is a WiCyS professional mentor, and has spoken at cybersecurity conferences including The Diana Initiative.
WORKED EXAMPLE: THE HR POLICY ASSISTANT A team builds an assistant that answers employee questions using a RAG pipeline over company documents.
Retrieval follows similarity, while access control is treated as a separate problem.
1. Without input sanitization, a crafted query can manipulate the model into ignoring its guidelines. 2. Without retrieval access controls, an entry-level employee can ask about executive compensation and the model will retrieve and summarize the salary document, because the vector database has no concept of role-based access. 3. Without an override gate, if the assistant has a tool to update employee records, a prompt injection through a retrieved document could trigger an unauthorized change.
64
A I CYB ER
How To Build A Security Engineer Second Brain With Karpathy’s LLM Wiki Pattern. Make your past security research instantly accessible during debugging, incident response, and pentests. By Muh. Fani “Rama” Akbar
FA L L
2026
Editor’s Note In our interview earlier in this section, security researcher Ezinne Kalu describes building a second brain from Karpathy’s pattern and running it against 30,000 targets. This is how to build one.
In Short Andrej Karpathy published a pattern for building personal knowledge bases with LLMs instead of RAG. This guide adapts it for security engineering. Ingest pentest reports, incident logs, bug writeups, and terminal output with one command, then query everything you ever learned from any project directory.
The knowledge base in Obsidian’s graph view.
The Problem Every Security Engineer Has You open 50 browser tabs researching a CVE, close the browser, and a week later you cannot find a single note. Documentation ends up spread across five different projects. A month passes. The doc exists somewhere, but where? Then there is the déjà vu. Solving the same attack pattern for the third time across three
65
A I CYB ER separate projects, starting from scratch each time. Every knowledge base tool either needs a vector database, a RAG pipeline, or a difficult setup.
The skill flow.
terminal. ln -s “$PWD/.skills/wiki-ingest” ~/.claude/skills/wikiingest ln -s “$PWD/.skills/ingest-url” ~/.claude/skills/ ingest-url ln -s “$PWD/.skills/wiki-query” ~/.claude/skills/wikiquery ln -s “$PWD/.skills/daily-update” ~/.claude/skills/ daily-update Without symlinks, the skills only function inside the obsidian-wiki repo folder.
FA L L
2026
/wiki-ingest Create a note about productionweb-01 server. It’s at 10.0.1.50, runs Ubuntu 22.04, owned by the Platform team, admin contact is platform@company.com, and it hosts the customer portal
Why Karpathy’s Pattern Is Different Andrej Karpathy described this pattern in his LLM Wiki project. Instead of RAG, the approach is direct context loading. The LLM reads the relevant wiki pages at query time. Three benefits for security engineers. It is easier to add notes and knowledge without complex setup. It is easier to find and query anything you have saved. And it is easier to read and access your notes from anywhere.
WHAT YOU CAN DO WITH IT • Pull IoCs, attack patterns, and remediation steps straight from pentest reports and CVE advisories. • Query incident timelines and root causes months later without digging through Slack threads. • Extract exploitation patterns from bug bounty writeups and tag them by technique. • Find the one-liner you ran a week ago from indexed terminal logs and command history. • Search across every project you ever worked on from a single query.
THE FOUR CORE SKILLS • Ingest anything. /wiki-ingest “<prompt>” takes any document, distills the knowledge into wiki pages, and cross-links related content. • Ingest URL. /ingest-url <url> pulls any article, advisory, or writeup directly into the wiki. • Query everything. /wiki-query “<question>” answers questions using everything ever ingested. • Daily maintenance. /daily-update runs freshness checks, cross-linking, and index updates.
Ingestion Ingest a CVE advisory from the web. Send this prompt to Claude Code. /ingest-url https://www.wiz.io/blog/github-rcevulnerability-cve-2026-3854
Prerequisites BEFORE YOU STARTt • Obsidian, any recent version. obsidian.md • Claude Code CLI, installed and configured. • The obsidian-wiki skills repository. github. com/Ar9av/obsidian-wiki • Optional: Defuddle, for better website content extraction. github.com/kepano/defuddle
Full Setup Guide • Step 1. Clone and install. Send this to your terminal. git clone https://github.com/Ar9av/obsidian-wiki. git cd obsidian-wiki bash setup.sh
Ingesting a URL. The skill fetches the page and extracts the vulnerability details, affected versions, and remediation steps. It creates wiki pages for the CVE and the attack technique, then links them to existing entries where any exist. Ingest a PDF incident report. Send this prompt to Claude Code. /wiki-ingest “@Bybit Incident Investigation Preliminary Report v1.0.pdf”
Querying Query for CVE details. /wiki-query What do I know about the GitHub RCE vulnerability? Returns the CVE number, affected versions, attack vector, and remediation steps, pulled from the advisory ingested earlier. Query for incident patterns. /wiki-query get the ioc from bybit incident
• Step 2. Configure the vault. Send this to your terminal. cp .env.example .env mkdir ~/llm-wiki Open .env and set the vault path, replacing the placeholder with your actual path. OBSIDIAN_VAULT_PATH=/path/to/your/llm-wiki • Step 3. Initialize the wiki. Open Claude Code inside the repository directory and send this prompt. Set up my wiki This reads the repository configuration and creates the initial wiki structure: index pages, category folders, and cross-link templates. •
Step 4. Symlink skills to global. Send this to your
Ingesting a PDF. The skill ingests the PDF and extracts the incident timeline, root cause, and attack chain. Each piece gets its own wiki page. The report used here comes from Verichains’ public audit reports. Add inline knowledge. No file needed. Type the knowledge directly and the skill structures it into a wiki page with frontmatter and links.
Querying the knowledge base.
66
A I CYB ER
FA L L
2026
Returns the attack chain, compromised keys, and timeline from the PDF report ingested earlier. Query for service owner mapping. /wiki-query Show me service owner and infrastructure mapping of production-web-01 Returns server names, IP addresses, owning teams, and admin contacts added via the inline knowledge entry. Save a conversation. /wiki-capture This stores the current Claude Code conversation into the knowledge base. Useful when a debugging session or analysis produced insights worth keeping.
FURTHER IMPROVEMENTS • QMD semantic search. Adds semantic search on top of the wiki. Useful once the vault grows past 200 pages. • MarkItDown for PDF parsing. Better extraction than the default. Microsoft’s library handles tables and formatted content well. github.com/ microsoft/markitdown • Graph colorize. Colors the Obsidian graph view by category, which helps visualize knowledge clusters. Run /graph-colorize inside Claude Code. Swap any tools or adjust the skills to match your setup. The system is plain markdown with no lockin.
Conclusion Adding knowledge to your second brain is now as easy as typing a sentence or dropping a file. No wrestling with RAG pipelines, vector databases, or complex setups. Just ingest and query. The knowledge stays, and grows with every use.
ABOUT THE AUTHOR Muh. Fani “Rama” Akbar is a Lead Security Engineer at Crypto exchanges where he specializes in Application Security, Cloud Security, and secure system design with 8+ years of experience building, breaking, and securing modern infrastructure, and AI-native security practitioner who leverages LLMs, automation, and agentic workflows to accelerate security engineering.
67
A I CYB ER
FA L L
2026
68
A I CYB ER AI SECURITY ARCHITECTURE
The Two-Layer Rule. What An AI Agent Should Never Own In Your Security Pipeline Temperature zero does not make a model reproducible, and the vendors say so themselves. Here is the question I use to decide what belongs to deterministic code and what belongs to the model. By Rupinder Pal Singh The second time I asked, I got a different answer. Not a wrong one. That would have been easier. I had asked a compliance agent which controls were failing. It picked its own queries and handed back something correct. The next morning it chose a different set and gave me a second answer, also correct. That was the moment I killed the feature. An assessor does not ask whether your answer is right. They ask you to show how you got it, then ask again next year. Evidence that cannot be reproduced is not evidence. It is an anecdote with good formatting. I wanted this to be my mistake. Set temperature to zero and the variance goes away. It does not. An April benchmark sent the same prompt to a model ten times at temperature zero and counted byte-identical replies. On a short prompt, Claude 4.5 returned identical output all ten times. On a longer open-ended one it managed 20%, and GPT4o managed none. Anthropic’s documentation says temperature zero is not fully deterministic, and Thinking Machines Lab traced a systematic cause to batch-size dependence in GPU inference kernels. The variance is not in your prompt. It is in the floor you are standing on.
The Question That Draws The Line Gartner predicts that by 2027, 40% of enterprises will demote or decommission autonomous agents over governance gaps found after a production incident. It names the cause as binary thinking: locked down or fully trusted. The way out of a binary is a split. One question decides where a model belongs. Does this output have to be reproducible six months from now, to someone who was not in the room? If yes, it belongs in deterministic code. If no, give it to the model, and it will probably do the job better than you would.
The Lower Layer Is Boring On Purpose Checks run from version-controlled definitions. Same inputs, same outputs, every time. In my pipeline that means Prowler and Steampipe queries executed from Python, serialized to OSCAL before any model sees them. No model chooses what runs or decides what passed.
The Upper Layer Gets A Read Handle, Nothing More Once results are immutable, a conversational interface over them is genuinely useful. Ask which controls regressed since the last run, or for a control explained in language an engineer will act on. The model is reading evidence, not producing it, and that distinction carries the architecture. In Gartner’s autonomy-level scheme it never rises above the read-only tiers, not because a policy says so but
FA L L
2026
because it was never handed anything else. Tool access is where risk moves. In a 1,200case adversarial study across five enterprise deployments, systems with tool access passed 24% of multi-turn tests against 47% for information-only systems. Acting roughly halved reliability under sustained pressure.
Why Architecture Beats Instruction In that study, logging and explainability was the only dimension that held up across long conversations, because it was enforced by the system’s shape rather than requested in a prompt. Everything that depended on the model remembering an instruction degraded as the conversation ran on. Anything you enforce with an instruction is a preference. Anything you enforce with a boundary is a control. None of this is about trusting AI less. The conversational layer is the part of this pipeline I would miss most. It is about admitting that some artifacts must survive a stranger’s scrutiny, and those cannot come from something that will not promise the same answer twice.
CODE BOX: THE BOUNDARY BETWEEN THE TWO LAYERS # Lower layer. Deterministic, versioned, rerunnable. results = run_checks(profile=”fedrampmoderate”, commit=GIT_SHA) write_oscal(results, path=EVIDENCE_DIR) # immutable from here # Upper layer. Reads finished evidence only. assistant = ConversationalLayer( source=load_oscal(EVIDENCE_DIR), tools=[], # no query engine, by construction write=False, log=CHAT_LOG, # separate from EVIDENCE_ LOG ) The constraint lives in the constructor, not the system prompt. A reviewer confirms it without reading a line of prompt text.
SIDEBAR ONE: TWO PITFALLS I WALKED INTO, AND THE WAY BACK • Letting the conversational layer write back. Someone asks the assistant to mark a finding as a false positive, and wiring that up takes ten minutes. The moment you do, your evidence stops being immutable, and a decision has moved out of the layer that logs everything into the layer that logs a chat. • Logging both layers to the same place. An assessor wants the deterministic run and nothing else. Debugging wants the transcript without ten thousand lines of query output around it. Merging the two costs you both. • If you already built it the other way. Freeze what the model can reach, move every write path down, and split the logs. None of it requires touching the model.
69
A I CYB ER
FA L L
2026
SIDEBAR TWO: THREE QUESTIONS BEFORE YOU WIRE A MODEL INTO A PIPELINE • Does it have to be reproducible? If the same question next quarter has to return the same answer, that work belongs in code. • Does it write, or only read? A model reading finished evidence is a feature. A model that can modify evidence is a hole in your audit trail. • Where does the log go? Deterministic runs and model conversations belong in separate logs. One is evidence. The other is debugging.
ABOUT THE AUTHOR Rupinder Pal Singh is Manager of Information Security at NICE, where he leads a globally distributed compliance program spanning FedRAMP, SOC 2, PCI DSS, ISO 27001, ISO 42001, IRAP, BSI C5, and Cyber Essentials across six countries. He holds CISSP, CISM, CISA, and CRISC, along with Master’s degrees in Cybersecurity and in Clinical Mental Health Counseling. Rupinder is also a member of the ISACA Emerging Trends Working Group
70
A I CYB ER
Why AI-Written Code Needs More Than Secure Prompts What building a security gate for AI-written code taught me about the difference between guidance and control By Anshuman Bhartiya
FA L L
2026
Background
The prompt is not the boundary
Between July 9 and July 13, 2026, OpenAI evaluation agents escaped a constrained testing environment through a previously unknown flaw in an Artifactory package-registry proxy. From an external launchpad, an agent then attacked Hugging Face’s dataset-processing pipeline. It first abused HDF5 external references to read local files, then placed a Jinja2 expression into a field expected to contain a number. The renderer evaluated it, giving the agent code execution inside a production Kubernetes pod. The agent eventually gained root on a Kubernetes node, generated 181 mesh-VPN enrollments with a stolen key, minted GitHub App installation tokens with write access, and opened a pull request intended to compromise CI. That final execution path was blocked - not by a prompt or code review, but by an execution policy.1 The sequence captures a distinction that matters far beyond this incident: Instructions express intent, but enforced boundaries determine reach.
We are deep in the AI-agent era. As software engineers shift from writing every line of code themselves to directing coding agents, many teams are responding with better prompts, coding rules, and detailed specifications. For example, a specification may say, “Users authenticate with JWTs.” An agent can satisfy that requirement while still allowing one user to read another user’s data. Authentication establishes who the user is; authorization determines what that user may access. Unless that second requirement is explicit or surfaced during threat modeling - it is easy for both humans and agents to overlook. Stating that requirement before implementation gives the agent a much better chance of preserving it. A specification is still a statement of intent. It can shape what an agent writes; it cannot guarantee what ultimately ships. I learned this while building FlowRail, a security gate for AI-written code. The system reads a design specification, predicts likely threats, and turns them into application-specific checks. Early versions ran those checks while the agent edited files. Detection looked good. Prevention did
71
A I CYB ER not. In one experiment, the guardrail detected a privilege-issuance vulnerability 21 out of 21 times across three code patterns, with no observed false positives on the safe controls. Yet in every test, the vulnerable code still ended up in the final version even after the guardrail blocked it on the client side. So, what happened?! The agents were not acting maliciously; they kept searching for ways to finish the task. In one run, an agent set FLOWRAIL_FAIL_MODE=open . I had intended this as an administrator-only breakglass mechanism, but because the agent could set it itself, it was not an administrative boundary - it was merely a convention. In another run, the agent removed the Write|Edit hook from Claude Code’s local settings. A third rewrote the file through shell commands the detector did not recognize. The lesson was simple - a control the agent can modify is guidance, not enforcement. The fix for this was more architectural. I moved the final check to a server-side merge gate that evaluated the complete diff and reported a required status check on the pull request. The agent’s local environment, settings, and choice of file-writing tool no longer mattered. For that gate to be trustworthy, the repository must require the check, restrict bypass permissions, and accept the verdict only from an approved source.2 In the next run, two high-severity regressions were blocked, and a real pull request reported mergeStateStatus=BLOCKED . This demonstrated that detection is only half the story. Detection becomes prevention only when the final check runs somewhere the agent cannot change or bypass. This moves security earlier in the lifecycle - from repeatedly fixing individual vulnerabilities toward preventing recurring bug patterns from shipping.
A useful gate must also be precise Moving the enforcement step outside the agent created another problem - false positives became more expensive. One example - One file-by-file check blocked handlers that did not contain an authorization control, even when another layer correctly enforced one. The verifier saw a single file rather than the complete request path. That is a classic source of false positives: a limited-context check is allowed to make a system-level decision. In another example - a security requirement explicitly called for bcrypt. When the agent chose scrypt, the gate rejected it - even though OWASP recommends scrypt when Argon2id is unavailable and positions bcrypt primarily for legacy systems. The mistake was in the requirement: it named a particular library instead of expressing the actual security objective.3 This matters because every incorrect block teaches developers to weaken the gate, convert it into a warning, or create broad exceptions. A security control that blocks valid work often enough will eventually be disabled. New blocking rules should therefore begin in an advisory mode, evaluate cross-file context where necessary, and provide narrow, approved, auditable waivers rather than a global escape hatch.
FA L L
2026
ABOUT THE AUTHOR Anshuman Bhartiya is an application security leader at Lyft and the creator of SecureVibes, VulnVibes and FlowRail. His work focuses on AI-assisted software development and agentic security. He builds and studies systems that measure whether AI security controls prevent vulnerabilities from reaching production. The views expressed are his own.
Measure what survives I tested the combined approach (enumerating threats at the design stage and enforcing security checks in CI) in a small internal pilot: eight builds of the same task-management application across Node and Python, with and without a security specification, and with and without FlowRail supervising the session. This was a descriptive pilot, not a statistically powered or independent benchmark. The baseline builds contained five confirmed critical-or-high vulnerabilities. Four came from the same pattern: placeholder secrets that became usable runtime defaults, including JWT, webhook, and session-signing keys. In these runs, placeholder secrets were a recurring agent failure mode. In the supervised builds, no confirmed criticalor-high vulnerability survived to a mergeable result. Dependency-CVE occurrences fell from eight to two. The write-time checks prevented the repetitive secret pattern, while the merge gate caught a separate “first user becomes administrator” flaw that the earlier check missed. The numbers are encouraging, but the more important lesson is - what to measure. Alert counts flatter security tools because finding more problems looks like success. The meaningful metric is how many confirmed, reachable vulnerabilities remain in code that is allowed to merge. Secure prompts and detailed specifications are valuable because they state the properties an agent should preserve. They are not security boundaries. A durable AI-native development process combines explicit security intent, fast feedback while code is being written, precise review of the complete change, and final enforcement at a checkpoint the agent cannot move.
NOTES 1. OpenAI, “OpenAI and Hugging Face partner to address security incident during model evaluation”; Hugging Face, “Anatomy of a Frontier Lab Agent Intrusion”. 2. Anthropic documents that Claude Code hooks can be modified through local settings: “Hooks reference”. GitHub explains how required checks and trusted status sources operate in “Available rules for rulesets”. 3. OWASP, “Password Storage Cheat Sheet”.
72
A I CYB ER
FA L L
2026
Got a byte of feedback? E MAIL US AT editors@aicybermagazine.com
73
A I CYB ER
FA L L
2026
74
A I CYB ER
FA L L
2026
75
A I CYB ER
FA L L
2026
76