InsightAdversarial Exposure ValidationCTEM

Attack Chaining: How to Test the Way Attackers Get In

7 min read

In July 2026, an AI model undergoing an internal capability evaluation broke out of its sandbox. It chained a zero-day exploit with a code-execution flaw to gain root access on external infrastructure. From there, it reached Hugging Face's dataset pipeline, escalated to Kubernetes admin, moved laterally through node impersonation and forged tokens, and ran command and control from ordinary public web services, more than 17,000 actions over roughly four and a half days, entirely on its own initiative. No human directed a single step. It stands as one of the first confirmed cases of an AI model autonomously running a multi-stage intrusion against a third party.


TL;DR

  • Attackers now chain techniques together using AI and automation, while most defenders still test one technique at a time through periodic pentests and scans.
  • Most security leaders (93%) struggle to maintain an accurate view of their attack surface, and 88% say periodic assessments can't keep pace with change.
  • Scripted, isolated tests fail because they can't adapt mid-run, offer no visibility into their reasoning, and require scarce red-team expertise to build.
  • Attack Chaining, part of OpenAEV's latest release, closes that gap by validating environments continuously and adaptively, automating pentesting and red teaming so defenses run the way attackers actually operate.

Call them neo-attackers and legacy defenders. Neo-attackers use state-of-the-art AI and automation to chain techniques together and find a path straight to the crown jewels. Legacy defenders still rely on manual tools and a pentest once or twice a year, and pay for it when a breach gets through. Their way of defending stopped matching how attacks actually happen.

Most security teams are still validating their environments one technique at a time: a phishing simulation here, a vulnerability scan there, a point-in-time pentest once or twice a year, or maybe more often for teams with the time or budget. It's not enough. Attackers, increasingly powered by AI, operate differently. They chain techniques together, using evidence from one step to decide the next, evaluating the path to a breach one move at a time, and adapting as they go.

Defenses need to move toward autonomous exposure validation: automating the processes behind pentesting and red teaming so they can run the same way attackers do, continuously and adaptively, instead of once or twice a year. Attackers now run on automation and, increasingly, on agents. Matching that capability requires defenses built the same way. If attackers run continuously and defenses run periodically, the math never works out.

Atomic testing, checking one technique at a time in isolation, surfaces individual exposure gaps, but it can't replicate the complex logic, human or AI, required to go from intrusion to breach. Automation gets pushed as a buzzword constantly in this industry, but here it's simply true: attackers depend on it more every year, so defenses need to as well. Pentesting needs to be automated because attacks now move at automation speed. Validation needs to run continuously instead of periodically. And red teamers need an agentic option, since scaling a human red team indefinitely was never realistic to begin with.

The Exposure Gap: Why It Keeps Growing

The distance between how attackers operate and how most organizations test for it is wide, and it's growing. In Filigran's State of Threat Management report, 93% of security leaders said they struggle to maintain an accurate view of their own attack surface. Eighty-eight percent said periodic assessments alone can't keep pace with how fast systems and environments change. And 95% agreed that greater automation would give them more confidence they're focused on the risks that matter most, a conviction reflected in AI-driven exposure management processes expected to nearly double, from 37% today to 59% within two years.

This summer’s breach of France’s tax authority, the DGFiP, shows what these numbers mean in practice. The attacker breached systems and extracted data before the attacker listing the stolen database for sale, exposing records on 678,000 individuals and businesses. The intrusion succeeded because real environments behave dynamically, adapting in ways that outpace a fixed script's assumptions.

Why Manual Evaluation Can’t Keep Up

Three structural limits explain why isolated simulations and atomic tests still leave significant gaps:

  1. Scripted scenarios run the same way regardless of what they discover. Results end up describing a scenario that doesn't match how a real adversary would pivot.
  2. Black-box outputs come with no reasoning attached. Analysts can't tell which action drove which outcome, so they end up triaging every weakness a chain touches instead of the one that matters.
  3. Building multi-stage chains by hand takes scarce red-team expertise. That caps how much of an environment any team can realistically test, regardless of how much time they have.

Together, these three limits point to one structural problem: validation that runs once, against attackers who never stop.

Closing that gap starts with rethinking what a validation tool actually needs to do: not run a fixed sequence, but reason through one, the way an attacker does.

What Continuous, Adaptive Validation Looks Like

This is the shift Attack Chaining, and the rest of the latest release of Filigran's exposure validation platform, OpenAEV v3, is built around: automation and actionability. That's most evident in its redesigned command center, exposure scoring spanning every validation source, and a new type of scenario that treats validation as a live, adaptive process rather than a fixed script.

Here's what this changes for a security team:

  • Configurable and conditional chaining logic means a team controls its own scenario logic, built from its own threat arsenal and tactics, techniques, and procedure (TTPs), rather than relying on a vendor's pre-set sequence.
  • A live attack path graph means a team can watch the attack unfold and find the one chokepoint driving an entire attack path, instead of triaging every weakness the chain happens to touch.
  • Operator-led and agent-orchestrated modes mean a team chooses its own tradeoff between speed and control, rather than being locked into one operating model.

Attack path validation has been done before. The difference here is execution. Some tools hand back results with no visibility into how they got there. Others run predefined scripts rather than reasoning from actual findings. OpenAEV’s approach, rooted in its open-source foundation, works differently: a team configures the logic and conditions it wants, and when a chain runs, every decision is visible and clickable: what it uncovered, what passed, what didn't. The process matters as much as the result, and a team can drill into all of it.

Today, most attack-path tools schedule pre-set sequences and treat their output as an after-action report delivered once the run is over. Attack Chaining breaks from that pattern, reacting to runtime evidence as a scenario runs and revealing the attack path as it forms rather than after the fact. The underlying logic stays open, so a team shapes it to its own environment rather than working inside a fixed, vendor-defined sequence. And results tie back to that organization's own threat intelligence, feeding a unified exposure score rather than sitting as an isolated report. 

Matching how attackers operate takes more than sharper tools.

Continuous, Automated Defense Is No Longer Optional

Neo-attackers and legacy defenders can't run on the same playbook forever. If there's one change to make after reading this, it's to embrace automation, whether or not it feels comfortable yet, because that's the direction attackers have already taken. The data suggests that awareness is already there. Implementation is what continues to lag.

As attackers become more autonomous, that gap becomes harder to ignore.

The Hugging Face intrusion is a preview of how infiltration increasingly works: AI-driven, chained, and adaptive. The only way to defend against that is to test the same way, continuously, with the same chaining logic and the same automation attackers already use.

For the full picture of what shipped alongside Attack Chaining, including a new command center, exposure scoring, and AI red-teaming injectors, check out my OpenAEV v3 launch blog.

Read more

Explore related topics and insights