97% of security teams cannot tell if their exposures are exploitable. Are you one of them?Read the report
Filigran

What Is Automated Red Teaming?

Automated red teaming uses software to continuously simulate real-world adversary TTPs against your live environment - replacing periodic manual engagements with a repeatable, scheduled, or continuous testing cadence. This guide defines the discipline, separates it from the adjacent terms it gets confused with, and covers how to evaluate and measure a platform.

By the Filigran team13 min read

Reviewed and maintained by Filigran’s product and content teams.

TL;DR:

  • Automated red teaming uses software to continuously simulate real-world adversary TTPs against your live environment, replacing periodic manual engagements with a repeatable, scheduled, or continuous testing cadence.
  • It is a different discipline from AI/LLM red teaming (which tests AI models for jailbreaks and unsafe outputs) - this article covers enterprise security validation only.
  • Automation speeds up execution, but most tools still rely on generic scenario libraries; the real differentiator is automating relevance, which is where threat intelligence (via OpenCTI's Prioritized Intelligence Requirements) feeding OpenAEV changes the equation.
  • Automated red teaming complements manual red teaming rather than replacing it - each covers a different part of the validation problem.

"Automated red teaming" gets used interchangeably with breach and attack simulation, continuous penetration testing, and even AI red teaming. That confusion costs practitioners time before they've even started evaluating a platform.

Automated red teaming, defined

Automated red teaming uses software and automation to continuously simulate real-world adversary tactics, techniques, and procedures (TTPs) against an organization's live environment. It replaces or supplements periodic manual red team engagements with a repeatable, scheduled, or continuous testing cadence, so security teams can measure detection and response effectiveness on an ongoing basis rather than once or twice a year.

That's the whole concept in one paragraph. The rest of this article covers how it actually works, where it fits next to the red teaming you already run, and what separates a genuinely useful platform from one that just automates the clicking.

How it relates to BAS, continuous pentesting, and other adjacent terms

Four neighbouring terms account for most of the confusion, and cadence is a poor way to tell them apart, because every vendor in the category now claims to be continuous. The distinction that actually holds is narrower: what decides which tests run.

TermWhat it isWhat decides what gets tested
Penetration testing
Human-led, scoped assessment of a defined target across a fixed window.A scope document agreed before the engagement begins.
Continuous penetration testing (PTaaS)
Penetration testing delivered on subscription, with testers returning on a rolling basis instead of once a year.A human tester's judgement, applied more often. The input hasn't changed, only the frequency.
Breach and attack simulation (BAS)
Automated replay of predefined attack scenarios to check whether detection logic fires.A vendor-maintained scenario catalog, selected from by hand.
Automated red teaming
Continuous automated emulation of adversary TTPs against the live environment, covering technical controls and human process alike.Whatever feeds scenario selection: a catalog by default, threat intelligence where the platform supports it. This is the whole variable.
AI (LLM) red teaming
Adversarial testing of AI models themselves for jailbreaks, prompt injection, and unsafe output.The model's own failure modes. Not an enterprise validation discipline at all, despite the shared vocabulary.

Read down that third column and the category sorts itself. Penetration testing and PTaaS are bounded by what a human scoped in advance; BAS is bounded by a catalog someone else maintains. Automated red teaming is the only row where the input is genuinely open, which is why it is also the only row where the choice of input decides whether the results mean anything.

Automated red teaming vs. AI red teaming: don't confuse these

Automated red teaming is a term that's rarely explained clearly, so a quick clarification before going further: automated red teaming (the subject of this article) tests an organization's infrastructure, people, and processes against real-world attacker behavior. It's a security validation discipline for CTI analysts, SOC teams, and CISOs.

AI red teaming (sometimes called AI/LLM red teaming) is a different, unrelated discipline that tests AI models themselves for harmful, biased, or unsafe outputs - things like jailbreaks, prompt injection, and model safety failures. Both fields use the word "automated" loosely, which is exactly why they blend together in search results and in casual conversation. They shouldn't: different buyer, different pain point, different tooling. If you came here looking for AI model safety testing, this isn't that article - though it's worth noting that Filigran's own roadmap for OpenAEV lists AI security posture validation as an upcoming capability, which will eventually let exposure validation platforms test AI systems as an asset class too. That's a distinct, separate use case from what follows here, not a current feature.

Automated vs. manual (traditional) red teaming

Automation doesn't replace the human red teamer. It changes what you use each approach for.

DimensionAutomated red teamingManual (traditional) red teaming
Cadence
Scheduled or continuousPoint-in-time, typically annual or per-engagement
Scale
Broad, repeatable across the environmentDeep, narrow, human-led
Cost & skill profile
Lower marginal cost per run; requires platform setup and scenario curationHigher cost per engagement; requires specialized offensive talent
Best suited for
Continuous validation of known TTPs, detection coverage tracking, regression testing after control changesCreative, human-adversary-driven engagements that chain novel techniques a scenario library won't anticipate
Where it's still required
Regulatory and insurance requirements that mandate a named human tester of record; engagements needing social engineering and physical access testing

Neither one makes the other obsolete. Manual red teaming still catches what a scenario library can't imagine: creative attack chains, physical access components, and the kind of lateral thinking that comes from a skilled human trying to break something specific. Automated red teaming catches what manual engagements structurally can't: whether your defenses still hold three weeks after the last engagement, when a new tool got deployed or a detection rule got quietly disabled. Filigran's own positioning on this is explicit: automation and manual red teaming aren't rivals, they're complementary mechanisms for continuously testing and improving your defenses.

How automated red teaming works

Most automated red teaming platforms follow the same basic process, regardless of vendor:

  1. Scoping and target definition. Define which assets, asset groups, or environments are in scope for testing.
  2. Scenario and TTP selection. Choose which adversary techniques to simulate, drawn from a scenario library, MITRE ATT&CK mappings, or custom-built chains.
  3. Automated execution against the live environment. The platform runs the selected TTPs, either through its own native agent or, in platforms offering Bring-Your-Own-Agent support (available in OpenAEV's Enterprise Edition), through existing EDR/agent infrastructure already deployed on the endpoint.
  4. Detection and response measurement. The platform checks whether security controls detected, alerted on, or blocked each simulated action.
  5. Reporting and remediation guidance. Results get translated into gaps, mapped back to specific controls or detection rules that need attention.
Key Insight

Step 2 is where most vendors quietly stop innovating. Automating steps 3 through 5 is a solved problem - most platforms in this space can execute a scenario quickly and report on it cleanly. The harder, less-automated problem is step 2: picking the right scenarios in the first place.

Why threat intelligence should drive scenario selection

Most automated red teaming tools automate execution speed, not relevance. They ship with a static or semi-curated scenario library: a catalog of attack techniques you select from manually, or that gets assigned based on generic risk scoring. That confirms a control works. It doesn't confirm the control was tested against what attackers are actually doing to organizations like yours.

Filigran's own State of Threat Management report puts a number on that gap. Breach and attack simulation and penetration testing are both widely used, but both are still snapshots, and only a minority of organizations feed live intelligence into a continuous, fully automated validation process.

88%
of security leaders and practitioners agree periodic assessments can’t keep pace as systems and environments change
44%
use breach and attack simulation
41%
use penetration testing
38%
use threat intelligence within a continuous, fully automated validation process
Source: Filigran’s State of Threat Management report

That's the actual bottleneck. Most teams already run automated testing; few connect it to live intelligence to decide what to test.

OpenCTI's Prioritized Intelligence Requirements (PIRs) and STIX 2.1-structured threat data close that gap. Instead of manually mapping "which threat actors target my sector" to "which scenarios should I run," OpenCTI's PIR-to-scenario mapping (an Enterprise Edition capability) converts threat intelligence directly into OpenAEV attack scenarios, with no manual step in between. The scenarios end up reflecting the organization's current threat landscape instead of a generic technique library. For a concrete example of this pattern in practice, see how MITRE ATT&CK-mapped threat-informed defense connects threat intelligence to hunting and validation workflows using OpenCTI, OpenAEV, and Splunk ESCU.

The major commercial automated red teaming platforms don't connect scenario selection to threat intelligence this way, at least not through a standardized, STIX-based intel layer. Their automation is scenario-library automation: pre-built or AI-generated attack chains pulled from a catalog. That's a legitimate approach, but it speeds up the wrong step if your real problem is figuring out which of the 40,000+ vulnerabilities disclosed each year actually matter to your organization.

What to look for in an automated red teaming platform

Before evaluating vendors, it's worth knowing what actually separates a platform that's useful day-to-day from one that looks good in a demo. A short checklist:

Integration with existing EDR/SIEM/XDR

Bring-your-own-agent support (available in OpenAEV's Enterprise Edition) matters more than it sounds like it should. Platforms that require deploying new, dedicated agents add infrastructure overhead and a new attack surface of their own.

Threat-intelligence-driven scenarios

Ask whether the platform can build scenarios from live threat intelligence, or whether "automation" only means faster execution of a fixed catalog - not just static libraries.

Technical and human/process coverage

Infrastructure and detection testing matters, but so does testing whether your people and processes hold up under pressure. Look for platforms that also support tabletop exercises to test organizational readiness, not just technical simulation.

Depth of remediation guidance

A finding that a control failed is only half the output; the platform should point toward what to fix and how.

Openness of the underlying engine

Closed-box platforms decide what "realistic" means on your behalf, and you can't audit or extend the scenario logic. An open-source engine means you can see exactly what's being tested and modify it.

Confidence in exploitability comes from testing the scenarios that matter, which is why "threat-intelligence-driven scenario generation" belongs on this checklist as a first-order requirement, not a nice-to-have.

Attack path management - understanding not just whether an individual technique succeeds but how a chain of exposures connects to a business-critical asset - is a related capability worth asking vendors about directly. It's a natural extension of automated red teaming once you're past validating individual controls and into validating end-to-end attack feasibility.

How to measure an automated red teaming program

Most programs report volume: scenarios executed, pass rate, mean time to detect. Those numbers describe how efficiently the platform runs. None of them answer whether the scenarios were worth running, and a program can improve on all three while testing the same irrelevant catalog faster every quarter.

Key Insight

A program that executes 300 scenarios flawlessly, twelve of which correspond to actors targeting your sector, does not have a 300-scenario result. It has a twelve-scenario result and 288 scenarios of reassurance.

Four measures separate the two, and none of them require a new dashboard, only a denominator drawn from intelligence rather than from the platform's catalog.

MeasureHow to read itWhat it exposes
Threat-relevant coverage
Of the TTPs your prioritized intelligence requirements identify as relevant, the share validated in the last 90 days.A scope set by what the platform shipped with rather than by who is actually targeting you.
Relevance ratio
Of the scenarios actually executed, the share that map to threat actors or campaigns you track.High execution volume disconnected from your threat landscape: the default catalog running on schedule.
Time to first validation
Elapsed time from intelligence landing - when a campaign is newly attributed to an actor you track - to a scenario running against your environment.A manual translation step still in the loop. Measured in weeks, an analyst is hand-mapping reports into scenarios.
Re-validation drift
Of the scenarios that passed 90 days ago, the share that still pass today.Silent regressions: the detection rule someone disabled, the control that was reconfigured, the agent that quietly stopped reporting.

The first two are only computable if threat intelligence is structured well enough to serve as a denominator, which is the practical argument for keeping it in STIX 2.1 and expressing priorities as PIRs rather than as a slide. The third is the measure that moves when the manual mapping step disappears. The fourth needs nothing but a schedule, and it is the one that most often surfaces a real problem in the first month.

OpenAEV as an automated red teaming tool

OpenAEV supports automated red teaming as part of a broader exposure validation platform, not as a standalone, single-purpose product. Its Enterprise Edition ingests OpenCTI threat intelligence (via Prioritized Intelligence Requirements) to build relevant scenarios automatically and executes through existing EDR agents via Bring-Your-Own-Agent; Community Edition runs the same breach and attack simulation and tabletop exercise workflows using OpenAEV's own native agent. Either way, technical (breach and attack simulation) and human/process (tabletop exercise) validation run in the same platform.

For a detailed look at how this plays out against a traditional red team engagement - specifically why the two aren't competing approaches but complementary ones - see Traditional Red Teaming vs. OpenAEV.

See how OpenCTI's Prioritized Intelligence Requirements turn threat intelligence into ready-to-run OpenAEV scenarios with a free 30-day OpenAEV Enterprise Edition trial.

Explore OpenAEV

Frequently asked questions

Is automated red teaming a replacement for manual red teaming?

No. Automated red teaming handles continuous, repeatable validation of known TTPs at scale. Manual red teaming still catches creative, human-driven attack chains and satisfies regulatory or insurance requirements for a named human tester. Most mature security programs run both.

How often should automated red teaming run?

This depends on environment change velocity, not a fixed calendar. Organizations with frequent infrastructure or control changes benefit from continuous or weekly automated runs; the point of automation is that "how often" stops being a resourcing constraint the way it is with manual engagements.

What's the difference between automated red teaming and breach and attack simulation (BAS)?

They overlap significantly. BAS typically refers to scheduled testing of specific, predefined attack scenarios focused on technical control validation. Automated red teaming is a broader umbrella that can include BAS-style technical testing plus continuous, threat-informed scenario generation. The industry is converging these terms as platforms combine breach simulation with automated workflows, and breach and attack simulation is increasingly being folded into the broader category of adversarial exposure validation.

Do I need a dedicated red team to use automated red teaming tools?

No. That's part of the point. Automated platforms lower the skill and staffing threshold for running adversary emulation, which is why SOC and blue teams, not just dedicated offensive security staff, increasingly operate these tools directly.

Is automated red teaming the same as AI red teaming?

No. Automated red teaming validates an organization's security controls, people, and processes against real-world attacker TTPs. AI red teaming tests AI/LLM models for unsafe, biased, or manipulable outputs. They're unrelated disciplines that happen to share vocabulary.

Is FireCompass's "CART" the same thing as automated red teaming?

FireCompass markets a specific product under the name Continuous Automated Red Teaming (CART). It describes their particular implementation of the broader automated red teaming category covered in this article, not an industry-standard term.

Validate your defenses against real-world attack scenarios.

Get 30 days of full OpenAEV Enterprise Edition access: threat-informed scenario generation from OpenCTI intelligence, Bring-Your-Own-Agent execution, and AI-powered remediation guidance.