AI Agents Push Organizations Toward Continuous, Autonomous Red Teaming
Artificial intelligence agents are emerging as both a powerful offensive tool and a new security challenge for organizations. Unlike conventional software, these systems can take actions across connec...
Artificial intelligence agents are emerging as both a powerful offensive tool and a new security challenge for organizations. Unlike conventional software, these systems can take actions across connected environments, potentially giving attackers automated help with reconnaissance, vulnerability discovery, exploit development and the identification of sensitive data.
Matt Hartman, a former acting head of cyber at the US Cybersecurity and Infrastructure Security Agency, said organizations should treat every agent with access to sensitive systems as a privileged identity. Agents create additional integration points and non-human accounts that may be difficult to track, while static security policies can fail to reflect their changing behavior.
AI is also improving conventional criminal techniques. Highly personalized phishing, convincing impersonation and automated reconnaissance can make traditional trust indicators less reliable. Hartman pointed to phishing-resistant authentication, strong identity controls, behavioral monitoring and zero-trust practices as important defensive measures.
Red teaming at machine speed
Security leaders increasingly argue that organizations should use AI agents to test their own systems before adversaries do. Former National Security Agency cyber chief Rob Joyce has warned that companies will be assessed by attackers regardless of whether they commission an exercise themselves; the difference is who receives the findings.
This approach, often described as agentic red teaming, aims to provide continuous or frequent testing rather than the periodic, human-led assessments traditionally delivered as consulting engagements. The systems still require human oversight, training and carefully defined boundaries, particularly when testing production infrastructure or sensitive environments.
Armadin, a company founded by former Mandiant executives, and security operations provider Tenex.ai recently described a three-day exercise for an unnamed global organization. The companies said an attacker swarm performed 17 million offensive actions, identified 38 validated attack paths and generated 238 findings. Tenex.ai reported processing more than 100,000 alerts and reconstructing the activity across 231 billion raw events. The vendors estimated that a conventional five-person team would have needed roughly four months for comparable work.
Such figures come from vendors promoting this technology and should be evaluated accordingly. Still, the underlying trend is clear: as AI lowers the cost and time required to probe systems, defenders may need automated testing to maintain comparable visibility. Safe authorization, isolation, monitoring and human approval remain essential to ensure that defensive simulations do not become real incidents.
