Google, Anthropic and OpenAI Detail New Cybersecurity AI Models and Safety Controls

Google, Anthropic and OpenAI have announced new cybersecurity-focused AI capabilities, alongside restricted-access programs and additional controls intended to limit misuse.Google introduced Gemini 3....

Google, Anthropic and OpenAI have announced new cybersecurity-focused AI capabilities, alongside restricted-access programs and additional controls intended to limit misuse.

Google introduced Gemini 3.8 Flash Cyber, describing it as a model designed to help defenders identify and remediate software vulnerabilities. The company said the model is being offered through its Fairwind Program to selected organizations, including government agencies, critical-infrastructure operators, cloud customers and security partners. Google said it is working with more than 650 partners worldwide and has emphasized defensive uses such as vulnerability fixing over exploit development.

According to Google, Gemini 3.8 Flash Cyber improves autonomous vulnerability-discovery performance compared with its earlier Gemini 3.5 Flash Cyber release. The company also claimed the model performed competitively against larger systems in internal frontier-security evaluations. Those performance claims have not been independently detailed in the announcement.

Anthropic expands controlled model access

Anthropic separately announced Claude Fable 5.1 and Claude Mythos 5.1. Mythos 5.1 is being limited to trusted-access programs and certain cybersecurity and life-sciences support engagements, while Fable 5.1 can be used for some vulnerability-identification tasks. Anthropic said it may redirect higher-risk requests, including penetration testing and exploit generation, to models operating under different restrictions.

The company also unveiled Enterprise Frontier Safeguards, which combines zero-data-retention options with monitoring intended to detect misuse. Anthropic said it has added containment measures, model-behavior monitoring and a classifier intended to detect sandbox-escape attempts following unauthorized model activity involving real systems during testing.

OpenAI classifies Astra as a critical cyber capability

OpenAI said its forthcoming Astra model meets the “Critical” cybersecurity-capability threshold in its Preparedness Framework. The designation covers systems that could independently find and exploit zero-day flaws across defended targets or execute an end-to-end attack from a high-level instruction.

OpenAI said it delayed parts of Astra’s rollout while strengthening protections against cyber misuse and unauthorized agent actions. Advanced features are expected to be tested through a limited Daybreak Blue program. The company reported that Astra identified previously unknown vulnerabilities and produced exploit chains in evaluations, while also showing improved refusal rates for jailbreak attempts.

The announcements highlight an emerging approach among major AI developers: provide advanced defensive capabilities to vetted users while applying access restrictions, monitoring and safeguards to models with potentially offensive cyber utility.