Anthropic expands cyber verification tiers, eases AI safety blocks
Anthropic has restructured its Cyber Verification Program into three tiers, each granting vetted security teams progressively fewer restrictions on cybersecurity research using its models. The change appears to integrate a previously separate initiative into the broader program, which may have been intended to balance dual-use risks with defensive security needs.
The lowest tier retains conservative classifiers that block certain queries, while the highest tier removes most barriers for approved researchers. The move follows reports suggesting that Anthropic’s default safety filters—applied to models like Claude Sonnet 5.5—could be seen as overly restrictive for legitimate security work. The company’s IPO prospectus, filed last month, positioned AI as an economic revolution while also acknowledging potential risks, highlighting the challenges of aligning safety with commercial goals.
The expansion arrives as Anthropic maintains its 2026 IPO timeline despite ongoing debates over AI safety. Earlier coverage noted the company’s balancing act: framing its models as transformative tools while recognizing their possible dangers. This policy update may reflect pressure to address concerns from enterprise customers and security researchers seeking fewer guardrails.
What remains unclear is how Anthropic will define eligibility for the highest tier—or whether the changes will satisfy critics who argue that safety controls should not be tiered but standardized. Observers will watch whether the move accelerates adoption among cybersecurity firms or draws regulatory scrutiny.
Sources: siliconangle.com
“Anthropic’s shift signals growing tension between AI safety controls and commercial demand for unrestricted security testing.”
Read the original reporting
The outlets below did the original reporting.
Related briefs
This brief was drafted automatically from the sources above and published under our editorial policy. Spotted an error? Tell us.