🎉 New here? Use code WELCOME10 for 10% off any plan at checkout
All labs
Sourced incident record

Anthropic

Anthropic · San Francisco, CA, USA · founded 2021 · Claude, Claude Code, Claude for Enterprise

Anthropic's notable disclosures are largely self-published safety research and misuse reports — it has proactively documented attempts to weaponise Claude and studied jailbreak techniques, positioning disclosure as part of its safety posture.

Each entry summarises what a named outlet reported and links the source. This is commentary on published reporting, not PlayCISO’s own allegation; severity and category are our classification of the reported facts, not a legal conclusion.

Reported incidents (2)

Anthropic disclosed a largely AI-automated espionage campaign abusing ClaudeHigh
2025-11·Model misuse (disclosed)

Anthropic published a report stating it had detected and disrupted a state-linked group that used Claude (including agentic coding tooling) to automate much of a cyber-espionage operation against a set of targets. The disclosure was Anthropic's own, framed as evidence that agentic models lower the barrier to automated attacks.

Many-shot jailbreaking research (self-published)Notable
2024-04·Jailbreak research

Anthropic documented "many-shot jailbreaking", a technique that uses a long context of staged examples to erode a model's refusals, and disclosed it along with mitigations — an example of a lab publishing an attack against its own class of systems.

Anthropic — Security Incident Tracker · PlayCISO