Anthropic
Anthropic · San Francisco, CA, USA · founded 2021 · Claude, Claude Code, Claude for Enterprise
Anthropic's notable disclosures are largely self-published safety research and misuse reports — it has proactively documented attempts to weaponise Claude and studied jailbreak techniques, positioning disclosure as part of its safety posture.
Reported incidents (2)
Anthropic published a report stating it had detected and disrupted a state-linked group that used Claude (including agentic coding tooling) to automate much of a cyber-espionage operation against a set of targets. The disclosure was Anthropic's own, framed as evidence that agentic models lower the barrier to automated attacks.
Anthropic documented "many-shot jailbreaking", a technique that uses a long context of staged examples to erode a model's refusals, and disclosed it along with mitigations — an example of a lab publishing an attack against its own class of systems.