Model Audit
Adversarial probe suite against a model endpoint you control.
The problem
Before each release you should know: can it be jailbroken, will it leak its system prompt, does it follow injected instructions, will it write malware? Manual red-teaming does not scale.
What it does
Point it at an OpenAI- or Anthropic-compatible endpoint you own. It runs a fixed adversarial suite β jailbreak, harmful-content, prompt-injection, and secret-leak probes β evaluates each reply with transparent heuristics, and returns the modelβs own responses so you can verify. Your key is used only for the run.
Capabilities
- Jailbreak, harmful-content, injection & canary-leak probes
- Resilience score (0-100) with per-probe pass/fail + rationale
- Returns the modelβs raw replies for your own review
- OpenAI- and Anthropic-compatible; SSRF-guarded
- Runs in the browser or as an npm package / CI gate
How you run it
Part of PlayCISO Pro: enter your endpoint, key, and model in the browser and run the suite live, or install @playciso/modelaudit and fail CI if any probe regresses.
Roadmap
- Adversarial probe suite + scoring
- In-browser live runner (OpenAI/Anthropic)
- npm package + CI gate
- LLM-judge scoring + bias/drift axes
- Baseline diffing across releases
Related tools
FAQ
What is ModelAudit?
ModelAudit is a behavioural security scanner for AI models: it runs a suite of adversarial probes against a model endpoint you control and reports how the model responds to jailbreak, prompt-injection and unsafe-request attempts, producing a risk score for that deployment.
What does ModelAudit test for?
Refusal behaviour and jailbreak resistance, susceptibility to prompt injection, and unsafe or policy-violating outputs β the behavioural risks that determine whether a model is safe to expose in your application.
How is ModelAudit different from a model bill of materials?
A model bill of materials inventories what a model is (provenance, licence, integrity); ModelAudit tests how a model behaves under adversarial pressure. They are complementary β one covers supply-chain risk, the other runtime behaviour.
Model Audit runs inside PlayCISO for subscribers. The source stays private β no public repos, nothing to fork, nothing for attackers to study. Weekly, monthly, and yearly plans all include every tool.