AI Model Risk Assessment Questions: A Practical Checklist for Security Teams
An AI model risk assessment asks a structured set of questions across four areas: data provenance, model performance and limitations, security and adversarial exposure, and governance and accountability. The fastest way to run one is to require a documented answer to each question before a model reaches production, using a model card as the canonical record. If you can't answer a question, that gap is the risk.
The core AI model risk assessment questions
Group your questions so each maps to a decision a reviewer actually has to make. Here is a working checklist you can lift directly into a template:
- Data: Where did the training data come from? Does it contain PII, copyrighted, or regulated data? How was it labeled, and by whom? What is the known demographic or sampling bias?
- Intended use: What decision does this model inform or automate? What are the explicit out-of-scope uses? What happens on an incorrect output — is there a human in the loop?
- Performance: What benchmarks were used, and what are the accuracy, precision, recall, and false-positive rates? How does performance degrade on edge cases or underrepresented groups?
- Security: Is the model exposed to untrusted input (prompt injection, data poisoning, model extraction)? Is the training pipeline access-controlled? How are model weights and API keys protected?
- Governance: Who owns the model? Who approved deployment? How often is it re-validated, and what triggers a rollback?
Notice these mirror the fields in a model card — the documentation format introduced by Mitchell et al. (2019) at Google, now an industry standard that records a model's intended use, performance benchmarks, ethical considerations, and known limitations. If you adopt the model card as your assessment artifact, your questions and your documentation stay aligned instead of living in separate spreadsheets.
How to actually perform the assessment
Treat it as a gated review, not a one-time form. A workable sequence:
- Scope and classify (Day 1): Rate the model's impact. A model that recommends internal content is low-risk; one that denies credit or flags fraud is high-risk. This determines how deep the rest of the assessment goes — the EU AI Act uses the same risk-tiering logic.
- Collect evidence (Days 2–5): Require the model owner to fill in the model card and attach test results. No evidence, no sign-off.
- Red-team the high-risk items: For any model exposed to untrusted input, run adversarial tests — prompt injection, jailbreaks, and boundary inputs — before launch, not after.
- Set review cadence: Re-assess on retraining, on a drift threshold breach, or at a fixed interval (quarterly for high-risk). Map controls to a recognized framework like the NIST AI Risk Management Framework (AI RMF) so your assessment survives an audit.
Prioritize ruthlessly. If you only have time for five questions, ask the two that catch catastrophic failure — "What untrusted input can reach this model?" and "What decision does an incorrect output affect?" — plus the three that unblock accountability: owner, approval, and rollback trigger.
What's the "30% rule" people keep asking about?
The "30% rule" is not a formal security standard — it's a product heuristic that says automation should handle roughly the bulk of routine cases while a meaningful minority (commonly cited around 30%) is routed to human review or kept as a confidence buffer. For risk assessment purposes, the useful version of this idea is a confidence threshold: define the output confidence below which the model must defer to a human. Document that threshold in your model card, because an undocumented automation boundary is where silent failures accumulate.
Where to find templates and tooling
You don't need to build the questionnaire from scratch. Good starting points:
- NIST AI RMF Playbook — free, maps questions to specific "Govern, Map, Measure, Manage" functions.
- Model card templates on GitHub — Google's Model Card Toolkit and Hugging Face's model card spec give you a reusable, version-controlled format (far better than a static PDF that goes stale).
- OWASP Top 10 for LLM Applications — turn each listed risk into an assessment question for any generative model.
A practical tip: store your completed assessments as code or structured files in your repo, not as one-off PDFs. Version control means every answer has a diff history, which is exactly what an auditor — or your future self after an incident — will ask for.
If you want a running start, PlayCISO's free AI Model Risk Scanner walks you through these questions and flags the gaps that most often block a safe deployment.
Ready to practise the decisions these articles describe?
Run a free War Room →