Enterprise AI Risk Assessment · Methodology v1.3
How we score AI risk
Most AI risk checklists count what you run. We score the gap between what your AI can do and what your controls can catch, because that gap is where incidents come from.
Why
What we believe
Using a lot of AI isn’t bad in itself. A company running 200 well-tested, monitored agents can be safer than one running three unmonitored ones. So exposure alone is never the verdict: we measure exposure, measure controls against the exposure you actually have, and report what’s left over.
A score only helps if it’s honest. We don’t tell you to add controls you already have, we hold the residual score until every question is answered so a half-finished assessment never shows a misleading verdict, and we don’t penalize low-risk programs for having few formal requirements.
Scores
Three numbers, one story
Inherent risk
The risk before any controls
Even a single AI system carries it. Set by what your AI touches, what it can do, and what happens if it fails. Higher means more at stake, not worse.
Control readiness
How well you can govern it
7 control areas, weighted by how relevant each is to your environment. Higher is better.
Residual risk
What your controls don’t cover
Inherent risk reduced by your controls. The headline number and the band you see first.
01
Inherent risk
Each of the five risk-surface questions produces a 0–100 value: the highest-risk option you select, plus a small increment for each additional selection. The five values are blended with these weights.
| Input | Question | Weight |
|---|---|---|
| Impact | If one of these AI systems failed or behaved incorrectly, what could reasonably happen? | 28% |
| Autonomy | What can your most capable AI systems do? | 22% |
| Data access | What information can these systems access or process? | 20% |
| Use cases | Where is AI being used in your organization? | 18% |
| Footprint | Approximately how many AI systems are in use or development across your organization? | 12% |
Low-risk programs score low. An organization still exploring, using public data only, generating content, with minor failure impact and one to five systems, lands around 9 out of 100.
02
Control readiness
Each control question scores the controls you consistently have, weighted by how much risk each one removes. Answers like “varies by team” or “we only have logs” cap the area, because a control that isn’t applied everywhere only partly protects you. Each area then counts for more or less depending on your exposure.
| Area | What it measures | Base weight | Counts for more when |
|---|---|---|---|
| Discovery | How completely you can see AI use across teams, vendors, and products. | 1× | Employee tools, third-party AI, or unknowns are present |
| Registration | What is consistently recorded about each AI system. | 1× | More than three use cases, or sensitive data |
| Testing | How systems are validated before deployment or major changes. | 1× | Customer-facing or agentic use, regulated data, fully autonomous systems, or severe impact |
| Observability | What you can monitor once systems are live. | 1× | Customer-facing or agentic use, or systems that communicate, modify records, or transact |
| Security & Guardrails | Which protections are enforced while systems run. | 1× | Personal data, health data or credentials, consequential actions, or severe impact |
| Cost Controls | How well you can attribute and limit AI consumption. | 0.4× | Agents, customer-facing or embedded AI, high autonomy, or 51+ systems |
| Compliance | The requirements you face and how ready your evidence is. | 1× | Personal, health, financial, or regulated data, or two or more named requirements |
Cost starts at a lower base weight: uncontrolled spend on one internal chatbot isn’t the same order of risk as uncontrolled spend on a fleet of production agents.
Compliance. Requirements awareness is 20% and evidence readiness is 80%. Naming any requirement earns full awareness credit, and so does saying none applies, for example a low-risk internal routing agent. Only “I’m not sure” scores zero. Evidence that is automatically mapped to requirements scores below evidence that takes some manual work, because evidence scattered per requirement is hard to produce when an auditor asks. Continuous tracking scores 100.
| Area score | Label |
|---|---|
| 0–24 | Limited |
| 25–49 | Developing |
| 50–74 | Established |
| 75–89 | Advanced |
| 90–100 | Comprehensive |
03
Residual risk
residual = inherent × (1 − 0.7 × readiness/100) + 0.15 × (100 − readiness)
Strong controls remove up to 70% of inherent risk, never all of it. A small baseline term means weak controls still carry some risk when exposure is modest, because AI use tends to grow. Example: inherent 70, readiness 40 gives 70 × 0.72 + 9 = 59, which is High.
- Low0–24
- Moderate25–49
- High50–74
- Critical75–100
As you answer
Why scores wait for every answer
No score shows while you answer. Inherent risk, control readiness, area scores, and residual risk appear together in your report, once all 13 questions are answered.
Readiness renormalizes over the areas answered so far, so an early composite could move from Moderate to Critical on the last few answers. We would rather show you nothing than a number that is about to change.
Findings
How findings read
A finding appears when an area guards real exposure and scores below its threshold, for example consequential agent actions with a Security & Guardrails score under 70. Each finding says which of four situations you are in, because the right advice differs.
| State | When | What we say |
|---|---|---|
| Control gap | A headline control for this risk isn’t in place | Names the missing controls, most important first |
| Right controls, low score | The headline controls are in place, but supporting ones are missing | Says so plainly, explains why the area still scores low, and recommends only what’s missing |
| Uneven coverage | The controls exist but vary by team or application | Recommends applying existing controls consistently, never re-adding them |
| Visibility gap | You weren’t sure what’s in place | Scores it as a gap: you can’t govern what you can’t see |
Rules
What we don’t do
- We don’t show any score until every question is answered, so a half-finished assessment never shows a misleading number.
- We don’t recommend controls you already have. Recommendations name only the controls you didn’t select; if you have them all but apply them unevenly, we recommend applying them consistently.
- We don’t pad the list. If no area has a material gap, we say so and suggest ways to keep controls current instead.
- We don’t treat “not sure” as neutral. Not knowing is scored as a gap.
- We don’t ask for an email to see your results. Answers stay in your browser unless you share a link.
Limits
What this is and isn’t
A self-reported estimate designed to start a useful conversation. It is not a security audit, a compliance determination, a benchmark against other companies, or legal advice. Scores reflect your answers, not an inspection of your systems.
Changes
Version 1.3
- 1.3. Findings distinguish a missing control from an area that scores low despite the right controls. No scoring weights changed.
- 1.2. Impact and footprint added to inherent risk; scores held until the assessment is complete; recommendations based on the controls you selected; revised compliance scoring.
