Platform
Evaluate risk before AI acts
Score how consequential a proposed action is by what it touches and whether it can be undone, then classify it from low through to critical before it runs.
Four dimensions
- What does it touch?35%
- How much does it move?25%
- Can it be undone?30%
- How many subjects?10%
Four, and no more. Each extra dimension makes a score marginally more accurate and materially harder to argue with.
Before anything runs
A classification is only useful if you can disagree with it.
Every risk score here comes with the inputs that produced it. That is the difference between a control an owning team maintains and a number they learn to ignore.
Risk evaluation
- What does it touch?
- Razorpay
- Contribution: high
- How much does it move?
- 74% of the ceiling
- Contribution: high
- Can it be undone?
- No — it has left the building
- Contribution: high
- How many subjects?
- One subject
- Contribution: low
Held for a named co-signer. Expires rather than proceeding.
Reversibility is the one teams forget.
It carries the second-largest weight on purpose. At the same amount, against the same system, an action that can be undone within the hour is a different proposition from one that has left the building — and the second is the one that turns an error into an incident.
The refund above scored 0.85 and classified critical. It cleared, because high risk is not a veto: it means the rule governing it had to be evaluated clause by clause against a specific amount, a specific order and a named accountable owner.
- What does it touch?35%
- A write to a system that holds money, a ledger or a customer record is a different act from a read. The target is the single strongest signal and it is knowable before anything runs.
- How much does it move?25%
- Scored against the bound, not in absolute terms. A refund at 90% of its ceiling is near the edge of what anyone agreed to; the same amount under a larger ceiling is not.
- Can it be undone?30%
- The input teams forget and usually the one that decides the outcome. An action that can be reversed within the hour is a different proposition from one that has left the building.
- How many subjects?10%
- One order is one problem. A dataset export is every row in it, and the difference is not a matter of degree.
What each band causes
A classification is a decision about how much scrutiny, not about whether.
Four bands, and what matters is the right-hand column: each one causes something specific and mechanical to happen. A band that does not change behaviour is a label.
- Risk: Low
Reads, lookups and reversible changes inside a bound nobody disputes.
Authorized and recorded. No person is asked.
- Risk: Medium
A write that matters but is reversible, or an amount well inside its ceiling.
Authorized, recorded, and surfaced in the daily review.
- Risk: High
Money moves, or a system of record changes irreversibly. Authorized only on a clean pass.
Every clause evaluated individually. Sealed either way.
- Risk: Critical
Irreversible, at scale, or above the ceiling its owner set. A person decides, or nothing happens.
Held for a named co-signer. Expires rather than proceeding.
Three things risk is not.
- Not autonomy
- Autonomy is what an agent may do at all; risk is how consequential this one action is. A level-1 agent can request a critical action and cannot commit one. Two axes, and collapsing them produces uniform lockdown.
- Not a veto
- A high classification does not stop an action. It determines how carefully the rule is evaluated and whether a person has to co-sign. Most high-risk actions here are authorized.
- Not a model output
- Scored as a pure function of the action, so the same inputs give the same band every time — which is what makes it explainable afterwards and replayable against a proposed change.
The estate, classified
Every action the estate attempts, with its band.
Sorted by score. The useful reading is the spread: an estate where everything is critical has not scored anything, and one where nothing is has not looked.
Scored actions
9 action types| Action | System | Amount | Score | Classification |
|---|---|---|---|---|
| post.journal_entry | Zoho Books | ₹3,43,472 | 0.84 | Risk: Critical |
| initiate.payout | RazorpayX | ₹87,159 | 0.77 | Risk: Critical |
| issue.refund | Razorpay | ₹6,517 | 0.73 | Risk: High |
| adjust.stock | Postgres · inventory | — | 0.40 | Risk: Medium |
| close.ticket | Zendesk | — | 0.31 | Risk: Medium |
| update.record | Salesforce | — | 0.31 | Risk: Medium |
| read.contract | Box · legal | — | 0.21 | Risk: Low |
| send.email | Gmail · customer | — | 0.21 | Risk: Low |
| create.vendor | NetSuite | — | 0.21 | Risk: Low |
Low
3
authorized on a clean pass
Medium
3
authorized on a clean pass
High
1
authorized on a clean pass
Critical
2
held for a co-signer
IllustrativeScored from the OpsAI sample estate by the same function the product uses, not hand-assigned for the page. Follow one through its decision.
Scored before the rule runs, not after the action does.
Risk is evaluated on the proposed action, which is what makes it useful. A score computed after execution is an incident report; the same score computed a millisecond earlier is a control. Everything downstream — how many clauses are checked, whether a person co-signs, how long a hold lasts — reads from it.
What critical causes
Approvals
Held for a named co-signer, with a clock. It expires rather than proceeding unattended.
What is evaluated
Policies
The rule the score decides how carefully to check, owned by the team carrying the consequence.
The scale dimension
Data access
One subject or every row in the set. Where an export is governed, and why the sink is the control point.
Where to start
Score the action you would not let a new joiner take unsupervised.
That instinct is already a risk model. Writing down which of the four dimensions produced it is usually a twenty-minute conversation, and it is the one that makes the rest of the controls proportionate.