Governance that fits in a week
Metrics That Actually Bite
Accuracy won't tell you if agency is leaving the building. Four leading indicators that detect trust withdrawal while you can still do something.
Dr. Sarah Dyson·August 19, 2026·4 min read·813 words
Share this note

We're drowning in metrics that can't embarrass us. Accuracy, latency, adoption, tickets touched, "hours returned to the business." All of them can be green while a team has quietly stopped believing they're allowed to halt a bad output. Green, in that case, is a lie with a date stamp.
Ethical governance needs measures that bite: numbers that move when trust is being withdrawn, when agency is leaving the room, when the conscientious people are absorbing load they won't be able to hold. If a metric can't get you into an uncomfortable conversation by Friday, it's decoration.
I'm not against accuracy. I'm against accuracy as a proxy for "we're fine."
Four leading indicators
1. Dissent latency. Hours between a questionable output and the first challenge that made it into a shared channel. Rising latency signals status threat, exhaustion or a learned belief that challenge is punished. Track it on one system, weekly, for a month before you set a target. The number itself is less important than the direction and the story.
2. Exception ownership. Percentage of overrides that have a single human name in the log. If this is below 80%, you don't have a governance process. You have a fog. Names are how you keep the human as the moral agent after the fact. They're also a fairness signal: if the same two people own every exception, you have a bottleneck and a caste.
3. First-voice ratio. In meetings with a model artifact on the table, how often does the person with the decision right speak before the artifact is treated as the decision? This is qualitative with a simple count. It's also the cheapest way to see whether The Status Threat Nobody Named is currently running the room.
4. Review depth versus volume. Minutes of human attention per artifact, against artifacts generated. When volume rises and minutes fall, you're not becoming efficient. You're thinning the only control you actually have. Set a floor, not a ceiling. If you can't staff the floor, cut generation. That sentence belongs in the business case, as The Human Overhead of Agentic Systems argues.
What not to add
Don't add a composite "ethics score." Composites are how uncomfortable news is averaged into a pastel. Don't add a sentiment bot that asks the team how safe they feel and then reports the average to the steering group that makes them unsafe. Don't add model-confidence as a governance metric. Confidence is the system's. Governance is yours.
A note on explainability metrics: "percent of decisions with a documented rationale" is only useful if the rationale would survive Explainability Is a Leadership Skill, four sentences a person will sign. A pasted SHAP plot isn't a rationale. Count those as zero.
How to keep the numbers from becoming theater
- Owner, not a working group. One leader's name on the weekly read-out.
- Direction first, targets later. Four weeks of observation before anyone is allowed to game a goal.
- Read them next to the trust ledger. Trust Capital on the Balance Sheet is the narrative, these are the sensors. A withdrawal with no metric movement means your sensors are in the wrong place.
- Kill a vanity metric for each one you add. Capacity is finite, including the capacity to care.
Fairness After the Spreadsheet reminds you that some of the most important movement will never be a KPI: who sets frames, who polishes. Use the stretch-assignment list as a fifth, analog indicator. Write it down quarterly. If you refuse to, you already know what it would show.
A weekly fifteen minutes
You can run this without a platform. Friday, with the four-sentence ritual:
- Dissent latency this week: __ hours. Story: __
- Named exceptions: __ / __
- First voice: decision-right holder first in __ of __ meetings
- Review minutes per artifact: __ (floor is __)
If the room laughs, you have a culture finding. Laughter is often status threat wearing a joke. Write that down too.
Keep the human as the moral agent
Metrics don't govern. People do. The point of a biting metric is to put a fact under a person who can still change the system this month. If your measurement stack can't name that person, it's an exhibit, not a control.
Don't wait for an industry standard. Standards will arrive late and be easier to game than these four. Start with dissent latency. It's the one that tells you whether anyone still believes they're allowed to be a moral agent in public.
This week
Add dissent latency to one live system's review. Measure it for four weeks. Don't set a target. Set a conversation. If the number only moves when you're in the room, that's the finding.
Carry the questions that sit under the numbers with the Ethical AI Leadership Decision Toolkit. Stay in the practice with EI Leadership Insights.
Filed under
Share this note
Related in this journal
- Trust Capital on the Balance Sheet
Trust is deposited, withdrawn and called. If you can't date the last withdrawal, you're managing a mood, not a system.
- Explainability Is a Leadership Skill
Non-technical leaders don't need the weights. They need a translation they can stand behind when someone asks why.
- Fairness After the Spreadsheet
Once the confusion matrix looks clean, fairness work isn't over. It has moved into voice, opportunity and who still gets to be seen as original.
EI Leadership Insights
Biweekly notes by email. One practice.