Skip to content
☕ Buy me a coffee → Any donations go towards API costs and time spent on the website. Thank you!
Evidence & Practice · 16 min read

Why Medium Risk Is the Safe Option — and How Do We Learn to Accurately Grade Risk With Confidence?

On a domestic abuse call, medium is the grade we reach for when we are not sure — it feels safer than standard, which triggers nothing, and less exposed than high. But a grade chosen to protect the officer protects no victim. Understanding what standard, medium and high actually mean is how we learn to grade with confidence, and act on the medium-risk victims we are currently missing.

N

Nathan Tracey

Illustration for “Why Medium Risk Is the Safe Option — and How Do We Learn to Accurately Grade Risk With Confidence?”

Audio edition

≈ 16 min · narrated

On the risk assessment in front of me, in the box where the grade is supposed to read standard, medium or high, an officer has written “medium/high”. I have seen “med/high” more times than I can count, and once, in earnest, “Hedium” — a word that does not exist, doing a job the three real words apparently could not. I have written something like it myself. It feels responsible at the time. It is the sound of someone who does not want to commit, attending an incident where committing is the whole point.

We almost never grade a domestic standard. On the front line the working scale runs from medium to high, with standard reserved for a thin band of cases nobody loses sleep over, and the result is a measure that has quietly lost its bottom. When the lowest grade you will actually use is medium, medium stops meaning there are identifiable indicators of risk and starts meaning the lowest I dare put. So when a case turns up that genuinely sits above that floor but below high, there is nowhere on the scale for it to go, and out comes the slash, the “med/high”, the invented word. The fix is not a fourth grade. It is being willing to use the first one.

The fix is not a fourth grade. It is being willing to use the first one.

The problem is older than policing

Before this is a domestic abuse problem it is a language problem, and a very well-documented one. People hear the same probability word and picture wildly different odds.

The clean origin of the idea sits in the CIA. In 1951 a national intelligence estimate warned of a “serious possibility” of a Soviet attack on Yugoslavia. The analyst who drafted it, Sherman Kent, had meant something like a 65% chance. When he asked colleagues what they had understood by the same phrase, the answers ran from around 20% to 80%. One form of words; the whole spread of the dice. Kent spent years afterwards arguing that estimative language should be pinned to numbers, precisely so that the reader and the writer were talking about the same thing.

It is not a Cold War curiosity. In 2018 Andrew and Michael Mauboussin ran the experiment at scale for the Harvard Business Review, asking roughly 1,700 people to put a number on everyday probability words. “Likely” drew answers spanning about 55% to 90%. “Real possibility” — the kind of phrase that would sit comfortably in any officer’s rationale — spanned roughly 20% to 80%, which is to say it meant “probably not” to some readers and “more likely than not” to others. This is the territory a recent episode of Radio 4’s More or Less went into, with the epidemiologist Adam Kucharski setting listeners a quiz on exactly these words. Most of us are confident we know what “likely” means. Most of us are wrong about what the person next to us thinks it means.

Medicine learned this the hard way. European regulators decided that a drug side effect occurring in 1 to 10% of patients should be labelled “common”. When researchers tested how patients read that word, the average person took “common” to mean around a 50% chance — off by a factor of ten or more, in a leaflet meant to inform consent. The academic work behind all of this, going back to Wallsten and Budescu in 1986, points to one finding that should make every risk-assessor uneasy: the disagreement is worst for the middle terms. We broadly agree on “certain” and “impossible”. We fall apart in the range between, which is exactly where medium lives.

What standard, medium and high are supposed to mean

Here is the part that should change the conversation: these grades are not a matter of feeling. They are defined, and the definitions are not new. They come down to policing through the Offender Assessment System used by probation, carried into domestic abuse work via the SafeLives DASH checklist and now sitting behind the College of Policing’s DARA, the Domestic Abuse Risk Assessment that frontline officers complete at the scene.

The three risk grades and their definitions. Standard: current evidence does not indicate likelihood of causing serious harm. Medium: there are identifiable indicators of risk of serious harm; the offender has the potential to cause serious harm but is unlikely to do so unless there is a change in circumstances. High: there are identifiable indicators of risk of serious harm; the potential event could happen at any time and the impact would be serious.
The definitions are not new and they are not vague. The difference between medium and high is not how worried you feel — it is whether serious harm could happen at any time, or only if circumstances change.

Read them slowly, because almost nobody does. Standard means current evidence does not indicate likelihood of causing serious harm. Medium means there are identifiable indicators of risk of serious harm; the offender has the potential to cause serious harm but is unlikely to do so unless there is a change in circumstances. High means there are identifiable indicators of risk of serious harm; the potential event could happen at any time and the impact would be serious.

Notice what the line between medium and high actually turns on. It is not the strength of your unease. Both medium and high accept that the indicators of serious harm are there. The hinge is timing and trigger: high is could happen at any time; medium is unlikely unless something changes. That is a real, answerable question about a real case — is there a trigger on the horizon, a court date, a pregnancy, a separation, a release from custody — and it is a far better question than the one we usually ask ourselves, which is some version of “how bad does this feel and how exposed am I if I get it wrong?”

Why we will not say standard

We will not say standard because standard feels like doing nothing, and doing nothing about a domestic feels like the start of a Domestic Homicide Review with your name in it. The grade has become a confession of how seriously you took the call rather than a description of the risk, and so it drifts upward. Nobody was ever criticised in the immediate moment for grading too high. The cost of over-grading is paid later, by somebody else, and diffusely — so we do not feel it, and we keep reaching for the higher word.

But the cost is real, and it is the victims I am most concerned about who pay it. A scale that only runs from medium to high has two settings, and a two-setting scale cannot do the one thing a risk grade exists for: to show where the risk actually lies, so we can aim our effort at reducing it. If nearly everything is at least medium, then medium triggers nothing beyond the routine, because the system downstream learns to read it as the norm — the noise floor. The genuine medium, the case with identifiable indicators and a foreseeable trigger that the definition was written for, then gets the same safeguarding as every other medium logged that day, most of which should in truth have been standard. We have spent all our discrimination at the bottom of the scale and have none left in the middle, where it counts most.

Turn it around. If we were willing to grade a case standard when the evidence genuinely does not indicate a likelihood of serious harm — and to mean it, accepting that standard is a decision to do no further safeguarding — then medium would immediately weigh more. It would sit one clear step above “nothing further”, which is what it is meant to be, and that step would be a reason to act: a safeguarding referral, a follow-up, a marker, a conversation with the offender management or the multi-agency side. Using standard honestly is what gives medium its teeth. The grade we are most afraid of is the one that would make the others work.

The uncomfortable part: the grade is a weak predictor too

I am not going to pretend the answer is simply “grade more people standard and you have fixed it”, because the evidence will not let me. The largest European study of police domestic abuse risk assessment, by Turner, Medina-Ariza and Brown in 2019, found the DASH tool only weakly predictive of who would be harmed again. In the Greater Manchester data they examined, the great majority of victims who went on to suffer repeat violence had been graded standard or medium, not high — close to nine in ten of them. A tool that misses most of the people it is meant to catch is not a tool you want to lean on harder by simply pushing more cases down a grade.

So the argument is narrower, and I think stronger for it. The point is not to grade down to save work. The point is to make each grade mean a definite thing, so that the grade carries information instead of carrying the officer’s anxiety. That is also exactly why the College of Policing built DARA in the first place: not to add a form, but to improve the chance that two officers at the same kitchen table reach the same answer. Its pilot reported a 38% increase in the proportion of officers reaching the same risk decision as a domestic abuse expert. Consistency is the whole game. A grade only protects anybody if it means the same thing on Tuesday as it did on Monday, and the same thing in my hands as in yours.

Policing already knows how to do this

The reassuring part is that we are not the first part of the public sector to notice that probability words leak. The people who assess national security threats sorted this out years ago, and they did it in the way Sherman Kent recommended back in 1964: they pinned the words to numbers.

Two probability scales compared. On top, the words 'possible', 'real possibility' and 'likely' shown as wide, overlapping bands stretching across much of the 0 to 100 percent range, reflecting how differently people interpret them in surveys. Below, the PHIA Probability Yardstick, which fixes each term to a defined band: remote chance around 5 percent, highly unlikely 10 to 20, unlikely 25 to 35, realistic possibility 40 to under 50, likely or probably 55 to 75, highly likely 80 to 90, almost certain 95 percent or above.
Left to themselves, probability words overlap and drift. The UK intelligence community’s Probability Yardstick fixes each one to a band — and the College of Policing already points analysts to it.

The UK intelligence community uses the Probability Yardstick, drawn up by the Professional Head of Intelligence Assessment. It is not complicated. “Realistic possibility” means 40% to just under 50%. “Likely” or “probably” means 55% to 75%. “Highly likely” means 80% to 90%. When an assessment says “highly likely”, everyone in the chain reads the same odds, and an analyst can be held to them later. This is not a niche document. The College of Policing’s own guidance on delivering effective analysis points analysts to exactly this scale. The climate scientists did the same thing: in IPCC reports, “likely” is defined to mean a 66 to 100% probability, every time, by rule.

We have, in other words, already accepted inside policing that a word like “likely” needs a number behind it before it can be trusted. We simply have not carried that habit across to the kitchen table, where the stakes are a living person and the words are standard, medium and high. There is no reason a force could not say plainly what likelihood of serious harm separates standard from medium, and anchor the grade the way the Yardstick anchors an intelligence judgement. The definitions already gesture at it. We just have to be brave enough to make them say it out loud.

Grading with confidence: the National Decision Model

None of this works if the grade still comes from the gut and the definitions are something we glance at afterwards to justify it. What turns a grade from a feeling into a decision is having a method for reaching it — and policing already owns the method. It is the National Decision Model, the College’s framework for any decision an officer makes, with the Code of Ethics sitting at its centre.

The National Decision Model shown as five stages arranged in a ring around a central Code of Ethics. Stage one: gather information and intelligence. Stage two: assess threat and risk and develop a working strategy. Stage three: consider powers and policy. Stage four: identify options and contingencies. Stage five: take action and review what happened. Each stage is annotated for a domestic abuse grading decision.
The National Decision Model is built for exactly this: a defensible judgement under pressure. Run the grade through it and the answer to ‘why medium and not high?’ is written down before anyone asks.

Run a DARA grading decision around that wheel and it stops being a guess. Gather information and intelligence: the account, the history, the markers, the children, the previous calls — including the absence of things, which is evidence too. Assess threat and risk and develop a working strategy: hold the picture against the actual definition. Are there identifiable indicators of serious harm? If yes, is there a trigger that means it could happen at any time — which makes it high — or is serious harm unlikely unless circumstances change, which makes it medium? If the evidence does not indicate a likelihood of serious harm at all, that is standard, and standard is a legitimate, defensible answer, not a failure of care. Consider powers and policy: the Domestic Abuse Act, protective orders, positive action, the force’s own thresholds. Identify options and contingencies: the safeguarding the grade should trigger, and what you will do if you are wrong. Take action and review: do it, write down why, and let a supervisor test the reasoning rather than just the box.

That last point matters, because the supervisory review is where over-grading either gets challenged or gets rubber-stamped, and the inspectorate has been blunt that in too many forces the review is the latter. The value of the National Decision Model here is not bureaucratic. It is that it produces a reason. “Medium, because there are identifiable indicators but no foreseeable trigger, and here is the safeguarding I have put in place” is a sentence you can defend to a supervisor, to a court, to a review years later. “Medium/high” defends nothing. It is the absence of a decision wearing the costume of one.

The word we are missing is not a fourth one

“Hedium” is a symptom, and a revealing one. It is what a scale produces when its users have stopped trusting their own lowest setting — the linguistic residue of a hundred small acts of defensive grading, each individually reasonable, collectively corrosive. The instinct behind it is decent. No officer writes “med/high” out of laziness; they write it because they care about being right and are frightened of being wrong, which is the correct way to feel at a domestic.

But the people that instinct is meant to protect are not served by a grade that means nothing. They are served by a medium that has weight because standard exists beneath it, by a high that means now because medium means not yet, and by a rationale that can be read back and acted on. The three words we already have are enough. We just have to mean them — and the next genuine medium-risk victim, the one currently lost in the noise we made by refusing to ever say standard, is the reason it is worth the courage.

See more in safeguarding and violence against women and girls, and the More or Less Policing strand.


Sources and further reading

If you or someone you know is affected by domestic abuse, the free, 24-hour National Domestic Abuse Helpline is on 0808 2000 247. In an emergency, always call 999.

Share this article

Rate this article

domestic abuse safeguarding DARA risk assessment violence against women National Decision Model College of Policing decision-making

Related reading