Before Serious Misconduct: Can Police Data Identify Harmful Team Cultures Earlier?
Can complaints, use of force, driving standards and misconduct data identify harmful police team cultures earlier? Research on peer effects and Operation Hotton suggests forces should test the possibility.
Audio edition
≈ 8 min · on-device voice
Police misconduct systems are built to decide whether an individual officer breached professional standards. They are less well designed to detect a team in which poor behaviour is becoming normal. Research on peer effects, misconduct networks and the IOPC’s Operation Hotton investigations gives forces a reason to widen the unit of analysis.
The argument is not that a team with more complaints or more use of force has a bad culture. Operational exposure, recording practices and chance can explain individual differences. The stronger signal is repeated overlap: the same team appearing unusually often across several independent measures.
Complaints, driving standards, use of force and misconduct are already recorded for different organisational purposes. Bringing them together could identify a concentration of risk that remains invisible when each dataset is reviewed separately.
Culture develops close to the team
Police culture is often discussed at force or profession level. Officers experience much of it more locally. They spend repeated shifts with a relatively small group, watching how colleagues deal with the public, how supervisors respond to poor decisions, which behaviours attract approval and what happens when somebody challenges the group.
Quispe-Torreblanca and Stewart’s 2019 study provides unusually strong UK evidence for taking peer influence seriously. They analysed records for about 35,000 Metropolitan Police officers and staff between 2011 and 2014, using changes in peer groups to help distinguish influence from the tendency of similar people to work together.
They estimated that a 10% increase in peers’ prior recorded misconduct allegations increased an individual’s later recorded allegations by 8%. The authors described this as evidence of a causal peer effect. Quispe-Torreblanca and Stewart, Nature Human Behaviour.
The result comes from one force and one period, so the size of the effect should not be assumed elsewhere. It does provide evidence that patterns in recorded allegations cannot always be understood independently of the people an officer works alongside.
The College of Policing’s evidence review supporting the 2024 Code of Ethics also identifies peers, supervisors and organisational expectations as influences on ethical decision-making. College of Policing, Code of Ethics.
Operation Hotton shows what an isolated micro-culture can become
The IOPC’s Operation Hotton investigations provide a documented example from England. They identified bullying, harassment, discriminatory behaviour and failures to challenge improper conduct among officers based at Charing Cross Police Station. IOPC, Operation Hotton recommendations.
The organisational conditions are important. A permanent night-duty team policing the West End operated relatively separately from other teams, including taking briefings and breaks apart. Officers described difficult work involving intoxication, aggression and violence. The IOPC reported that the team rarely saw colleagues above sergeant rank. One inspector was responsible for six teams.
The IOPC concluded that isolation and inadequate supervision may have allowed conduct problems to become more widespread and remain unchallenged. It recommended stronger supervision, welfare arrangements and quality assurance.
By the time serious misconduct reaches a formal investigation, the organisation may be looking at behaviour that has been reinforced for months or years. Team-level analysis offers a way to look for earlier concentrations without lowering the threshold for misconduct findings.
Four datasets can ask a different question
Most forces already hold the information needed for a basic test.
A complaint is normally reviewed as a complaint. A collision or poor driving event is considered through driving standards. A use-of-force report is assessed within operational governance. Misconduct information sits with professional standards.
Each process is legitimate. The organisational question changes when Team X keeps appearing across them.
One high-use-of-force team may simply make more arrests. A team with more vehicle collisions may complete more emergency drives. A team receiving more complaints may police a difficult night-time economy. None of those counts should be treated as a proxy for integrity.
If the same team is unusual for complaints, use of force, driving standards and misconduct after reasonable adjustment for exposure, however, the force has a defensible reason to examine why. The appropriate response is review, not judgement.
The Venn diagram is a better mental model than an integrity score. The individual circles retain their different meanings. The centre asks which officers, supervisors or teams repeatedly sit in the overlap.
Look at relationships, not only totals
Research from the United States shows how misconduct data can be analysed as a network rather than a list of individuals.
Wood, Roithmayr and Papachristos reconstructed misconduct networks using 16,503 complaints involving 15,811 Chicago police officers over six years. Their work showed substantial connectivity through officers being co-named in complaints and argued that the social networks between individual and organisational explanations deserve closer attention. Wood, Roithmayr and Papachristos, Socius.
The US context is different from policing in England and Wales, so those findings should not be imported as proof of the same network structure here. The method is transferable. A force knows which officers worked together, attended the same incidents, used force together and shared supervisors. It can test whether adverse indicators cluster around recurring combinations of people.
That produces two useful questions: which officers repeatedly appear, and which officers repeatedly appear together?
Use of force and driving need fair comparisons
Raw totals will produce misleading results.
Forces in England and Wales and British Transport Police recorded 812,447 use-of-force reports in the year ending 31 March 2025. The Home Office warns that changes in recorded force partly reflect recording practice and operational exposure. Home Office, police use of force statistics.
A response team dealing with violent incidents should use more force than a low-contact function. Comparisons therefore need denominators such as arrests, incidents attended, violent incidents or officer deployment hours, depending on role and data quality. Similar teams should be compared with each other and with their own historical pattern.
Driving requires the same discipline. A single collision says little. Repeated preventable collisions, speeding or other driving-standard concerns may provide another independent signal, but only after accounting for how much and under what conditions officers drive.
There is not currently evidence establishing poor police driving as a predictor of misconduct. Its proposed value is as a separate measure of behaviour and risk-taking that can be tested against the other circles rather than assumed to be connected.
Officer assaults can be examined without treating assaulted officers as a risk
Recorded assaults on officers are potentially informative but require care. Assaults are genuine occupational harm and officers should not be discouraged from reporting them.
College of Policing material drawing on Lee Johnson’s research in Lincolnshire describes differences in how officers understand and record assaults, including the influence of seriousness, intent, administrative burden and occupational culture. College of Policing, Assaults on police: culture, legitimacy and risk.
The research does not put forward evidence that officers record assaults to protect themselves from allegations about their own use of force. Operational experience does, however, give reason to test a related question. Some officers appear quick to arrest for assaulting an emergency worker and are also quick to resort to physical force, when better communication might have controlled the encounter without anybody getting hands-on.
That observation does not establish a causal relationship. It creates a testable operational hypothesis: whether unusually high assault recording and unusually high use of force sometimes cluster around the same officers or teams, and what those incidents look like when reviewed in context.
Assault recording should therefore sit outside the four core indicators and be used as contextual analysis. A high rate should never count against a team by itself.
Follow the pattern when people move
Team membership changes create another analytical opportunity.
If complaints, force or driving indicators change after several officers join a team, analysts can examine the shift. If a pattern reduces after a supervisor leaves one team and appears after that supervisor joins another, the association can be tested rather than inferred from anecdote.
This follows the logic of the Metropolitan Police peer-effects study, which used movement between peer groups to estimate changes in behaviour.
Longitudinal analysis could help distinguish three different organisational problems: risk concentrated around one officer; risk associated with a recurring peer group; or risk associated with the supervision and working environment of a particular team. Those explanations should lead to different interventions.
Test it before building an early-warning system
A force does not need a new predictive algorithm to start.
Professional standards and analytical teams could retrospectively combine several years of team membership with complaints, driving standards, use of force and misconduct data. They could compare similar operational teams, adjust for exposure and identify whether unusual overlap occurs more often than would be expected from workload and chance.
Analysts could then examine the underlying incidents through body-worn video, complaint outcomes, supervisory records and operational context. Staff surveys or interviews could test whether statistically unusual teams also report different local cultures.
Only if that retrospective work shows useful predictive or explanatory value should a force consider prospective monitoring. A live system should flag a pattern for human review, not assign an integrity score or make a conduct finding.
The evidence supports taking peer influence and local culture seriously. What has not yet been demonstrated is whether the information UK forces already collect can reliably identify harmful micro-cultures earlier.
Forces can answer that empirical question with their own historical data. The useful output would be evidence about whether several weak signals, when they repeatedly converge on the same group, can identify organisational risk early enough for supervision and support to alter the trajectory.
Discussion questions
- 01
Which team-level indicators in your force could reveal something that individual misconduct records miss?
- 02
If the same officers repeatedly appear together across complaints, force and driving events, who should review the pattern?
- 03
Could your force tell whether a risk pattern moved when officers or supervisors changed teams?
































