An NHS trust recently discovered that a performance figure submitted to NHS England was wrong. Oxford Health NHS Foundation Trust had reported that 36.1 per cent of urgent community response referrals were seen within two hours. The correct figure was 94 per cent.

That was not a rounding error, statistical noise, or a difference of interpretation. It described a substantially different reality. The incorrect figure contributed to Oxford Health being ranked 35th among 61 mental health and community providers and placed in segment three of the NHS Oversight Framework. The Trust said the correct data should have placed it in segment two. Because the league table was comparative, its error also altered the published positions of other providers.

There is an obvious lesson here about data quality. There is also a more important one about how organisations know what they think they know.

Something happens. Somebody observes it. The observation is recorded, classified, aggregated, submitted, calculated, compared, interpreted and reported. By the time the result reaches a Board, regulator, court or public authority, the original event may sit several transformations behind the proposition now being treated as fact.

A reported position is not necessarily the same thing as the reality it purports to describe.

When representation becomes reality

Oxford Health illustrates the problem neatly because we can see the stages. Operational performance was about 94 per cent. The submitted representation was 36.1 per cent. That number entered a comparative calculation and helped produce a ranking of 35th. The ranking contributed to a segment-three classification.

Had the error remained undiscovered, perfectly serious people could have discussed Oxford Health as a “segment-three Trust”. The words would have accurately described the published classification while inaccurately describing the organisation’s actual performance on the measure that helped produce it.

This distinction becomes important whenever an abstraction acquires institutional authority. A number becomes a score. A score becomes a category. A category becomes a description. Before long, the description begins to stand in for the thing itself.

We need abstractions. Large institutions could scarcely function without records, scores, categories, models and summaries. The problem begins when we forget that these are representations produced through methods, assumptions and human choices.

Precision does not cure that problem. In fact, precision can disguise it. “Thirty-five out of sixty-one” sounds authoritative. “36.1 per cent” looks exact. Yet a decimal point adds precision, not truth. A calculation may be mathematically flawless while its conclusion remains materially wrong because the input was wrong.

The Post Office: from number to character

The Horizon scandal demonstrates what happens when this process moves from organisational description to human judgement.

For years, apparent accounting shortfalls recorded through Horizon were treated as evidence against individual sub-postmasters. A system reported a discrepancy. That could become evidence that money was missing, then that an individual had caused the loss, and ultimately that the person was dishonest.

Those propositions are not equivalent.

“The system records a shortfall” is a statement about a computer record. “There is a genuine financial shortfall” is a statement about events. “This person caused it” is a statement about causation. “This person is dishonest” is a judgement about conduct and character.

Each step requires additional evidence. Yet institutional confidence in the first proposition can lend borrowed authority to all those that follow.

That is one of the dangers of reification: an abstraction ceases to be treated as a representation and starts to be treated as reality. Once that happens, people can find themselves required to disprove what a system has asserted about them.

Provenance, in circumstances like these, is not an obscure technical concern. It becomes a matter of justice.

A model is not a student

The 2020 examination grading crisis exposed the same problem in another form. With public examinations cancelled, statistical methods helped determine results. Historical performance and other evidence were used to standardise grades, producing outcomes that in some cases differed materially from teachers’ assessments.

The wider lesson was not simply that a particular model caused controversy. It was that population-level patterns and individual truth are different things.

A statistical model may estimate what tends to happen among people who share certain characteristics. An individual person is not what tends to happen to people like them.

That distinction is easy to lose when the result emerges as a crisp grade. A calculated B, C or D looks like a fact about the student. In reality, it is an output produced by evidence, assumptions, comparative data and rules.

The number or grade may be useful. It may even be the best available estimate. What it must not do is conceal what kind of knowledge it represents.

Windrush: when absence of evidence becomes evidence of absence

Windrush shows an inverse form of the same error. Some people had lawful rights to live in the United Kingdom but lacked documents later expected of them. In some cases, records confirming their status had not been retained or had never been supplied to them.

An evidential gap then acquired a meaning it could not support.

“We do not possess a record proving your status” can easily slide into “you do not possess that status”. Yet absence of a record and absence of a right are entirely different propositions.

This is why words matter as much as numbers. Terms such as “unverified”, “failed”, “non-compliant”, “illegal”, “high risk” or “ineligible” may sound merely descriptive. Each can conceal a sequence of assumptions about what has been observed, what evidence is available, what standard has been applied, and what conclusion somebody has drawn.

Once attached to a person, those words can determine employment, liberty, access to services, reputation or legal standing.

The record has become more powerful than the reality it was supposed to record.

Provenance means following the journey

Provenance is sometimes reduced to a simple question: where did this number come from?

That is not enough.

For any proposition important enough to influence a consequential decision, we should be able to follow its journey. What was originally observed? Who recorded it, when, and under what definition? Was the information final or provisional? What was absent? What was aggregated? Which assumptions entered the analysis? What calculation transformed the source data? What comparisons changed its meaning? Who interpreted the result? What description was then attached to it? Which decisions have subsequently relied upon that description?

An authentic source can still contain an error. A correct calculation can still rest upon faulty inputs. A classification can follow its rules perfectly while misrepresenting a particular case. A report can reproduce an authoritative publication accurately even when the underlying publication is wrong.

Provenance does not magically establish truth. It allows us to examine how a proposition acquired authority and whether that authority remains justified.

That is a rather more demanding discipline than simply asking whether somebody can produce a source.

AI makes an old problem faster

Artificial intelligence makes this problem more pressing, not because it invented it, but because it can compress the chain between source and conclusion.

Until recently, many transformations were visible. Information moved through spreadsheets, analysts, reports, meetings and papers. Each stage left some trace of the human process involved.

Generative AI can read dozens of documents, reconcile language, summarise disagreement and produce a polished account within seconds. Used well, this is extraordinarily valuable. Used carelessly, it can hide precisely the distinctions directors, professionals and public authorities most need to see.

Which statement came directly from a source? Which was inferred? Which source was current? Which qualification disappeared during summarisation? Which contradiction did the model resolve for itself? Which confident sentence rests on weak evidence?

Fluency makes the problem harder. Poor information once had a reasonable chance of looking poor. AI can make poor information read beautifully.

The resulting text may become more coherent while the provenance of its propositions becomes less visible.

The question is not merely whether the number is right

Boards and senior leaders do not need to become statisticians or forensic technologists. They do need to remain suspicious of the authority that polished representations acquire.

Four questions are particularly useful.

Is it accurate? In other words, does the proposition correspond sufficiently with what we are trying to understand?

Where did it come from? Can we follow the source, transformations, qualifications and interpretation?

What does it mean? Are we looking at an observation, estimate, proxy, classification, prediction, inference or professional judgement?

And what happens to people if we are wrong?

That last question restores proportion.

A wrong performance figure changes a league-table position. A computer shortfall can become an accusation. A statistical estimate can become a young person’s examination grade. An absent record can become a challenge to somebody’s right to remain in the country they call home.

Numbers matter because institutions act on them. Words matter because institutions use them to describe what the numbers mean.

People are neither.

The task is not to abandon measurement, classification or technology. It is to remember what they are: instruments through which human beings attempt to perceive a reality that remains richer, messier and more particular than the representations we construct of it.

A sound Board should never ask only, “What does the report say?”

It should also ask: What happened between reality and this report, and how justified are we in believing what now appears before us?