Kapari
The bench, taken apart

We bolted an expert AI onto our engine. Then we removed it.

The idea was appealing: once the simulation is done, a language model reads the whole file and delivers the read of a seasoned practitioner. We built it, then put it up against our arithmetic, blind, on eight decisions. It surfaced nothing the computation had not already found. We deleted the code.

iAt Kapari the numbers and the verdict are computed. The language model gives a voice to the panel profiles. It decides nothing.
What research has established since 1954

A consistent rule beats expert judgment, and this is not news.

In 1954 the psychologist Paul Meehl compared the clinical forecasts of human experts with those of simple statistical formulas. The formulas won. The result was so counterintuitive that it was replayed for fifty years. In 2000, Grove and colleagues pooled 136 studies: the conclusion held. Robyn Dawes showed along the way that a linear model with crude weights is enough to beat expert intuition.

The reason is not that the expert is ignorant. It is that the expert is inconsistent. The same file, read twice, does not prompt the same reaction. A rule never varies. Kahneman gave this invisible cost a name: noise.

A language model is an expert in Meehl's sense: brilliant, well read, and inconsistent. Two passes over the same file do not produce the same answer.

How it is built

The computation supplies. The model dresses.

In Kapari, the spread of reactions, the dispersion, the camps that form, the breaking point, the reception verdict: all of it comes out of deterministic arithmetic. No language model sits anywhere on that path.

The language model has one precise job: giving each panel profile a voice. It writes what a 52-year-old floor supervisor would say about this announcement. It adds nothing up, concludes nothing, grades nothing.

The same rule governs the signals: a family breaking ranks, a driver no one carries, opposing camps converging on a single grievance. These are computations over the run data, with fixed wording. A model writing those sentences could invent one that sounds right and is wrong.

The measurement that settled it

Eight decisions, two blocks, a jury blind to who wrote what.

We wanted to know whether an artificial expert added anything. The protocol: on eight decisions, two reading blocks, the one produced by our computations and the one produced by the expert. An automated jury, blind to the origin of each block, had to say which one surfaced material the other lacked. The adoption threshold was fixed in advance, at five cases out of eight.

Zero cases out of eight

The expert surfaced no rare material the computation had not already delivered. On six cases the two blocks were equivalent, each with its own angle.

Three slips out of eight

The expert drifted into forecasting three times, despite the guardrails built to stop it. Prediction travels through verb tense, and overflows any list of banned words.

Verdict: the code was removed. It lives on in the repository history, alongside the jury archives. This is the doctrine that has governed every new idea for a smart read ever since: a computation over the run data first, the model at most to dress it, never to find.

What this changes for you

A number that shifts from one day to the next cannot be defended in a boardroom.

The same file returns the same verdict. You can reopen a run three months later in front of your board: it will say the same thing. A tool whose output depends on a model's mood leaves you with no answer the day someone reruns it and gets something else.

No one can throw hallucination at you. That is the first objection raised against any AI-produced result, and it is often fair. Here the path to the verdict is arithmetic and can be traced line by line.

And you can check it. We sit a public exam on 16 famous US decisions, against 48 objections people actually raised, graded objection by objection. The good paper and the bad one are shown side by side.

Sources

The work this choice rests on.

Meehl, P. E. (1954), Clinical versus Statistical Prediction. Grove, W. M. et al. (2000), "Clinical versus mechanical prediction: a meta-analysis", Psychological Assessment. Dawes, R. M. (1979), "The robust beauty of improper linear models in decision making", American Psychologist. Kahneman, D., Sibony, O. and Sunstein, C. (2021), Noise.

Other work is implemented elsewhere in the engine: Gary Klein's premortem, triggered automatically when agreement runs too wide; the conditions for valid expertise from Kahneman and Klein (2009), which surface a warning when a decision has no precedent; and the mediating assessments protocol of Kahneman, Lovallo and Sibony (2019), which grades a decision on its process.

The bench, taken apart

More from this series

When everyone agrees, the engine raises a flag

A good decision can end badly. So what do you grade?