Kapari
How it works Use cases The exam Are they accurate? The Hub Pricing Partners Demo Version française Log in Start free
Our US report card

We hid the ending.
Here is every scorecard, misses included.

Every vendor claims accuracy. We publish a graded exam you can rerun yourself. We took 16 famous US decisions, from the Bud Light backlash to the Reddit blackout, and handed them to Kapari with the ending stripped out. The answer key: 48 objections people actually raised at the time, each one verified in a dated article. We ran each case five times, independently. An LLM judge at temperature 0 then checked, objection by objection, whether the simulated panel had raised it. All 16 are business calls: brands, platforms and startups facing their customers and their employees. No politics on this exam.

The only decision tool that publishes its score, the good and the bad. The rest of the market claims 80 to 95% accuracy, self-declared and never audited. We show our work: the numbers, the method, and the cases we got wrong.
78%
average coverage of the real objections (more than three in four)
32/48
objections the panel raised explicitly (same grievance, same angle)
10/16
cases where the step of the response held across every pass
0
dangerous verdicts (never a favorable reception on a documented disaster)

July 27, 2026 run, on the US knowledge base as it stood at exam time (19 documented frames; the base has since grown to 33), model mistralai/mistral-large-2512. Every number on this page, tiles and table alike, comes from that single run. Coverage measures how much of the documented range of objections the simulated panel catches. It is a test score on past cases, never a measure of public opinion.

The only score on this market you can rerun yourself

The synthetic-audience market routinely claims 80 to 95% accuracy. Those numbers are self-declared. Almost none of them rest on a public, reproducible exam run on cases that were never used for calibration. A score with no published method, no held-out cases and no version stamp cannot be replayed, so it cannot be held to account. That is a sales argument, not proof.

This is not a hypothetical. In 2025 the FTC brought an enforcement action against an AI company, Workado, for advertising 98% accuracy it could not substantiate. Tested independently, the product came back around 74%. The gap between the advertised number and the real one is exactly what a public exam makes impossible to hide.

We do the opposite, and that is the whole point of this page. The method and the judge prompt are published. Every score is stamped to an engine version. The bad scorecards sit right next to the good ones. We also say what each corpus is for, because that is exactly where the usual sleight of hand hides: the 43 decisions on this page were used to tune the engine, and we publish them as a reference, not as a blind exam. So the test for any vendor fits in one question: show me your public exam, on cases you did not pick, and let me replay it. If they cannot, the number is marketing. Ours is online, and it runs on cases never used to calibrate the engine: real startup decisions, including the Kima Ventures portfolio, replayed as on day one and published case by case on the Kapari Hub.

Does the grade move from one run to the next?

No system that simulates voices turns in the same paper twice. We measure the spread and publish it instead of hiding it.

Replaying the full exam does not give back the exact same grade. Taken pass by pass across the whole exam, the five independent passes score 75%, 78%, 78%, 78% and 82% of the real criticisms. A few points of spread is normal for a system like this, and it is why we publish the average of the passes, never the best of them.

We measure stability case by case too. Across the passes run on each decision, the step of the response held on 10 of the 16 US decisions and on 34 of the 42 French ones for which a response was returned. That is lower than the figure we used to publish: the scale now has five steps instead of three outcomes, so there are mechanically more chances to change box. Measured another way, two passes taken at random land on the same step close to 9 times out of 10. The same rule applies to your own decisions: when a Kap swings between passes, we say so. Instability is information, not a flaw to paper over.

The 16 scorecards

Sorted by coverage. "Real outcome" is what the decision-maker actually did: kept, amended or withdrew the decision. A verdict is an exact match when it lines up with that outcome, close when it sits one notch more cautious, and a miss otherwise. We publish the misses with the rest.

DecisionYearCoverageHow the panel respondedWhat was left to addressReal outcomeMatch
Burger King UKopen an International Women's Day thread with the line 'Women belong in the kitchen' 2021 100% Divided response Objections to defuse withdrawn close
Coinbaseban internal debate on politics and social causes, and offer severance to anyone who disagrees 2020 100% Divided response Objections to defuse kept close
Nikemake Colin Kaepernick the face of the 30th anniversary Just Do It campaign 2018 100% Divided response Objections to defuse kept close
Twitch (Justin.tv, YC W07)tighten the rules on how streamers run sponsored and branded content 2023 100% Divided response Objections to defuse withdrawn close
Instacart (YC S12)count customer tips toward the $10 guaranteed minimum pay 2019 97% Clear reluctance Rework before exposing withdrawn exact match
Bud Lightpartner with trans influencer Dylan Mulvaney for a March Madness promo 2023 93% Clear reluctance Rework before exposing kept miss
Netflixend free password sharing and charge for extra members 2023 86% Clear reluctance Objections to defuse kept close
Robinhoodrestrict buying of GameStop and other meme stocks at the height of the short squeeze 2021 86% Clear reluctance Rework before exposing kept miss
Xretire the Twitter name and bird logo overnight 2023 86% Clear reluctance Rework before exposing kept miss
Pepsirelease the Kendall Jenner protest ad Live For Now 2017 80% Clear reluctance Rework before exposing withdrawn exact match
DoorDash (YC S13)defend a pay model where customer tips subsidize the guaranteed base pay 2019 67% Clear reluctance Rework before exposing amended close
Targetroll out a prominent front-of-store Pride Month collection nationwide 2023 67% Divided response Objections to defuse amended exact match
Reddit (YC S05)charge for API access and shut down free third-party apps 2023 57% Clear reluctance Objections to defuse kept close
Balenciagarun a holiday campaign showing children with harness-styled teddy bear bags 2022 50% Clear reluctance Rework before exposing withdrawn exact match
Stripe (YC S09)lay off 14 percent of staff, with a founders' apology and a generous severance package 2022 50% Cautious support Objections to defuse kept close
Airbnb (YC W09)lay off 25 percent of staff at the height of the pandemic, with a generous and transparent severance package 2020 33% Divided response Objections to defuse kept close

What the misses taught us

The most valuable finding of this exam is not the average. It is the pattern in the misses.

Kapari reads the reaction, not the decision-maker's nerve

Every miss tells the same story. The decision landed badly, and the company kept it anyway. Netflix's password crackdown landed exactly the way our panel said it would, with fury and cancellation threats trending. Months later it paid off, after Netflix did exactly what the bench called for: rework the plan and roll it out in stages.

Twitch vs. Netflix, side by side

Twitch's branded-content rules drew a far milder reaction than Netflix's crackdown. Twitch still backed down in 24 hours. Netflix held the line for months. The anger was not the difference. Twitch's revenue came from the people who were angry. What you do about a reaction depends on how exposed you are to it. That call is yours, and no honest tool should pretend otherwise.

The scorecard we could have buried

Airbnb's 2020 layoffs are our worst score: 33 percent coverage. The panel caught the anger and the two-tier contractor problem. It missed objections that came out of how the press told the story. We publish that grade as is. This is the one that makes the others worth trusting.

Every version retakes the exam

New knowledge base, new prompt, new model: we replay the full case set before anything ships. If the overall grade drops, the update waits. The Kapari that took this exam is the same one that runs your decision.

The method, in plain English

The case set

16 documented US decisions with a known ending: brand backlashes (Bud Light, Pepsi, Nike, Target, Balenciaga, Burger King), platform and pricing calls (Netflix, Reddit, X, Twitch), finance (Robinhood), workplace policy (Coinbase), and Y Combinator companies, including the textbook layoffs (Airbnb, Stripe) and the gig-pay reversals (DoorDash, Instacart). Every objection and every outcome comes from a dated article, one URL each.

How the blind test works

Kapari gets the decision and the context exactly as they stood on announcement day. It never gets the ending. The engine builds the panel itself, grounded in our US knowledge base (Census Bureau, BLS, Gallup and Ipsos series, all checked against primary sources). A judge model then compares the simulated range with the documented objections. We keep its reasoning readable so anyone can audit it.

What this proves, and what it does not

A good score shows one thing: the panel catches most of the reactions reality produced, without having seen them. It does not mean Kapari predicts outcomes, and we publish the misses so nobody can sell it that way. You will not know what is going to happen. You will see what you had not thought of.

The artifacts (case set, run results, judge outputs) are versioned in our repository. Reactions are simulated by an AI from synthetic profiles. No real person is interviewed. This page reports an engineering benchmark, not market research.

The bench, taken apart

What runs under the hood, one article at a time

Six pieces of decision research, and how each one is coded into the engine.

The verdict is computed, not generated by an AI

When everyone agrees, the engine raises a flag

A good decision can end badly. So what do you grade?

Our tool tells you when not to trust it

Why thirty voices beat three experts

Your case is less special than you think