We hid the ending.
Here is every scorecard, misses included.
Every vendor claims accuracy. We publish a graded exam you can rerun yourself. We took 16 famous US decisions, from the Bud Light backlash to the Reddit blackout, and handed them to Kapari with the ending stripped out. The answer key: 48 objections people actually raised at the time, each one verified in a dated article. We ran each case five times, independently. An LLM judge at temperature 0 then checked, objection by objection, whether the simulated panel had raised it. All 16 are business calls: brands, platforms and startups facing their customers and their employees. No politics on this exam.
July 27, 2026 run, on the US knowledge base as it stood at exam time (19 documented frames; the base has since grown to 33), model mistralai/mistral-large-2512. Every number on this page, tiles and table alike, comes from that single run. Coverage measures how much of the documented range of objections the simulated panel catches. It is a test score on past cases, never a measure of public opinion.
The only score on this market you can rerun yourself
The synthetic-audience market routinely claims 80 to 95% accuracy. Those numbers are self-declared. Almost none of them rest on a public, reproducible exam run on cases that were never used for calibration. A score with no published method, no held-out cases and no version stamp cannot be replayed, so it cannot be held to account. That is a sales argument, not proof.
This is not a hypothetical. In 2025 the FTC brought an enforcement action against an AI company, Workado, for advertising 98% accuracy it could not substantiate. Tested independently, the product came back around 74%. The gap between the advertised number and the real one is exactly what a public exam makes impossible to hide.
We do the opposite, and that is the whole point of this page. The method and the judge prompt are published. Every score is stamped to an engine version. The bad scorecards sit right next to the good ones. We also say what each corpus is for, because that is exactly where the usual sleight of hand hides: the 43 decisions on this page were used to tune the engine, and we publish them as a reference, not as a blind exam. So the test for any vendor fits in one question: show me your public exam, on cases you did not pick, and let me replay it. If they cannot, the number is marketing. Ours is online, and it runs on cases never used to calibrate the engine: real startup decisions, including the Kima Ventures portfolio, replayed as on day one and published case by case on the Kapari Hub.
Does the grade move from one run to the next?
No system that simulates voices turns in the same paper twice. We measure the spread and publish it instead of hiding it.
Replaying the full exam does not give back the exact same grade. Taken pass by pass across the whole exam, the five independent passes score 75%, 78%, 78%, 78% and 82% of the real criticisms. A few points of spread is normal for a system like this, and it is why we publish the average of the passes, never the best of them.
We measure stability case by case too. Across the passes run on each decision, the step of the response held on 10 of the 16 US decisions and on 34 of the 42 French ones for which a response was returned. That is lower than the figure we used to publish: the scale now has five steps instead of three outcomes, so there are mechanically more chances to change box. Measured another way, two passes taken at random land on the same step close to 9 times out of 10. The same rule applies to your own decisions: when a Kap swings between passes, we say so. Instability is information, not a flaw to paper over.
The 16 scorecards
Sorted by coverage. "Real outcome" is what the decision-maker actually did: kept, amended or withdrew the decision. A verdict is an exact match when it lines up with that outcome, close when it sits one notch more cautious, and a miss otherwise. We publish the misses with the rest.
| Decision | Year | Coverage | How the panel responded | What was left to address | Real outcome | Match |
|---|---|---|---|---|---|---|
| Burger King UKopen an International Women's Day thread with the line 'Women belong in the kitchen' | 2021 | 100% | Divided response | Objections to defuse | withdrawn | close |
| Coinbaseban internal debate on politics and social causes, and offer severance to anyone who disagrees | 2020 | 100% | Divided response | Objections to defuse | kept | close |
| Nikemake Colin Kaepernick the face of the 30th anniversary Just Do It campaign | 2018 | 100% | Divided response | Objections to defuse | kept | close |
| Twitch (Justin.tv, YC W07)tighten the rules on how streamers run sponsored and branded content | 2023 | 100% | Divided response | Objections to defuse | withdrawn | close |
| Instacart (YC S12)count customer tips toward the $10 guaranteed minimum pay | 2019 | 97% | Clear reluctance | Rework before exposing | withdrawn | exact match |
| Bud Lightpartner with trans influencer Dylan Mulvaney for a March Madness promo | 2023 | 93% | Clear reluctance | Rework before exposing | kept | miss |
| Netflixend free password sharing and charge for extra members | 2023 | 86% | Clear reluctance | Objections to defuse | kept | close |
| Robinhoodrestrict buying of GameStop and other meme stocks at the height of the short squeeze | 2021 | 86% | Clear reluctance | Rework before exposing | kept | miss |
| Xretire the Twitter name and bird logo overnight | 2023 | 86% | Clear reluctance | Rework before exposing | kept | miss |
| Pepsirelease the Kendall Jenner protest ad Live For Now | 2017 | 80% | Clear reluctance | Rework before exposing | withdrawn | exact match |
| DoorDash (YC S13)defend a pay model where customer tips subsidize the guaranteed base pay | 2019 | 67% | Clear reluctance | Rework before exposing | amended | close |
| Targetroll out a prominent front-of-store Pride Month collection nationwide | 2023 | 67% | Divided response | Objections to defuse | amended | exact match |
| Reddit (YC S05)charge for API access and shut down free third-party apps | 2023 | 57% | Clear reluctance | Objections to defuse | kept | close |
| Balenciagarun a holiday campaign showing children with harness-styled teddy bear bags | 2022 | 50% | Clear reluctance | Rework before exposing | withdrawn | exact match |
| Stripe (YC S09)lay off 14 percent of staff, with a founders' apology and a generous severance package | 2022 | 50% | Cautious support | Objections to defuse | kept | close |
| Airbnb (YC W09)lay off 25 percent of staff at the height of the pandemic, with a generous and transparent severance package | 2020 | 33% | Divided response | Objections to defuse | kept | close |
What the misses taught us
The most valuable finding of this exam is not the average. It is the pattern in the misses.
Kapari reads the reaction, not the decision-maker's nerve
Every miss tells the same story. The decision landed badly, and the company kept it anyway. Netflix's password crackdown landed exactly the way our panel said it would, with fury and cancellation threats trending. Months later it paid off, after Netflix did exactly what the bench called for: rework the plan and roll it out in stages.
Twitch vs. Netflix, side by side
Twitch's branded-content rules drew a far milder reaction than Netflix's crackdown. Twitch still backed down in 24 hours. Netflix held the line for months. The anger was not the difference. Twitch's revenue came from the people who were angry. What you do about a reaction depends on how exposed you are to it. That call is yours, and no honest tool should pretend otherwise.
The scorecard we could have buried
Airbnb's 2020 layoffs are our worst score: 33 percent coverage. The panel caught the anger and the two-tier contractor problem. It missed objections that came out of how the press told the story. We publish that grade as is. This is the one that makes the others worth trusting.
Every version retakes the exam
New knowledge base, new prompt, new model: we replay the full case set before anything ships. If the overall grade drops, the update waits. The Kapari that took this exam is the same one that runs your decision.
The method, in plain English
The case set
16 documented US decisions with a known ending: brand backlashes (Bud Light, Pepsi, Nike, Target, Balenciaga, Burger King), platform and pricing calls (Netflix, Reddit, X, Twitch), finance (Robinhood), workplace policy (Coinbase), and Y Combinator companies, including the textbook layoffs (Airbnb, Stripe) and the gig-pay reversals (DoorDash, Instacart). Every objection and every outcome comes from a dated article, one URL each.
How the blind test works
Kapari gets the decision and the context exactly as they stood on announcement day. It never gets the ending. The engine builds the panel itself, grounded in our US knowledge base (Census Bureau, BLS, Gallup and Ipsos series, all checked against primary sources). A judge model then compares the simulated range with the documented objections. We keep its reasoning readable so anyone can audit it.
What this proves, and what it does not
A good score shows one thing: the panel catches most of the reactions reality produced, without having seen them. It does not mean Kapari predicts outcomes, and we publish the misses so nobody can sell it that way. You will not know what is going to happen. You will see what you had not thought of.
The artifacts (case set, run results, judge outputs) are versioned in our repository. Reactions are simulated by an AI from synthetic profiles. No real person is interviewed. This page reports an engineering benchmark, not market research.