Kapari
Understanding the market

Synthetic audiences: a map of the market, its five core concepts and its three open debates

An AI-simulated audience, or synthetic audience, is a panel of profiles generated by a language model that react to a question, a product, or a decision the way people answering a survey would.

iWritten by Kapari, a player in this market: every figure carries its source in the sentence, including the ones that do not flatter us.
What this is about

The subject in four minutes.

An AI-simulated audience, or synthetic audience, is a panel of profiles generated by a language model that react to a question, a product, or a decision the way people answering a survey would. No one is actually surveyed: the answers come out of the model, grounded in real data or not. The founding idea, silicon sampling, dates to Argyle et al. (Political Analysis, 2023): a model conditioned on real profiles can approximate the response distributions of human subgroups. The market that grew out of it is young and well funded: Simile has raised $100 million, Aaru more than $50 million, and most players claim 80 to 95% accuracy, self-reported and never audited through a reproducible public exam. Peer-reviewed research measures something else: on a synthetic panel whose averages look correct, 48% of the statistical relationships differ from reality, including 32% with the sign flipped (Bisbee et al., Political Analysis 2024), and the error against a real poll runs from 4 to more than 23 points depending on the topic (Verasight reports, 2025-2026). The Pew Research Center, the reference institution in polling, rejects silicon sampling and surveys only real people (2026). The thesis of this page fits in two sentences. The market is racing toward the precise wrong number: an opinion percentage produced by a simulated panel feels reassuring and sells, but the map of opinions beneath it is wrong nearly half the time. The honest value lies elsewhere: exploring the range of plausible reactions to a decision, the objections, the fault lines, before reality answers. This page maps the subject: three families of players, five concepts, three debates, ten questions. It is written by Kapari, a player in this market that chose exploration and says so: its test bench for decisions is neither a poll, nor an opinion measurement, nor a prediction.

The market

Three families of players, a market finding its feet.

The market for AI-simulated audiences falls into three families. The first sells synthetic polling: a panel of generated respondents that promises an opinion percentage at the price and speed of a query. The second builds digital twins of customers (and, for the same gesture applied to a whole decision, the digital twin for decisions) from a company's own data, available for questioning at will to test a product or a message. The third, Kapari's, is the test bench for decisions: a simulated panel reacts to one specific decision to surface the objections and fault lines before the announcement, with no opinion number. The first two families are bought by marketing, product, and insights teams looking for speed; the third belongs in the office of the executive and their leadership team, to see the objections before announcing. The market is young and well funded: Simile has raised $100 million and Aaru more than $50 million, while most players claim 80 to 95% accuracy, self-reported and never audited through a reproducible public exam. Regulators have started sorting the field: the FTC took action against Workado in 2025 over an unsubstantiated accuracy claim, a product that fell to around 74% once tested. A market that raises this much money and proves this little is still finding itself: sellers race toward the exact number, and buyers win by demanding the method.

Concept 1 / 5

Silicon sampling (simulated respondents)

Silicon sampling is the technique that replaces the human respondents of a study with characters generated by a language model: you describe a profile to the model, then ask it the question as you would a person. The founding paper (Argyle et al., Political Analysis 2023) shows that a model conditioned on real profiles can approximate the response distributions of human subgroups. The market's entire offering descends from this idea, and from its limit: what the machine returns is what its training data says about those profiles, not what real people would answer today.

It is the word that opens both the scientific literature and the sales decks, and it is the one the polling world's reference institution rejects: the Pew Research Center refuses silicon sampling and surveys only real people (2026). Knowing where the technique comes from lets you ask a vendor what they have added, or not, to the founding paper.

Source: Argyle et al., "Out of One, Many," Political Analysis, 2023; Pew Research Center, 2026.

Concept 2 / 5

Opinion flattening (algorithmic fidelity and homogenization)

Opinion flattening is a language model's tendency to pull the voices it simulates toward an average opinion: minorities and sharp positions disappear from the picture (Santurkar et al., ICML 2023; Wang, Nature Machine Intelligence 2025). The trap goes deeper than the average: on a synthetic panel whose averages are correct, 48% of the statistical relationships differ from reality, including 32% with the sign flipped (Bisbee et al., Political Analysis 2024). A correct average can therefore dress up a false map: who thinks what, and why, slips away while the total reassures.

This is the number one buying trap, and the most counterintuitive result in the field. A buyer who validates a tool on its averages alone will see nothing, and the conclusion can flip. Algorithmic fidelity, a model's ability to reproduce the answers of a given group, is judged on the structure of the answers, never on the total.

Source: Santurkar et al., ICML 2023; Wang, Nature Machine Intelligence, 2025; Bisbee et al., Political Analysis, 2024.

Concept 3 / 5

Demographic persona vs. grounded profile

A demographic persona is a profile the model fills in from imagination based on a label ("woman, 45, executive"); a grounded profile rests on real data: interviews, published studies, sources a third party can check. Stanford measured both ends of the scale: an agent built from a 2-hour interview with a real person recovers 85% of that person's answers when the General Social Survey is asked again (Park et al., 2024, on 1,000 people), while agents built from bare demographic descriptions come out more biased and more stereotyped.

"Where do your profiles come from" is the first question to ask a vendor: it separates the tool that invents its respondents from the one that documents them. And the 85% holds for the fidelity of one individual interviewed for two hours; it does not carry over to a panel simulating the stakeholders of a decision, whom no one has interviewed.

Source: Park et al., "Generative Agent Simulations of 1,000 People," Stanford, 2024.

Concept 4 / 5

Reaction range vs. point estimate

A reaction range is an output that shows the spread of plausible reactions, the camps that form, and the objections that keep coming back, where the point estimate compresses everything into one approval percentage. The range owns up to what the single number hides: against a real poll, an LLM panel's error runs from 4 to more than 23 points depending on the topic (Verasight reports, 2025-2026).

This is the dividing line between the market's promises. A tool that returns a single number is judged as a measuring instrument, and the literature says it fails that test; a tool that returns a range is judged on the quality of its exploration. Knowing which of the two promises you are buying keeps you from paying for a measurement that is not one.Our page on refusing the single number explains why Kapari chose the range.

Source: Verasight reports, 2025-2026.

Concept 5 / 5

Computed verdict vs. generated verdict

A computed verdict comes from fixed rules applied to the panel's responses: same case, same verdict, replayable in front of a third party. A generated verdict is a text written by the language model itself, different from one run to the next. Decision science has known this divide for a long time: a simple rule applied consistently holds its own against expert case-by-case judgment, because the rule does not tire and does not change its mood (Meehl, 1954; Dawes, American Psychologist, 1979), two works presented on our page about the decision theories that power the engine.

Replayability is the only check a buyer can run alone: without it, no public exam is possible, and any claimed accuracy remains a declaration. Facing two tools, asking which one returns the same verdict twice on the same case settles the choice faster than any brochure.

Source: Meehl, "Clinical versus Statistical Prediction," 1954; Dawes, American Psychologist, 1979.

Debate 1 / 3

Explore or measure: can a simulated audience replace the poll?

The replacement camp has three serious arguments. Cost and speed, first: a simulated panel puts research within reach of decisions that never had any. The founding result, next: in Argyle et al. (Political Analysis 2023), a model conditioned on real profiles approximates the response distributions of human subgroups. The engineer's bet, finally: models improve fast, and the gap with the field will close. Capital is following that bet: Simile has raised $100 million, Aaru more than $50 million. The exploration camp answers with measurement: on a synthetic panel whose averages are correct, 48% of the statistical relationships differ from reality, including 32% with the sign flipped (Bisbee et al., Political Analysis 2024); against a real poll, the error runs from 4 to more than 23 points depending on the topic, and subgroup breakdowns are close to unusable (Verasight reports, 2025-2026). A correct average is not enough when the map of opinions beneath it is wrong. What the peer-reviewed literature says is not "this is worthless"; it is "this is not a measurement": fragile as a numerical substitute for a poll, defensible as a tool for exploring plausible reactions. The Pew Research Center drew the conclusion for its own trade: it rejects silicon sampling and surveys only real people (2026). The heart of this debate is covered in our sourced answer on the reliability of synthetic audiences; this paragraph sums it up without repeating it. Kapari's position fits in one sentence: its test bench for decisions chose exploration, it shows the range of plausible reactions before you announce, and it is neither a poll, nor an opinion measurement, nor a prediction.

Debate 2 / 3

What is a self-reported accuracy score worth?

The score camp argues from use: a buyer needs a simple benchmark to compare tools, a seriously built internal score beats no information at all, and demanding third-party audits from a newborn market would freeze innovation before it proves anything.

The method camp answers that an unverifiable number does not inform, it reassures: most players claim 80 to 95% accuracy, self-reported and never audited through a reproducible public exam. And when someone does test, the gap shows: the FTC took action against Workado in 2025 over an unsubstantiated accuracy claim, and the product fell to around 74% once tested.

What the public record says: no standard third-party audit exists in this sector, and the FTC precedent (2025) is the only public comparison point between a claimed number and a tested one. Bisbee et al. (Political Analysis 2024) give the deeper reason: a global score can look accurate while 48% of the panel's statistical relationships differ from reality. An accuracy percentage is too short an answer to a question that requires three: accurate at what, measured how, verifiable by whom.

Kapari's position: refuse the percentage contest and defend a single criterion, is the method published and replayable by a third party;it applies that criterion to itself with a public exam against reality where every grade is posted, the worst one included.

Debate 3 / 3

Does simulating human groups mean caricaturing them?

The critical camp holds the best-documented position: models crush minorities and extreme positions toward an average opinion (Santurkar et al., ICML 2023; Wang, Nature Machine Intelligence 2025), and agents built from bare demographic descriptions are more biased and more stereotyped than agents backed by interviews (Park et al., Stanford, 2024). Simulating "a 45-year-old female executive" would then amount to animating a cliché, and to letting that cliché speak for real people.

The grounding camp answers that the stereotype is a construction defect, not a fate: the same Stanford study shows that grounding in real data reduces bias across groups, and a tool can refuse to simulate individuals and work only with documented types, limits stated on screen. The problem is not simulating; the problem is simulating from nothing.

What the literature says: both camps cite the same work and draw different rules from it. The established point: without grounding in real data, the simulation of a human group drifts toward its cliché, and the drift hits minorities first. Grounding softens the defect without erasing it.

Kapari's position: take the critique seriously, every voice in its panels is grounded in published, verified sources, no voice claims to replicate a real person, and its limits are stated on screen.

Question 1 / 10

What is an AI-simulated audience?

An AI-simulated audience, or synthetic audience, is a panel of profiles generated by a language model that react to a question, a product, or a decision the way people answering a survey would. No one is actually surveyed: the answers come out of the model, grounded in real data or not. The founding concept, silicon sampling, dates to Argyle et al. (Political Analysis 2023). Three families of tools descend from it: synthetic polling, the digital twin of customers, and the test bench for decisions. The difference between them is the promise: measure an opinion, replicate a customer, or explore the reactions to a decision.

Question 2 / 10

How is a simulated panel built?

Everything starts with the profiles: either the model fills them in from imagination based on a demographic label (a demographic persona), or they are grounded in real data, interviews or published sources a third party can check. The difference is measurable: at Stanford, agents built from bare demographic descriptions come out more biased and more stereotyped than agents backed by interviews with real people (Park et al., 2024). Profile construction decides the panel's value: before comparing sizes or prices, ask the vendor where theirs come from.

Question 3 / 10

What does a synthetic audience do well?

Explore. Before an announcement, a simulated panel brings out the range of plausible reactions to a decision: the objections, the fault lines between groups, the blind spots no voice defends, everything a polite meeting leaves under the table. On that ground, the tool is useful as long as it claims neither to measure an opinion nor to predict an outcome: it prepares the decision before reality answers. That structured work differs from a free-form conversation with an assistant:our comparison of a simulated panel and a chatbot details the gap.

Question 4 / 10

What can a synthetic audience not do?

Measure. On a synthetic panel whose averages are correct, 48% of the statistical relationships differ from reality, including 32% with the sign flipped (Bisbee et al., Political Analysis 2024). Against a real poll, the error runs from 4 to more than 23 points depending on the topic, and fine-grained subgroup breakdowns are rated close to unusable (Verasight reports, 2025-2026). A percentage produced by a simulated panel is not a data point, it is a hypothesis: treating it as a measurement means deciding with a thermometer that has never been calibrated. The usage rule follows: anything that requires a defensible opinion number, market share, voting intention, satisfaction score, belongs to real people being surveyed, not to a simulated panel.

Question 5 / 10

How many voices does a simulated panel need?

Size buys nothing back: multiplying the voices of one model adds no information, flattening pulls them toward the average (Santurkar et al., ICML 2023). What matters is coverage of the camps, the lesson Tetlock draws from twenty years of measuring expert judgment: diversity of angles protects against the misstep better than the depth of a single expertise (Expert Political Judgment, 2005). Thirty or so contrasting voices, each grounded in a documented profile, cover a decision better than three experts who agree with each other;our page on panel size walks that reasoning to the end.

Question 6 / 10

How do you verify a vendor's claimed accuracy?

By demanding the method, not the number. Three questions are enough: is the exam public and replayable by a third party; were the test cases used to calibrate the tool; are the bad grades published along with the good ones. Most players claim 80 to 95% accuracy, self-reported and never audited through a reproducible public exam. The precedent exists: the FTC took action against Workado in 2025 over an unsubstantiated accuracy claim, and the tested product fell to around 74%. A vendor who refuses these three questions has answered.

Question 7 / 10

What are the red flags with a synthetic audience vendor?

Four flags: an accuracy percentage with no published method; a promise to predict a behavior or an outcome; fine-grained subgroup breakdowns sold as reliable, when the Verasight reports (2025-2026) rate them close to unusable; and no limits stated anywhere on screen. A serious vendor says what its tool cannot do; one that answers everything with the same confidence is rehearsing the claim the FTC went after at Workado (2025).

Question 8 / 10

What personal or confidential data does this kind of tool handle?

It depends on the family of tools. A panel of simulated profiles requires no personal data about your customers: the voices are generated, no one is surveyed. Digital twins of customers, on the other hand, are built from company data, which moves the question into data-protection law. In every case, the risk lives in what you feed the tool: the decision you describe, the documents you attach, the confidential context. Three questions for the vendor before signing: where are submitted texts stored, are they used to train its models, and who can read them?

Question 9 / 10

How much does a simulated audience cost compared with a traditional study?

The cost and speed gap is the market's central argument, and it is real: a simulated panel launches the same day, with no respondents to recruit and no fieldwork to organize. No standard price holds; each vendor announces its own. The right calculation is therefore sequential: simulate to prepare, spot the objections and sharpen the questions, then measure with real people when the stakes demand it. One prepares, the other settles the question.

Question 10 / 10

What is the Stanford study's 85% fidelity actually worth?

The study (Park et al., Stanford, 2024) interviewed 1,000 real people for 2 hours each and built one agent per interview; each agent recovers 85% of its own person's answers when the General Social Survey is asked again. That is individual fidelity, obtained through interviews: it does not carry over to a panel simulating the stakeholders of a decision, whom no one has interviewed. When a sales pitch cites this number, ask whether the vendor's method also rests on real interviews; if the answer is no, the number belongs to Stanford, not to the product.

What next

The referee is a real dossier.

Now that you have the map, try the instrument: run a real decision through the test bench for decisions, for free, and read the range of reactions before you announce.

This map of the market is written by Kapari, a player in that market: every number here carries its source in the sentence, its limits are stated in the same place as everyone else's, and its panel remains what it is, simulated voices that survey no one, measure no real population, and predict nothing; judge us with the questions on this page.