Kapari
Understanding the market

What is a customer digital twin, and what is actually proven?

A digital twin of the customer is a simulated replica of a real customer, built by an artificial intelligence from data about that person, that you question in the customer's place.

iWritten by Kapari, a player in this market: every figure carries its conditions in the sentence, including when they limit the promise.
What this is about

The digital twin of the customer, in plain terms.

A digital twin of the customer is a simulated replica of a real customer, built by an artificial intelligence from data about that person, that you question in the customer's place. The term comes from industrial engineering: the digital copy of a piece of equipment, fed by its sensors, on which an intervention is simulated before anyone touches the real thing. Transposed to the customer, the promise is simple: question your customers at will, without re-contacting anyone, and replace repeated studies with a permanent asset. The vendors who carry it form one of the three families on the market map of AI-simulated audiences, alongside synthetic polling and the test bench for decisions. The strongest scientific result in the field is real, and it is quoted with its conditions: an agent built on a 2-hour interview with a real person recovers 85% of that same person's answers when the General Social Survey is asked again (Park et al., Stanford, 2024). Every word of that sentence matters: a two-hour interview, the same person, a questionnaire asked again. One rule protects a research budget better than any other: what has been proven was proven under precise conditions, and a score bought outside its conditions describes nothing. The rest of this page gives you what you need to check which of those conditions an offer actually meets: where the word comes from and why it reassures, what these vendors sell, what the science establishes and does not establish, the five questions to ask a vendor, and the difference between a twin and a decision panel.

Where the term comes from

From turbines to customers.

The words "digital twin" come from a field where they are earned: industrial engineering. Engineers build the digital copy of a turbine, an engine or a factory, fed continuously by sensors, to simulate an intervention before touching the real equipment. In that field, the copy is recalibrated against reality at every deviation, and the practice has proven itself. That reputation is what the vocabulary carries when it moves from the turbine to the customer: buying "a twin" is buying the idea of an exact copy, and the word imports the solidity of an engineering discipline into the sales pitch, without bringing along the measurement apparatus that made it solid. Yet what the transfer hides is precisely what separates a machine from a human being. A machine obeys stable laws and is measured by sensors; a customer changes their mind without warning, contradicts themselves from one day to the next, and answers differently depending on context, three movements no frozen copy captures. The name of a technology says nothing about its proof. The useful question is not "is this a twin?" but "what was measured, on whom, and when?"

The promise

What this family of vendors sells.

What this family of vendors sells fits in one sentence: a replica of your customers you can question at will. No more fieldwork to organize, no respondents to recruit: the twin answers the same day, as many times as you ask, and the one-off study becomes a permanent company asset. For a CMO or a head of insights, the appeal is real and deserves to be said without irony: the cost and the turnaround put research within reach of decisions that never had any, and you test as many messages as you want where you used to test one. It is the promise of the poll without the fieldwork, of opinion measurement without respondents. But a shift happens right in front of the buyer, and the buyer has to see it: the promise turns a question of method (what is the answer worth?) into a question of subscription (how many questions can I ask?). The entire value of the offer then rests on a single pillar: the replica's fidelity score to the person replicated. Where that score comes from, what it proves and what it does not prove: that is what the next two sections cover.

What is proven

What the research establishes.

The best paper in the field deserves to be stated plainly, because it is solid. Park et al. ("Generative Agent Simulations of 1,000 People," Stanford, 2024) interviewed 1,000 real people for 2 hours each, then built one AI agent per interview. The result: each agent recovers 85% of that same person's answers when the General Social Survey is asked again. That is individual fidelity, to the same person, obtained through interviews. It should be acknowledged without reservation: a faithful twin of an individual is possible when you pay the price of the interview, and it is the most solid result in the entire family. But every word of the conditions matters. Two hours of interview: not a purchase history. The same person: not a segment, not a group. A questionnaire asked again: not a business decision. Remove a single one of those conditions, and the number no longer describes anything published. When a sales pitch quotes 85%, the question to ask fits on one line: does your method also rest on two hours of interview per real customer? For the reader who wants the field's full body of evidence, the reliability question examined in depth gathers it, dated sources included.

What is not

What the research does not establish.

Between the Stanford protocol and the twin sold off the shelf, the sales pitch skips a chain of transpositions, link by link. First link, passive data: a twin built on CRM records, purchase histories or demographic labels has interviewed no one for 2 hours; the 85% Stanford measured was measured on agents built from 2-hour interviews, faithful to the same person's answers when the General Social Survey is asked again (Park et al., 2024): it belongs to that protocol, not to those twins. Second link, aggregation: as soon as you move from the individual to the group, the average can look right while the map falls apart; on a synthetic panel with correct averages, 48% of statistical relationships differ from reality, including 32% with the sign flipped (Bisbee et al., Political Analysis 2024). Third link, the gap to the field: against a real survey, a simulated panel's error runs from 4 to more than 23 points depending on the topic, and subgroup crosstabs are judged nearly unusable (Verasight reports, 2025-2026). Fourth link, flattening: models pull minorities and sharp positions back toward an average opinion (Santurkar et al., ICML 2023; Wang, Nature Machine Intelligence 2025); the profession's reference institute drew its own rule from this, and the Pew Research Center refuses silicon sampling and interviews only real people (2026). One limit of principle remains, and no technical progress will lift it: a twin answers "what would this customer say"; it does not answer "how does this announcement land on the whole ecosystem", customers, employees, partners and payers at once. None of this says a twin does not work. It says exactly what has not been proven, and where the proof you are being sold ends.

The comparison

Digital twin or decision panel?

The right dividing line is not the tool, it is the question being asked. A digital twin of the customer is the right tool when the question is: what would this segment I have genuinely documented say? It assumes fine-grained knowledge of a customer you have the means to anchor in real data, ideally through interviews, the only path for which science has validated individual fidelity (Park et al., Stanford, 2024, agents built on 2-hour interviews with the same person). A decision panel is the right tool when the question is: how does this announcement, this price increase, this reorganization land on my stakeholders? There, no one is cloned: you compose contrasting voices, customers, employees, partners, payers, and you put one specific decision in front of them; the size and composition of a panel are reasoned, they are not decreed. Two moves, two questions: clone a person and question it, or compose a panel and watch it react. Neither replaces the other, and buying one for the other's question is paying for a score out of context. Kapari's position fits in one honest sentence: Kapari does not sell individual digital twins; its test bench for decisions composes a panel of stakeholders anchored in documented data to surface the objections and fault lines before the announcement, and it is neither a poll, nor an opinion measurement, nor a prediction.

Question 1 / 5

Where does my twin's data come from: real interviews or labels?

The only proven individual fidelity in the field comes from agents built on 2 hours of interview per real person, matching 85% of that same person's answers when the General Social Survey is asked again (Park et al., Stanford, 2024); the same study observes that agents built from plain demographic descriptions come out more biased and more stereotyped. A twin fed on labels or passive data has no such anchoring. If the answer is "our proprietary data," ask which data, how old it is, who collected it from the customer directly, and whether a third party can inspect it.

Question 2 / 5

Was your accuracy score measured on my customers, or on an academic paper?

Stanford's 85% measures an agent's fidelity to the answers of the same person, interviewed for 2 hours, when the General Social Survey is asked again (Park et al., 2024): that number belongs to that protocol, not to a product that does not reproduce its method. An unverifiable in-house score reassures without informing: the FTC sanctioned Workado in 2025 for an unsubstantiated accuracy claim, with the product falling to around 74% once tested. Demand a score measured on your customers, with a protocol written before the test and published with the number.

Question 3 / 5

What happens to the twin when my customer changes their mind?

A turbine recalibrates continuously through its sensors; a customer changes their mind without telling anyone, and that is the whole difference from the industrial equipment the word was borrowed from. Ask what triggers an update of the twin, how often, with what fresh data, and what re-interviewing costs. A twin that is never refreshed answers with yesterday's opinion in today's confident tone, and it is today's customer you are deciding about.

Question 4 / 5

What limits do you put in writing?

The peer-reviewed literature documents at least three: under aggregation, 48% of the statistical relationships in a synthetic panel with correct averages differ from reality, including 32% with the sign flipped (Bisbee et al., Political Analysis 2024); against a real survey, the error runs from 4 to more than 23 points depending on the topic (Verasight reports, 2025-2026); and models flatten minorities and sharp positions (Santurkar et al., ICML 2023; Wang, Nature Machine Intelligence 2025). A vendor who writes none of these limits on screen or into the contract leaves you to discover them on your own decisions.

Question 5 / 5

Can a third party replay your method?

Without replayability, same file, same result, verifiable by someone else, any claimed accuracy remains a declaration. The polling profession's reference institute has settled the question for its own craft: the Pew Research Center refuses silicon sampling and interviews only real people (2026). Ask for the published exam, for test cases never used to calibrate the tool, and for the bad grades displayed next to the good ones. A refusal on this point is already an answer.

What next

The referee is a real dossier.

If your question is a decision to make react rather than a customer to copy, read the spirit of the digital twin applied to a whole ecosystem, the way Kapari uses it.

If the question that brings you here is an announcement to test rather than a customer to clone, the test bench for decisions can be tried for free on a real decision. Its method is backed by a public exam against reality, method published, every grade displayed, the worst included. You will judge on evidence, not on promise.

The test bench for decisions surfaces the objections and fault lines before the announcement; its panel remains simulated, it interviews no one, measures no real opinion and predicts nothing.