Asking whether synthetic respondents are reliable is like asking whether a thermometer is reliable. For body temperature? Probably. For measuring customer loyalty? Please put the thermometer down.

Reliability depends on the domain, target quantity, prompt, model, scoring method, reference data, and decision. One promising result can't bless the entire category. One ugly failure can't prove every narrow use is worthless either.

The short answer

Key takeaways

  • Judge reliability for one defined task, population, setup, and metric.
  • The SSR purchase-intent preprint is encouraging but narrow; a newer cross-domain preprint finds serious individual- and segment-level failures.
  • Use synthetic evidence for low-cost screening, not high-stakes substitution without local validation.

Reliable at what, exactly?

ClaimEvidence you'd need
Stable across rerunsRepeated runs under the same frozen setup
Matches an aggregate human distributionHeld-out human benchmark in the same domain
Gets individuals rightIndividual-level prediction against real responses
Finds real segment differencesHuman data showing the same gaps
Supports this decisionA test of downstream decision error and cost

Sources: NIM: Leaving Insight to Digital Twins?, Kantar: What is synthetic sample?

The encouraging SSR result

The 2025 SSR preprint tested 57 personal-care surveys with 9,300 human responses. It reports that free-text elicitation plus semantic scoring reached 90% of human test-retest reliability and produced realistic response distributions in that setting.

That's worth paying attention to. It's also a preprint, one domain, a particular dataset, and a particular method. It doesn't validate PaperPMF or every vendor that says “SSR” on a landing page.

Sources: Maier et al.: LLMs Reproduce Human Purchase Intent via SSR, PyMC Labs: semantic-similarity-rating package, PaperPMF methodology

The evidence that ruins the easy story

A July 2026 cross-domain preprint tested four models on U.S. social attitudes and cross-cultural values. Under the protocols studied, no model beat the strongest non-LLM baseline at individual prediction. The models also exaggerated how much demographics predicted attitudes and often pointed segment decisions in the wrong direction.

That doesn't directly test product concepts or SSR. It does kill the lazy idea that a better model plus demographic prompting automatically creates a trustworthy population.

Sources: When Synthetic Users Fail: A Cross-Domain Benchmark

Run this audit before using the result

  • Define the exact outcome: ranking, distribution, theme, individual answer, or segment gap.
  • Freeze model, prompt, persona method, scoring, and exclusions.
  • Repeat the run and inspect how much the result moves.
  • Compare against relevant held-out human data if the decision matters.
  • Check for compressed variation, excessive positivity, stereotypes, and invented subgroup differences.
  • Set a human or behavioral escalation rule before seeing the result.

Sources: NIM: Leaving Insight to Digital Twins?, Kantar: What is synthetic sample?, When Synthetic Users Fail: A Cross-Domain Benchmark

How to use PaperPMF without fooling yourself

PaperPMF exposes a versioned synthetic procedure and respondent-level evidence in its full report. That's good for inspection. It still hasn't turned generated people into sampled customers.

Use weak or mixed output to find revisions and questions. Use strong output to justify the next test, not the launch. For pricing, health, safety, large inventory, or segment targeting, bring in humans and behavior.

Sources: PaperPMF methodology

Sources and verification

Product details are based on official documentation reviewed on August 4, 2026 unless noted. Features and pricing can change; verify them with the provider before making a purchase.

  1. Maier et al.: LLMs Reproduce Human Purchase Intent via SSR Underlying SSR preprint and scoped benchmark; reviewed 2026-08-03.
  2. PyMC Labs: semantic-similarity-rating package Open-source algorithm implementation; reviewed 2026-08-03.
  3. NIM: Leaving Insight to Digital Twins? 2026 marketing-domain supporting and adverse evidence; reviewed 2026-08-03.
  4. Kantar: What is synthetic sample? Commercial-industry adverse findings; reviewed 2026-08-03.
  5. When Synthetic Users Fail: A Cross-Domain Benchmark 2026 preprint outside product purchase intent, included for general failure modes; reviewed 2026-08-03.
  6. PaperPMF methodology First-party methodology and configuration; reviewed 2026-08-03.