A bigger synthetic panel gives you more model outputs. That's useful. It can make the aggregate less jumpy, expose more generated reasons, and support richer inspection. It doesn't summon representativeness out of thin air.

PaperPMF uses 10 generated respondents for the free preview and a separate 100 for the full report. The panels don't overlap. So when the two disagree, the answer isn't to average them into a 110-person super-panel.

The short answer

Key takeaways

  • Ten is a quick aggregate screen; 100 gives more respondent-level synthetic evidence and filterable output.
  • The preview and report are separate generations, not one expanding sample.
  • Human survey margin-of-error formulas don't transfer automatically to LLM-generated personas.

What each PaperPMF tier contains

Feature10-response preview100-response report
PanelGenerated preview panelSeparately generated full panel
Main resultOne aggregate five-point PMFAggregate PMF, mean intent, and top-box share
Individual evidenceNot shownRespondent PMFs and two reactions each
FiltersNoneAge, gender, location, occupation, and income
ThemesNoneFixed themes computed for the full panel
Best useQuick first lookAudit and deeper synthetic analysis

Sources: PaperPMF methodology

What 100 generated respondents can improve

More generations can smooth random variation in the aggregate, reveal more kinds of generated objections, and make rare response patterns easier to inspect. Respondent-level records also let you see whether a tidy mean is hiding disagreement.

Filters can recompute the displayed PMF for exact demographic fields in the full panel. The saved themes don't recompute with those filters, so don't present a filtered chart and full-panel themes as one subgroup analysis.

Sources: PaperPMF methodology

What sample size can't repair

  • A biased or stereotyped persona-generation method.
  • A loaded product concept or audience brief.
  • Missing lived experience and real budget constraints.
  • Prompt, model, anchor, or embedding choices that distort the score.
  • A claim that the generated panel represents a real population.
  • A gap between stated—or generated—intent and actual purchase.

Sources: NIM: Leaving Insight to Digital Twins?, When Synthetic Users Fail: A Cross-Domain Benchmark

Why the familiar human formulas don't slot in

Classical sampling error assumes a probability-sampling story and independent observations from a defined population. Generated personas come from a shared model, prompt, training data, and construction process. Their errors can move together.

So a synthetic N of 100 doesn't automatically earn the margin of error attached to a probability sample of 100 humans. Treat vendor confidence language carefully unless the statistical model and validation support that exact claim.

Sources: AAPOR Standard Definitions, 10th edition, AAPOR: Data Quality Metrics for Online Samples

When preview and report disagree

  1. Check that the prepared concept and audience brief are the same.
  2. Remember that the panels were generated separately.
  3. Inspect the 100 respondent records for polarization, strange personas, and repeated language.
  4. Repeat the test only under a frozen method if stability matters.
  5. Escalate the uncertainty to human research instead of cherry-picking the nicer result.

Sources: PyMC Labs: semantic-similarity-rating package, PaperPMF methodology

Pick the tier by the next decision

Use 10 when you want a quick clue and aren't ready to inspect individual evidence. Use 100 when the generated reactions, filters, and method record will change what you do next.

If the next decision is costly, neither number should be the final witness. Bring in real people or real behavior before the money gets serious.

Sources and verification

Product details are based on official documentation reviewed on August 4, 2026 unless noted. Features and pricing can change; verify them with the provider before making a purchase.

  1. PaperPMF methodology First-party methodology and report details; reviewed 2026-08-03.
  2. PyMC Labs: semantic-similarity-rating package Open-source PMF translation and aggregation implementation; reviewed 2026-08-03.
  3. NIM: Leaving Insight to Digital Twins? 2026 evidence on positivity and compressed variation; reviewed 2026-08-03.
  4. When Synthetic Users Fail: A Cross-Domain Benchmark 2026 preprint on general synthetic-user failure modes; reviewed 2026-08-03.
  5. AAPOR Standard Definitions, 10th edition Professional probability/nonprobability sample definitions; reviewed 2026-08-03.
  6. AAPOR: Data Quality Metrics for Online Samples Professional guidance on sample quality and reporting; reviewed 2026-08-03.