Validation & Surveys

Score a completed CrowdProof run against a survey or shares you already have. Survey alignment and known-outcome backtests. Unscored runs stay unscored.

Why Validate?

When the crowd finishes you get a report: how the crowd reacted to your idea. A stopped or failed run has no report. Scoring is optional. When you score a finished run against real survey data, you get an alignment percentage. Unscored runs stay unscored.

When you score a finished run, it can feed the public record at /accuracy. Unscored runs stay unscored.

The Survey Lifecycle

If you already have real-world Positive, Neutral, and Negative shares, enter them on the same report page and click Compare to Simulation. You do not need to collect five CrowdProof survey responses first.

To recruit new respondents through CrowdProof, attach a public survey:

  1. Open the report page for a completed simulation
  2. Click Create Survey Link in the validation section
  3. Share the public survey URL
  4. Once 5+ responses are collected, click Score Against N Responses
  5. The alignment percentage is stored on the simulation. Score again from the report when more survey responses arrive, or use Update Real Results to replace owner-entered shares.

What the Survey Asks

The public survey asks up to four questions, each mapped to a corresponding simulated distribution.

1. Overall Reaction

"What is your overall reaction?" with three choices: Positive, Neutral, or Negative. This question is on the survey. When you score, real answers are compared against the simulation's final stance distribution (the share of agents that ended positive, neutral, or negative). Unscored runs stay unscored.

2. Closest Position

"Which of these positions is closest to yours?" with choices built from the simulation's final factions. This appears only for completed sims that produced named factions. When you score, real picks are compared against the simulated faction share distribution. Unscored runs stay unscored.

3. Purchase Intent

"How likely would you be to try it?" with three choices: Likely, Maybe, or Unlikely. When you score, real answers are compared against an engagement-weighted intent distribution derived from agent stances and activity levels. Unscored runs stay unscored.

4. Argument Resonance

"Which of these reactions sounds most like yours?" showing up to 4 of the strongest reactions from each of the camps that formed. Spare slots fill from the engagement ranking. This is a question only an emergent social simulation can ask. When you score, real picks are compared against those posts' simulated engagement shares. Unscored runs stay unscored.

Demographics

An optional age-band question collects which age group they are in. It does not change the headline alignment score. A band with enough respondents can be scored against the simulated people of the same age. A band the simulation has too few members in falls back to the whole crowd.

Open-Text Comments

An optional free-text box ("Anything else?") collects qualitative feedback. Once you have 3 or more open-text comments, the report page can cluster them into themes. Fewer comments stay unanalyzed.

How Alignment Is Scored

Each scored question compares the real distribution of answers against the corresponding simulated distribution using a distance metric.The headline alignment percentage is weighted by the number of respondents behind each scored question, so an optional question with fewer answers cannot count as much as the overall-reaction question. A score of 100% means the simulated crowd matched real answers exactly; lower scores indicate divergence.

Minimum sample size is 5 responses. Per-age-band and per-wave breakdowns require 5 responses in each segment.

Tip: Share the survey link broadly. More responses produce a more reliable alignment score. The dashboard shows collection progress so you know when you have enough.

Collection Waves

You can collect responses in multiple waves. A wave with 5 or more responses can show that wave's alignment. Wave-over-wave deltas show when two or more waves qualify.

  1. In the report page's survey section, click Start Wave N
  2. Share the same survey URL (responses are stamped with the current wave number)
  3. When you score, a wave with 5 or more responses can show that wave's alignment. Wave-over-wave deltas show when two or more waves qualify.

A wave with 5 or more responses can show that wave's alignment on the report and on published case study pages. When the crowd finishes, the Markdown export includes those scores, when any did. A stopped or failed run has no Markdown.

Comment Analysis

Once you have 3+ open-text comments, the report page offers an "Analyze N Comments" button. This calls an LLM to cluster comments into themes and flags each theme as either:

  • Anticipated: the simulated crowd raised similar points
  • Missed: the simulated crowd did not surface this theme

The report and a published case study open that list with how much of what people said the crowd anticipated, weighted by mentions. A theme one person raised does not count the same as one half the room raised. One comment can touch more than one theme.

When comments span multiple collection waves, each theme is also tagged as "new" (only appeared in the latest wave) or "recurring" (raised across waves).

Where Scores Appear

When you score a finished run, that alignment can show in these places. Unscored runs stay unscored.

  • Report page: headline alignment. Question-level scores when any did. A band with enough respondents can be scored against the simulated people of the same age. A band the simulation has too few members in falls back to the whole crowd. A wave with 5 or more responses can show that wave's alignment.
  • Dashboard: A Validated NN% badge, when any run is scored. Unscored runs stay unscored. An empty dashboard has no badge.
  • Markdown export: When the crowd finishes, download the report as Markdown with the validation tables. A stopped or failed run has no Markdown.
  • Public share links: a scored card. Survey-collected scores name real people. Owner-entered scores name real-world responses.
  • Case studies: alignment. Question-level scores when any did. A band with enough respondents can be scored against the simulated people of the same age. A band the simulation has too few members in falls back to the whole crowd. A wave with 5 or more responses can show that wave's alignment.
  • Accuracy page: /accuracy survey alignment averages scored runs, and known-outcome backtests score past launches against what actually happened

Publishing Case Studies

Validated simulations can be published as public case studies from the report page. A published case study appears on the /case-studies index page and includes:

  • The tested stimulus and audience
  • Alignment score with simulated-vs-real comparison bars
  • Question-level scores when any did.
  • Per-age-group and per-wave alignment (when available)
  • Comment themes with mention-weighted coverage, plus anticipated or missed flags
  • The executive summary, camps that formed, the strongest points for and against, and recommendations, when any did.

Published case studies are public receipts for survey-scored runs. The /accuracy page also shows known-outcome backtests. They help prospective users evaluate CrowdProof's predictive quality before running their own simulations.

Exporting Raw Data

The report page's survey section has an Export CSV button that downloads every survey response with columns for submission time, wave, reaction, purchase intent, faction choice, argument pick, age band, and open-text comment. Use this for analysis in external tools.

API Access

On Enterprise, scoped keys create, retrieve, and export. Live stream and inject stay in God Mode.

  • POST /api/sim/{id}/survey creates a survey link
  • GET /api/sim/{id}/survey/results returns response counts, shares, and comments
  • POST /api/sim/{id}/validation/from-survey scores alignment
  • GET /api/sim/{id}/survey/responses.csv downloads the raw CSV
  • GET /api/stats/accuracy returns the public accuracy aggregate

See the API Reference for full details.

Next Steps

  • Simulations - Create and configure a run. When the crowd finishes you get a synthesis report. A stopped or failed run has no report
  • Accuracy - Survey alignment and known-outcome backtests. Unscored runs stay unscored
  • Case Studies - Published case studies, when any did. Unscored runs stay unscored. An empty catalog has none

Pay once for one full simulation: 500 people, 20 rounds, 3 platforms. Simulations, Accuracy, and Case Studies stay above.