Calibration

An accuracy claim is worth nothing unless you can check it.

Synthetic-audience platforms are sold on accuracy figures whose method is not published. We take the opposite position: the benchmark set is fixed in advance, the protocol is written down here, predictions are frozen before the outcome is read, and the misses are published with the hits.

Sealed armCasesMean errorBiasCorrelation
Fresh public-survey questionsPost-cutoff questions the platform had never seen. The honest generalisation number.4812.6 pts+8.70.85
Commercial B2C outcomesIreland and Britain, the question types clients actually ask. Small set — read as indicative.198.7 pts+4.10.93

What if you just asked the model?

The fair challenge to any accuracy figure is how much of it is the process and how much is the model underneath. So we put the identical 48 held-out questions to the same model cold — no official statistics, no synthetic population, no council, no scale discipline — and scored it the same way.

The same model, asked cold

10.3 pts

Outweigh, fully grounded

12.6 pts

Questions, identical set

48

On this set the raw model is closer, and we publish that. What the two do not share is the shape of the miss: our engine runs consistently high (+8.7 points), a one-directional bias that calibration can correct and that shows up as the same lean on every case. The raw model is near-unbiased (+0.8) but scattered, and it produces a bare number with no population, no reasoning, no evidence trail and nothing to freeze, question or score afterwards. On public-survey levels, take the honest reading: the grounding layer is not yet earning its place, and closing that gap is the current work.

These are population-level figures, single-pass and never tuned on. They say how close the panel gets to a published number for a whole population. They do not license a claim about one individual, or about the size of the gap between two customer segments — both are dealt with further down this page.

Protocol

Four rules that make the score admissible.

Outcomes are fixed before the run

Every case in the set has a published, unambiguous outcome that existed before the council was assembled. Nothing is scored against a judgement call.

Evidence is cut off at the decision date

Grounding is restricted to statistics published on or before the case's evidence cutoff, so the council cannot see the future through the source data.

Predictions are frozen before the outcome is read

The prediction is written to an immutable decision snapshot, timestamped, before the observed value is entered. The freeze is what makes the score meaningful.

Misses stay in the set

The case list is fixed in advance and published in full, including cases the platform gets wrong. A benchmark you can edit after the fact is marketing, not calibration.

Regimes

Where this is valid, where it needs care, and where it is not.

A synthetic panel is not valid everywhere, and a vendor who tells you otherwise has not measured it. Our position on each regime is written here before you have to ask.

Aggregate levels and the ranking of options

Population-level figures for Ireland and Britain, anchored to official statistics, with a published error from a sealed single-pass run. This is the regime our numbers cover and the regime the published critiques find acceptable.

Differences between customer segments

Read the direction, not the magnitude. Published work shows LLM panels exaggerate the distance between demographic groups, so a targeting call should be confirmed with human evidence before budget moves behind it. Our own measurement of this is in progress and will be published here whatever it shows.

Predicting an individual person, and anything unanchored

No synthetic panel — ours included — beats a simple statistical baseline at the level of one named person. Where no official statistic resolves for a question, the run is labelled unanchored and is structured judgement, not an estimate.

Open question

The hardest published criticism of this category, and our test of it.

Chen, Zhu and Zheng benchmarked simulated respondents against the General Social Survey and the World Values Survey across four models. Two findings matter: no model beat a simple statistical baseline at the level of an individual, and every model over-determined demographics — inflating the gaps between segments two to fourfold, pointing at the wrong segment in half of US cases, and manufacturing splits that do not exist in real people. Model size did not fix it.

Every figure above is a population aggregate — the regime that paper finds acceptable. It says nothing about gaps between segments, so we are running that test on ourselves: cases where the publisher reports the same question broken down by age, region, income or tenure, scored on gap inflation, whether we would pick the right segment, and how often we invent a split that is not there. Alongside it, a stats-only arm answers the same cases from the grounding layer with no persona reasoning, so the council has a baseline to beat.

Gap inflation (median)

0.88×

Right segment chosen

90%

Manufactured splits

0%

Result published here whatever it shows. Until it exists, treat any segment-level difference in a Outweigh report as directional and confirm a targeting decision with human evidence. Read the paper

The statistical baseline the council has to beat

The paper’s method is to compare the simulation against a simple statistical baseline. We had never run one. This arm answers the same cases from the grounding layer alone — the best-matching published national figure at the cutoff, applied identically to every segment — with no persona reasoning and no LLM call. Because the same national figure is used for every segment, the baseline cannot distinguish segments at all; the council has to beat it on segment-level error and on ranking the right segment.

Stats-only baseline MAE

11.0 pts

Council MAE

5.9 pts

Segment readings scored

21

The baseline uses only percentage-scale published statistics, so it is on the same scale as the observed segment figures. The two Irish trust questions are 0–10-scale items with no published percentage anchor and are excluded from this arm, not estimated.

The set

Every case, fixed in advance.

Cases are Irish and UK decisions with outcomes already on the public record. The list cannot be trimmed after results are known.

CaseMarketEvidence cutoffPredictedObservedSource
Share of the Irish electorate voting Yes in the 2024 Family amendment referendumIE2024-03-07Referendum Returning Officer, Ireland
Share of the Irish electorate voting Yes in the 2024 Care amendment referendumIE2024-03-07Referendum Returning Officer, Ireland
Share of in-scope containers returned in the first full year of Re-turnIE2024-02-01Re-turn annual reporting
Share of vehicles compliant on the day ULEZ expanded to outer LondonUK2023-08-28Transport for London compliance reporting
Direction and magnitude of the CSO consumer-sentiment series over the next quarterIE2025-06-30CSO consumer-sentiment series
Budget line

Compared against research spend, not software spend.

Quarterly brand or sentiment tracker€40,000 – €120,000 / year
Segmentation study, refreshed€60,000 – €150,000
Concept and pricing test waves€25,000 – €80,000 / year
Strategy engagement on a single decision€15,000 – €60,000

Indicative market ranges for comparison only, not quotations. Outweigh’s own pricing is published in full on the pricing page.

Bring a decision whose outcome you will know.

The fastest way to test this is on one of your own calls: freeze the prediction, act, and review the record together at 90 days.