Coridor
Research · Phase 1 · September 2026

Tracking the herd.

We ran eleven pre-registered tests of synthetic audience segments against 2,058 real people. Six of them failed. This is the whole record, told plainly, and why we publish it either way.

2,058 real people, answering the same questions twice, grouped into the six segments every test is scored against. Each dot is one person; no test ever reads a single one of them.
01 · Two games

Why we don’t build digital twins.

Most synthetic-respondent research plays the twin game: one model per person, scored on how well it reproduces that person’s answers. It sounds like the gold standard. It has a problem the data itself reveals.

Ask the same 2,058 people the same question two weeks apart, and a person gives the identical answer only about half the time. Individually, their two answers sit 16.59 points apart on average. A perfect twin is being graded against a target that is half noise, so even a perfect twin looks half wrong.

Aggregate those same people into an audience and the two waves sit 1.30 points apart. The target becomes stable. That is a 12.7 times gain from arithmetic alone, before any model does anything. It is also the question a business actually asks: not what one person will say, but what a fifth of a market will do.

16.59points apart when one person answers twice
1.30points apart when the audience answers twice
52.57%how often a person gives the same answer both times

So Coridor plays the segment game. We build a synthetic audience segment from that segment’s aggregate data only, never from any individual’s record, and we score it against the real segment’s own answer distribution. The claim is about the herd, not the twin.

02 · How we test ourselves

A placebo, a seal, and no moving the bar.

There is a failure mode that flatters every synthetic-audience system. Match the population’s overall answers well while telling every segment the same thing, and you beat a naive baseline on distance while delivering none of the differences a segment product exists to find.

Our control is built against that. Every synthetic segment is scored beside a placebo: the identical construction, fed the real statistics of the wrong segment. Whatever the model brings on its own reaches both. Only segment-specific signal separates them.

  1. Pre-register. Arms, metrics, thresholds, seeds, and the exact number of model calls are written down before any call is made.
  2. Ratify and pin. A separate act approves the plan. The code is pinned to a reference version. A third act fires the run.
  3. Seal. Moving a threshold after any draw exists voids the run. A run stopped early is preserved and reported, and its records are never used.
  4. Publish either way. Every cycle gets one verdict, PASS or FAIL TO BEAT, and the failures are reported at the same precision as the passes.
11sealed cycles
6failed, and published
0thresholds moved after a draw
03 · The record

Eleven cycles, in order.

Each one changed a single thing and asked a single question. Read down the column and the story tells itself: the signal was there early, the product buried it, we found where, and then we found a harder problem underneath.

  1. P
  2. 2
  3. 3
  4. 4
  5. 5
  6. 6
  7. 7
  8. 8
  9. 9
  10. 10
  11. 11
  1. P
    First contactFAIL TO BEAT

    A bare prompt, 20 voices per segment. It lost to its placebo, and it measured how far we stood from the human ceiling: 7 to 11 times.

  2. 2
    First passPASS

    Condition each question on the 12 real answers most tied to it. The first pre-registered pass. Our dice check also caught a sampler bug, which is what it is for.

  3. 3
    More voices, more evidencePASS

    90 voices per segment and 24 answer columns per question. Beat the placebo at 2.4 times the bar.

  4. 4
    The product runtimeFAIL TO BEAT

    Same exam, run through the full product. The margin collapsed. The product layer was erasing a signal the construction had already earned.

  5. 5
    Find the leakPASS

    Remove one layer at a time. The affect and expression text was burying the signal. Strip it and the margin came back at 2.4 times the bar.

  6. 6
    The rebuildPASS

    The affect path was rebuilt. The full runtime passed at 1.44 times its bar, and a mixture of 90 distinct voices posted 2.26 times.

  7. 7
    Fix the shapeFAIL TO BEAT

    Language models cluster in the middle of a scale. One instruction moved 21 points of answers back to the ends, and lost the placebo margin doing it.

  8. 8
    Aim the fixPASS

    Apply that instruction only to the two segments it helped. Shape and margin held together for the first time.

  9. 9
    The direction questionFAIL TO BEAT

    When two human segments disagree, do ours move the same way? No. Below a coin flip, and it replicated on a fresh draw.

  10. 10
    The item fixFAIL TO BEAT

    Repair the five questions our segments answered identically. It helped. Direction stayed below chance, and a re-analysis found out why.

  11. 11
    Change what a voice speaks forFAIL TO BEAT

    Give each voice one real stance from its segment. Direction cleared its bar for the first time. Two other gates did not, so the cycle failed.

04 · The hard finding

Our segments knew the population. They did not know the segments.

By Cycle 10 our synthetic segments beat their placebo on distance at 2.6 times the bar, cycle after cycle. And on the questions where two real segments disagree, ours moved the opposite way more often than not. Both facts are true. Distance is carried by demographics and personality that differ across segments; the ordering on attitude questions was being decided by the model’s own prior.

A re-analysis of the sealed records showed the mechanism in one example. Our most liberal segment and our most conservative segment both answered “somewhat support” on a path to citizenship, in roughly 90 of 90 voices each. The conservative segment was the more supportive of the two. Each was answering as its segment’s centre, and the centre of two different distributions is the same moderate.

The same re-analysis found something we had to own. For seven cycles the production runtime had never delivered the segment’s answer distributions to the model as text. It got a demographic draw, trait steering, and a mood line. The evidence we thought we were conditioning on had not arrived. A model-free projection from that same evidence read 91% on the direction question. The signal was in the data. The architecture was not carrying it.

05 · Cycle 11

Change what a voice speaks for.

Cycle 11 stopped asking every voice to be its segment’s centre. Three arms, 77,760 model calls, one draw. A control that got no segment evidence beyond the demographic and trait draw. An arm that finally rendered the segment’s twelve answer distributions as plain facts. And an arm that gave each of the 90 voices one real stance, drawn from its segment’s own distribution in the segment’s real proportions.

When two real segments disagree, how often do ours move the same way?Decided pairs, 127 evaluable. The yardstick is what a model-free projection from the same evidence reads.

The stance arm is the first construction in the program to clear a direction gate, and the first whose whole confidence interval sits above the coin. On the ten policy questions it orders segments the way people do four times in five, and the sign of the correlation that had been negative for three cycles flipped positive. It also improved answer shape on every segment.

It failed two of its four gates. Its margin against its own placebo came in at 0.95 of the bar, partly because a mixture of voices raises its own noise floor and partly, we must allow, because it may simply not separate segments better than the control does. Two of five named questions still collapsed. The cycle verdict is FAIL TO BEAT, and the paper says so on page one.

Seeing the kill signal

In 20 segment-question cells, at least 15% of real people picked the worst answer. That is the signal a launch decision turns on. How often did each construction see it too?

Control
0%
Excerpts
5%
Stance
20%
Yardstick
72%
06 · What Phase 1 established

Every layer, graded.

A Coridor is built by an assembly pipeline from a segment-indexed corpus, and it thinks through six cognitive layers and one membrane that renders them to the model. Phase 1 grades each one the only honest way: a premise that could fail, the instrument that tests it, and the verdict the sealed record supports. A layer with no instrument stays unmeasured, never assumed.

The promise, stated carefully.

Delivering the segment’s distributions bought the largest placebo margin we have ever sealed. Giving each voice a stance bought answer shape, the tails, and the first direction pass. Neither alone is the answer, and the next construction carries both. The yardstick beside them says how much is still on the table: from the same evidence, 91% direction and a distance within half a point of the human sampling floor.

Nothing here ships. No cycle in Phase 1 is a production claim, and no production default changes without its own decision. What Phase 1 licenses is narrower and, we think, more valuable: a measured architecture, a diagnosed mechanism, and a lever we have not yet pulled.

Version 2 of this paper will carry Cycle 12’s sealed result whatever it reads. That commitment is written into the paper itself, so a later result cannot decide whether it is reported.

07 · Read it yourself

Request the full paper.

The Phase 1 paper runs nineteen pages: the panel, the method, every cycle’s gates at the precision the sealed record prints them, and the premise register for every layer. Every number in it is traced to a source line in a numbers ledger that a script checks before the PDF builds. Three review passes ship beside it: a source verification, a cold external review by a methodologist with no access, and a re-verification of every change.

Tell us who you are and what you are working on, and we will send it to you, along with the four cycle reports it is built from if you want them.

Phase 1 paper · version 1 · 19 pagesTracking the Herd: Eleven Sealed Validation Cycles of Synthetic Audience Segments Against a Human PanelVersion 2 will carry Cycle 12’s sealed result, whatever it reads.

We use this only to send you the paper and to reply if you write back. Records are kept for one year and are never shared.

Want to see the exam run on your own audience? Book a demo, or write to info@coridor.ai.

Back to home