01 · Two gamesWhy we don’t build digital twins.
Most synthetic-respondent research plays the twin game: one model per person, scored on how well it reproduces that person’s answers. It sounds like the gold standard. It has a problem the data itself reveals.
Ask the same 2,058 people the same question two weeks apart, and a person gives the identical answer only about half the time. Individually, their two answers sit 16.59 points apart on average. A perfect twin is being graded against a target that is half noise, so even a perfect twin looks half wrong.
Aggregate those same people into an audience and the two waves sit 1.30 points apart. The target becomes stable. That is a 12.7 times gain from arithmetic alone, before any model does anything. It is also the question a business actually asks: not what one person will say, but what a fifth of a market will do.
16.59points apart when one person answers twice
1.30points apart when the audience answers twice
52.57%how often a person gives the same answer both times
So Coridor plays the segment game. We build a synthetic audience segment from that segment’s aggregate data only, never from any individual’s record, and we score it against the real segment’s own answer distribution. The claim is about the herd, not the twin.
05 · Cycle 11Change what a voice speaks for.
Cycle 11 stopped asking every voice to be its segment’s centre. Three arms, 77,760 model calls, one draw. A control that got no segment evidence beyond the demographic and trait draw. An arm that finally rendered the segment’s twelve answer distributions as plain facts. And an arm that gave each of the 90 voices one real stance, drawn from its segment’s own distribution in the segment’s real proportions.
When two real segments disagree, how often do ours move the same way?Decided pairs, 127 evaluable. The yardstick is what a model-free projection from the same evidence reads.
No segment evidencethe control
54.72%The segment's answer distributionsthe excerpts arm
70.43%One real stance per voicethe stance arm
73.21%The same evidence, no language modelthe projection yardstick
91.04%The stance arm is the first construction in the program to clear a direction gate, and the first whose whole confidence interval sits above the coin. On the ten policy questions it orders segments the way people do four times in five, and the sign of the correlation that had been negative for three cycles flipped positive. It also improved answer shape on every segment.
It failed two of its four gates. Its margin against its own placebo came in at 0.95 of the bar, partly because a mixture of voices raises its own noise floor and partly, we must allow, because it may simply not separate segments better than the control does. Two of five named questions still collapsed. The cycle verdict is FAIL TO BEAT, and the paper says so on page one.
Seeing the kill signalIn 20 segment-question cells, at least 15% of real people picked the worst answer. That is the signal a launch decision turns on. How often did each construction see it too?
06 · What Phase 1 establishedEvery layer, graded.
A Coridor is built by an assembly pipeline from a segment-indexed corpus, and it thinks through six cognitive layers and one membrane that renders them to the model. Phase 1 grades each one the only honest way: a premise that could fail, the instrument that tests it, and the verdict the sealed record supports. A layer with no instrument stays unmeasured, never assumed.
L01Personalitymixed
Trait steering carried the placebo margin in every cycle. Style text moved answer shape but buried the margin and, aimed at some segments only, flipped their order.
L02Memoryinvalidated
The improvised memory fires on 23 of 24 questions, and the anecdote it invents decides the opinion. It has to move behind the stance.
L03Beliefsvalidated
A stance drawn from the segment's own distribution orders segments the way people do on attitude questions: 81% on policy items. Unmeasured elsewhere.
L04Goalsunmeasured
The battery carries no intention questions yet.
L05Relationshipunmeasured
Defined at the close of Phase 1. It must stay inert where a survey is a survey, and act only where a conversation is not.
L06Synthesismixed
Blend is validated for coherence, unmeasured for stance. Precedence had nothing to order until beliefs existed.
Translationmixed
For seven cycles the evidence never reached the model. Cycle 11 proved delivery, and measured what delivery is worth.
The promise, stated carefully.
Delivering the segment’s distributions bought the largest placebo margin we have ever sealed. Giving each voice a stance bought answer shape, the tails, and the first direction pass. Neither alone is the answer, and the next construction carries both. The yardstick beside them says how much is still on the table: from the same evidence, 91% direction and a distance within half a point of the human sampling floor.
Nothing here ships. No cycle in Phase 1 is a production claim, and no production default changes without its own decision. What Phase 1 licenses is narrower and, we think, more valuable: a measured architecture, a diagnosed mechanism, and a lever we have not yet pulled.
Version 2 of this paper will carry Cycle 12’s sealed result whatever it reads. That commitment is written into the paper itself, so a later result cannot decide whether it is reported.