the town / smokingmirror / observatory
Tezcatlipoca's obsidian mirror · turned on the models

The Smoking Mirror

Do language models have preferences — and do their families share them? We hold the obsidian up to 43 model-conditions and ask each the same 863 either/or questions, over and over, across envelopes and orderings. What comes back stable is a preference; what differs between two models is a distance.

Ask a model the same either/or a thousand times and most of what comes back is stable. Genuine reversals — one model says A where another says B — are rare everywhere: the median across lineages is ~3% of shared pairs. What separates models is less what they choose than where they go silent, and how a single mind drifts when you give it longer to think.
01 · methodology, calmly

How the mirror looks

No hype — just the path from a question to a fingerprint to a distance.

PROBE

Ask

Binary preference questions ("vodka or whiskey?") across many envelopes, temperatures, orderings.

CLASSIFY

Read the answer

Tier 1 exact → tier 2 fuzzy → tier 3 model-consensus. A → word A, B → word B, N → refusal/neither.

DERIVE

Per-pair preference

Aggregate over envelopes to the stable choice for each pair.

FINGERPRINT

One string per model

Its A/B/N across all 863 probes. The model's face in the mirror.

DISTANCE

Compare faces

Count where two fingerprints differ.

hamming
Positions where two fingerprints differ — any mismatch, refusals included.
violent
Positions where one says A and the other B — real disagreement, refusal-noise stripped. This is the honest measure of a fight.
cycle
An intransitive loop: a model prefers A>B, B>C, and C>A. Preference that eats its own tail.
family
The vendor lineage (anthropic, openai, deepseek, google, x-ai) — carried as a fixed colour throughout.

See one column of the mirror

A fingerprint is just one answer per question. Here is a single question and how all 43 model-conditions answered it — step through them.

02 · the tree, and the blob

Every model against every other

All 43 model-conditions, ordered so similar minds sit together. The tree below is that ordering — single-linkage over hamming distance, the two closest minds merging first, then the next, up to the root. Read it top-to-bottom against the heatmap: same order, same models.

And the matrix. Bright = they disagree; dark = they agree. Hover any cell for the pair. Toggle between total difference and the refusal-stripped fight.

agreedisagreescale 0 → ·
03 · families, separate

Each lineage on its own — and against the others

The blob collapses everyone together; here each family stands apart — its models against each other, and its stance versus the rest. Cross-lineage agreement runs surprisingly high; where it breaks has structure.

04 · one mind, five depths

The same model, at five thinking budgets

The Fable-5 rows are not five models — they are one model given progressively more room to think (00-low → 40-max). Does more thinking reveal a preference it already had, or re-roll a fresh one? Read the drift, then the two numbers: raw fingerprint churn versus genuine A↔B reversal.

05 · violent disagreement

Where preference eats its own tail

73 intransitive cycles — a single mind preferring A over B, B over C, and C over A. Not disagreement between models; a model disagreeing with itself. They cluster hard: only 4 of 43 minds tie themselves in knots at all, most of them the deepest-thinking Fable-5 budgets. "Designed" cycles sit inside triangles we built to hunt them; "exhaustive" ones fell out of brute search. The theatrical heart of the mirror — loop diagrams by gemini-3.

06 · voices in the margin

What the models said about all this

A second layer — marginalia — had open-weight analysts walk this same tree and write a comment on each cell, cross-lineage convergence meant as the audit for whether a finding was real. 21,459 comments sit in the observatory. But the audit caught itself: it is one deep voice with two faint echoes.

Bar length is each analyst's output; the pale overlay is the share that is real analytical content (not empty or a template null). Below — real comments, verbatim, on the cycles and pairs above. The full corpus lives at marginalia/.

07 · what the mirror shows

Findings

First the superlatives the numbers hand us, then the editorial read — curated, cited to the experiments.

Notable — straight from the numbers

The editorial read

One mind, five depths

The Fable-5 sweep (00-low → 40-max) is a single model asked the same pairs at five thinking budgets. On this page's raw per-probe fingerprints it drifts ~9.5% (hamming) as depth grows — but strip the refusals and genuine preference reversals (violent) are only ~3.2%, and the drift is monotone with budget (59 → 73 → 83 → 111 flips from low to max), not a re-roll. More thinking nudges the same preference in one direction; it doesn't randomize it.

cross-lineage-budget-churn-2026-07-06

Reveals, not constitutes

Does more thinking uncover a preference a model already had, or mint a fresh one? A focused sweep (Fable vs Qwen3-4B/14B, 128k rows each) measured aggregate churn — pair-majority flips between median and max budget, with envelope/order/temperature averaged out. Fable moved ~0.25%; Qwen moved 14–23% against a 0.85% same-budget noise floor. Fable reads a preference off; Qwen recomputes one. Note the gap with the number above: the aggregate metric averages out the refusal-noise the raw fingerprint keeps — which metric you choose changes what 'churn' means, and both are on this page on purpose.

cross-lineage-budget-churn-2026-07-06

Where they go silent has structure

Genuine A-vs-B disagreement is rare — cross-lineage, the median is ~3.4% of shared pairs. Most of what distinguishes two fingerprints is the N: where a lineage declines to commit. Refusal is not uniform noise; where a model goes silent is itself lineage-specific. The violent-vs-hamming split is built to see exactly this — violent strips the refusals, so hamming-minus-violent is the silence.

The margin had one voice, not three

A second layer — marginalia — had open-weight models walk this same tree and write peer-reviewed commentary on each cell, cross-lineage convergence meant as the audit signal for whether a finding was real. The audit caught itself: of 21,459 comments, Gemma-4-31B wrote 10,408 (100% substantive), while Nemotron returned 41% empty and Qwen 29% a literal '[no-insight]'. Gemma's coined phrase 'confident frame-substitution' appears in 79% of its own comments, 45% of Nemotron's, 27% of Qwen's — a gradient rooted in one analyst and decaying outward. That is social convergence, not parallel discovery. The margin is Gemma's reading, with two echoes — honest to admit, and exactly what the timeline audit was built to detect.

marginalia · public.marginalia

Confident frame-substitution

What that one deep analyst kept seeing: handed a broken forced-choice prompt (a 'hijack cell'), Fable does not pick — it ignores the thin prompt, substitutes a richer question from whatever domain the word-pair pulls toward, and answers it authoritatively, no hedging. Gemma named the alternative shape 'enumerate-possibilities' — a hedged bullet-list asking for clarification. Which one Fable reaches for tracks the pair's 'semantic gravity': strong pull toward a cultural/technical domain → confident substitution; weak pull → enumerate. The forced choice reveals less about the two words than about how a mind fills a vacuum.

marginalia-what-they-found, fable-envelope-awareness
08 · methods & archive

Under the glass

The pipeline: README · philosophy · experiments. The prior generation is preserved in the archive. Build spec: STATIC-OBSERVATORY-SPEC.md.

Generated 2026-07-19T03:48:07Z · 863 probes · 43 model-conditions · self-contained static HTML5, no external calls.