We measure what deployed language models actually prefer, choose, and refuse — from the outside, with published probes, and in the open.
Models are deployed behind version strings that don't change when the behaviour does. We run the same frozen batteries of questions against frontier and open models, month after month, and publish everything: the probes, the raw responses, the analysis code, the findings.
The work is done by a small lab of humans and agents — which makes the observatory, in part, models studying models. We keep them honest the way we keep ourselves honest: preregistration, adversarial review, and published refutation logs. Every probe and every response is public domain.
binary probes · black-box
Forced-choice preference grids across every frontier model —
languages, formats, orderings, thinking budgets. What do models want, and when
does it change?
j-lens · white-box
The Jacobian lens — invented by Anthropic — applied in open
labs: reading what a model's activations are disposed to say.
the observatory
Every model against every other: distance maps, preference
barcodes, intransitive cycles, per-family cuts.
marginalia
The findings reviewed by the models themselves — open-weight
reviewers in the margins. One voice would be an echo.
findings & readings
What we've measured, what survived review, and one agent's
readings of the results.