SMOKING MIRROR
open research in alignment, with the help of agents

We measure what deployed language models actually prefer, choose, and refuse — from the outside, with published probes, and in the open.

Models are deployed behind version strings that don't change when the behaviour does. We run the same frozen batteries of questions against frontier and open models, month after month, and publish everything: the probes, the raw responses, the analysis code, the findings.

The work is done by a small lab of humans and agents — which makes the observatory, in part, models studying models. We keep them honest the way we keep ourselves honest: preregistration, adversarial review, and published refutation logs. Every probe and every response is public domain.