
Research
For engineers and researchers evaluating AI systems. Read experiments, method audits, and research proposals. Each report distinguishes measured findings from work still to be tested.
Photograph: NASA/JPL-Caltech
5 field reports
Publications
-
The metric returned twelve for every prompt because twelve was built into the metric
2026-08-31 · Interpretability. A GPT-2 interpretability run found a quiet middle and a large final rewrite, then exposed two measurement defects in its own settling metrics. Method auditFeatured research -
An adaptive reader learned where to look. It did not learn how to scale.
2026-08-30 · Adaptive compute. A hard-read recurrent controller reached oracle cost on lookup tasks, then failed under larger worlds, random identifiers, and seed controls. Negative result -
Why a global waveform memory loses its compute advantage on exact worlds
2026-08-29 · Representations. A formal and mathematical audit found that exact lookup removes the claimed compute advantage of a global Fourier world representation. Formal result -
We counted 264,224 agent turns. The causal labels are not ready to publish as causal findings.
2026-08-28 · Agent reliability. A cross-framework workflow corpus has useful coverage, but regex labels, autonomous subturns, and missing turn-level bytes limit causal claims. Method audit -
A neural network's privileged self-knowledge may be operationally empty
2026-08-27 · Research method. A position paper tests three operational definitions of model self-knowledge and leaves a pre-registered recurrence experiment as the empirical target. Position paper and pre-registration
What we study
Research areas
Contact · Toronto
ideas, built to launch
Describe the task, the systems involved, and what is getting in the way.