the research · proved, measured, and where it failed
Detection — catching an agent’s drift — is a theorem, and it is what we sell. Whether returning then helps, and costs less, is the open question. We tested it three ways, each written down and frozen before it ran and built to lose. It is not established in any of them, and where the evidence is legible it leans the other way. The whole record is below, nulls beside the proofs — because a claim is only worth the boundary printed next to it.
Can drift from a goal be detected at all — and can an agent do it by watching only itself?
A proof over agent states in a metric space, with the ground state as the goal as first spelled.
A fixed, unchangeable, findable reference is necessary and sufficient to detect when an agent has left the ball around where it started, in O(1) memory. And no monitor that compares an agent only to a window of its recent states can be both sound and complete — for any window, any rule. So “the reference must never change” is not a stylistic choice; the proof forces it. This is what laserbrain sells.
Is the reference a genuine measure, or hand-waving?
The JSON grammar the agent spells its state into carries a pseudometric.
Displacement Φ is well-defined, deterministic, and symmetric — same input, same number, every time — so “distance from ground” is a real quantity, not a vibe. Which grammar is rich enough for a given agent’s reasoning is a separate, open question; the theorem blesses a fixed reference, never a particular vocabulary.
Does an agent with the harness return sooner than one left to run?
Open-ended tasks, control vs. harness, replicated.
Yes. The harnessed agent returned to its goal in about half the steps (median 5 vs. 10). That is the step-count half, and it holds. The honest caveat travels with it: fewer steps is not fewer tokens — steps do not charge the monitor for its own calls, and a run can be short and expensive.
Does the early return keep the answer as good as running longer?
A blind, stronger judge scored every pair in both orders; a win had to survive the swap.
The preregistered rule read “supported” — but the judge disagreed with itself on 42% of pairs. Judging open-ended answers that have no right answer is near a coin flip. Of the five pairs where the harness acted: three ties (same quality, less cost) and two clean losses (the return produced a worse answer). Consistent with the hope, dominated by noise, not a result.
And on tasks that DO have an objective answer?
Debug/loop coding tasks with hidden unit tests — Pass@1, no judge needed.
The harness did not help. It matched or trailed the control at ~4× the tokens, and every ceiling failure was a run where it intervened. Theory-consistent: a task with its own built-in criterion does not need an outside reference, and the nudge derails a run that was fine. So we say it plainly — laserbrain is not for well-specified, test-backed work. Its domain is open-ended work, which is exactly where quality has no ground truth.
So we ran it only where it should help, with the best measure funding could buy.
Criterion-absent tasks, a three-judge panel, each pair double-order, a rule fixed before any data: if the judges cannot agree (Fleiss κ < 0.4), the verdict is inconclusive — full stop.
The panel agreed at κ = 0.10. So the honest output is inconclusive — the pilot’s problem confirmed at scale: criterion-free answers cannot be reliably judged even by three models, and the κ-gate refused to read a verdict from a broken measure. It was not neutral, though: the harness acted on only 2 of 16 runs, and both of those, where the judges could agree, went to the control. Cost was swamped by noise.
the ledger · what may be said, and what may not
we may say
we may not say
The rule of thumb: claim detection, not cure. The first is a theorem we own outright; the second is a study we ran three times and did not win. The same discipline that would have let us claim a victory is why the null is honest — every test was frozen before it ran. If a better way to measure quality on open-ended work appears, we will run it and report whatever it says.