the research · proved, measured, and where it failed

Returns sooner. We could not show it returns as good.

Detection — catching an agent’s drift — is a theorem, and it is what we sell. Whether returning then helps, and costs less, is the open question. We tested it three ways, each written down and frozen before it ran and built to lose. It is not established in any of them, and where the evidence is legible it leans the other way. The whole record is below, nulls beside the proofs — because a claim is only worth the boundary printed next to it.

detection · a theoremproven

The fixed reference

Can drift from a goal be detected at all — and can an agent do it by watching only itself?

A proof over agent states in a metric space, with the ground state as the goal as first spelled.

A fixed, unchangeable, findable reference is necessary and sufficient to detect when an agent has left the ball around where it started, in O(1) memory. And no monitor that compares an agent only to a window of its recent states can be both sound and complete — for any window, any rule. So “the reference must never change” is not a stylistic choice; the proof forces it. This is what laserbrain sells.

the metric · by constructionbuilt

A real displacement

Is the reference a genuine measure, or hand-waving?

The JSON grammar the agent spells its state into carries a pseudometric.

Displacement Φ is well-defined, deterministic, and symmetric — same input, same number, every time — so “distance from ground” is a real quantity, not a vibe. Which grammar is rich enough for a given agent’s reasoning is a separate, open question; the theorem blesses a fixed reference, never a particular vocabulary.

coverage · H2 · N = 18measured

It cuts the recursion

Does an agent with the harness return sooner than one left to run?

Open-ended tasks, control vs. harness, replicated.

Yes. The harnessed agent returned to its goal in about half the steps (median 5 vs. 10). That is the step-count half, and it holds. The honest caveat travels with it: fewer steps is not fewer tokens — steps do not charge the monitor for its own calls, and a run can be short and expensive.

benefit · H1 pilot · N = 12inconclusive

The judge could not tell

Does the early return keep the answer as good as running longer?

A blind, stronger judge scored every pair in both orders; a win had to survive the swap.

The preregistered rule read “supported” — but the judge disagreed with itself on 42% of pairs. Judging open-ended answers that have no right answer is near a coin flip. Of the five pairs where the harness acted: three ties (same quality, less cost) and two clean losses (the return produced a worse answer). Consistent with the hope, dominated by noise, not a result.

benefit · ground-truthed · N = 15negative

Where there is a right answer

And on tasks that DO have an objective answer?

Debug/loop coding tasks with hidden unit tests — Pass@1, no judge needed.

The harness did not help. It matched or trailed the control at ~4× the tokens, and every ceiling failure was a run where it intervened. Theory-consistent: a task with its own built-in criterion does not need an outside reference, and the nudge derails a run that was fine. So we say it plainly — laserbrain is not for well-specified, test-backed work. Its domain is open-ended work, which is exactly where quality has no ground truth.

benefit · powered re-run · N = 16inconclusive

The panel, in its own domain

So we ran it only where it should help, with the best measure funding could buy.

Criterion-absent tasks, a three-judge panel, each pair double-order, a rule fixed before any data: if the judges cannot agree (Fleiss κ < 0.4), the verdict is inconclusive — full stop.

The panel agreed at κ = 0.10. So the honest output is inconclusive — the pilot’s problem confirmed at scale: criterion-free answers cannot be reliably judged even by three models, and the κ-gate refused to read a verdict from a broken measure. It was not neutral, though: the harness acted on only 2 of 16 runs, and both of those, where the judges could agree, went to the control. Cost was swamped by noise.

the ledger · what may be said, and what may not

we may say

  • A sound-and-complete detector of drift — provably beyond what any self-watching agent can do.
  • The reference never changes, and we can prove why it must not.
  • Same input, same response — by construction.
  • On open-ended work it returns to ground in about half the steps.

we may not say

  • That it makes your agent better, or finishes the task — H1, tested three ways, not established.
  • That it is cheaper on tokens — measurable only against an unmonitored spiral, not in general.
  • That it helps on well-specified, test-backed work — measured, and it does not.

The rule of thumb: claim detection, not cure. The first is a theorem we own outright; the second is a study we ran three times and did not win. The same discipline that would have let us claim a victory is why the null is honest — every test was frozen before it ran. If a better way to measure quality on open-ended work appears, we will run it and report whatever it says.

← laserbrainattach it — freethe protocol line →