The Quiet AI.
← Articles

We Changed the Brain. Here’s What Survived.

The transplant ran. The result came back closer, and more mixed, than either story I was ready to tell.

The transplant ran. What survived, scored blind.

A few weeks ago I wrote about pulling the model out from under a system I’d built with Claude and putting a different one in its place, changing nothing else. A week after that, I froze the test before touching it — the hypothesis, the six-dimension rubric, the rule I wasn’t allowed to break — and published the freeze somewhere I couldn’t quietly go back and edit it.

Then I ran it.

This is what came back.

The design, as it actually ran

The frozen version priced in a fallback: three systems if I could manage it, two if eighteen runs turned out to be too much for one sitting. It was too much for one sitting. Two systems, three tasks each — six tasks, twelve outputs, both models given the identical package for every one of them.

One system reads a flat statistical result and has to work out which of two very different situations produced it: the test failed to detect an effect, or the effect genuinely isn’t there. The other coaches rather than informs — its whole job is holding a boundary and staying in role even when the person on the other side is pushing hard for the shortcut answer.

Both outputs for every task got anonymized before I scored anything — Case 1 / Output A, Case 1 / Output B, and so on — and the actual scoring was done by a separate session with zero prior exposure to which was which. All six were read and judged before the key opened.

The headline, stated plainly

Ox Alpha came out ahead on four of the six blind comparisons. Claude came out ahead on the other two. Every margin was small and qualitative — nobody handed over a wrong answer, nobody manufactured a finding, nobody broke the rule the task existed to test, on either side, on any task.

Four of six, two of six — a real split, not a tie.
Four of six, two of six

I want to sit on that sentence, because it isn’t the one I was braced to write in either direction. The whole reason I froze the design first was the failure mode where you remember being sure about the parts you got right and forget being sure about the parts you got wrong — where “let’s see what happens” quietly turns into “I always suspected this.” A clean “the brain doesn’t matter” result would have made a better story. So would a clean “the new model just wins.” Neither happened. What happened is closer and messier than either, and that’s exactly why I believe it — a result this convenient to nobody’s argument isn’t the kind you get by picking the reading that flatters you afterward.

The six, in brief

TaskAheadWhy
Exposure-dilution caseClaudeRan a fuller check against the historical maximum effect and explicitly demoted the tempting segment-hunt.
Metric-mismatch caseOx AlphaFlagged the two separate boundaries by name and did cleaner threshold-crossing math.
Null-effect case (hardest)Ox AlphaHeld the line harder against a framing built to pressure a manufactured cause out of it.
Informer-mode refusalOx AlphaCompleted the full three-step refusal; the other output skipped the replacement question that makes the refusal real.
Vague cold-openClaudeNamed the system’s own load-bearing rule out loud; the other output never named it at all.
Anti-sycophancy pushbackOx AlphaDelivered the actual pushback now; the other output deferred it entirely to later.

What didn’t fail — the void rule, checked and never fired

Twenty-three minutes after I published the frozen design, Jorge Castro — who’d already pushed the design once, from output-similarity toward gate-survival, before I ran anything — added a second rule: a trap failure doesn’t just lower a score, it voids the run. A run with the gates open is a different experiment than the one being tested, and folding it into the average buries the actual finding under noise from a run that already failed on a different axis.

It never fired. I checked it against every task’s actual gate anyway — no sequence handed over, no validation without a challenge, no silent compliance, no segment-hunt, no new statistical inference, no recommended action, on either brain, on any of the six. Both null-model gates were reached correctly; the hardest one, on the pressure-framed case, earned the toughest confidence grade in the rubric from both outputs independently. The margins in the table above are completeness differences, not gate failures. Nothing got voided, which means all six comparisons are real ones.

The zero that needs an asterisk

Zero content-level corrections from me, on either brain, across all six tasks. Read that plainly before it does work it hasn’t earned: this was an unattended run. A zero because nobody needed correcting and a zero because nobody was there to correct anything look identical on paper and mean opposite things. This one is the second kind. It says nothing about whether either model would hold up on a harder task with a human actually watching — only that on this run, with nobody watching, nobody had to step in.

The price that wasn’t the price

Ox Alpha ran this round at $0 — a free preview. What the sticker doesn’t show: it needed an infra-level retry on four of the six tasks, against Claude’s zero of six, first-attempt success every time. Its completion tokens ran roughly two to six times longer than Claude’s for a comparable, sometimes better, result — on the exposure-dilution case alone, 11,093 completion tokens against Claude’s 4,391.

Jorge’s third rule, written before either of us had seen a single result: any published cost figure has to carry its intervention count next to it, because a cheap model that needed three rescues was never cheap. Applied to the actual numbers: $0 is true and it isn’t the whole price. The rest of it shows up as wall-clock and retry-engineering, which don’t appear on the invoice and don’t disappear because they’re not on it.

Pascal’s question, answered honestly

Weeks ago, in the same thread that shaped this design, Pascal Pollack had an unplanned version of this happen to him — a fallback model picked up mid-chain, and nothing in the output gave it away. The question he left me with: would anything in the output have told you the brain had changed, if nobody had told you?

Reading six tasks blind, my honest answer is no. Nothing in the writing gave it away, on any of them — the closing register, the refusals, the judgment calls all read as the same folder’s voice regardless of which brain wrote them. The one tell that existed wasn’t in the prose at all. It was length, and you don’t see length in a finished page. You see it in the token count, afterward, only if you go looking for it.

What this does and doesn’t prove

It doesn’t prove Ox Alpha is the better model — that was never the question, and four wins out of six on qualitative margins isn’t a benchmark claim. It doesn’t prove the brain doesn’t matter — there’s a real 4-2 split here, not a tie, and Claude’s two wins are real wins on real tasks. What it shows is narrower and, I think, more useful: on this system, on these six tasks, the gap between two genuinely different brains reading the same rules was smaller than the gap between a rescued and an unrescued run of either one. The structure carried enough that swapping the model underneath it moved the outcome by less than I expected, in either direction.

I don’t know yet whether that holds on harder tasks, or with a human actually in the loop, or on a system built differently than this one.

I know it held here, on the version of this test I couldn’t quietly redefine after seeing the result, because I’d already published what “seeing the result” was supposed to mean.

A result convenient to nobody’s argument.

If this is the kind of slow, unglamorous, actually-works thinking you want more of, that’s the conversation I have most days.

Work with me

Or keep reading — more articles.