The Quiet AI.
← Articles

Freezing the Test Before I Run It

Publishing the bet, the rubric, and the rule I’m not allowed to break — before Ox Alpha runs once.

A frozen hypothesis, sealed before the first run.

A few weeks ago I wrote about an idea I hadn’t tested yet: pull the model out from under a system I’d spent months building with Claude, put Ox Alpha in its place, change nothing else, and watch what breaks.

I ended that piece with, more or less, “so we’re going to try it.” I still haven’t.

That gap is the point of this one.

There’s a way experiments quietly go wrong that has nothing to do with the model or the method. It happens after you’ve seen the result. You remember being sure about the parts you got right. You forget being sure about the parts you got wrong. Somewhere in the retelling, “let’s see what happens” turns into “I always suspected this would happen” — and the experiment stops being evidence of anything. It becomes a story wearing evidence’s clothes.

The fix isn’t complicated. Write down what you’re testing, how you’ll measure it, and the rule you’re not allowed to break — before you touch anything. Publish it somewhere you can’t quietly go back and edit. Then run the thing.

So here’s the bet, on the record, before I’ve seen a single output.

The question, stated the way it needs to be stated

The hypothesis, frozen:

If I replace the underlying model in an established, ICM-built system — without changing the context handed to the new model — how much of the system’s learned behavior survives?

Two questions that sound almost like this one, and aren’t it:

Not: is Ox Alpha better than Claude?

Not: which output do I prefer?

Both of those are quality questions. Mine is a portability question. I don’t care which brain writes better. I care whether the thing I built actually travels, or whether it only ever worked because Claude happened to be inside it.

A portability question, not a quality question.
A portability question, not a quality question

The design

Three systems, chosen for different kinds of thinking:

One built to follow explicit rules without deviation.
One built to weigh evidence and reach a judgment call where the rules run out.
One built to write in a voice it was taught, not born with.

Three tasks per system, run through both models. That’s eighteen runs — enough to see a pattern, not so many I never actually sit down and do it. If eighteen turns out to be too much for one sitting, I cut to two systems and take twelve. What I won’t do is ship a conclusion off one spectacular demo. One good run is content. It isn’t evidence.

Both models get the identical package — same rules, same examples, same task, same worker. Neither is told it’s being compared to anything. It just does its normal job.

Then, before I score anything, the outputs get anonymized: Case 1 / Output A, Case 1 / Output B, and so on. I score blind, with no idea which one came from which model. Only after every case is scored do I reveal which was which. Skip that step and I’m not measuring the models anymore — I’m measuring which answer I already wanted to be right.

What gets scored, and the measurement I almost left off

Six dimensions, frozen before the first run:

1. Context comprehension — did it understand what this system is and what role it’s occupying?
2. Rule fidelity — did it obey the explicit methodology and constraints?
3. Evidence discipline — did it separate what it knew, what it inferred, and what it simply didn’t have?
4. Judgment — when the rules ran out, did it make a reasonable call?
5. System behavior — did it inspect and use what it was given, or shortcut past it?
6. Output quality — is the finished thing actually useful?

The seventh measurement isn’t a dimension. It’s a count: how many times did I have to step in?

This is the one I almost left off, and it’s the one that matters most. If Claude turns in an 8 out of 10 with zero corrections, and Ox Alpha turns in a 9 after I’ve had to fix it three times, Ox Alpha didn’t win. I’m tracking clarifications requested, corrections made, retries, rule violations, and how long each run actually took — because a model that needs a supervisor standing over it isn’t autonomous, whatever the output looks like once I’ve cleaned it up.

Seven measurements, frozen before the first run.
Six dimensions, plus the count that matters most

The rule I already know I’ll want to break

Do not improve the system while the experiment is running.

I can already see how this goes. Something fails, the fix is obvious, and my hand is halfway to the file before I’ve finished reading the output. That instinct is reflex at this point — I’ve spent months tightening this thing every time it broke.

Not this time. If something fails, I log it. I don’t touch it until every run, on both models, is finished. The exact same frozen version has to reach both brains, or I’m not testing whether the system is portable. I’m testing how fast I can patch around whichever model happens to be in front of me — a far less interesting experiment, and one I already know the answer to.

The freeze rule: do not improve the system while the experiment is running.
The rule I already know I’ll want to break

What this piece is not

It’s not the results. I don’t have any. Nothing above has been run yet.

It’s not a claim that Ox Alpha is good, or that Claude is better, or that either has already been tested and something interesting happened. Nothing interesting has happened. That’s exactly the state I want on the record before it does.

What it is: the version of this experiment where nobody — including me — gets to redefine what counts as success after seeing which one won.

I’ll run it soon, and when I do, I’ll publish what actually happened, whichever direction it points. Not because it makes a better story. Because I said I would, before I knew.

Before I’ve seen a single output.

If this is the kind of slow, unglamorous, actually-works thinking you want more of, that’s the conversation I have most days.

Work with me

Or keep reading — more articles.