The Quiet Ai.
← Articles

The Ten-Second Test

A new AI model answers with a probability and nothing else. Reading up on it gave me two questions to ask before I reach for any model: what kind of work is it, and how many at once?

Two desks in one room. At the left, a small figure flicks index cards from a stack into one of two trays with barely a glance. At the right, a larger figure hunches over a typewriter, drafting a single page slowly.

A model came out on September 15 that never writes a word back. You give it a question in plain English, is this message urgent?, and it answers with a probability. That is all it says.

It’s called Jev, it comes from a company called TypeSafe, and according to Every, which wrote it up on September 23, it spent about two years in stealth before release. It reads text and structured data only for now (TypeSafe’s docs say no images yet), and it classifies instead of generating. Every lists its price at 4.2 cents per million tokens in, with output free. Every lists it beside other models; the small model I’m using as a yardstick is $1 in and $5 out.

I haven’t run it. The facts about Jev here come from Every’s write-up and TypeSafe’s own description, so weigh them that way.

What I did instead was apply a test. A launch like this invites a mistake first: new, cheap, fast, so surely something in my week could use a decision engine. That is shopping for a place to put a tool, and the tool can’t tell you where it belongs. Two questions can, and neither one mentions Jev.

Question one: what kind of work is it?

Every’s rule of thumb is that a task a person does in under ten seconds suits Jev, and a task that takes more reasoning goes to a language model. I’d put it as three verbs. Sort. Route. Flag. Is this message a refund request? Which of three inboxes does it belong in? Does this headline mention a competitor? Each is a yes, a no, or a pick-one, and a person would make the call at a glance.

Now put a different job next to those. Write the reply to that refund request so it sounds like me. No one makes that call in ten seconds, because it isn’t a call. It’s voice, and craft, and a decision about what to leave out. Jev doesn’t do it, by design, because it doesn’t write.

Two lanes. Top lane, “Decision model. Picks.”: five items go in on the left, and each comes out the right side with a YES or NO chip and a score bar. Bottom lane, “Language model. Writes.”: one prompt goes in and a block of new text comes out.

Every describes a method for the fuzzy jobs that sit in between, and it’s worth stealing whatever model you use. A question like is this urgent? is really several. Split it into separate yes/no checks (my examples: does it mention a deadline, does the sender say they’re blocked, is it from a customer), have the model score each one, and combine the scores in ordinary code. Then, before you trust any of it, run the whole thing on a sample you’ve already judged by hand and see where it disagrees with you. The judgment stays yours. The model answers small questions, and the code you wrote decides what the answers add up to.

Question two: how many at once?

My rule: the edge shows at hundreds to thousands of items in one run. Below that, a small language model you already have is cheap enough.

Do the arithmetic once, because it’s short. Say a piece runs about a thousand tokens, roughly 700 words. A thousand of them in one run is a million tokens: about a dollar of input at that small model’s price, about four cents at Jev’s. That’s about 96 cents saved per run: worth a new tool only if you run it constantly.

Five pieces a day is five thousand tokens. Half a cent on one side, two-hundredths of a cent on the other. You won’t feel either bill. You will feel the afternoon it takes to wire in a second kind of tool.

The example people are sharing sits squarely on the right side of that line. Elvis Sun, who sells PR tools, posted a demo of Jev reading 384 headlines and picking stories for 15 brands in 24.9 seconds for 19 cents. That is his own test, and I haven’t checked it. I’m pointing at the shape of the job, hundreds in one pass, not the result.

Where my own work lands

I ran both questions on my own work and it fails both. The work I’d want help with is writing, which is voice, so it fails the first question before the second comes up. And it arrives one piece at a time, not by the thousand, so it fails the second question too. I read up on Jev, I understood what it’s for, and I didn’t adopt it.

It isn’t a permanent answer. If my work ever shifts from single pieces to hundreds of small judgments in a run, I’ll ask the two questions again and expect a different result.

I came to the conversation about Jev through Jake Van Clief’s online community. Jake’s framework, ICM (the Interpretable Context Methodology, from his paper with David McDermott), is a way of setting up AI work as plain folders and text files that a person can review at each step and an AI can read. What I took from the conversation there is fundamentals first: don’t drop a tool into a workflow you don’t understand. The two questions are that idea made short.

A dark card with the rule set apart: “A ten-second decision, made a thousand times. Miss either half, and you don’t need a decision model.”

A ten-second decision, made a thousand times. Miss either half, and you don’t need a decision model.

What I can’t tell you

Whether Jev is accurate enough for your sort. A write-up can’t answer that, and I haven’t run it on anything of mine. It’s why Every’s method ends on the sample you’ve already judged: that test is the only one that answers the accuracy question for your data.

Try this today

Take one task from your week and finish this sentence: a person could decide this in ten seconds if ___. If you can’t finish it without slipping in “and then they’d have to think about,” it isn’t that shape, and a language model is the right tool. If you can, count how many you’d do in one sitting. Dozens, and a small language model already costs you cents. Hundreds or more, and it’s worth reading about tools built for exactly that.

More notes like this one are at thequietai.com.

Jev writes nothing back. The sentence that matters here is the one you write first: what the job actually is.

Less hype, more attention.

If this is the kind of slow, unglamorous, actually-works thinking you want more of, that’s the conversation I have most days.

Work with me →

Or keep reading — more articles.