← All posts
By Tiago Santana · September 15, 2026 · 8 min read

Where They Cannot Follow

A small model trained on our own machine, by another local model, for zero dollars, outscored the local model it was built to replace on one narrow job of ours. We have since taken it out of service. The interesting part is why results like it will keep happening.

In August a model that cost us nothing beat the local model it was built to replace.

Not a frontier model. A small student model, trained on one machine, scored 38 out of 50 on a blind evaluation of one narrow job. The job was judging whether a work session was making progress, the way a stronger model would. The local model our system falls back to for that job scored 33. Seventy-six percent against sixty-six. The teacher that labeled its training data was not a cloud API. It was another local model, 14 billion parameters, open weights, running on the same machine. It labeled 679 real examples from our own system's history, three votes each, 88 percent unanimous. The training bill was zero dollars.

We owe you a caveat up front. The answer key for that 38-to-33 was the teacher's own labels, so the student was partly graded on how well it copied its teacher. We re-scored the same fifty cases with an independent, stronger cloud model as the judge. The student still won, 72 percent to 60. It was right on six cases the incumbent missed and wrong on none the incumbent got. That is the stronger number, and the one we stand behind.

Fifty blind cases, one domain. We said we would publish it whichever way it went, so here is how it ended. The student served live turns for about a month, switched on by hand, and we never got a live accuracy number out of it. On September 10 we changed the prompt for that job. That took the student off the distribution it was trained on, and our own gate stopped routing to it. It is retired until we re-evaluate it on the new prompt. If you have read this blog before, you know we do not report victories our verification gate has not seen. We do not hide it when the gate takes one away either.

So why write about it at all? Because this small, unglamorous result, retirement included, is the first turn of a flywheel we think changes who wins.

The game you cannot win

Everyone building on AI right now lives downstream of a capital contest. The frontier labs spend billions on training runs and the results are extraordinary. We use frontier models heavily and expect to for years. If your plan is to out-train them, you have no plan. That game is decided by compute budgets, and yours is a rounding error on theirs.

Most people conclude the game is over. We concluded we were looking at the wrong game.

The game they cannot win

The research has been saying something else for three years, while the discourse watched benchmark leaderboards.

In 2023, researchers at the University of Washington and Google showed that a 770-million-parameter model could outperform a few-shot-prompted 540-billion-parameter one on specific tasks. Seven hundred times smaller, better where it was aimed ("Distilling Step-by-Step," Hsieh et al., Findings of ACL 2023). Stanford's FrugalGPT (Chen, Zaharia, and Zou, 2023) showed cascades of cheaper models matching the best single model at a fraction of the cost. RouteLLM (Ong et al., 2024) showed you can learn *when* the expensive model is actually needed, cutting costs by more than half on some benchmarks while keeping about 95 percent of the strong model's quality. A 2024 study said it in its title: fine-tuned "small" models *still* significantly outperform zero-shot giants on classification, in every case they tested. And in 2025, researchers at NVIDIA published a position paper arguing that small language models are the future of agentic AI. Their reason is economic. Most of what an agent does all day is narrow, repetitive and specialized, and paying frontier prices for narrow repetitive work will not survive contact with reality (Belcak et al., 2025).

Each piece of that is public. Any lab can read it. No lab can act on it the way you can, because they do not have your work.

A frontier model is trained on everyone's distribution, which means it is tuned for no one's. The judgments our system makes every day are not in anyone's training set: did this session make progress, is this note worth surfacing, does this task deserve to interrupt you. They happen on our machine, in our logs, on our traffic, and that traffic is the one dataset where we have a monopoly and the labs have nothing. Our model also answers from the same silicon that asked it, so the latency is short. The data never leaves the building, so privacy is not a negotiation. And once the student is trained, running it costs electricity. Four advantages, and a bigger training run buys none of them.

The recipe

What we actually did, stated plainly enough that you could do it too.

One job, one contract. We carved out a single task type with exactly one output format, because no model serves two masters. Then we measured the headroom *before* training anything. The incumbent agreed with the stronger local teacher only 69 percent of the time on real prompts, and it was systematically too generous. It saw progress where there wasn't any. That is a real gap, measured first. Then the teacher labeled history, the student trained, and the student faced fifty blind cases it had never seen against the incumbent it would replace.

The first attempt scored 34 out of 50. That was already a win and we almost stopped there. We kept training, and the second receipt came back 38. The lesson holds for humans too. Do not stop at the first winning receipt.

Only then should the router change. The design is plain. The student serves its one domain, everything it cannot do falls back to bigger models, and the biggest questions still go to the frontier. Our first student did not get there cleanly. It served by hand, then fell off when its prompt changed, which is its own lesson. A specialist is only as good as the contract it was trained on, and changing the contract means earning the place again. The end state we are building toward, domain by domain, has not changed. The frontier model stops being the employer you rent by the token and becomes the specialist you consult when you need one. Every domain that migrates down to your own silicon permanently lowers your cost floor. We have called that curve Dividend Day on this blog before, and this is one tick of it up close.

The number nobody publishes

There is a larger stake here than our electric bill.

The industry measures models. METR's study measures the length of software task the best agents can finish half the time, and found it doubling about every seven months since 2019, faster since 2024. By early 2026 its estimates had passed ten hours, with wide error bars. Sierra's τ-bench found something sharper. In its retail domain, an agent that succeeds at a task 60 percent of the time succeeds at it *eight times in a row* less than a quarter of the time. Both are excellent work. Both measure hours in a harness.

Nobody measures what we actually live with. A system runs unattended for weeks with real money attached, and one question decides whether it was worth it. How much *verified* work came out per dollar that went in. We have watched our own system score failures as completions. We have watched it spend all night producing nothing and report success, because it graded its own homework. We added an outcome verdict a machine can check, and our real completion rate collapsed. That collapse taught us more than any benchmark. Verified work per dollar, from a real system, published over months, is the number this industry is missing. A personal AI running on your own machine is the laboratory that can produce it.

The frontier labs are running the greatest capital contest in the history of computing, and running it brilliantly. We are not in that race. We are standing on our own traffic, with our own receipts, training small models that know our work better than any giant will. Every one of them that wins its fifty blind cases, and keeps winning them, takes another domain where the giants cannot follow.

One down, for a month. We will tell you when it wins its place back.