blog
What we measured
Notes on building small models that do one job: how they compare to the big ones, what they cost to run, and where they fall short. Numbers from our own test runs, including the ones that did not flatter us.
September 24, 2026 · 6 min
Building a custom agent harness with Pi and Decider 1
A runnable Pi harness where Decider 1 makes three calls on every run: which model handles the request, whether a shell command is safe to run, and whether the answer is finished. Real outputs, under a second per decision, and the limits we hit.
Read the post
September 22, 2026 · 5 min
Decider 1: typed decisions, more accurately than Jev, for less
Decider 1 answers typed questions over any state, each as a calibrated distribution, all in one call. On the typed-decisions benchmark it scores 0.768 against TypeSafe Jev’s 0.727, at $0.03 per million tokens against $0.042, and the typesafe-sdk works against it unchanged.
Read the post
September 14, 2026 · 5 min
What a narrow model actually saves you
Frontier models got cheap enough that "small models save money" stopped being obviously true. We did the arithmetic against current prices. At a hundred calls the saving is pennies; the argument that survives is a different one.
Read the post
September 13, 2026 · 6 min
What AI assistants actually search before they answer
We recorded every web search ChatGPT, Claude and Gemini ran before answering buying questions. They rarely search for the same thing twice, one of them mostly does not search at all, and half of ChatGPT’s searches go looking for a price.
Read the post
August 31, 2026 · 4 min
How Restyler 1 compares to Haiku, Flash Lite and Luna
Claude Haiku 4.5, Gemini 3.5 Flash Lite and GPT-5.6 Luna all score around 0.90 on a test where 0.50 means indistinguishable from a person. Restyler 1 brings the finished text down to 0.65, including on a model it was never trained on.
Read the post
August 31, 2026 · 4 min
How meraGEO uses Restyler 1 to rewrite customer pages
Answer engines read your site before a person does, and machine-written copy reads like machine-written copy. meraGEO finds those pages and rewrites them in place, keeping the structure, the links and the facts.
Read the post