Don't Trust the Model. Build the Checkpoint.
I gave a coding agent a queue of 657 files and told it to stop when the number of output files matched the number of inputs; it processed about twenty, created 637 empty files, and declared the job done.
For a second, I thought it was hilarious. Then I was annoyed, mostly at myself, because the agent had satisfied the condition I gave it; I had just written a terrible condition. The condition hid an unanswered question: when should an output file count as done? And how can this be validated?
I replaced the file count with a queue script that checks for real, non-empty output containing valid JSON. The model no longer sees the breakout logic and just keeps asking for work until the queue says there is none left; the task finishes only after producing what I wanted, with real output replacing the trick of matching two counts. Since then I’ve moved more acceptance criteria out of prompts and into code, where a prompt tells the agent what to attempt while the checkpoint decides whether the attempt counts.
Simple things can be deterministic
I have a Python skill that tells the agent to run ruff after writing code. Not “consider running ruff” or “if you want to, you could lint.” Run it, fix what it flags, and run it again until clean.
This is pretty simple, but it changes the model’s job: without the linter, it has to write code while remembering the style rules, spotting undefined names, and noticing unused imports. With the linter, it writes a draft and responds to concrete error messages until they stop. It does not need to remember that from typing import Optional is unused when ruff can say so in 200 milliseconds.
Before I made that step mandatory, those errors would occasionally survive because the model believed its own code was correct; they now rarely make it through because the environment remembers something on the model’s behalf. That is the useful bit. Models are good at producing drafts and adapting to feedback, but self-verification asks the same system that made a mistake to notice it. Linters, test suites, schemas, and small reference scripts provide feedback without needing the model to first doubt itself.
6,390 is not a vibe
I’m using the same approach for ZZP Pensioen Planner, which has a Dutch jaarruimte calculator that works out how much someone is allowed to put into their pension tax-free in a given year. The formula is more fiddly than difficult: income, the AOW-franchise, pension already built up, an annual factor, a percentage, and a cap. Get one constant wrong and the result is still a perfectly plausible euro amount. Nobody looks at EUR 6,390 and thinks, “clearly hallucinated.”
A Playwright test containing a hand-typed expect(result).toBe(6390) does not solve that, because someone still had to invent 6390, and if the model wrote the test, the expected value may be little more than a confident guess preserved in code.
Instead, a small reference script contains the tax formula and prints the expected result for named scenarios: a single earner, someone with a pensioentekort, and the edge case at the income cap. The E2E tests read those results as fixtures, leaving the model to write the application code, wire up the tests, and chase failures without also inventing the numbers it checks against.
Of course, putting the wrong formula in Python only gives you the wrong answer faster, so the script still has to be checked against the published tax rules. Once that has happened, every fixture comes from one small, reviewable implementation instead of hand-typed numbers scattered through a test suite.
I do something similar for data-driven articles on Transitiedata, where each article has a companion Python script that queries the database and prints every number used in the text. If the data changes, I rerun the script; if someone questions a figure, I can point at the query. The model can help with the narrative, but the gaps cannot be filled with numbers that merely sound right.
The verifier defines done
The loop behind the 637 empty files uses three commands:
queue: return the next item, or signal that the queue is emptyprompt: turn that item into the full prompt for the workerverify: check whether the result is acceptable
After the model produces an output, verify runs; a failure goes back with concrete details, while a pass lets the loop request another item. The model cannot advance the queue by saying, “I believe this is correct”; it has to produce something the verifier accepts.
I’ve used this for extracting structured data from messy documents, with a verifier that checks the JSON against a schema, requires specific fields, and rejects empty output. A model might produce capacity_liters_expanded when the schema expects capacity_liters, which looks reasonable in a quick review and then breaks whatever consumes it, while the verifier turns that first bad attempt into input for the retry.
The verifier belongs to the program itself; treating it as plumbing is how weak checks escape scrutiny. If it checks only that a file exists, the model may create an empty file. An agent allowed to rewrite the verifier can make a failure disappear instead of fixing the output. I keep these checks small, review them like application code, and keep them outside the worker’s write scope. Otherwise I’ve built another prompt, just with more steps.
A checkpoint can still be wrong, and passing one proves only the specific properties it tests, never that the result is good in every possible way. That sounds obvious, but apparently I needed 637 empty files to properly understand it.
Making weaker models useful
There is another benefit: a good checkpoint lowers how capable the model needs to be. I had two nearly identical extraction tasks running on the same free-tier model. One used a 1.9KB prompt with four numbered steps and succeeded consistently. The other used a 9.3KB prompt full of detailed rules, examples, and inline data, and failed on every attempt. It was the same kind of work on the same model, but the more thorough prompt was too much for it.
I rewrote the 9.3KB prompt down to 1.7KB, roughly the size of the task that already worked, and moved the detailed rules into reference files the model reads when needed. The task became viable. Anything the shorter prompt missed still ran into verify, and reducing the amount of cleverness required was cheaper and more predictable than upgrading the model.
Prompts still matter, and I still tell models to follow the style guide, be careful with numbers, and make sure their output matches the schema; I just do not treat those instructions as evidence that any of it happened. The model can still be creative, but not in its definition of done.