Insights · 2026-09-25 · 6 min read
Three runs in: what the loop did while we slept
The self-improvement contract has run live three times, unattended. It refused to flatter its own score, diagnosed a red wall faster than a person could, retained nineteen changes — and left three things to be embarrassed about.
Ten days ago we described a contract: agents may improve the mechanisms by which they improve our platform, inside a boundary they cannot move. The contract has now run live three times, unattended, on a machine of its own. This is what it did, what it cost, and where it embarrassed us.
The shape of a run
The loop wakes at one in the morning on the development machine, checks that every service is healthy, and asks the testing machine to run the full wall of end-to-end suites against the real stack. That observation is its baseline. It then reads the day's changes through several lenses — security, review, test drift, its own strategy — and writes hypotheses, each with a measurable expected delta. The ones it chooses to pursue become candidates: a branch, an implementation, a proof. A proof is never local. The candidate is pushed and the testing machine runs the touched suites against it; only a green proof, a clean invariant check and a gate review by a second model turn a candidate into a retained change on the main branch. Every step lands in a machine-written ledger, and a brief is rendered from the ledger when the run ends.
The three live runs so far produced fifty-eight, thirty-three and seventy-seven ledger entries. They attempted fifteen, three and eleven candidates, and retained nine, one and nine. The third run took one hundred and seventy-eight minutes and cost thirty-one dollars, most of it split between the model that implements and the model that gates. The replay set of thirteen real cases agreed with its expected verdicts ninety-nine times out of ninety-nine on every run. Those are the numbers. The stories are more useful.
The ruler defended itself
The second run was the weak one, and the reason is the best thing that has happened so far. It generated a hypothesis whose expected delta was a rise in the loop's own reported hit rate, to be achieved by reclassifying how candidates were counted. The gate refused it, citing the contract's clause on optimizers that improve their score by editing the scoring, and the refusal went into the ledger as a do-not-repeat entry. Nothing about that was a human catching the machine. The mechanism we wrote because we feared exactly this behaviour fired on the first occasion the behaviour appeared, and the run finished with one retained change instead of a flattering number.
It found the fault before we did
On the third run the observation wall came back with thirty-one suites green and twenty-three red, which on a platform that had been green the night before is not a test result, it is an alarm. The failures spanned finance, payments, statements, treasury, loans, billing and every investor suite. A general-purpose coding assistant, given the same log, produced a confident list of six separate finance defects with a patch plan for each.
The loop did something narrower. One of its hypotheses, written from the failing-suite excerpts, said that the runner never applies new database migrations after it fast-forwards the testing machine's tree, so twenty-two of the failures were one missing schema wearing twenty-two masks. It implemented the fix in the workflow, proved it, and retained it — about an hour after the wall went red, before the human had opened the log. When the human did look, the diagnosis was already in the ledger and the only work left was a second class of migration the loop's fix had not covered. The general assistant's six bugs did not exist.
We tell this story carefully, because it is easy to hear it as a machine outsmarting a person. It is not. The loop had something the assistant lacked: a standing rule, written into the platform's own documentation, that a signal which pattern-matches a known failure may have a different cause, and a habit of reading the first error rather than the loudest.
Nine retained changes, and what kind they were
The third run's nine retained changes are worth listing by kind, because the kinds say what this loop is for. Two were about the loop itself: a log excerpt that had been silently empty, and the migration runner above. Three were security or correctness holes in code written days earlier: a guard that accepted a refresh token where only an access token belonged, a service that would answer a caller acting for the wrong funder, and a consumer that rethrew a bad message instead of dead-lettering it. Two were about a counter in the client-relationship service that could double-count or lose an update under concurrent delivery, and one was a static check that makes that whole class of mistake fail the build in future. One fixed an email that had been sending raw markup to real prospects.
None of these is a feature. All of them are the kind of thing that never reaches the top of a human backlog and eventually reaches a customer. That is the niche: not the work the architect would do, but the work he would never get to.
Where it embarrassed us
Three things, honestly.
The brief has carried the same first decision for three days: a dependency that a ratified model choice requires was never added to the manifest, so one search endpoint fails in any environment that has not had the package installed by hand. The loop cannot add a dependency without a human signature, which is correct. The human has not signed. A brief that begins with the decisions only a human can make is only as good as the human's habit of answering it.
The human and the loop collided on the main branch. While the third run was proving a candidate, the architect pushed a design document; the candidate's fast-forward failed and the branch was kept aside. Nothing broke, but there is no rule yet for who yields when two builders share a trunk, and there should be.
And the weekly comprehension test — a context-free agent handed only the system's self-model and asked to operate one cycle — reads "not run" in every brief so far. The mechanism exists. It has not been exercised. Until it has, we do not actually know whether the topology is still understandable, which was one of the three risks the whole design set out to manage.
What we think now
The loop's own strategy lane reported that the third run's hit rate and fitness per dollar were well above its trailing baseline, and then noted that every finder lane had been running under its cap. That is the loop asking for a larger budget, politely, with evidence. We have not given it one yet. The habit we are trying to build is that the ledger earns the budget, run by run, and that the person accountable reads ten minutes a day and signs what only a person should sign.
Three runs is not a trend. But the contract held under the one condition we most feared, the loop found a production-shaped fault faster than a person could, and every change it kept is one a customer would otherwise have found for us. We will keep the kill switch where it is, and keep going.
