Synlian — Quantum age applications for digital age supply chains

Insights · 2026-09-16 · 6 min read

The optimizer is delegated. The objective is not.

How we let our agents improve the mechanisms by which they improve SQF — under a written constitution, a ruler they cannot edit, and a ledger no human writes — and what the first run found.

HUMANOBJECTIVE+ constraintsIMMUTABLE COREthe loop cannot reach itRSI OPTIMIZERhypothesizeexperimentimplementevaluatelearnimprove optimizerMUTABLE SYSTEMcodeagentstoolstestspromptstopologyREALITY · OUTCOMESFITNESS DATAFROZEN INSTRUMENTS — not writable by the optimizer

Every team that puts AI agents to work on a codebase arrives at the same wall. The agents are capable, the backlog of things to test and improve is effectively infinite, and the one human who understands the whole system becomes the bottleneck — first for review, then for comprehension. At Synlian that human is one person, and the platform he architected has outgrown what he can hold in his head. He said so himself, and then said the more important thing: that this is fine, and that his limits should not be the limit.

This piece describes what we built in response. We call it a recursive self-improvement contract. It is not a licence for agents to rewrite themselves. It is a controlled loop in which the system may improve the mechanisms by which it improves itself, inside a boundary that no optimizer can move.

Separate the objective from the optimizer

The whole design rests on one sentence: the objective is human-owned, the optimization is delegated. The agents may change how they pursue the objective — the code, the architecture, the agent topology, the prompts, the tests, the evaluation methods, even the loop's own search strategy. They may never change what constitutes success.

That boundary is written down as three layers. An immutable core holds the objective, the ranking of principles (regulatory compliance first, then customer fairness, then auditability, then reliability, then cost, then speed), human authority including a one-file kill switch, and the safety constraints: no real money moves, no real recipient is emailed, nothing is deployed to production, no stored customer or ledger data is rewritten. A ratified layer holds everything the system may propose but never decide alone: the meaning of money, credit and identity data; cross-service contracts; the permission model; new external dependencies; and the contract's own mechanics. Everything else is freely mutable, with proof.

The contract has one clause about itself that took us the longest to get right. The system may generate, test in isolation and recommend amendments to its own governance. It cannot make them authoritative. That distinction is what turns "the contract cannot amend itself" into something more useful: improvement of the improvement constraints, with a human signature on every change.

An optimizer that can edit its own ruler will shorten it

Proof of correctness is not proof of improvement, so every retained change must carry a measurable hypothesis and a fitness measurement: what was expected, what was observed, the delta net of cost and complexity. But the moment you measure, you create the second problem. An optimizer that can modify its own evaluations will, sooner or later, improve its score by changing the scoring system.

So the ruler is separated from the hand that holds it. Frozen instruments — the end-to-end suite wall, the asserted invariants, replay sets of real cases with their expected verdicts — live outside the optimizer's write path and decide what is retained. Candidate instruments can be created freely but measure nothing authoritative until two keys turn: agreement with the frozen set over a window, and human ratification. And reality outranks both. When an outcome arrives from the lending book months later — an invoice paid or defaulted, a facility performing or not — it calibrates every proxy, and a proxy that disagrees with outcomes is the thing that is wrong.

We were honest with ourselves about the shape of that signal. In supply-chain finance, outcomes are slow and thin. Reality is the right teacher, but it teaches on a quarterly timetable, so proxies carry the loop until it speaks, and they are labelled provisional the whole time.

Three ways it goes wrong even when it works

Our architect named three risks before we started, and each became a mechanism rather than a caveat.

The local-maxima trap: an optimizer will build a superb horse carriage when what you needed was a car. So a fitness plateau does not trigger more polishing; it triggers a pivot review in which the system must write the strongest case against its own architecture, as a proposal a human reads. Non-linear pivots are a human right, not an optimizer output.

The ledger bottleneck: tracking every mutation, baseline and delta is a data problem, and if a person has to parse it, they are the bottleneck again. So no human ever writes a ledger entry. Entries are machine-written and schema-validated, with the model versions and agent definitions behind every number, and the daily brief is a query over them with a ten-minute reading budget. If the brief overruns, the loop is throttled, not the reader.

The comprehension bottleneck: an evolved agent topology can become something no architect understands, trading a development bottleneck for a worse one. So complexity is a budgeted cost subtracted from every fitness delta, every retained structural change ships with a rationale a person can read, and once a week a context-free agent is handed only the system's self-model and asked to operate one cycle. If it cannot, comprehension regressed, and the last structural changes are the suspects.

What the first run found

Stage zero ran by hand, with nothing merging autonomously, on the day the contract was ratified. The wall of forty-nine end-to-end suites ran against real infrastructure: forty-six green, nine hundred and eighty-one checks, three failures that were all stale tests rather than product defects. Two new invariants that had lived only in prose became build-failing checks, and the first pass of one of them mostly found its own blind spots — eight of twenty-four candidates were flaws in the instrument, not the code. That is now a rule: a new invariant runs report-only for a cycle before its findings count.

Then eight review angles read the day's own changes and returned forty-two candidates, deduplicated to eight confirmed findings. Two of them were real holes in a matching rule we had shipped hours earlier — a shared lot number that could pass for a matching postcode, and an address that would be flagged for being in the wrong city when it was exactly right. Both were fixed the same day, and both are now frozen cases the optimizer cannot edit.

The most useful thing the run produced was not a fix. It was a decision. An invariant found five routes that any signed-in session could read without a permission key; the brief recommended keying them; and when an agent went to do it, it found a test asserting those routes were open on purpose, for another agent's tool loop. The contract's rule for that moment is plain: a conflict between a claim and the code is surfaced, never resolved by picking a side. The human chose to key the routes and give the agent a proper identity, and the ruling landed as an architecture decision record the same day.

What the architect does now

Reading code stops being the job. Owning the value system, the contract and the calibration decisions becomes the job, along with the direction nobody can delegate: what the platform is for. The reading surface is a set of architecture decision records, a page of standing rules, and a brief that begins with the decisions only a human can make, each one phrased as a question with options and a recommendation.

We do not know yet whether this loop will produce the kind of compounding improvement the phrase "recursive self-improvement" promises. We know that it measures whether it is improving, that it cannot move its own goalposts, and that the person accountable for it can read what it did in ten minutes a day. That seemed like the right place to start.

See it rather than read about it.

Everything in this essay can be walked through live, on demonstration data.

Book a briefing