Essay · September 20, 2026

Craft & engineering

Steering by intention, closing the loop

Delegation became real before validation caught up.

Under Ezkey, the real shift was not just better agents, but a more concrete mode of delegation that immediately exposed the need for explicit completion criteria.

The essential point

Under Ezkey, I recently experienced a real step up in capability. Not only because the bots are better, but because delegation became concrete: several specialized collaborators, each with its own memory, context, and working environment. For the first time, I could steer most of a task primarily through intention, even while travelling, from my phone rather than from my usual worktrees.

But that step up exposed its weak point immediately. A multi-agent system can analyze, split work, execute, and even perform a quality pass without actually knowing how to close the loop. As long as what counts as complete is not explicit, observable, and enforced in the cycle, the result can look good without truly being finished.

A real elevation, not just an impression

My normal mode, until now, remained tied to the workstation: one or several worktrees, the IDE open, constant back and forth, and an orchestrator role still fairly close to classic pair programming. What changed recently is that I decided to go all in on the experience with my bot team.

One of the things that made it credible is that each bot had its own context and its own execution environment, concretely separate Docker containers. At that point, delegation stops being a metaphor. It is no longer just a well-structured conversation. Work can actually split apart, move forward in parallel, and come back to a coordinator who keeps the overall direction intact.

I was travelling when I launched this experiment. A large part of the follow-up happened from the mobile app, with Grok Bot also open on my workstation. I steered the work through intention far more than usual. The real functional check on my PC only happened at the very end. For me, that was a real step up in the way I interact with software work.

The limit showed up exactly where it mattered

On the main piece of work, everything seemed to be going well. The task was the intentional introduction of Ezkey's opt-in base mode, where cryptographic keys are not introduced systematically, heartbeat is disabled, and the cryptographic integrity of checkpoints is turned off. The analysis happened, the work was distributed, the relevant bots produced what they were supposed to produce, and there was even a quality pass.

That is exactly where the real problem appeared.

At validation time, the check stopped at something like: this looks correct. But a feature that removes cryptographic guardrails is exactly the kind of work where that impression becomes dangerous. The gap was deeper than that: I had not defined clearly enough what needed to be observed in order to declare, credibly, that the work was truly complete.

In other words, the system knew how to move forward. It did not yet know how to close the loop.

Knowing is not enough

What makes this interesting to me is that I am not coming to these subjects from scratch. I have already written about the importance of testing. In my return to foundations, I already argued that it is not enough to have tests in the abstract; you need to know which ones to run, at what level, and what they should actually verify. I have also written about moving from code to specifications, then from specifications to intention, precisely to give evolving models a compass that sits above the immediate code.

I have also argued, in Giving AI back its wings, that we should not clip the probabilistic wings of the model. In other words: avoid micromanagement, keep paths open, and provide project values that help the system make better decisions.

I already knew all of that. And yet, in this new distributed mode of work, that knowledge did not translate into reliable behavior on its own.

That is the lesson. Knowing the principles is not the same thing as making them operational inside the system itself. A compass helps you choose a direction. It does not replace observable criteria that let you conclude the destination has actually been reached.

The correction was not more control, but a better frame

The right reading was simpler: what is missing in the system for distributed work to finish well without all my follow-up nudges?

So I ran a small retrospective with my coordinating agent, my generic right hand. Together, we looked at what had been missing: a more explicit definition of finished, more observable validation expectations, and a firmer way to impose those guardrails in the normal work cycle. The correction was then injected in two places: the documentation corpus, to state the rule, and the task descriptions of the relevant bots, so it would become a working reflex rather than an improvised reminder.

That first correction was about depth: being able to say the work is finished because it was observed, not because it looked right. The same piece of work, once actually tested, showed the other facet. Base mode turns processing off. A key rotation that can no longer complete must then be refused by the API, and the interface must stop offering it. Until the process, the API surface, and the screen tell the same story, the work is not complete. A software factory becomes more autonomous when it carries that kind of coherence, the blast radius of a choice, into its own rules.

For me, that may be the most interesting point of the whole experience. An agentic software factory is not just a team of bots that can talk to each other. It starts becoming serious when the work system learns about itself and corrects its own rules.

If I had to reduce the lesson to one line, it would not be: bots can collaborate now. It would be: work steered by intention only becomes reliable once you formalize what finished actually means.