Interrogate the plan
Before we assess anything, we establish what you are actually trying to build and why. That conversation sets the scope.
AI Factory is a four-week engagement that finds out where the failure in your agentic workflow actually sits, and builds the system that closes it.
Book an assessment callA coding agent's right to work unattended comes from one thing only: the quality of its oracle, the check that decides whether the work was done correctly. An oracle only grants autonomy when it is cheap, frequent, and unfakeable at the same time. Miss any one of the three and it cannot be trusted, no matter how good the model behind it is.
Most oracles fail quietly. The agent writes both the code and the proof that the code is correct.
The check passes on a version of the output the end user never sees. The process exits clean — but nothing actually ran.
A watchdog proves its own fix using a canary it launched itself. None of these show up as a crash — only as a production incident months later.
We audit your repository, and your working practice, against an eleven-pillar, five-level readiness model. The repository is scored at its weakest pillar, not its average, because that is where agents fail first.
The eleven pillars we assess:
A score for the repository as a whole, anchored on its weakest pillar.
For one specific class of change: may it update your main branch unattended, and under which limits.
We do not hand you a number and leave. We test your oracles with a negative control (disable the behaviour, run the checks, confirm the right ones go red) and design the nine landing gates that decide whether a change earns unattended release — from the check itself, to who controls it, to the separation between landing code and exposing it to users.
Before we assess anything, we establish what you are actually trying to build and why. That conversation sets the scope.
Every finding comes from the repository itself, not from an interview about it. We assemble the maturity score and the landing decision for one change class.
One class of change, one baseline, one metric with its direction written down, and an append-only log where a failure is a recorded row, not a deleted one. No class of change earns unattended landing before a pilot loop has run and the results have survived being read by someone else.
We size the change limit from your own repository's history, not a number from a book, install the rules where your agents will actually read them, and hand over a quarterly reassessment cadence.
On the classes of change that clear the gates, routine delivery stops waiting on a human review queue. The expensive part of routine work was never the typing. It was the waiting: for a reviewer, an environment, or someone who remembers why a module is shaped the way it is.
In similar engagements, clients have seen faster delivery, lower review and rework overhead, and fewer production issues on the change classes that cleared the gates. The exact impact varies by client and starting maturity, which is exactly what week three's pilot loop is for: it gives you your own measured baseline, on your own repository, before anything is allowed to land unattended.
This is not a promise of full autonomy. The honest outcome of an assessment is often "this one class of change may land unattended, and nothing else may yet." That is a useful, and safer, place to start.
Experience from real organisational consulting, applied to a new tool.
Tell us where to send the proposal. We reply within two working days with a scoped outline of the four weeks.