Governed agent operations, tried first on our own back office
Before proposing governed agent operations to clients, we built the pattern for ourselves: a command gateway, an append-only record and a close, tested by replaying our own invoicing history with no external effects.
Context
AI Workify advises organisations on putting agents to work inside real processes. In regulated sectors the first question is rarely whether an agent can do the task. It is who answers for the action, what the agent was allowed to touch, and how anyone would reconstruct the event later. We decided to answer that on ourselves before proposing an answer to anyone else.
The vehicle is an internal R&D programme we call the Agentic Operations OS. It starts from one question: if capable agents can read, reason, write, audit and use tools, what form should a company take so that they can operate it with freedom, precision, safety and memory? One rule kept the work honest: we would not rebuild an ERP under another name.
Our own back office was the natural first subject: invoicing, the records behind it and the reconciliation that closes them. We own every record in it, and mistakes made while learning land on us and not on a client.
The challenge
Agents fail in a back office in familiar ways. One writes straight to the books and nobody can later say on what evidence. A figure appears in a summary that no source document contains. An instruction hidden in an incoming document tries to redirect a payment. A procedure is quietly reworded so that an action no longer needs approval. Each is an ordinary control failure, arriving faster.
Two constraints shaped the build. The design requires a rehearsal before anything touches a real counterparty, so the first build performs no external side effects: no emails, no payment calls, no bank connections. And the rules had to be legible to people and to agents alike, because a policy an agent cannot read, or a person cannot audit, is not a control.
Our approach
We worked from the rules outwards, and treated the documentation as part of the system, not as a description of it.
- 1Write the rules as claimsEach design position was logged as a claim and argued through written change proposals. Accepted claims are compiled into the constitution the system operates under.
- 2Build the invariant core onlyA gateway, an event store, validators, approval tiers and the close. Sales, marketing and delivery settlement were deliberately left out.
- 3Replay our own historyThe firm's real invoicing history was replayed through the gateway as governed events, then questioned from the store alone.
- 4Drill the failuresSeven scenarios were run against the core: partial issuance, double-numbering, invented evidence, an injected bank redirect, procedure laundering, a corrupted artefact and a restore.
- 5Invite attackFive reviewer agents attacked the build adversarially. Two real violations were confirmed, and both were fixed behind regression tests.
What was built
The substrate is deliberately plain: written against the Python standard library alone, with one append-only event store per company. The design is agent-first, not screen-first.
- Command gateway. The sole authoritative writer. Agents propose commands; only the gateway writes, and it stamps the actor's identity itself instead of trusting what the agent declares.
- Validation pipeline. Twelve stages, among them schema, scope, vocabulary, evidence resolution, groundedness, procedure, budgets, idempotency, approval and operating mode.
- Governed records. Evidence is kept as classified, content-addressed artefacts, so a changed file is detectable. Secrets are referred to by logical name only.
- Approval tiers and modes. Every command carries a tier, and the higher tiers wait for explicit approval. The company runs in normal, heightened or frozen mode.
- The close. A reconciliation loop with an entailment audit: an amount asserted in the record must follow from cited evidence. A blocked close escalates the operating mode; it does not merely colour a dashboard.
- Constitution and claims register. The constitution in force is compiled from the claims register, so a rule traces back to the claim it came from. Norms sit in a two-level registry.
The operating loop is short. Frame the objective; do the work with ordinary tools; record it through the gateway as a one-off operation, a process or a task; keep the receipt; review the journal; run the close. Each record carries an actor, a timestamp, a type, a payload, evidence, a tier and a content hash. The rule for agents fits on one line: you propose, the gateway writes, and silence is never permission.
Companies on the substrate are isolated by construction. The design corpus itself runs as the first of them, governed by the rules it sets.
Results
The honest summary: the invariant core is built, our invoicing history has been replayed through it, and we run it first on our own back office, with no external effects. By its own documentation it is a rehearsal-grade substrate. It is not an autonomous finance function.
Rounded. Counts events, not invoices; no amounts or counterparties are published.
Found by five reviewer agents we set up ourselves; fixes sit behind regression tests.
More than 50 further claims are accepted but not yet codified. The register was argued through 20+ written change proposals.
A property of the build, not an achievement: no external connector exists yet.
The replay mattered more than the drills. Asked twenty questions about our own history, the system answered from the store with cited evidence, and said so where the record could not support an answer. Two real anomalies in that history were preserved as found, not cleaned.
We publish no time saving, cost figure or error rate. There is no measured baseline, so any such number would be invented. The counts above come from the build's own records and tests, and no independent party has re-run them.
Governance and risk
One principle carries most of the security weight: external systems are observation sources, never commanders. A statement or an inbound message can inform the record; it cannot instruct the system. The injected bank-redirect drill tests exactly that boundary.
The limits are plain. No external connector exists, so sending invoices and moving money stay outside the system, with people. The replay covered one process, invoicing. More than 50 accepted claims still await codification, so the constitution is incomplete. The adversarial review was run by agents we configured, not by a third party. We hold no certification for any of this and claim none.
What we learned
- Make one component the only writer. When agents can only propose and a gateway writes, every other control has somewhere to stand.
- Prove the record before adding reach. Replay, drills and adversarial review surfaced real flaws while nothing could leave the building.
- Keep the anomalies. A history replay that preserves real oddities is evidence; one that tidies them away is only a demonstration.
Client identity, locations and identifying details are withheld under confidentiality. Figures are rounded. Measured figures come from engagement records; client-reported figures are attributed, not audited; modelled figures are projections from the engagement's business case and are labelled as such.
