Designing a pre-review agent a sceptical accountant can check
Concept and technical design of an agent that would pre-review audited financial statements for a multilateral development bank. It would extract, reconcile and draft; a named specialist decides. Nothing has been built.
Context
A multilateral development bank lends to public bodies for infrastructure. To keep disbursing on a loan, it requires each executing agency to submit audited financial statements every year and progress reports twice a year. A handful of financial specialists read them, extract the figures, check them against the bank's own records and write a supervision report.
The documents arrive in a seasonal wave. As the problem was described to us, findings could take up to three months to reach the people deciding on the next disbursement and, as we understood it, the peak left little room for exhaustive cross-checks. Temporary reviewers would help for a season and leave no lasting capacity.
The challenge
We designed for difficult files without having seen any. The brief assumes native PDFs next to skewed photocopies, stamps over figures, pages out of order and statements of hundreds of pages with the relevant notes scattered through them. It expects tables broken across pages, line items that differ from the bank's taxonomy and more than one currency.
The user is more difficult, and should be. The reviewer is an accountant whose job is professional scepticism. A tool that shows numbers without showing where they came from gets re-checked line by line, which saves nothing, or gets trusted, which is worse. Fiduciary responsibility cannot move to a model.
Two constraints came from the institution. Earlier staff-built AI tools had been paused for months while they were brought into line with IT security policy, so security requirements came first. And the system would sit nearly idle for nine months of the year and run flat out for three.
Our approach
We treated this as a control-design problem first and a model problem second. One rule runs through the design: the language model reads, deterministic code counts, a person decides.
- 1Start from the controlWe started from what the specialist checks today and what a late finding means for the next disbursement. Source: one client-side sponsor's account, not observed files.
- 2Design the review screen firstEvery extracted value links to its exact place in the source document. The brief makes this mandatory.
- 3Size for the seasonA volume model sized the annual load and the peak. The architecture, as specified, queues documents, scales out in filing season and falls to near-zero cost afterwards.
- 4Phase it so it can fail cheaplyA short discovery, a one-month proof of concept on extraction scored per field, then a one-month MVP adding reconciliation, drafting and audit logs.
What was designed
The output is a technical brief for the engineers who would build the agent: an internal scope draft, summarised as one module of a wider proposal to the bank. It specifies five modules.
- Document clean-up. Automatic rotation, de-skewing, contrast correction and OCR before any language model sees a page.
- Accounting extraction. Tables rebuilt across page breaks, line items mapped to the bank's taxonomy, currencies detected explicitly. The audit opinion is classified; the notes are searched for defined risk signals such as litigation or covenant breaches.
- Cross-check. Extracted figures are reconciled with the bank's ledger within a stated tolerance and flagged. The statement's own arithmetic and fiscal year are verified too.
- Drafting. A supervision report on the bank's template and a management letter, both as drafts for a specialist to edit.
- Human review. A split screen with click-to-source on every field. Reviewers can overwrite any value; each change is logged as corrected by a human.
Cloud, infrastructure and model were left open, to be chosen per task by benchmark within the bank's IT requirements.
The same pre-review pattern was sketched, far more thinly, for procurement checks and for environmental and social safeguards. Those sketches are not part of this case.
Results
One sponsor's account, given in conversation before any measurement. We have not verified it.
Sizing assumption: annual statements and one round of progress reports land in the same quarter. Not a measured volume.
Target for the queue-based architecture in the technical brief. Untested.
Requirement from our later internal review of the brief, which had first asked for 50 or more files per document type.
There are no outcome figures in this case, because there are no outcomes. We did not model hours saved, full-time equivalents or return on investment: without a measured baseline of review time per file, any such number would be invented.
The proposal did state a direction: a review cycle moving from months to days. One version of our own material said hours. No baseline stood behind either, so neither appears among the figures above. Our later internal review flagged wording that commits to outcomes before anything is measured, and a working name that called the agent an auditor. The agent audits nothing; it prepares a file for a person to check.
That review changed our position in three ways. An adverse opinion was an automatic block in the first brief; it is now a recommendation that a person confirms. The brief's illustrative reconciliation tolerance of about 1% is gone: we fix no tolerance or accuracy threshold before measuring the baseline. And a firm quote waits for a stratified sample of real files.
What remains unproven is most of what matters: OCR quality on the bank's real scans, access to the ledger, and whether per-field accuracy is high enough that review takes less time than reading from scratch. The plan is optimistic too: a month each for proof of concept and MVP assumes every input arrives on time, and the wider proposal allowed four months for the same module.
Governance and risk
A named person decides. In the design, the agent recommends accepting or rejecting a statement and shows why. A specialist accepts, corrects or overrides, and approves. Flags inform that decision; they never replace it.
Everything is traceable. The specified audit log records who uploaded a file, what the model extracted, what a human changed, who approved and when.
Data stays inside the perimeter. The brief requires isolated tenancy and zero retention with any external model provider. We told the bank's IT function that client data would be used at inference only, not for training. IT, security and risk set their requirements before design choices are made. The proposal commits the code to the bank through technology transfer.
Known failure modes are designed for. Currency differences raise an alert instead of being silently normalised. A statement for the wrong fiscal year, or one whose totals do not add up, is flagged in its own right.
What we learned
- Design the review screen before the model. If a reviewer cannot click from a figure to its source, they re-check everything and the time saving disappears.
- Let language models read and deterministic code count. A reconciliation with an explicit tolerance gives the same flag every time, which is what an auditor asks for.
- Set no accuracy or tolerance threshold before measuring the baseline, including how often human reviewers disagree.
Client identity, locations and identifying details are withheld under confidentiality. Figures are rounded. Measured figures come from engagement records; client-reported figures are attributed, not audited; modelled figures are projections from the engagement's business case and are labelled as such.
