It is day four of month-end close. The reconciliation agent has run overnight and cleared 1,912 of 2,104 intercompany items, routing the remainder to a queue. The controller is pleased; four days of clerical work has become one.
Two months later the external auditor picks a single line — item 1,447 — and asks a simple question: what did the agent see, what did it decide, and who reviewed that decision?
Nobody can answer. The prompt was assembled at runtime and never stored. The model version was upgraded by the vendor mid-quarter. The only surviving artefact is the journal entry itself. Three weeks of manual re-performance follow, and the agent is quietly retired to a sandbox.
This is the shape of most agentic AI failure in the finance function — and it is not a technology failure.
| KEY TAKEAWAYS • An AI agent that reconciles, flags or posts is not a tool supporting a control. It is the control — which puts it in scope for ICFR and ITGC testing. • The AI agent audit trail is the binding constraint, not model accuracy. An agent at 99.4% that can't show its working is worth less than one at 96% that can. • COSO (February 2026), the amended PCAOB standards AS 2201 and AS 2101, and the EU AI Act now converge on the same requirement: reconstructable prompts, inputs, outputs, model version and human review. • Trust is the real adoption gap. Only 14% of mid-market finance leaders fully trust AI outputs even after review. • Start with a process that has a reconciliation point — AP, bank reconciliation or three-way match — not forecast commentary. • Private mid-market firms are not exempt. The same question arrives via statutory audit, lender covenants, PE reporting and acquirer diligence. |
| • | 54% of CFOs at $1B-plus companies rank integrating AI agents into finance among their top transformation priorities for 2026 — ahead of improving data quality itself (Deloitte, Q4 CFO Signals). |
|---|---|
| • | 56% of finance leaders now use AI tools in daily work, roughly double the 2023 figure (State of AI in Finance 2026). |
| • | Only 17% run AI inside core workflows, with 45% still in limited pilot mode (General Atlantic). |
| • | Capability, not budget. A Gartner survey in early 2026 found CFOs naming capability-building inside finance — not technology, not cost — as their most pressing near-term constraint. |
The conventional reading is that finance is conservative and slow-moving. That explanation is comfortable and mostly wrong. Engineering, marketing and customer success all reached full workflow automation without waiting for finance-grade rigour — because nobody audits a campaign brief.
Finance is not behind on AI because finance is cautious. It is behind because finance is the only function whose output has to survive an audit.
Agentic AI vs RPA in finance: what actually changed
The terms are used loosely enough to be useless in a procurement conversation. The distinction that matters is what happens at the exception.
| Rules-based automation (RPA) | AI agent | |
| Handles the normal case | Yes, reliably | Yes |
| Handles the exception | Breaks, or escalates to a human queue | Reasons about it and acts |
| Reproducibility | Deterministic — same input, same path | Non-deterministic by default |
| Control testing | Standard ITGC approach works | Standard approach breaks down |
| Evidence produced | Log of rule executions | Only what you designed it to capture |
| • | It acts rather than suggests. An assistant drafts variance commentary for a human to accept. An agent posts to the ledger, applies cash, or releases a match — it writes to systems of record. |
|---|---|
| • | It is bounded to a process, not a department. "An agent for finance" is not a scope. "An agent for unapplied cash under £50,000" is. |
| • | It is non-deterministic by default. Given identical inputs on two runs, a language-model-driven agent may take different paths — or reach a different answer. This is the property that breaks traditional control testing. |
| • | It is consequential. Its output changes a number that a lender, a board, an acquirer or a shareholder relies on. |
Is an AI agent a SOX control?
If a system flags unusual journal entries, automates reconciliations, or runs anomaly detection over financial data, it is not a tool that supports a control. It is the control — which brings it into ICFR scope, with IT general controls tested over the underlying system.
| • | COSO (February 2026) — internal control guidance for generative AI requires a complete audit trail: prompts, inputs, outputs, model and configuration versions, and evidence of human review, sufficient to reconstruct what the AI acted on. |
|---|---|
| • | PCAOB AS 2201 and AS 2101 — amended standards apply to audits of fiscal years beginning on or after 15 December 2026. ITGCs are tested first; if they fail, reliance on everything above them collapses. |
| • | EU AI Act — most high-risk obligations become enforceable from August 2026: risk management, technical documentation, logging, post-market monitoring. |
Why reproducibility beats accuracy
The binding constraint on a finance agent is not how often it is right. It is whether you can prove, twelve months later, how it got there. An agent that is 99.4% accurate and cannot show its working is worth less than one at 96% that can. Almost every struggling deployment we see has optimised for the first number.
"We're private — does this apply to us?"
Largely, yes. Most mid-market firms are not SOX registrants, but the same reconstruction question arrives through other doors: statutory audit, PE sponsor and lender reporting, acquirer diligence — where an unexplainable control is a valuation discount — and the EU AI Act, which applies regardless of listing status if you sell into the EU.
Can finance teams trust AI outputs?
Not yet, on the evidence. A 2026 Maximor survey of 100 middle-market finance leaders found 66% consider human oversight of agentic AI extremely or very critical; more than four in five had encountered a hallucination in a finance context; and only 14% completely trust AI outputs even after human review.
Read that last figure carefully. Review is happening and it still isn't producing trust — because an unstructured glance at an output tells a reviewer nothing about how it was reached. Trust is not a function of accuracy. It is a function of traceability.
Which finance process should the first agent handle?
Choose a process with a natural reconciliation point — somewhere ground truth exists and you can prove the agent right or wrong without opinion.
| Process | Ground truth? | Verdict as a first deployment |
| AP / invoice matching | Yes — PO and receipt | Best starting point. High volume, contained blast radius |
| Bank reconciliation | Yes — statement | Strong. Clean pass/fail measurement |
| Intercompany reconciliation | Yes — counterparty ledger | Strong, if entity data is clean |
| Cash application | Yes — remittance | Good, though remittance quality varies |
| FP&A narrative / forecast commentary | No | Avoid first. Impressive demo, unprovable output |
| 1. | Choose a process with a natural reconciliation point. Ground truth is non-negotiable for a first build. |
|---|---|
| 2. | Design the evidence schema before the agent. All thirteen fields, agreed externally, before build starts. |
| 3. | Pin the model version and treat upgrades as releases. A vendor-side upgrade mid-quarter invalidates the control unless detected, re-validated and documented. Version drift is the most common cause of a control being written off retrospectively. |
| 4. | Make human review a defined control step, not a courtesy. "A human checks the output" is not a control. "All items above £25,000 plus a 10% random sample below it are reviewed by the financial controller within two working days" is. |
| 5. | Map each touchpoint to a financial assertion. Existence, completeness, valuation, rights and obligations, presentation and disclosure — recording whether the touchpoint is in or out of ICFR scope, with the rationale. |
| 6. | Retain logs for a full audit cycle. At least 366 days for anything touching financial reporting. |
| 7. | Rehearse the auditor conversation before go-live. Reconstruct a random transaction end to end. If that takes more than ten minutes, the agent is not ready for the close, whatever the accuracy dashboard says. |




