Skip to main content
Finaisse
All articles

Agentic AI

What agentic AI gets wrong about finance

Agentic AI is sold on autonomy: how much the software can do without you. In finance, the question that matters is who decides where the line sits, and whether you can see what passed under it.

By Kris Subramanian·August 2026·7 min read

The matcher solved the cheap part

A bank reconciliation comes out $60K short on 340 statement lines. The matching engine handles 328 of them without anyone looking. That's what matching engines have done well for years.

The morning goes on the other twelve.

$48K turns out to be timing, though establishing that means checking what cash application did after cut-off. $12K is bank charges nobody booked. The last item is $600 and takes longest, because the receipt was applied to the wrong customer three weeks ago and nothing in the reconciliation knows that.

Then the entry gets posted, which takes ten seconds.

The automation worked exactly as designed. It automated the part that was never expensive. What's left isn't a matching problem, and running the matcher harder won't touch it: the twelve exceptions are unmatched precisely because the reason sits in another system.

The line exists. The argument is about who draws it.

The usual argument doesn't help much here. One side promises software that acts on its own, the other insists on a human in the loop, and neither describes how a finance team actually works. Plenty of finance work should complete without anyone looking at it. Some of it should never complete without a named person, no matter how confident the software is.

A receipt that matches an open invoice exactly, on the right customer, for the right amount, doesn't need a person.

An account that reconciles to zero and sits well below materiality doesn't need a controller to open it.

Insisting otherwise isn't control. It's ceremony — and it's how teams end up spending close week on things that were never going to matter.

The interesting question was never whether the software acts, but who sets the boundary — and whether they can see across it.

A vendor who decides that for you has taken something that belongs to your controller.

Three tests, not one

Where an item lands comes down to three things, and materiality alone isn't enough.

Materiality. Is it big enough to matter? A $400 variance and a $400K variance are not the same object. Treating them identically is how a checklist turns a trivial item into a queue position.

Confidence. Is the answer clear? A receipt matching one open invoice exactly is a match. A receipt that plausibly fits two invoices is a question. Ambiguity is a reason to escalate, not something the software should resolve on your behalf.

Class. Is this an operational action or a control point?

This is the test most discussions miss.

Applying a receipt is reversible: unapply it, reapply it, and nothing reaches the ledger. Posting a journal is different. It becomes part of the accounting record and can only be corrected through another entry.

Volume work that is reversible and unambiguous should complete on its own. The things a named person signs for shouldn't — unless the organisation has explicitly decided they can.

Identifying a deduction and coding its reason is operational. Writing one off is a control point, and it has an approval matrix behind it for good reasons.

And where something does complete on its own, the three tests aren't the last word. Validation is the gate.

Auto-posting isn't a model deciding it feels confident enough. It's a set of deterministic checks that pass or fail: the entry balances, the account and period are open, the amount sits inside the threshold you set, the support is attached, no policy exception applies. Any check fails and it doesn't post. Confidence decides whether the work is worth preparing. Validation decides whether it can complete without you.

Why “doesn't post” is the wrong promise

Which is also why any vendor telling you their agent never acts is either not describing their product accurately or has built something that will annoy your team by month three.

The useful commitment isn't that the software stays out of the ledger, but that the thresholds belong to you.

They can differ by account class and entity. Changing them is logged. And nothing crosses a control point that you haven't said can be crossed.

The goal isn't maximum autonomy. It's autonomy you can audit.

Automatic shouldn't mean invisible

Setting policies is the easy half. The claim only means something if you can inspect what they let through.

Knowing that a journal was posted correctly is the easy part. Every system can show you that.

The harder question, and the one an auditor actually asks when sampling automated work, is different:

Why didn't I see this one?

That requires the routing decision to be as traceable as the outcome.

Not simply a note that the system handled it, but the rule that applied, the confidence at the time, the threshold in force, and the policy that allowed it.

And crucially, those values have to be stamped onto the decision rather than looked up afterwards. A threshold raised in March tells you nothing about why something completed in January.

When that holds, “the system handled it” stops being an assertion and becomes a record. When it doesn't, the policy is a slide.

An agent can only prepare what it can see

The other constraint is context, and it decides whether any of the preparation is worth reading.

Go back to the $600 item. An agent with the ledger and the bank statement can tell you it doesn't match. It cannot tell you why, because the reason is a misapplied receipt in another system, and it has no way of knowing that collections has been chasing that customer for three weeks or that the item is holding up a close dependency.

It hasn't made a mistake. It prepared exactly what its evidence supported, and its evidence was a fraction of the situation.

Confident work built on a partial view is more dangerous than no work at all, because it arrives looking finished.

The same shape, everywhere

The pattern holds across the function.

In reconciliation, the matcher clears the routine volume and the agent works the residual: finds the cause, explains it, and drafts the adjustment. Below threshold on a low-risk account it can complete and leave a record; above it, a person approves.

In collections, the agent sees exposure change, assesses the accounts underneath, prioritises by what actually matters and prepares the recommended action. A person decides how to play it.

In close, the agent detects a blocked dependency, identifies what's blocking it, works out the downstream impact and prepares the next step. A person owns the call.

In each case the agent is preparing a decision rather than simply completing a task. That's a harder thing to do, and worth considerably more.

What this looks like in practice

This is the principle Finni is built on.

It assembles the work behind a decision and brings the evidence and the context with it. Where your policy permits it and every validation passes, it completes. Where either doesn't hold, it stops.

Everything else arrives finished and reviewable, with the support attached.

And everything that completed without you stays visible, with the reason it never needed you.

The measure

The wrong question is how much the agent can do without a person. The better question is whether it can explain the twelve items your matcher couldn't — and whether you can audit the 328 you never saw.

The threshold is yours. So is the record of what it let through.

See an agent that shows its work and leaves the right decisions to you.