Skip to main content
Ebook

The Tools Fallacy: 25 years of finance software automated everything except the work

AuthorPayhawk Editorial Team
Read time
14 mins
PublishedAug 14, 2026
Quick summary

New research across 1,520 finance leaders finds the best-tooled teams automate up to 90% of everything a system can record, track, or enforce, and still verify invoices by hand seven times out of ten. The gap has a cause, the cause has a history, and AI is heading straight for it.

  1. Foreword
  2. Key findings
  3. The number that survives everything
  4. Every road hits the same ceiling
  5. The work that was never written down
  6. A function priced out of its own questions
  7. The same mistake, at AI speed
  8. What the answer would have to look like
  9. METHODOLOGY
  10. SOURCES
  11. APPENDIX
GoogleAdd us as a preferred source
Payhawk - G2 4.6 rating (600+ reviews)
Get fresh finance & AI insights, monthly.
Unsubscribe anytime.

By submitting this form, you agree to receive emails about our products and services per our Privacy Policy.

Foreword

Every CFO I meet has bought the tools. Expense platforms, ERP integrations, approval flows, cards that code themselves. Ask whether the close got fundamentally easier and the answer usually comes with a pause.

We commissioned this study expecting to find a technology gap. We surveyed 1,520 finance leaders across seven markets and measured 25 concrete behaviours, from real-time visibility to automated matching. What came back was stranger than a gap. The tools work. Among teams that finished the tooling project, everything a system can record, enforce, or reconcile runs at automation rates approaching 90 percent. One category of work barely moved, and it is the same category at every company size, in every market, at every level of ambition.

I run a software company. It would suit me to tell you the answer is more software. The data says the answer is something our industry has never shipped and most finance teams have never written down. This report is about that missing artifact, what it has quietly cost, and why the AI wave has made the question urgent.

Hristo Borisov
CEO and co-founder, Payhawk

Key findings

  1. Among fully tooled finance teams, automation of recording, tracking, and controls runs at 81–90%. Automated invoice matching runs at 30.4% (Payhawk study, n=214).
  2. Automated two and three way matching is the lowest-adopted of all 25 behaviours measured (20.7%). The highest, at 42.1%, is agreement that simple, well-designed software drives adoption. Even the best-performing item on the list is a belief about tools rather than a process that runs without people (N=1,520).
  3. Nothing else closes the gap: matching automation reaches just 29% at the largest companies, 37% at the most structurally complex groups, and 25% among teams who say they are committed to automating everything they can. No cut of the data comes near the 80–90% those same teams achieve elsewhere.
  4. Finance is repeating the pattern with AI: even among AI-forward organisations, 78% use AI tools while 55% have set minimum AI rules; 41% of tool users have none. Externally, 7% of finance leaders prioritise governance over deployment speed (Avalara, July 2026).
  5. The market result: 84% of finance organisations have implemented or plan to implement AI. 7% report high or very high impact (Gartner, June 2026).

The number that survives everything

This month, somewhere in your company, a person will open an invoice, pull up the purchase order, find the goods receipt, and read all three. She will notice the quantities agree and the prices almost do, remember that this supplier ships partials in the first week, approve the difference, and move to the next one. Multiply her by every invoice you receive. That is the work this report is about.

We measured 25 finance behaviours across 1,520 finance leaders in seven markets, evenly quotaed from individual contributors to the C-suite and from 50 to 5,000 employees. Then we did something the industry rarely does with adoption data: we isolated the teams that had finished the tooling project, applying three criteria: their expense tools are deeply integrated with the rest of their systems; their spend flows into accounting without manual reconciliation; and all of their spend sits consolidated on a single platform. 214 companies clear all three bars, and for them the last two decades of finance software paid off in full.

Among these teams, 89.7% run a unified process across all entities, 88.8% have fully automated accounting, 84.6% track spend in real time, 80.8% apply proactive controls to all company spend, 68.2% can report adequately to stakeholders outside finance. Then comes automated two and three way matching: 30.4%.

Hold on to the shape of that chart. These are the same 214 companies on every bar, so no difference in budget, talent, or ambition separates the top of the chart from the bottom. Five outcomes land between 68 and 90 percent, one lands at 30, and whatever stops matching automation survives the best infrastructure money currently buys.

Nor is this a matter of some teams being further along than others. Score every company in the study by how many of those three criteria it meets, from none to all three, and accounting automation climbs without pause: 7 percent, then 30, then 55, then 89. Matching automation climbs too, from 13 to 25 to 32 percent, and then stops. The final step, the one separating a well-tooled finance function from a fully tooled one, adds thirty-four points of accounting automation and takes two points off matching.

Every road hits the same ceiling

If tooling completeness is not the answer, three other explanations remain, and each comes up whenever this number appears. The data rules out all three.

Scale

From companies of 50 to 100 employees up to companies of 1,001 to 5,000, that figure moves from 16.1% to 28.6%. A thirtyfold difference in scale, and the resources that come with it, buys twelve points. The largest firms in the study leave 71% of this work manual.

Complexity

Groups with more than 100 legal entities reach 36.8%, against 14.1% for single-entity companies. Necessity pushes harder than scale does, and still leaves nearly two thirds of the most complex organisations in the market checking invoices by hand.

Ambition

30.9% of teams describe themselves as committed to automating everything that can be automated. Within that group, matching automation stands at 24.9%. Sharper still: teams that are both fully tooled and committed reach 95.7% on automated accounting, 87.0% on real-time tracking, and 24.6% on matching (n=138). Willingness plus infrastructure moves everything except this.

The work that was never written down

If we rank all 25 behaviours, a pattern appears. At the top sit the outcomes a tool can deliver directly: simple, well-designed software that people actually adopt (42.1%) and proactive controls applied to all spend (40.3%). At the bottom sit automated matching (20.7%) and automated ESG reporting (23.1%), the two behaviours that require interpretation. The interface problems got solved and the rule problems got solved, while the judgment problems stayed exactly where they were.

Underneath a generation of finance software sits a single assumption: that every problem in the function is tool-shaped. Tools move data and enforce rules, and over two decades they have become superb at both, but neither capability lets a tool carry a sequence. It cannot hold the AP lead's map of which supplier double-invoices in December, which entity books freight differently, or which mismatch is harmless noise rather than the first sign of a duplicate payment.

That map is a procedure, a matter of what to check, in what order, against what, and when to stop, and nobody ever wrote it down because until recently nothing could run it except the person who held it. Call the assumption the Tools Fallacy, and the last twenty years of finance technology read as its faithful execution: spectacular progress on everything that could be specified in advance, and near silence on everything that turned on a person's judgment.

None of which means judgment work is beyond automation, because a minority of teams already do it. In Ardent Partners' operational benchmarking, best-in-class AP teams reach 49.2% touchless invoice processing at $2.78 per invoice while the average buyer pays roughly four times as much, and around 75% of AP departments already run some form of automation or AI tooling.

The tooling, in other words, is nearly universal, but the outcomes belong to a small cohort, and what that cohort did differently was codify the procedure exception by exception. Codification is slow, expensive, expert work that almost nobody pays for, so adoption splits between the few who invested in it and the long tail who never will at the current price.

It is worth pausing on the fact that our survey figures land in the same band as Ardent's operational measurements, because two unrelated methods arriving at one ceiling is harder to dismiss than either finding alone.

McKinsey arrived at the same place from the opposite direction. After analysing more than fifty agent deployments, its first lesson puts the workflow above the agent, noting that codified practice serves as both the training manual and the performance test, and that today those practices often exist only as "tacit knowledge in people's heads."

An unwritten procedure also tends to concentrate, since it lives in whoever has performed it longest, which is how a single resignation turns a routine close into a bad quarter. The accounting pipeline is genuinely recovering here, with enrollment up 8.9 percent in spring 2026 for the third straight year and first-time CPA candidates at their highest since 2018, set against roughly 124,200 US openings a year. That is good news that changes little, though, because hiring transfers headcount rather than knowledge, and a procedure held in one person's head moves only when someone finally writes it down.

A function priced out of its own questions

The ceiling reads like an operations detail until you price it. Only 31.1% of teams track spend in real time, so for the other two thirds, month-end is where the past first comes into view: the uncoded transactions, the missing receipts, the commitment someone made three weeks ago.

Only 26.1% say their tools let them report adequately to the rest of the business. A function consumed by verification answers slowly, and a business that gets slow answers learns to stop asking. The real cost of manual judgment work is every question that was never worth an analyst-day: the supplier worth renegotiating, the software nobody uses, the entity quietly drifting off policy.

The people inside the function know this. Asked whether expense management is treated as a genuinely strategic topic, 35.0% of the C-suite says yes, against 18.2% of the individual contributors who do the actual checking.

The distance between those two numbers is the distance between the org chart and the work. Strategic intent set at the top has not reached the people carrying it out, and the gap is not a communication failure but a capacity one: a function spending its hours on verification cannot feel strategic to the person doing the verifying, whatever the intent above them.

The same mistake, at AI speed

Now the industry is running the experiment again, faster. In June 2026, Gartner reported that 84% of finance organisations have implemented or plan to implement AI, while 7% report high or very high impact.

Its analysts locate the difference in discipline: the organisations getting value follow structured roadmaps that connect AI to outcomes, and the rest chase capability. Bain's April 2026 survey of 951 companies put the same finding in colder terms. Savings targets were broadly missed, the firm concluded that "the technology worked, the value didn't arrive," and 90 percent of the surveyed companies responded by raising their AI budgets again without pausing to diagnose why, buying more capability to remedy a surplus of it.

The strongest evidence on what does separate winners comes from McKinsey's global survey. Across 31 organisational attributes tested, fundamentally redesigning workflows sits among the strongest predictors of AI performance, and high performers are about three times likelier to have done it.

A second differentiator deserves more attention than it gets: the high performers had defined processes for deciding when a model's output needs human validation, which is to say they wrote the checking procedure down. That discipline is rare against the backdrop of the wider numbers, where 88 percent of companies use AI somewhere but only 39 percent can point to any earnings impact, and where inaccuracy leads the list of AI failures teams have actually experienced.

Our own data shows the same sequencing error at work inside finance. Among the AI-forward organisations in the study, the minority already deep into adoption, the pattern descends in a clear order: 78.3% use AI tools and run skills initiatives, 68.9% have committed budget, 64.0% have implemented integration measures, and only 55.3% have established minimum rules for AI in their policies.

The tool comes first, the money second, the rules last. Fully 41% of the AI tool users have set no minimum rules at all, even as 65% of that same group expects AI agents to be supporting their employees before long. The vanguard is acquiring capability faster than it is writing the procedure to govern it, which is the Tools Fallacy under a newer name.

Tax compliance platform Avalara's July 2026 study of 1,505 finance leaders found 7% prioritising governance over deployment speed, 71% under pressure to move fast, 30% yet to update internal controls for AI agents that take or recommend actions, and 44% only somewhat confident they could explain an agent's actions to an auditor or a regulator. Nearly a quarter say accountability for a serious agent error is unclear or belongs to no one.

Beneath the official deployments runs the unofficial one: 66% of office professionals have used AI at work while believing their policy forbade it. Waiting is a deployment decision too. It deploys the ungoverned version.

A single root cause is now producing two different failures. The unwritten procedure is what lets a single resignation destabilise a close, and it is also why AI pilots stall in finance: the model arrives ready to work and finds nothing to follow.

The question for this quarter:
Which of our procedures exist anywhere outside a person's head, in a form something other than that person could run?

What the answer would have to look like

Take an inventory against that question and most finance teams will find the honest answer is: very few. A close checklist points at the real sequence without capturing it, because the sequence that matters, the one carrying all the exceptions and the stopping rules, lives in people rather than in any document. Whatever eventually fills that gap, this research and the sources around it already describe the properties it would need.

It treats the procedure as an artifact

The sequence that once lived only in the AP lead's head becomes explicit, ordered, and runnable: something a team can inspect, improve, and hand to someone else. This is what McKinsey means when it describes codified practice as both the training manual and the performance test.

It drafts, and lets people check

The judgment call is attempted first and reviewed second, inside the approval flow finance already trusts, so the human moves from producing the answer to validating it. Defined rules for when an output needs that validation rank among the strongest differentiators of AI high performers in McKinsey's data, which is to say the checking itself is a procedure, and the teams that pulled ahead had written it down

It computes rather than estimates

Finance is the one function where approximately right is professionally disqualifying, and the survey data explains why it matters here more than elsewhere: inaccuracy already leads the list of AI failures companies report, with explainability close behind and still largely unaddressed. What the function needs is an exact answer with its working shown, not a confident guess.

It acts inside permissions, and leaves a trail

Asked what would raise their confidence in AI agents, Avalara's 1,505 finance leaders named the same three things: agents that operate within existing systems of record, outputs grounded in verified data, and an audit trail behind every action. The specification, in the end, was written by the people who would have to buy it.

For as long as finance has existed, writing a procedure down and putting it to work were two different things: the document sat in a drawer while the real work stayed in someone's head. That has just changed. A written procedure is now a runnable one, and the teams that grasp what that makes possible will spend next year codifying the work that only lives in people, while everyone else evaluates another platform.

METHODOLOGY

  • The Payhawk Spend Management Study surveyed 1,520 finance leaders across seven markets spanning EU and the US, across four seniority tiers (individual contributor, director, VP, C-suite), five company-size bands from 50 to 5,000 employees with different number of entities.
  • Agreement means top-2 box (6 or 7) on a 7-point scale.
  • The ranked list covers 25 behaviours; a separate block of self-rated optimisation items is reported independently.
  • "Fully tooled" means top-2 agreement on all three of: deeply integrated expense tools, spend-to-accounting integration with no manual reconciliation, and a single consolidated spend platform (n=214).
  • AI-specific questions were routed to a subset of respondents with high self-rated AI maturity (n=405 to 451; mean 7.5/10 against 5.4 for the full sample) and are reported only for that base.
  • Crosstabs describe associations, and the report's argument rests on the pattern across them.
  • As a validity check, our self-reported matching-automation rates fall in the same range as Ardent Partners' operationally measured straight-through processing benchmarks.

SOURCES

  1. Payhawk Spend Management Study, 2026. N=1,520 finance leaders, seven markets.
  2. Gartner, press release, 8 June 2026: CFOs need structured finance AI roadmaps. Survey of 183 CFOs.
  3. Bain & Company survey of 951 companies, April 2026, via CFO.com and Bloomberg.
  4. Avalara / Censuswide, Agents of Change, 21 July 2026. N=1,505 finance leaders; fielded 15 to 22 June 2026.
  5. McKinsey & Company, One year of agentic AI: six lessons, 12 September 2025.
  6. McKinsey & Company, The state of AI, 5 November 2025. N=1,993 respondents, 105 countries.
  7. Ardent Partners, AP Metrics That Matter 2025, via Corpay, July 2026.
  8. AICPA / National Student Clearinghouse, enrollment data, 9 June 2026.
  9. US Bureau of Labor Statistics, Occupational Outlook: Accountants and Auditors, 2024 to 2034 projections.
  10. PagerDuty / Wakefield Research, Shadow AI survey, 11 June 2026. N=1,250; fielded 9 to 20 April 2026.

APPENDIX

The 25 behaviours, ranked

Every respondent rated each statement on a 7-point agreement scale. The figure shown is top-2 box agreement (6 or 7) across the full sample of 1,520. The two automation behaviours that anchor this report's argument, automated matching and automated ESG reporting, sit at the foot of the list.

# Behaviour Domain Agreement
1 Simple UX acknowledged as a driver of adoption and reduced resistance Tooling & adoption 42.1%
2 Proactive controls ensure all spend follows internally agreed policies Visibility & control 40.3%
3 Unified process across all entities ensuring visibility and standardisation Process & automation 39.5%
4 Spend traceable, enabling clear financial KPIs and progress tracking Visibility & control 38.5%
5 Streamlined approval workflows with appropriate controls maintained Tooling & adoption 38.0%
6 Culture has capacity to embrace technological advancement Culture & mindset 37.4%
7 Agile mindset cultivated, with people-focused leadership Culture & mindset 36.7%
8 Automations consolidate data from different sources for better decisions Process & automation 34.2%
9 People enabled to understand the implications behind spending decisions Culture & mindset 33.9%
10 Spend management system helps negotiate better supplier terms Visibility & control 33.0%
11 People have autonomy and tools for situationally appropriate decisions Culture & mindset 32.8%
12 Expense handling digitalised as early as possible in its cycle Process & automation 32.5%
13 Spend consolidated in a single platform for quick, accurate insights Visibility & control 31.6%
14 Real-time spend tracking at any point, for agile decisions and audits Visibility & control 31.1%
15 Committed to automating processes as far as possible within evolving standards Governance & matching 30.9%
16 Governance structured to give context to incurred spend Governance & matching 30.7%
17 Spend management integrates with accounting, eliminating manual reconciliation Tooling & adoption 30.3%
18 Accounting process fully automated, supporting a lean finance function Process & automation 30.3%
19 Expense tools deeply integrated, allowing highest levels of customisation Tooling & adoption 30.1%
20 Spend structured almost entirely through POs or budgets Governance & matching 29.3%
21 POs treated as a high priority for elevated visibility Process & automation 26.4%
22 Digital technologies deliver adequate reporting to non-financial stakeholders Tooling & adoption 26.1%
23 Expense management already treated as a strategic topic Culture & mindset 25.2%
24 ESG reporting automated, ensuring proactive regulatory compliance Governance & matching 23.1%
25 Automated 2/3-way matching to increase oversight of purchases and payments Governance & matching 20.7%

Behaviours 24 and 25, are the two the report treats as judgment-shaped work. Behaviour 23 (expense management treated as a strategic topic) is a perception item rather than an automation behaviour, and is not part of that claim.