AI Bookkeeping Accuracy: Why 95% Tells You Nothing
Almost every AI bookkeeping tool claims 95% accuracy or better, and the number is close to meaningless. In our own German ledger, a program that always answers "19% VAT" would score about 87%, and half the money sits in the largest one percent of expense lines. Here is what to measure instead.
- Category
- General
- Updated
- Author
- Stan Kharlap
Every accounting product sold in 2026 has an accuracy percentage in its marketing, and almost all of them land between 85% and 99%. Treat that number the way you treat a restaurant calling itself authentic. It is not a lie. It is just not a measurement, because nobody tells you the denominator.
Here is why, using our own ledger, and then the questions that actually separate products.
What does "95% accurate" actually count?
An accuracy figure is a fraction, and the fraction has to be over something: lines, documents, fields, euros. Vendors almost never say which. The default assumption, and the flattering one, is per transaction line, unweighted, so a three-euro coffee counts exactly as much as a forty-thousand-euro invoice.
That choice does most of the work in producing a good-looking number, because real ledgers are lopsided in two directions at once, and both favour whoever is doing the counting. Norman does bookkeeping and tax filing for businesses in Germany, and across a couple of hundred thousand categorised expense lines so far this year, our median line is about thirty euros.
Always answer "19% VAT" and you are already 87% accurate
Start with the VAT rate, since that is the field with direct tax consequences. Germany has a standard rate of 19%, a reduced rate of 7%, and a set of zero-rated and exempt cases. In our expense lines this year, roughly seven in every eight carry 19%.
So a program consisting of a single line, return 19, scores about 87% on VAT rate accuracy. No model, no OCR, no learning. It would beat several published accuracy claims, and it would be wrong on every case where getting the rate right requires knowing anything at all.
The same trick works on categories. We have a few dozen expense categories in active use; five of them cover well over half of all lines. A classifier that has memorised the head of that distribution and guesses the most common category whenever it is unsure posts a strong per-line number while being useless precisely where a human wanted help. That is the expected behaviour of anything tuned to maximise per-line accuracy, because the cheapest available accuracy is always in the head.
Half the money is in one percent of the lines
Now weight by euros instead of by rows, which is what your tax return actually does.
In our data, the largest one percent of expense lines by amount carry about half of all expense value. The smaller half of all lines, put together, account for under two percent of the money.
Read those two sentences together and the problem becomes hard to unsee. You can be right about 99% of lines and wrong about half the money, or wrong about 20% of lines and financially near-perfect. On a skewed ledger, per-line and euro-weighted accuracy are barely correlated. Nobody publishes the euro-weighted figure, and it is the only one that maps onto what you owe.
The hard cases fit inside the error budget
The third problem is the sharpest. A 95% claim grants a 5% error budget. Ask how big the genuinely difficult population actually is.
For us: about one expense line in fifty is a reverse-charge booking, where the supplier charges no VAT and the buyer accounts for it themselves. Cross-border suppliers of any kind are under one line in ten.
Reverse charge alone fits inside a 5% error budget twice over. A product could get every single reverse-charge decision wrong, in both directions, and still truthfully advertise 95% per-line accuracy. Those are also the bookings most likely to produce a filing error a tax office notices, because they change which boxes of the VAT return move, not just the amount inside one box. The headline metric is structurally blind to the cases it is being used to reassure you about.
What to ask an AI bookkeeping vendor instead
Ask for the denominator, and ask for the tail. None of these are hard to answer if the answer is good:
| Metric | What it measures | Where it misleads |
|---|---|---|
| Category accuracy, per line | Share of lines with the right category | Dominated by a handful of common categories; a head-guesser scores well |
| VAT rate accuracy, per line | Share of lines with the right rate | A constant "19%" already scores about 87% on a German ledger |
| Euro-weighted accuracy | Share of the money booked correctly | The number that tracks your filing. Almost never published |
| Reverse-charge recall | Hard cross-border cases caught | Around 2% of lines, so invisible in any headline figure |
| Abstention rate | How often it declines to guess and asks you | Usually not published at all; a low number is a warning, not a feature |
The last row inverts the usual sales conversation, and I would weight it most heavily. A system that abstains on cases it cannot determine is more useful than one that is a few points "more accurate" and silently confident. Confidence without abstention is not accuracy, it is unreviewed work, and you sign the filing.
One more question: which decisions go to a model at all? A lot of what gets sold as AI judgement is a lookup. VAT rates come from a table, reverse charge follows from the supplier's country and the nature of the supply, thresholds come from statute. Where a rule exists, running it is not just cheaper than asking a model, it is auditable, which matters when a tax office asks why a line was booked the way it was. Where we draw that line is in how we built an agent that files VAT returns, and the loop for the ambiguous remainder is in what our categorisation learns from corrections.
So what is Norman's number?
Fair question, and it would be cheap of me to end without one.
Of the machine categorisations our pipeline made this year, about fifteen in sixteen were never changed by hand. Roughly one in sixteen gets corrected by the person whose books they are. That is our closest analogue to the figures in everyone else's marketing, and by the argument above you should not be impressed by it: an unweighted per-run count, on a ledger with all the skew I just described.
Here is the part I cannot give you. Our correction record stores what changed and why, but not which transaction it belonged to, so we cannot currently weight our own override rate by euros. I went looking for that number while writing this and it is not there. That is the same gap I am criticising, sitting in our own telemetry, and linking a correction back to the amount is now on the list.
One thing the corrections do tell us: when the machine gets a category wrong, it is almost always the reading of the document rather than the rule that mapped it. Only about one correction in sixty traces back to a rule. That is the empirical case for keeping the deterministic half deterministic.
Does the EU AI Act make vendors prove accuracy claims?
No, and this is worth understanding before you assume regulation has solved it for you.
The AI Act's transparency obligations under Article 50 became applicable and enforceable on 2 August 2026. They require you to be told that you are dealing with an AI system, and they require synthetic content to be marked. Article 50 says nothing about accuracy, performance, or how a vendor may describe either.
The obligations that do touch technical documentation and risk, the high-risk regime, were pushed back by the AI Omnibus, Regulation (EU) 2026/1744, in force since 27 July 2026. Stand-alone Annex III systems now have until 2 December 2027, and AI embedded in regulated products until 2 August 2028.
So the law obliges a bookkeeping vendor to tell you an AI is involved, and permits it to describe that AI's accuracy however it likes. The buyer does the diligence. The buyer's guides know it, too: one 2026 roundup tags competitors' figures "[vendor claim]", says outright that such claims are mostly unproven, and presents its own numbers without that tag.
Frequently asked questions
How accurate is AI bookkeeping really?
On routine domestic transactions, modern systems genuinely are strong, and claims of 85% to 99% are plausible for that population. They are close to meaningless as product comparisons, though, because they are almost always unweighted per-line figures on a heavily skewed ledger. Ask for euro-weighted accuracy and for performance on unusual cases specifically, and expect most vendors not to have those numbers.
What does 95% accuracy mean in accounting software?
Usually: 95% of individual transaction lines got the right category, measured on the vendor's own sample, with no weighting by amount and no separate reporting for difficult cases. Because a few common categories and one common VAT rate cover most lines in a typical ledger, a trivial rule-based guesser scores in the high eighties. The percentage tells you very little about the remaining, expensive minority.
Can AI handle reverse-charge VAT correctly?
It can, but only if the product treats it as a rule rather than a guess. Reverse charge follows from facts, chiefly the supplier's country and the type of supply, so the right design is a decision tree over those facts, with the model used to extract them from a document rather than to infer the outcome. Because these bookings are a small share of lines, a headline accuracy figure will not reveal whether a product gets them right. Our guide to reverse charge sets out the rules.
Which AI is best for taxes?
The wrong framing, because the model is rarely the differentiator. Any current frontier model can read a receipt and add up VAT. What separates products is everything around it: whether deterministic rules handle the decisions that have deterministic answers, whether the system abstains when a required fact is missing, whether every booking leaves an inspectable audit trail, and whether a human approves the filing.
Should I ask a vendor for their error rate on large transactions?
Yes, and it is the single most informative question on the list. Because value is concentrated so heavily in the largest few percent of lines, error rates on that slice dominate your actual financial exposure while contributing almost nothing to a per-line accuracy figure. A vendor who has measured this will tell you. A vendor who has not will re-quote the headline number.
The number I would want on a spec sheet
If I could standardise one figure for this category, it would not be accuracy. It would be the share of bookings the system completed without asking, next to the euro-weighted error rate on that same set. Two numbers, reported together, over a stated period on a real ledger. That pair cannot be gamed by guessing the head of the distribution, because guessing inflates the first number and wrecks the second. We can publish the first today and not yet the second, which is a fair measure of how early all of this still is.
Until something like it exists, treat every percentage in this market as a claim about the vendor's confidence rather than about your books. Ask what happens to the forty-thousand-euro invoice and the one purchase from an Irish software company. That is where your filing is decided, and it is nowhere in the 95%.
Norman handles the operational finance work behind the scenes
From invoicing to bookkeeping, Norman keeps recurring finance work organized so you can stay on top of deadlines with less manual effort.