Back to Technology

AI Financial Reports: The Refund Test

AI financial reports need reproducible calculations. A Norman replay shows how refunds, subtotal rows and CSV formatting can change what a number means.

Category
General
Updated

You ask an AI assistant how much you spent on office supplies. There was a €500 purchase and a €200 refund. It returns a tidy table, a short explanation and a number in bold. Before reading the explanation, look at that number. Is it €300 or €700?

AI financial reports should make the calculation reproducible. A fluent explanation cannot rescue a total that adds a refund to the original expense. I want to be able to change one input, predict the effect and see that same effect in the report and its export.

This is a useful test precisely because it is small. Nobody needs a complicated dashboard to understand the expected answer. Yet the example forces a system to preserve several things that a list of amounts can lose: the booking side, the meaning of a refund and the relationship between a total and its components.

What should an AI financial report actually deliver?

On September 22, 2026, Accrual launched Arc, describing work delivered for review, including Excel workbooks and client reports. Source: Accrual. On September 15, Workiva introduced Agent Studio for building AI agents around company knowledge and governed workflows. Source: Workiva.

My reading of these announcements is that the useful unit of work is moving beyond an answer in a chat. People are being asked to review an artifact they may use in another process. That raises a practical question: what travels with the number when the conversation is gone?

At minimum, I want the reporting period, currency, calculation rules and a way to inspect what contributes to the result. A paragraph saying that spending fell is much less useful if the exported file cannot explain which spending it means.

Our earlier article on an AI accounting agent discusses the wider workflow. Here, we focus on something smaller: the deterministic calculation underneath a report in an AI accounting product. We replayed Norman's report arithmetic and CSV formatting locally using fictional data. This is not a demonstration of an AI creating a workbook.

How should an AI report handle a refund?

Start with one expense category. Put a €500 purchase in it, then return €200 of that purchase. For this example, VAT is zero and business use is 100%. Those assumptions keep the experiment about refunds, rather than mixing several accounting questions into one result.

The replay uses the unchanged calculation functions extracted from Norman's reporting code. It takes the amount's magnitude, applies the booking direction and reverses the contribution for a refund. The resulting expense is €300. Running the same calculation with no refund gives €500; a full €500 refund brings it to zero.

Synthetic local replay: a 500 euro purchase less a 200 euro refund leaves 300 euros of expense.
Fictional amounts, replayed through Norman's report arithmetic. Zero VAT and full business use; this is an explanatory diagram, not a customer screenshot.
Synthetic casePurchaseRefundNet expenseAdding absolute amounts
No refund€500€0€500€500
Partial refund€500€200€300€700
Full refund€500€500€0€1,000

The last column is a deliberately wrong alternative, not a measured model output or a claim about a previous Norman bug. It shows why a plausible-looking operation, adding the numbers after removing their signs, fails this specific test.

We fixed the category in advance. The replay therefore says nothing about whether a model would classify the purchase correctly. Separating that question makes a failure easier to diagnose: first establish the calculation you expect, then evaluate whether the system supplies the right inputs.

Why is the amount's sign not enough?

A positive number can look like income if you have only a bank-style list in mind. But a stored transaction amount and its booking side are separate pieces of information. In this example, the supplier refund belongs with an expense. Treating every positive amount as revenue would answer a different question.

There is a production reason to test this distinction. A read-only aggregate check of Norman's active transaction records found a few hundred non-refund records whose stored amount sign differed from their booking side. The check covered a ninety-day value-date window ending before September 28, 2026. We retrieved counts, not individual transactions or descriptions.

That is a statement about stored representation. It is not a count of incorrect reports, and it does not tell us why those records have that representation. The population also differs from the narrower set of business transactions used by a particular report.

In the local replay, a positive stored amount of €500 with an expense booking side still contributes €500 of expense. That is the behavior we can demonstrate. The production count explains why testing it is worthwhile; it does not turn the synthetic example into a production accuracy benchmark.

Can a subtotal be counted twice?

Now give the report two operating-cost components: €400 of rent and €300 of office costs after the refund. The operating-cost total is €700. Expand the detail and all three numbers can be visible at once.

Add every visible amount and you get €1,400. Nothing went wrong with addition. The mistake was treating a parent row and its children as independent expenses. A person can make that mistake in a spreadsheet, and an agent can make it when turning a rendered table back into data.

Norman's CSV row builder labels summary rows with level zero and operating-cost details with level one. That preserves a useful relationship in the export. It does not mean that all level-zero rows can safely be added together: the full report also contains calculated subtotals and results.

For our test, the invariant is narrower and clear: the operating-cost children sum to their €700 parent, and the parent is counted once. This is the same kind of distinction we explored in transaction matching scores. A value becomes useful only when the reader knows what it represents.

Should the CSV match the screen?

The reviewed Norman screen, CSV export and PDF export share a report-row builder. That gives them a common source for report values while allowing different layouts. It also gives us one place to test the relationship between the €700 subtotal and its components.

We passed the fictional report through the actual CSV functions. The English numeric format writes 700.00 with comma-separated fields. The German format writes 700,00 with semicolon-separated fields. Both encode the same amount.

You can inspect the English-format CSV and the German-format CSV. Both use English demonstration labels. These are outputs from the local replay, not exports from a customer account. We checked their cells and delimiters; we did not test opening them in a particular spreadsheet application.

Formatting deserves a test because a file is another interface. Once someone downloads it, the explanatory chat may no longer be available. Period, currency and row meaning need to survive that handoff. Our discussion of AI agent tool results makes the same point about data passed between tools.

What should you test before trusting the report?

Write the expected answer before asking the system for one. Start with the three refund cases above, then add the parent-and-child case. Keep the input records fixed while changing the export format. The amount should stay the same even when its printed representation changes.

Then test the boundaries this small replay leaves open: category selection, missing amounts, date filters, currency conversion and legacy refund records. A zero produced by known inputs is different from an answer produced when an input was unavailable. The interface should make that distinction inspectable rather than ask the reader to infer it from a blank cell.

Those are evaluation recommendations, not claims that this replay covers the entire reporting application. It makes no model call, runs no full reporting endpoint and measures no accounting accuracy rate. Its value is that the demonstrated arithmetic and exported cells can be checked independently.

For an AI financial report, that is where I would start a review. Take one purchase, one refund and one total. Follow them into the file. If the system cannot preserve their meaning through that small journey, a longer explanation will not make the result easier to trust.

Frequently asked questions

How do you check an AI financial report?
Start with a small dataset whose expected result you can calculate independently. Keep the period, currency and inclusion rules fixed. Check a purchase, a partial refund, a full refund and a subtotal with child rows. Then compare the calculated values with the exported cells. This tests specific behavior, not overall accounting accuracy.
Should a refund reduce expenses in an AI report?
In the synthetic supplier-refund example here, a €500 purchase and a €200 refund produce €300 of net expense. That result assumes a fixed expense category, zero VAT and full business use. The report must preserve the meaning of the refund. A positive stored amount alone does not establish that an operation is revenue.
Why do CSV totals differ from a financial report?
One possible cause is adding a subtotal and the detail rows already included in it. Another is interpreting a decimal separator or delimiter incorrectly. Our local replay checks a €700 operating-cost subtotal and its €400 and €300 components in two CSV formats. It does not test how a particular spreadsheet application imports the files.

Norman handles the operational finance work behind the scenes

From invoicing to bookkeeping, Norman keeps recurring finance work organized so you can stay on top of deadlines with less manual effort.