Blog/AI workflows
AI workflows

Before you let AI read your receipts, decide this one thing

Every tool can "read receipts with AI" now. What actually saves time is knowing which fields to look at again.

inform.praxis Oct 3, 2026 · 8 min read
What you will take away

Have two different models read each receipt and check only the fields where they disagree. Measured on 38 synthetic receipts and 1,778 fields.

The same receipt, read twice. Fields where both readings match turn green; the one that differs turns amber, and that is the only field a person checks.

01. The problem

Receipts pile up fast. To close the month you copy the date, store, net amount, tax, and total into a spreadsheet. Today an AI can read the photo for you. That part is solved.

The hard part comes next: can you put that AI-made table straight into your books? It is right most of the time, but not always, and it never tells you where it went wrong.

Since any field might be wrong, you end up comparing every receipt with the table, line by line. You brought in AI to save time, and the checking time is still there.

How well the AI reads mattered less than knowing which fields to look at again.

02. Read it twice

The idea is simple. Two different AI models read the same receipt separately. Then every field is compared.

  • Fields where both readings match go straight into the table.
  • Fields where they differ get a “check” flag. A person compares only those with the receipt.

The receipt in the video above is a real case from the test (a Korean grocery receipt). One model read the store name as 소담바구니, the other as 쇼담바구니 — one character apart. Dates and amounts matched. So instead of the whole receipt, a person checks one field. And that field was indeed the wrong one.

The second reader: different, but not weaker

Pick the second model carelessly and the method breaks, in one of two ways:

  • Two models from the same vendor tend to make the same mistakes in the same places. A value they both get wrong passes as a “match”.
  • A much weaker model is wrong so often that almost every field gets flagged. People start rubber-stamping the first reading.
Ten fields read by three pairings. Same-vendor models miss a shared mistake; a weak model flags nearly everything.

So I tested seven candidates on the same receipts and kept the pair with similar accuracy but different failure patterns: 97.4% and 94.9% field accuracy, from different vendors. The second reader is also very cheap — reading 40 receipts twice cost about one US dollar.

No flag does not mean correct

One caveat matters. A flag means “the two readers disagreed”, not “this is wrong”. And no flag does not guarantee the value is right: if both readers make the same mistake, it passes silently. The only way to know how often that happens was to test it.

03. I tested it

Checking the results needs ground truth, and real receipts don’t come with an answer key. So I generated 40 synthetic receipts whose correct values I knew, mixing four conditions from clean to crumpled, tilted, and blurred. For the 38 that finished processing, I compared all 1,778 fields with the truth.

Field accuracy97.1%

of 1,778 fields

Money fields99.7%

2 wrong out of 756

Wrong fields that got flagged98.0%

50 of 51

Wrong and not flagged1

0.06% of all fields

The result that matters: 50 of the 51 wrong fields were flagged. The fields a person has to check went from 1,778 to 84, and almost every error sat inside those 84.

Each dot is one field. Of 1,778, only 84 need a check, and 50 of those are real errors.

The one silent miss

Exactly one field was wrong without a flag: an item name both models misread by the same character (제주풍 → 제주품, think “Jeju-style” vs “Jeju-made”). Same mistake on both sides, so it passed as a match.

Both readers got the same character wrong. It was part of an item name, not a number, so the totals were unaffected.

The cost: some flags are false alarms

Of the 84 flagged fields, 50 were actually wrong. In the other 34, only one reader differed and the value in the table was correct. So about 4 in 10 flags turn out fine when you check.

That is a deliberate trade. Instead of 1,778 fields you check 84, and nearly every error is among them. The checking drops by more than 95%; some of it is a wasted glance.

Where the readers disagreed

FieldWrongFlaggedNote
Date · time · approval no.00Correct on every receipt
Net · tax · total00Correct on every receipt
Line unit price · amount22Every error flagged
Store name44Every error flagged
Address1822Most disagreement
Item name2550The one silent miss is here

The fields your books depend on — date, net, tax, total — were never wrong. Line prices had one error each, both flagged. The rest clustered in text fields like addresses and item names, where look-alike characters are common and every store formats receipts differently.

04. What I learned

1. “Where to look” beats “more accurate”

Pushing accuracy from 97% to 98% is hard and expensive, and if you still don’t know where the remaining 2% is, people still check everything. Pointing at the likely errors cut the checking to about one twentieth.

2. Test on data with known answers first

“It seems to work” is not enough. With synthetic receipts and a known answer key, every model change could be compared the same way — and no real customer data was needed.

3. Leave the judgment to people

This doesn’t remove the person. It moves their time to the fields that need judgment. Repetition to tools, judgment to people — that’s what this blog is about.

Tested on 40 synthetic receipts I generated (38 fully processed). Results on real receipts will vary with condition and format.