Before you let AI read your receipts, decide this one thing
Every tool can "read receipts with AI" now. What actually saves time is knowing which fields to look at again.
Have two different models read each receipt and check only the fields where they disagree. Measured on 38 synthetic receipts and 1,778 fields.
01. The problem
Receipts pile up fast. To close the month you copy the date, store, net amount, tax, and total into a spreadsheet. Today an AI can read the photo for you. That part is solved.
The hard part comes next: can you put that AI-made table straight into your books? It is right most of the time, but not always, and it never tells you where it went wrong.
Since any field might be wrong, you end up comparing every receipt with the table, line by line. You brought in AI to save time, and the checking time is still there.
How well the AI reads mattered less than knowing which fields to look at again.
02. Read it twice
The idea is simple. Two different AI models read the same receipt separately. Then every field is compared.
- Fields where both readings match go straight into the table.
- Fields where they differ get a “check” flag. A person compares only those with the receipt.
The receipt in the video above is a real case from the test (a Korean grocery receipt). One model read the store name as 소담바구니, the other as 쇼담바구니 — one character apart. Dates and amounts matched. So instead of the whole receipt, a person checks one field. And that field was indeed the wrong one.
The second reader: different, but not weaker
Pick the second model carelessly and the method breaks, in one of two ways:
- Two models from the same vendor tend to make the same mistakes in the same places. A value they both get wrong passes as a “match”.
- A much weaker model is wrong so often that almost every field gets flagged. People start rubber-stamping the first reading.
So I tested seven candidates on the same receipts and kept the pair with similar accuracy but different failure patterns: 97.4% and 94.9% field accuracy, from different vendors. The second reader is also very cheap — reading 40 receipts twice cost about one US dollar.
No flag does not mean correct
One caveat matters. A flag means “the two readers disagreed”, not “this is wrong”. And no flag does not guarantee the value is right: if both readers make the same mistake, it passes silently. The only way to know how often that happens was to test it.
03. I tested it
Checking the results needs ground truth, and real receipts don’t come with an answer key. So I generated 40 synthetic receipts whose correct values I knew, mixing four conditions from clean to crumpled, tilted, and blurred. For the 38 that finished processing, I compared all 1,778 fields with the truth.
of 1,778 fields
2 wrong out of 756
50 of 51
0.06% of all fields
The result that matters: 50 of the 51 wrong fields were flagged. The fields a person has to check went from 1,778 to 84, and almost every error sat inside those 84.
The one silent miss
Exactly one field was wrong without a flag: an item name both models misread by the same character (제주풍 → 제주품, think “Jeju-style” vs “Jeju-made”). Same mistake on both sides, so it passed as a match.
The cost: some flags are false alarms
Of the 84 flagged fields, 50 were actually wrong. In the other 34, only one reader differed and the value in the table was correct. So about 4 in 10 flags turn out fine when you check.
That is a deliberate trade. Instead of 1,778 fields you check 84, and nearly every error is among them. The checking drops by more than 95%; some of it is a wasted glance.
Where the readers disagreed
| Field | Wrong | Flagged | Note |
|---|---|---|---|
| Date · time · approval no. | 0 | 0 | Correct on every receipt |
| Net · tax · total | 0 | 0 | Correct on every receipt |
| Line unit price · amount | 2 | 2 | Every error flagged |
| Store name | 4 | 4 | Every error flagged |
| Address | 18 | 22 | Most disagreement |
| Item name | 25 | 50 | The one silent miss is here |
The fields your books depend on — date, net, tax, total — were never wrong. Line prices had one error each, both flagged. The rest clustered in text fields like addresses and item names, where look-alike characters are common and every store formats receipts differently.
04. What I learned
1. “Where to look” beats “more accurate”
Pushing accuracy from 97% to 98% is hard and expensive, and if you still don’t know where the remaining 2% is, people still check everything. Pointing at the likely errors cut the checking to about one twentieth.
2. Test on data with known answers first
“It seems to work” is not enough. With synthetic receipts and a known answer key, every model change could be compared the same way — and no real customer data was needed.
3. Leave the judgment to people
This doesn’t remove the person. It moves their time to the fields that need judgment. Repetition to tools, judgment to people — that’s what this blog is about.
Tested on 40 synthetic receipts I generated (38 fully processed). Results on real receipts will vary with condition and format.