Guides › OCR for receipts and screenshots

Getting OCR to actually work on receipts and screenshots

Published June 5, 2026 · Updated September 28, 2026 · thisuseful

Optical character recognition has gotten good enough that people expect it to just work, then get a garbled result from a crumpled gas station receipt and assume the tool is broken. Usually it isn't — the input is. OCR engines are pattern-matching against character shapes, and a few specific problems break that matching more than people expect.

What actually ruins accuracy

Thermal receipts fade and go gray-on-white within weeks, which kills contrast — photograph them soon after purchase, not at tax time. Screenshots taken at less than 100% zoom compress text into fewer pixels than the engine needs to distinguish similar letterforms like "rn" from "m". And a photo taken at an angle introduces perspective distortion that a straight-on scan wouldn't have; if your phone's camera app has a document mode, use it instead of a regular photo.

Receipts specifically

Long thermal receipts photographed in one shot usually fail on the bottom third, where curl and shadow are worst. Flatten the receipt under a book for thirty seconds first, or photograph it in two overlapping halves and run each through separately. Totals and dates — the two fields people actually need — sit in the sections most prone to fold damage, so this is worth the extra thirty seconds.

Screenshots and mixed layouts

A screenshot with a sidebar, a chat bubble, and a code block in one image gives the engine three different text densities to reconcile at once, and it often merges or skips lines at the boundaries. Crop to just the text block you need before running it — narrower input consistently outperforms a wider image the engine has to segment itself.

After conversion, before you trust it

Numbers and dates are the highest-risk output — "8" becomes "3", "1" becomes "l" more often in OCR than any word-level error, because the engine is choosing between visually similar shapes with no language context to catch the mistake. Check every digit in a total or date field manually; don't spot-check the paragraph and assume the numbers came along for free.

If the image contains a signature, ID number, or anything you wouldn't want cached somewhere, use a converter that runs entirely in the browser and delete the source image afterward — server-side OCR tools generally retain uploads longer than their privacy pages suggest.