How to extract text from images with OCR

In short: How to shoot a page so recognition succeeds, what to expect from tables and handwriting, and how to get from a photo to editable, searchable text.

Text recognition has become good enough that the limiting factor is almost never the software — it is the photograph. Most disappointing OCR results come from an image that was never going to work: shot at an angle, half in shadow, or with the page curving away from the lens. Getting the capture right takes a few seconds and changes the result more than any setting.

Shoot the page properly

Recognition works line by line, so anything that distorts the shape of a line of text — perspective, curvature, shadow across the middle of the page — costs accuracy directly. The camera wants a flat, evenly lit, front-on rectangle.

Books are the common hard case, because the page curves near the spine. Pressing the page flat, or shooting the two halves separately, fixes most of it.

  • Hold the camera parallel to the page, not at an angle
  • Use even, indirect light; avoid a hard shadow crossing the text
  • Fill the frame with the text block, not the whole desk
  • Press curved pages flat, or capture each half separately

What OCR handles well — and what it does not

Clear, upright, well-lit printed text is close to solved. Ordinary body text in a book, a slide, a receipt, or a printed handout usually comes through cleanly enough to edit rather than retype.

Three things remain unreliable and it is worth knowing which before you rely on the output. Complex tables lose their structure even when the individual cells are read correctly; handwriting varies enormously by hand; and dense multi-column layouts can be read in the wrong order.

  • Reliable: printed body text, slides, signage, receipts, printed handouts
  • Unreliable: handwriting, complex tables, stylised or decorative type
  • Check carefully: numbers, codes, and anything you will act on

On-device recognition and what it means for privacy

Where text recognition runs on the device, the image does not have to leave it to become text. For material that is sensitive — a contract, a medical letter, a page from an internal document — that is a meaningful difference rather than a technicality.

MemoFlow performs text recognition on the device for supported captures. Whatever tool you use, it is worth knowing which side of that line a given feature sits on before photographing something confidential.

Clean up in the right order

Recognised text arrives as a block, and the fastest cleanup is not to read it top to bottom. Fix structure first, then scan for the errors that concentrate in predictable places.

Numbers deserve individual attention. A misread digit is invisible in prose — it looks exactly like a correct one — which is why any figure you intend to use should be checked against the image.

  • Restore paragraph breaks and headings first
  • Then check numbers, codes, and proper nouns against the photo
  • Delete artefacts: page numbers, headers, stray marks read as characters

Keep the image with the text

The photograph is the evidence for the transcription. Keeping both in one note means a suspicious figure can be checked in a second, and a quote can be verified before it is used.

It also preserves what OCR cannot capture: the diagram, the layout, the handwritten margin note next to the paragraph you extracted.

Key points

  • Most OCR failures are photography failures — shoot flat, front-on, and evenly lit.
  • Printed body text is reliable; handwriting and complex tables are not.
  • On-device recognition means sensitive pages do not have to leave the device.
  • Fix structure first, then verify numbers and names against the image.

Be the first to know when MemoFlow is ready.

MemoFlow captures from the camera or your photo library, recognises text on supported devices, and drops the editable result straight into a note.