← Home









Getting Data Out of Documents
OCR on PDFs, scans, and tables: getting accurate text out of messy documents.
Posts
Loading posts…
Examples
Sixty-second worked examples, one document at a time.

When OCR Moves a Cell Into the Header
2026-07-23

Extract Handwritten Tables With Structure Intact
2026-07-22

Check OCR Labels and Values Together
2026-07-21

How to Preserve Grouped Headers in a Scanned Table
2026-07-20

How to Extract a Scanned Table Without a Text Layer
2026-07-19

Why Datasheet OCR Needs Product-Specific Evals
2026-07-18

How to Extract IRS Tax Tables Without Losing the Formula
2026-07-17

Plain Text Loses the Structure in Tesla’s Revenue Table
2026-07-16

Why Plain Text Breaks Financial Tables
2026-07-15

Avoid PyMuPDF Licensing Issues
2026-07-03