Isaac Flath
BlogAboutSubscribe
← Home

Getting Data Out of Documents

OCR on PDFs, scans, and tables: getting accurate text out of messy documents.

Posts

Loading posts…

Examples

Sixty-second worked examples, one document at a time.

When OCR Moves a Cell Into the Header

When OCR Moves a Cell Into the Header

2026-07-23

Extract Handwritten Tables With Structure Intact

Extract Handwritten Tables With Structure Intact

2026-07-22

Check OCR Labels and Values Together

Check OCR Labels and Values Together

2026-07-21

How to Preserve Grouped Headers in a Scanned Table

How to Preserve Grouped Headers in a Scanned Table

2026-07-20

How to Extract a Scanned Table Without a Text Layer

How to Extract a Scanned Table Without a Text Layer

2026-07-19

Why Datasheet OCR Needs Product-Specific Evals

Why Datasheet OCR Needs Product-Specific Evals

2026-07-18

How to Extract IRS Tax Tables Without Losing the Formula

How to Extract IRS Tax Tables Without Losing the Formula

2026-07-17

Plain Text Loses the Structure in Tesla’s Revenue Table

Plain Text Loses the Structure in Tesla’s Revenue Table

2026-07-16

Why Plain Text Breaks Financial Tables

Why Plain Text Breaks Financial Tables

2026-07-15

Avoid PyMuPDF Licensing Issues

Avoid PyMuPDF Licensing Issues

2026-07-03

Get new posts by email.

An email when I have something worth sending. Unsubscribe anytime.

BlogAboutCommunityYouTubeGitHubXLinkedIn
© 2026 Isaac Flath