Document Processing with AI: What Works Today

If you are looking for the place AI most reliably earns its cost in an ordinary business, it is here. Turning incoming documents into structured data is repetitive, high volume, language-heavy, and — crucially — verifiable.

What works well today

Extracting fields from varied layouts. This is the core case. Invoices from two hundred suppliers, each with a different layout, all needing supplier, date, number, net, tax and total. Traditional template-based systems required configuration per layout and broke when a supplier redesigned. Current approaches handle unfamiliar layouts because they read the document rather than matching positions.

Classification and routing. Deciding what an arriving document is, and which process or department it belongs to. Reliable and immediately useful, particularly where a shared inbox is manually triaged today.

Finding specific content across a set. Which of these contracts contains an automatic renewal clause, and what notice period does each require. Work that was previously not done at all because it was too slow.

Summarising and comparing. Reducing a long submission to its key points, or identifying where a supplier's terms deviate from your standard.

What still causes trouble

Poor scans. Skewed pages, low resolution, photographs taken at an angle, faint print. The general rule holds: if a person struggles to read it, the system will too.

Handwriting. Better than it used to be, still unreliable for anything consequential. Handwritten amounts should always be confirmed.

Complex tables. Merged cells, multi-page tables with repeating headers, and totals that need to reconcile. Extraction often looks correct and is subtly wrong — a row misaligned, a subtotal read as a total.

Documents that reference other documents. Terms incorporated by reference to an attachment that was not supplied. The system will answer from what it has and will not know what it is missing.

The verification advantage

Document processing has a property that makes it unusually safe: much of the output can be checked automatically.

Line items should sum to the net. Tax should be a defensible percentage. The supplier should exist in your system. Dates should be plausible. An invoice number should not already be recorded.

These checks catch a substantial share of extraction errors without human involvement, and they let you route only genuinely doubtful cases for review. Any serious document system should have this layer — ask about it directly, because a system without it is sending everything to a person or nothing.

How to start

Pick one document type with meaningful volume. Gather a realistic sample — including the bad scans and the odd suppliers, not just clean examples. Define which fields matter and what a wrong value would cost for each.

Then set a target that is not one hundred percent. Aim to handle the clear majority automatically with verification, and route the rest to a person with the document and extracted values side by side. That is a system that works in production, rather than one that works in a demonstration.

What to expect

A well-built system on a common document type will handle most items without correction, flag its uncertain cases, and make the remainder faster to process by hand. That is a substantial improvement. Providers promising complete automation across all document types are describing something we have not seen survive contact with a real post room.

All Articles
Let’s Talk

about the process
AI should run.