Human in the Loop: Designing Review Steps That Work

Every responsible AI deployment includes human review. Most of them include it badly.

A review step that people click through without reading is worse than no review step, because it manufactures the appearance of control while providing none — and it moves accountability onto someone who was never given a real chance to catch the problem.

Why review steps decay

When a system is accurate most of the time, reviewers stop expecting errors. Approving becomes reflex. This is not carelessness; it is the predictable result of asking a person to look for something that is almost never there.

Paradoxically, the better the system, the faster the review step decays. A system that is right ninety-nine percent of the time will have reviewers approving without reading within weeks, which means the one percent goes through unexamined.

Make the review a real task

The fix is to give the reviewer something to do rather than something to approve.

Instead of "here is the extracted invoice data, approve or reject", show the source document alongside the extracted values with the relevant passage highlighted, and ask the reviewer to confirm the three fields that matter most. Now they are performing a check rather than authorising an outcome.

Where practical, require an action that cannot be performed without engaging — entering a value, selecting between genuine alternatives, marking which source supports the conclusion.

Route by uncertainty, not uniformly

Reviewing everything equally spends the same attention on the obvious and the doubtful. Better systems distinguish the two.

Straightforward cases matching an expected pattern can go through with light checking. Cases that are unusual, incomplete, high-value or where the system's own signals are weak get full review. This concentrates human attention where it can actually change the outcome, and it keeps the review population small enough that reviewers stay alert.

Give reviewers the time

A review step assigned to someone with no allocated time will be performed at the speed that fits the time available, which is to say instantly.

If a process produces sixty items daily requiring a minute each, that is an hour of someone's day. Either it is in their workload or the review is fictional. This is a planning decision, not a detail.

Close the loop

When a reviewer corrects something, that correction should go somewhere. If it disappears, three things follow: the same error recurs, the reviewer concludes that reporting is pointless, and you lose your only production quality signal.

Collect corrections, look at them periodically, and use them to improve the system. Then tell the reviewers what changed. Visible response is what keeps reporting alive.

Measure whether it works

You can test this directly. Periodically introduce a known error and see whether it is caught. If it passes review consistently, your review step is decorative and you should know that before a customer discovers it.

This sounds adversarial but it is ordinary quality assurance, and it is far kinder than discovering the same fact through an incident.

Be explicit about responsibility

State in writing what the reviewer is accountable for. Confirming three specific fields is a duty someone can discharge. "Approving the output" is not, and it leaves the reviewer carrying risk they cannot manage.

All Articles
Let’s Talk

about the process
AI should run.