Human-in-the-Loop

The system proposes; a person commits.

Keeping a person inside an automated decision by design rather than by accident. The pattern names where human judgment sits — approving every action, supervising a stream of them, or absent altogether — and what each choice costs.

Origin

The phrase grew out of control engineering and military simulation, where man-in-the-loop named a system whose operation depends on a human operator: a missile steered by a person at the controls, or a flight simulator whose realism comes from a real pilot flying it. There is no founding paper. The term spread as the name for a design decision rather than a discovery.

Thomas Sheridan and William Verplank gave the idea a spine in a 1978 MIT report, grading automation on ten levels that run from a computer offering no help at all to one that acts alone and tells no one. Machine learning borrowed the phrase in the 2010s for labeling and correction workflows, and generative systems turned it into a product question: which actions may a model take by itself.

The Pattern

Keep a person inside the loop wherever an action is costly to reverse, hard to verify, or answerable to someone. The system proposes and the person commits, and the handoff between them is designed rather than assumed.

Origin
Thomas Sheridan and William Verplank (1978)
Source
Levels of automation, from control engineering
In practice
Gate on consequence, not on confidence
When to Use
How to Use
Postures 01 / 10

In, On, or Out of the Loop

Three postures: a person approves every action, supervises a stream of them, or is absent entirely. Which one a feature uses is a decision to write down, not a property of the model.

Best for
Scoping an automated feature
Use when
Autonomy is being decided
Avoid when
The action is trivially reversible
Approval 02 / 10

Propose, Then Commit

The model drafts and a person commits. The draft saves the work; the commit keeps the accountability with someone who can be asked why the action was taken.

Best for
Irreversible actions
Use when
A mistake costs money or trust
Avoid when
Volume turns review into a formality
Routing 03 / 10

Confidence Thresholds

Send the confident cases straight through and route the rest to a person. The threshold is a product decision tuned to the cost of a wrong answer, not a model setting left at its default.

Best for
Classification at volume
Use when
Confidence scores are calibrated
Avoid when
Scores are uncalibrated guesses
Handoff 04 / 10

Escalation with Context

When the system hands a case to a person it must hand over the context too: what it tried, what it is unsure about, and what happens if nobody acts on it.

Best for
Support and moderation queues
Use when
Cases arrive mid-flow
Avoid when
The reviewer already sees everything
Learning 05 / 10

Corrections as Training Data

Every correction a reviewer makes is a labeled example. Capture it deliberately, with the reason attached, or the loop teaches the system nothing and the same case returns next week.

Best for
Models that keep learning
Use when
Reviewers already fix output
Avoid when
Corrections cannot be attributed
Failure mode 06 / 10

The Rubber Stamp

A reviewer who approves nine hundred correct suggestions will approve the next one without reading it. Review that is never wrong has stopped being review.

Best for
Auditing a loop that looks healthy
Use when
Approval rates approach 100 percent
Avoid when
Sampling already catches the errors
Cost 07 / 10

Review Does Not Scale

A person in every loop sets the ceiling on throughput. Measure seconds per review and queue depth before promising a human check on everything.

Best for
Capacity planning
Use when
Volume grows faster than the team
Avoid when
Volume is small and steady
Accountability 08 / 10

Someone Owns the Outcome

Log what the system proposed, who approved it, and what they could see at the time. Without that trail the loop cannot be audited and nobody can answer for the result.

Best for
Regulated workflows
Use when
Decisions affect people
Avoid when
The action carries no consequence
✦

Human-in-the-Loop in the Age of AI

Agents act in sequences, so the question moves from approving an output to approving a step, and from reviewing everything to choosing what may run unattended.

✦ AI Era 09 / 10

Approval Gates for Agents

An agent that can send, buy, delete, or deploy should stop at those steps and ask. Gate on the consequence rather than the model's confidence, and make each pause cheap to clear.

Shift
Output review → step approval
Use when
Actions leave the product
Watch for
Prompts so frequent they are dismissed
✦ AI Era 10 / 10

Supervising by Sample

When volume outgrows review, move from in the loop to on the loop: sample the stream, watch the error rate, and pull the thread the moment it moves.

Shift
Every case → a sampled case
Use when
Throughput outruns reviewers
Watch for
Rare failures a sample never sees
Further Reading