Human in the loop
Oversight is worth what the person can actually check. Attach it to irreversible actions, with real arguments shown.
Part of the Agents track on lAItest.
The question is not whether a human approves. It is which step, and what they are shown at the moment they do.
A common misconception
Commonly believed: Human in the loop means a person watches the agent work and stops it if something goes wrong.
Actually: Watching does not scale, and people rubber-stamp what they cannot read. Oversight that works is narrow: one approval, attached to one irreversible action, showing the exact arguments, immediately before it runs.
Sort actions by reversibility
Reading is cheap to get wrong. Writing is recoverable if the old version was kept. Sending a message, moving money and deleting data are not recoverable at all. A sane agent runs the first class freely, logs the second, and stops at the third for a person who is shown exactly what is about to happen.
Elicitation is the protocol version of this
An MCP client can offer elicitation, which lets a server ask the user for a missing detail during a task instead of letting the model invent one. Same principle as an approval gate: when the information or the authority is not inside the loop, get it from the human rather than guessing plausibly.
Which approval design is most likely to catch a genuine mistake?
Answer: A confirmation showing the exact arguments of one irreversible action, immediately before it runs. A plan approved up front describes intentions, not the arguments that eventually get sent, and a log arrives after the damage. Approval works when it is close in time to the action and specific about what that action will do.
In one sentence
Approval is worth exactly what the person can check, so attach it to irreversible actions and show the arguments that will really be sent.