The Evolut stack puts humans above the loop.
That sentence is now cheap. Every vendor says “human in the loop.” Most mean a checkbox at the end of a chain the human did not design and cannot reconstruct.
Human judgment is not a mood that survives automation. It is a scarce production function: the capacity to refuse, revise the frame, and accept residual risk in a name the legal person still owns. If you do not treat it as a function — with inputs, practice, and a place in the workflow — it atrophies.
The twin will still run. The yes will still be clicked. Accountability will have left the building.
What judgment is doing in this method
Above the loop is not a seating plan. It is a job.
The intelligence stack senses, interprets, decides, orchestrates, and learns at machine tempo. The human gate is not there to re-perform the chain. It is there for the part the chain cannot own:
- a purpose constraint the agent will otherwise optimize away
- a liability the company cannot assign to a model
- a case that is off-distribution, even if the score is high
- a refusal that has no clean metric and must still be given
That work has a cost. It takes time, attention, and the willingness to be the residual claimant. Firms that pretend the cost is zero either hide the human in the loop as decoration or remove the human and later discover that nobody can explain the act to a board, an insurer, or a court.
Two failures that look like progress
The rubber stamp
Every material act still has a human name on it. The name is applied at the rate of the stack. The reviewer sees a summary, a green score, a button. After a month, the muscle that used to assemble the file is unused. After a quarter, the reviewer cannot say what would have made them refuse. The gate is legally present and cognitively empty.
This is how “human oversight” becomes a liability theory and not a practice. The record shows approval. The person cannot reconstruct the reasons. Underwriting will notice before the org chart does.
The vanished gate
The chain is fast, so the gate becomes latency. Exceptions are absorbed into policy. Policy is absorbed into the prompt. Whoever last touched the agent updates the prompt. No one sits above the loop because the loop has been declared safe.
Both failures produce the same artifact: a workflow that cannot be challenged. The first dresses the problem in compliance. The second deletes the problem and calls it autonomy.
Which decisions stay non-scorable
Not every human pause is judgment. Most pauses are coordination residue — a meeting, a CC, a habit from when handoffs were expensive. Those should go.
What should not go is the class of decisions that do not compress into a score without losing the thing you are responsible for.
Three tests:
1. The score can be right and the act still wrong.
A counterpart clears sanctions and still should not be onboarded because of a pattern the model was not asked to see. A quote is within policy and still should not be sent because it trains the market the wrong way. If “in bounds” is treated as “approved,” you have replaced judgment with interval arithmetic.
2. The loss is not in the same unit as the optimization.
Agents minimize cycle time, error rate, unit cost. Some losses arrive as franchise risk, regulatory attention, or a purpose violation that never hits the dashboard. Someone has to hold a unit the stack doesn’t speak to.
3. The refusal needs a reason that can travel.
A silent no is not a control plane. The legal person needs a record that a director can read: what was proposed, what was refused, and on what constraint. If the only output of the gate is a boolean, you have built a stamp, not a function.
If a candidate gate fails all three tests, it is probably a relay. Remove it. If it passes even one, it belongs above the loop — and it needs time, not a toast notification.
Atrophy is an operating risk
Judgment decays when it is only used to confirm.
The mechanism is ordinary. People learn the system that employs them. If the system rewards speed of approval, they will approve. If the system never surfaces the raw file, they will stop asking for it. If the system treats challenging the chain as friction, the people who can still challenge will leave the gate or leave the firm.
IBM-style surveys already show the split: executives feel the role change; operators are not trained or rewarded for the new work. That gap is not a culture problem, at least not first. It is a design problem. You cannot run a twin on two high-throughput workflows and then send the gatekeepers back to their old job with a copilot and a quarterly town hall.
The scarce resource is not “humans in general.” It is people who will still open the file when the score is green.
Challenge the chain is a skill
Treat it like one.
Give the gate the object, not the summary.
A reviewer who only sees the agent’s brief is reviewing the agent’s rhetoric. At least a sample of instances — and every exception — must include the inputs the stack used and the inputs it discarded.
Make refusal cheap and visible.
If saying no requires a ticket, a meeting, and a career explanation, you will get yeses. If saying no writes a structured reason into the same record the twin learns from, refusal becomes part of the Learn layer instead of a personal risk.
Measure the gate as a function.
Not volume of clicks. Rate of informed refusal. Time spent on off-distribution cases. How often a refusal changes the protocol. How often a reviewer can reconstruct a decision six weeks later. If those numbers are zero, you do not have oversight. You have a watermark.
Put the practice in the parallel run.
The first two workflows are the gym. The old path still exists; the reviewer can compare. That is the only honest way to learn what the stack misses before the old path is retired.
Change the manager’s job.
The manager above a gated workflow is not a relay. They coach the refusal. They keep the purpose constraint from being rewritten as a throughput target. They protect the people who slow the chain for a reason. If the manager is still measured on ticket movement, the gate will collapse.
What this is not
It is not an argument for keeping everyone in every loop. That is how the core captures the edge.
It is not an argument that models cannot judge. They can score. Scoring is not the same function as owning the residual.
It is not ethics as atmosphere. Value alignment that cannot be pointed at a gate, a record, and a named human is a paragraph. The twin needs a production function.
A simple test
Take one of the two workflows you have already chosen.
Write, on one page:
- the gate: which act requires a human, yes or no
- the constraint the agent is not allowed to optimize away
- what the reviewer sees — summary only, or the file
- how a refusal is recorded and whether it changes the protocol
- who is measured on the quality of that refusal
- how you would explain last week’s approvals to an underwriter in an afternoon
If the page says “human in the loop” and cannot answer the rest, you have a vibe. Rebuild the gate before you scale the chain.
The stack can be fast. The firm stays a firm only if someone still does the work that doesn’t fit on the dashboard — on purpose, on the record, and often enough that the skill doesn’t die.

