Evaluation suites, logs, rollback, review queues. Agents as supervised juniors — not an ethics appendix.
We do not treat agents as colleagues who have already earned autonomy. We treat them as junior staff: evaluated, logged, reversible, queued for review. That is not caution for its own sake. It is how a company remains a moral and legal actor while its tempo becomes inhuman.
Value alignment that arrives after deployment is decoration. Value alignment as guardrails around the stack is the work.
In 2026, that sentence stopped being philosophy. It became underwriting.
The supervision gap is now a balance-sheet problem
Gartner’s prediction is blunt. By 2027, 40 percent of enterprises will demote or decommission autonomous agents because governance gaps showed up only after production incidents.
Shiva Varma, who wrote the note:
Enterprises are treating AI agent governance as binary, either locked down or fully trusted, and that is the root cause of failure.
Agents operate at different autonomy levels and across different trust boundaries. When the same controls are applied indiscriminately, organizations encounter two common failure modes: over-restriction of simple agents, which slows delivery and drives shadow development, or under-restriction of more autonomous agents, which increases operational, security and compliance risk.
Uniform governance is not prudence. It is how you get both a black market of shadow agents and an unsupervised executive.
The legal default has not moved with the stack. Art Gilliland, to CSO:
If a company builds a system and that system causes damage, the company should own the outcome. The alternative, where nobody is responsible because ‘the system did it’ is a loophole big enough to drive a truck through.
Insurers have noticed. From January 2026, ISO’s CG 40 47 and sister forms let carriers carve generative AI out of ordinary general liability. Cyber markets are rewriting language around agents that act as designed and still cause loss. A twin without reconstructable records is not a strategic asset. It is an uninsurable one.
Singapore’s IMDA, in the Model AI Governance Framework for Agentic AI, put the same point in official prose: bound the agent’s powers up front; make humans meaningfully accountable; put technical controls in the lifecycle; do not pretend continuous human oversight of every step will scale.
Boards are next. Mehtab Barkat names the standard taking shape — cognitive adequacy. A director who cannot describe an agent’s failure modes, reward hijacking, or drift is not exercising care. Periodic slides after the fact are not oversight of a system that closes an OODA loop in hours.
What you can hold accountable is the governance around it: who approved its scope, what controls existed, and whether the deployment matched the risk.
Not ‘Did the agent do something bad?’ but ‘Did you have review, escalation, and containment for agent behavior before you deployed it?’
That is the product.
What EU law already requires — and what it does not yet
This is the file in September 2026. There is no European “agent act.”
Live now. The AI Act (Regulation (EU) 2024/1689), as amended by the Digital Omnibus (Regulation (EU) 2026/1744). Prohibitions and the Article 4 literacy duty since 2 February 2025. GPAI-provider rules since 2 August 2025. Article 50 transparency — disclose AI interaction; mark synthetic content — since 2 August 2026 (marking grace to 2 December 2026 for systems already on the market). The AI Office and national authorities enforce this from that date. Official reading, not statute: Commission guidelines on prohibitions and on who is a GPAI provider; the GPAI Code of Practice; the Code of Practice on marking AI-generated content.
Article 4 is the board clause already in force: staff must be literate enough to authorize the system. That is the statutory cousin of cognitive adequacy.
On published dates. Annex III high-risk (employment, credit, education, essential services, biometrics, certain justice and migration uses), including Article 14 human oversight, from 2 December 2027. Annex I embedded high-risk from 2 August 2028. A twin that scores people or access is on that clock. An internal invoice twin usually is not.
Liability without an AI liability directive. The Commission withdrew that proposal in 2025. No replacement. What does arrive is the revised Product Liability Directive (EU) 2024/2853: to be transposed by 9 December 2026; software and AI are products; manufacturers can be strictly liable for defects from updates they control and from continuous learning under their control. Harm is still the PLD’s category — persons, not every commercial loss. National fault law fills the rest.
Still in force. GDPR on personal data the agents touch, including logs. NIS2 on cybersecurity for essential entities; an agent with credentials is a non-human identity inside that perimeter.
The Act does not say “lease.” What it does say: literacy now; transparency now; high-risk oversight on a calendar; strict product liability for learning systems from December 2026; no EU-wide fault shortcut. The guardrail product is how you meet those duties before the record is empty.
Five things the guardrail product has to be
1. Inventory and a named owner
No owner, no agent. Every running agent — including the ones a business unit stood up last Tuesday — has a human whose name sits on purpose, access, and review. Visibility first. Guesswork is not governance.
No legal personhood for the agent!
2. Leased authority, not permanent power
Purpose, scope, and time. An agent may exist as a persistent identity. Its right to act should expire. Authority is borrowed for a job, re-checked at each consequential step, and allowed to lapse while the agent itself remains. Permanent credentials on a junior who never sleeps is how scope becomes a career.
3. Proportional controls by autonomy — and by access
Gartner’s useful split: autonomy is what the agent can do; scope is what it can reach. Risk rises with either. A read-only summariser and a payment agent do not share a policy. One-size-fits-all produces the two failures above. Classify. Then govern the class.
4. Circuit breakers and reconstructable records
Evaluation suites before production. Logs that can answer who authorized, on whose behalf, with what minimum privilege, and what actually happened. Rollback that does not require a war room to remember why the work started. If you cannot reconstruct the chain, you cannot defend it, insure it, or learn from it.
5. What the board must be able to understand
Not the model card. The failure modes. Where human approval is mandatory. Override rates and time-to-contain. Which agents hold which leases. Cognitive adequacy is not fluency in transformers. It is the ability to refuse a system whose drift you cannot see.
IMDA’s caution belongs on that board pack: meaningful human control has to live in the lifecycle — and continuous oversight of every token is not the design. Humans stay above the loop: checkpoints, not theatre.
How this sits in the twin
The Evolut twin is a digital operation at the edge, inside the firewall, reporting to the CEO. That mandate is worthless if the twin cannot be insured, audited, or rolled back.
Governance is not a committee that visits the twin. It is the band around the intelligence stack: Sense, Interpret, Decide, Orchestrate, Learn — each layer with an evaluation, a log, a human yes/no on material acts, a way back.
Agents as supervised juniors. Not unsupervised executives. Ethics by design, not bolted on.
Build the guardrails as software you ship with the first two workflows. Then the parallel run can compare cycle time, quality, cost, and risk. Without the last number, “deprecate the loser” is a slogan.
A simple test
Could we explain this agent to an underwriter, a director, and a court — from the same record — in an afternoon?
If no, the twin is not ready to run in parallel. Inventory, lease, class, breaker, board pack. Then turn it on.
If yes, keep the lease short. The junior does not get tenure because last week’s score looked good.
The picture is not “slow the agents down.” The picture is: keep the judgment, keep the name that can be sued, keep a record that can be read — and stop asking an ethics paragraph to do the work of a product.

