The model can surprise you. The firm cannot.

Today’s LLMs — and most of what is sold as agentic AI — are non-deterministic.

Same inputs do not guarantee the same outputs. Plans drift. Tool calls wander. Multi-agent stacks multiply the branches.

That is useful in a lab. It is a liability in a firm that has to be insured, audited, and held to account.

Do not put that liability on the client

The Evolut Strategy does not stand up unsupervised agents inside the core and hope governance catches up.

We build a digital twin at the edge — still inside the firewall, reporting to the CEO. The twin is a deterministic system with a precisely defined set of potential outcomes.

Agents do not have agency. They do not have digital personhood. The human stays above the loop. Always.

Use the model. Bind the commit.

We use models and agents. We use them inside a guarded envelope, mostly for automation: classification, extraction, routing, drafting inside a schema, orchestration of known steps.

We do not expect the stack to invent a new firm. We expect direct effects on cost and cycle time. And indirect effects only humans can spend: judgment, ideation, strategy, the next line in the portfolio.

Determinism lives in the shell, not in the model

The model remains stochastic. Determinism is the contract around it.

A workflow is a chain: input arrives → information is gathered → a decision is taken → an action is committed → a record is left.

The twin rebuilds that chain as an intelligence stack — Sense, Interpret, Decide, Orchestrate, Learn — with a finite menu of admissible next states at each gate.

The model may propose. The harness accepts only what matches the schema, the policy, and the outcome set. Anything else is refused, queued, or rolled back. On material acts, commit is a human yes or no.

An agent that can originate authority is a junior executive without a contract. An agent that can only fill a slot in a verified workflow is automation. The first is a personhood claim. The second is a tool.

Neighbor one: harness the run, don’t hope the model behaves

Dhage’s 2026 study of harness engineering starts from an operational fact. LLM agents show large run-to-run variance on identical tasks, tools, and environments. Acceptable for exploration. Unacceptable in finance, compliance, or audit.

A first-pass harness — finite-state execution, forced tool selection, validation, retry — was not enough.

Adding structured planning changed the result. The plan must validate against a fixed schema before invoking any tool. Reproducibility and a determinism index reached 1.000 across 100 repeated runs in every model–task cell tested. Task success rose with it.

That is the technical cousin of the Evolut twin. The model is not asked to be reliable. The execution contract is.

Where we differ: those tests measure isolated pipelines. We put the same idea inside a CEO-owned replica of a live operation, run it beside the legacy path, and only deprecate the old workflow when cycle time, quality, cost, and risk force the comparison.

Neighbor two: industrial twins that validate before they act

The Digital Twin Consortium’s Industrial AI Agent Manifesto states the same law for plants and grids: deterministic validation and execution; process-aware constraints; emergency stop; human override; complete provenance.

The twin is the place where state, limits, and audit already live. An agent query is checked against current constraints before actuation.

Kraft Heinz is the factory picture. Their manufacturing twin produces line-by-line recommendations. The published operating model is staged: first the human still decides; then the human is in the loop and can still refuse; only later would machine talk to machine. The named effects are throughput and waste — not surprise.

Close. Still not the Evolut object. An industrial twin governs a physical process. Ours governs a firm workflow — the offer factory, the diligence pack, spare-parts identification, the onboarding file — and it reports to the chief executive, not to a plant system of record.

The same split, in software terms

The Convergent AI Agent Framework treats the harness as an asset: domain invariants in a registry, a deterministic assertion interface, state locking so the model cannot silently relax a red line. Humans authorize policy. They do not debug token by token.

Formal methods work on orchestration, bluntly. You cannot prove what the model will say. You can prove that the orchestrator terminates, respects budget, never grants an unprivileged tool call, and never deadlocks. Verify the part you control.

Sense and Interpret may use a model. Decide and Orchestrate commit only through gates. Learn writes back into the twin under review — not into unsupervised memory that later becomes policy.

What we refuse

Open agent ecosystems that treat software as a colleague. Digital personhood is a category error with legal consequences. Liability attaches to the firm.

If the system cannot name who authorized the act, which outcome set it belongs to, and how to roll it back, it is not ready for the edge — let alone the core.

Human-in-the-loop as theatre. A checkbox after the fact is not a control plane.

Above the loop means the C-suite designs the mandate, owns the reporting line, and sits at every material gate to approve or refuse. The agent is a supervised junior. Fluency in a demo does not make it an unsupervised executive.

Novelty as the success metric. Non-deterministic systems are good at surprise. A company is not.

Score the line. Rebuild two chains. Compare.

We score the 70% question — the high-margin line two people and an agent stack could copy in 60 to 90 days — then rebuild two high-throughput workflows inside a bounded twin.

California Management Review describes the fracture this design sits on. Older systems could behave deterministically at the governance layer. Generative stacks broke output-verifiable GRC.

The durable pattern they report in reinsurance is the one we start from: map the process as humans already run it, give the model the lower-risk stages, keep irreversible commits human-led, and compare against a champion path.

We add the rest of the Evolut move. The path lives in a twin at the perimeter. The mandate is written before the immune system attacks. Deprecation is earned in parallel, not declared in a steering committee.

Direct effects, then indirect ones

Direct: cycle time, unit cost, error rate, and risk on the two workflows. Comparable, because old and new see the same inputs.

Indirect: the capacity that falls out when coordination is no longer a management layer. That capacity is not “the agent’s idea.” It is time returned to work only people should do — and to the next offer in the portfolio.

That is how the edge grows into the core. Not how a model surprises the org chart.

If you want a conversation that scores the line and names the first two workflows — no deck, not a sales meeting — write to contact (at) evolut.ing with the high-margin chain you are actually worried about.

Send the notes!

Short pieces on AI-native work: workflows, guardrails, humans above the loop.