The twin is only as safe as the object it represents.

AI transformation safety is not a model problem. It is a binding problem: the digital twin must prove it is talking about this vial, this wafer, this implant – and not a plausible substitute.

A company can stand up an intelligence stack that senses, interprets, decides, orchestrates, and learns. It can put humans above the loop. It can log every decision and roll back every action.

None of that is enough if the stack cannot tell a genuine part from a counterfeit, a sterile lot from a diverted one, a calibrated instrument from a lookalike.

The failure is not in the agent. The failure is at the seam between bits and matter.

That seam is where most “AI transformation safety” programs go quiet. They govern prompts, policies, evaluations, and access. They do not govern whether the object in the room is the object in the record. In software-only businesses, you can postpone this. In engineering and health, it cannot.

An agent that releases a batch, signs a work order, or recommends a dose is not supervising text. It is supervising a physical consequence.

There is now a concrete industrial answer to that seam. Merck KGaA in Darmstadt, Germany, built it after centuries of treating product quality as a physical fact rather than a spreadsheet column. They call the platform M-Trust™. The useful lesson is not the brand. It is the method.

Two twins, one risk

Evolut talks about a digital twin of the firm: a governed replica of high-throughput workflows, run in parallel with the core, retired only when the numbers force it.

M-Trust describes a digital twin of the object: a decentralized identifier bound to a physical good so authenticity, quality events, and custody can be verified as the object moves.

These are not the same twin. They become dangerous when people treat them as if they were.

An organizational twin that is fed unverified physical data will learn the wrong world at machine speed. A product twin that is never connected to the decision loop will sit in a compliance drawer while agents act on ERP fields that anyone could have typed.

Safety in an AI-native operation is the join between the two. The firm-twin must be allowed to act only on object-twins whose binding to matter has been proven — not assumed, not inferred from a barcode that can be photocopied.

Merck’s own language is exact about the problem they set out to solve: product safety, traceability, and counterfeiting; quality control that still depends on human inspection as processes become fully machine-executed; the need for machines to identify objects reliably and securely. The platform is a cyber-physical trust layer that immutably links the physical and digital worlds and turns human quality-control capability into less error-prone machine-to-machine interaction.

That last clause is the one AI programs should steal. Agents do not need more dashboards. They need counterparts in the physical world that can answer, under cryptographic and material proof: yes, this is the thing you think it is.

Why the physical connection is not optional in safety work

An AI transformation safety project has three jobs that look digital but aren’t.

Sense must be about the world, not the warehouse system. The first layer of an intelligence stack is Sense. If Sense reads only from internal databases, the stack is a hall of mirrors. Counterfeit APIs, diverted semiconductors, swapped implants, and relabelled reagents all look perfect in software. They only fail when someone is harmed, a line stops, or a regulator arrives. In health and engineering, that failure is the safety incident.

Decide must not outrun evidence. Agents compress OODA. That is the point. Compression without a physical check means a wrong release decision can propagate across a value chain before a human notices the smell, the color, the torque, the batch discrepancy. Human checkpoints above the loop are necessary. They are not sufficient if the human is reviewing the same unanchored record the agent used.

Learn must not train on laundered reality. Partnerships around M-Trust already name the AI consequence. Zebra and Merck framed their collaboration as a way to raise product safety and traceability and to supply high-quality data for training and testing AI models — data whose authenticity and origin can be trusted. That is the correct order. First bind the object. Then let models learn from it. The reverse order produces confident systems that have memorized fraud.

Regulated industries already know the paper version of this problem. Certificates travel. Objects do not always travel with them. Digital Product Passports, serialization, and track-and-trace try to close the gap. They still fail if the identifier can be cloned, the document can be copied, or the scan never touches a property of the object itself.

A trusted connection to the physical world is therefore not a late integration. It is a precondition for letting agents anywhere near a release gate.

Ways to create trust between the digital and the physical

Trust at this boundary is a stack of bindings. Each can be attacked. Each must fail closed.

1. Give the object an identity that is not a sticker.
A serial number, QR code, or RFID tag is a label. Labels can be copied. A serious binding starts from something hard to separate from the object: a security pigment in the material or coating; a microscopic transponder embedded in the product, label, or pack; a physical unclonable function; even a machine-learned fingerprint of surface structure.

M-Trust calls these crypto anchors. They choose the right anchor for the security need rather than forcing one form factor.

2. Derive the digital twin from the anchor, not the other way around.
The important direction of travel is physical → digital. Merck’s anchoring work generates reproducible, collision-resistant hashes from an anchor’s fingerprint, optionally mixing in ambient context—time, location, temperature, pressure—so the twin isn’t a free-floating document later pointed at a box. In the platform, the twin is a DID document. Resolving it returns verification methods, not a marketing page.

3. Read with devices that are part of the trust system.
An anchor that cannot be interrogated in the field is a lab curiosity. M-Trust ships readers for the current families of anchors: a SEC-Reader for security inks, an IMP-Reader for microtransponders, and a station to keep them alive. The developer surface is a console, Flutter SDKs, and APIs. Verification is a first-class act: scan, check, green or not. The Companion flow exists so that act can be rehearsed before it is wired into production.

4. Treat verification as a credential, not a screenshot.
A scan that lives in a PDF attached to an email is not a safety control. A scan that becomes a signed event — a verifiable credential, a ledger transaction, an update to the twin — can be checked later by a different party, a different machine, a regulator, or an agent.

The Gemini partnership of Merck and Google to accelerate Agentic AI enterprise transformation creates and resolves twins and issues credentials; the transaction layer writes to a ledger when that is the chosen persistence. Web 3.0 here is not décor. It is how two machines that do not share a boss can still share a fact.

5. Keep the physical and digital representations in lockstep.
The binding is only useful if it stays true as the object changes state: tested, released, shipped, installed, recalled. Updates to the twin are permissioned, keyed, and attributable. Hardware and software are one portfolio, not a middleware afterthought. Merck’s differentiation claim is the patent surface covering both sides of the seam — crypto-chemicals and crypto-objects on the material side, hash-entropy methods on the digital side.

6. Design for the regulation that is already coming.
The EU Digital Product Passport is the obvious forcing function. The deeper requirement is older: in pharma, devices, chemicals, food, automotive, and electronics, you must be able to show what a thing is, where it has been, and whether it is still the same thing.

None of these steps replace human judgment. They give judgment something that can count as evidence.

Engineering and health: where an unbound agent becomes a physical event

Health
A counterfeit or diverted medicine, a misidentified implant, a reagent that is not the reagent, a cold-chain claim that was never measured on this shipper — these are not data-quality issues. They are patient-safety issues that happen to be recorded in software. An agent that drafts a batch disposition, flags a shortage alternative, or pre-fills a recall scope will act at a speed no pharmacovigilance meeting can match.

If the pack’s identity is a printable code, the agent is supervising a forgery surface. If the pack carries an anchor that can be read, hashed, and bound to a twin, the agent is supervising a fact. Materials safety and supply-chain transparency are not adjacent to care. In modern manufacturing of therapies and devices, they are care.

Engineering
A wafer, a safety-critical fastener, a calibrated sensor, a composite coupon, a spare that will go into an aircraft or a plant — the same logic applies. Semiconductor work already treats provenance as a security problem: chips are rerouted, altered, counterfeited. The cyber-physical layer is what sits above materials science: the component gets a record that is not merely logged but tamper-evident at each handoff.

In capital equipment, “the digital thread” has been promised for a decade. Most threads break at the first contractor who reprints a traveler. An anchor that lives in the part does not.

In both domains, the safety case for AI is easy to state and hard to implement:

  • Agents may recommend. They may not release.
  • Release requires a fresh physical verification against the twin, not a cached status bit.
  • Training data for the next model is accepted only if it originates from an anchored event.
  • Recall, containment, and audit walk the same chain the agent walked — object, anchor, twin, credential, human checkpoint.

If any link is a shared password and a CSV, the safety case is theatre.

What this does to an Evolut stack

The Evolut Strategy already insists on a twin at the edge, two high-throughput workflows, parallel run, humans above the loop, evaluation, logs, rollback.

Add a sixth constraint for any workflow that touches matter:

Sense may not treat an unbound identifier as an observation.

Practically:

  • When you score the 70% question in a health or engineering line, ask whether a two-person agent team could copy the paperwork of the line or the object. If they can only copy the paperwork, the exposed margin is fake, and so is the safety story.
  • Pick first workflows that close on a physical gate: incoming identity check, in-process quality hold, release, install, returns. Those workflows already have inspectors. They are the right place to replace heroic inspection with anchored machine-to-machine checks plus a human yes/no.
  • Put the object-twin in the system of record for that workflow. The firm-twin orchestrates. It does not invent identity.
  • Treat every successful scan as training-grade data and every failed scan as an incident precursor, not as a UI state.
  • Do not let the intelligence stack “learn around” a missing reader. Absence of a physical check is a stop, not a low-confidence score.

This is slower than a chatbot on the warehouse laptop. It is the only version that deserves the word safety.

Merck did not wait for AGI to take product identity seriously. They treated the physical object as the root of trust because, in their industries, it always was. AI does not repeal that. It raises the cost of ignoring it.

The digital twin of the company will be evolut_ing into the core only if the twins of the objects it handles can survive contact with the world. Build the binding. Then let the agents work.

Sources used for the M-Trust facts: mtrust.io, docs.mtrust.io/guides, Merck’s January 2025 launch materials, the Kugler Q&A on anchoring patents, the Zebra collaboration. The piece is not sponsored copy; M-Trust is the worked example of a method, not a product recommendation.

Send the notes!

Short pieces on AI-native work: workflows, guardrails, humans above the loop.