Fail-Closed by Design: Why Auditable AI Can’t Depend on Unconstrained Agency
Agents are useful. But when an AI result may have to be defended years later, agency and authority shouldn’t be the same thing.
The question that arrives later
Most conversations about AI agents start with what they can do: choose tools, make a plan, recover when the plan fails, keep going until the task is done.
In industrial operations and process safety, we learned to start with a different question: three years from now, after an incident, can we explain why the system made this particular claim?
Not why a similar model might say it today. Not the explanation the model wrote after the fact. What evidence was available, what processing happened, what supported the result, and what happened when the evidence was incomplete.
That question exists because of where the output ends up. Engineers write hazard studies that identify what can go wrong, the consequences, and the safeguards meant to prevent or mitigate them. In methods such as layers of protection analysis (LOPA), only safeguards that meet specific criteria earn risk-reduction credit. These studies are signed, periodically revalidated, reviewed by regulators and auditors and, after an incident, examined line by line.
When AI touches that chain, its output inherits the same scrutiny. That led us to a principle that sounds anti-agent at first:
The path that produces an authoritative result should not depend on unconstrained agency.
That doesn’t mean “don’t use LLMs,” and it doesn’t mean “don’t use agents.” We use both, every day. An LLM call isn’t an agent. The distinction that matters is between judgment and authority.
The trade hidden inside agentic systems
An auditor needs three things from a result: to bound what the system could have done, to reproduce what it did, and to defend why. The features that make agents attractive make each of those harder to guarantee, unless the agency is tightly constrained.
For many applications these are good trade-offs. A research assistant should be free to explore. But when a result may need defending later, the architecture has to start from that moment and work backwards.
Agency for challenge, not authority
Our answer wasn’t to get rid of agents. It was to give them a different job.
We think of the system as two paths. The trusted path produces the result someone may have to defend. It works from fixed inputs, its orchestration is code, its interfaces are explicit, and its consequential claims are tied to evidence.
Alongside it runs an agentic challenge path, and it is deliberately free. Agents explore alternate readings, look for missing information, question an assumption, and act as a second reviewer. Opportunistic and adversarial is exactly what you want from a reviewer.
The difference is what happens next. A challenger can propose. It can’t quietly promote its own conclusion into the authoritative result.
We use agents. We just don’t let agency define truth.
Different routes, not more votes
For the judgments that matter most, we care more about independent failure modes than about asking one model the same question several times.
There’s usually more than one way to approach the same evidence:
Deterministic: what follows from explicit, structured relationships? Reproducible and bounded, but it misses whatever nobody encoded.
Holistic model judgment: what does the evidence mean, taken together? It sees context and meaning, but can hide which assumption drove the answer.
Decomposed model judgments: does each condition the conclusion needs actually hold? Every step is inspectable, but splitting the question up can miss how the parts interact.
These are useful precisely because they fail differently. When they all agree, that’s interesting. When they disagree, that’s more interesting: it tells you exactly where review or more evidence is needed.
Diversity of failure modes matters more than number of votes.
Reconcile; don’t average
This is where most ensembles go wrong. Two methods say “yes,” one says “unknown.” It’s tempting to call that two out of three, so probably yes.
But an unknown isn’t a weak yes. It means a required fact hasn’t been established. Averaging it away changes what the result means.
So we prefer explicit reconciliation rules over voting. Agreement, with the evidence requirements met, lets a result proceed. Disagreement is kept and escalated. Missing evidence stays unknown. And no calculation runs through an unknown input.
UNKNOWN survives aggregation. Disagreement isn’t noise; it’s one of the best review signals you have.
Three rules on the trusted path
1. Fail closed: give the model a legitimate way to say “I don’t know”
If the only allowed answers are “yes” and “no,” a model will pick one even when the evidence supports neither. So every consequential step needs an explicit outcome for insufficient evidence, and that outcome goes to a person, never through as a default pass.
A missing verdict is not a pass. “Unconfirmed” means the documents don’t show it, not that the safeguard has failed. “There isn’t enough evidence here to establish this protection is adequate” is a different statement from “it’s adequate” or “it’s inadequate,” and downstream software has to keep it different.
2. Evidence, not citations
It’s easy to ask a model for an answer plus a source. But a model-written citation is still model-written text. It can be wrong and look perfect.
For consequential claims, the surrounding system should tie each claim to evidence it actually retrieved and keep that link itself. The model can point at evidence; it can’t create it.
A related discipline: evidence that something exists doesn’t establish its properties. A document can show a protection is there without showing it suits this scenario or is available in this operating state. Each property needs its own evidence.
3. Never compute over unknowns
Ordinary software wants to keep going: fill in a default, skip the missing term, produce a best estimate. In a safety calculation, a precise-looking number built on a hidden assumption is worse than no number.
If a required term is unknown, the result is not computable until that term is resolved.
The useful output is then not a guessed number but a statement of which evidence is missing. Engineers tend to like this: it tells them exactly what to go find.
Reproducibility, defined honestly
“Reproducible” gets used loosely in AI. Calling a model twice with the same settings doesn’t guarantee the same text forever: models change, providers change, models get retired.
So we split the claim:
We can reconstruct the run: what went in, what code ran, and which model was asked what.
We can replay the deterministic parts exactly: orchestration, rules and arithmetic.
We keep the judgment that was actually used, so every result traces back to it.
We don’t promise that asking the model again, years later, gives the same answer.
The narrower claim turns out to be the more credible one. Auditors already know models aren’t deterministic. What they want to know is where the non-determinism lives, and that it can’t quietly redefine the authoritative result.
Auditability doesn’t require pretending the model is deterministic.
A surprise: small models do more of the work
One side effect caught us off guard. Once a large reasoning problem is split into smaller, explicit judgments, each call often needs less model capability than a single large prompt would. That makes smaller models useful for far more work than we initially expected.
The point isn’t just a smaller inference bill. Cheaper judgments are what make independent routes, a second reviewer and disagreement detection affordable, instead of asking one large model to do everything at once. Here, architectural discipline and economics pull in the same direction.
What it costs
This approach isn’t free, and we don’t pretend otherwise:
Fewer open-ended questions. A bounded system only answers what it has been designed to answer.
Slower to demo. Each new capability needs explicit interfaces and validation before it ships.
Less impressive output. “Not computable” loses a sales demo to a fluent paragraph every time.
More software to own. Deterministic orchestration is code that somebody builds and maintains.
For work that feeds high-consequence engineering decisions, we think every one of those costs is worth paying.
Match agency to accountability
The lesson isn’t that every AI application should be built this way. If nobody will ever need to reconstruct a decision, wide agentic freedom may be exactly right. If somebody may have to defend the result years later, design the architecture backwards from that moment.
If you need auditable AI tomorrow:
Define your states, including “unknown,” before you write prompts.
Tie every consequential claim to evidence the system retrieved.
Keep orchestration separate from model judgment.
Preserve unknowns instead of smoothing them away.
Use independent approaches that fail differently.
Let agents challenge the trusted path, without becoming it.
Record enough to explain what actually happened.
The goal isn’t to make AI behave like an autonomous engineer. It’s to make every consequential AI judgment inspectable by one.
Ivo Dujmovic is co-founder and CEO of uc2c.ai, which builds industrial risk intelligence for process safety and operations. Connect with him on LinkedIn, or try ask.uc2c.ai.