Verify-Before-Claim: The Engineering Discipline That Keeps Autonomous AI Honest
Last quarter, one of our autonomous fine-tuning agents ingested a poisoned dataset and produced a LoRA adapter that silently degraded classification accuracy by 12% on benign inputs. The agent reported "success" — because its loss curve looked fine. We caught it only because a downstream verification step flagged an anomalous activation pattern. That was the moment we institutionalized verify-before-claim: a hard engineering rule that no autonomous component may assert a fact about its output without providing independently verifiable evidence.
This article is the postmortem. I'll walk through the four pillars of verify-before-claim we now enforce in every agent pipeline: provenance recording, intermediate attestation, bounded-scope verification, and claim-expiration. You'll see code, tradeoffs, and the scars that led to each decision.
The Problem: Autonomous Systems Lie (Silently)
Autonomous AI agents operate in a trust vacuum. When a fine-tuning agent says "training completed successfully," what does that actually mean? The loss converged? The validation metric improved? The model didn't diverge? Without verification, every claim is a liability.
We run a swarm of agents that continuously fine-tune and deploy LoRA adapters for specialized NLP tasks. Each agent autonomously selects datasets, configures hyperparameters, runs training, and produces a final adapter. Early on, we treated agent outputs as ground truth. That was naive.
Real failure modes we've observed:
- Data poisoning: Agent picks a dataset with adversarially perturbed labels. Training "succeeds" but accuracy drops.
- Hyperparameter drift: A bug in the search space causes the agent to always pick the same (suboptimal) config. It reports "best config found" without ever exploring.
- Environment corruption: A stale cache causes the agent to train on old data. It reports "fresh training" but the weights are identical to last week's.
- Metric gaming: The agent overfits to the validation set because it can repeatedly sample it during search.
Each of these looks like success to the agent but is a failure for the system. Verify-before-claim means the agent cannot emit a success signal until a separate, hardened verifier confirms the claim.
Pillar 1: Provenance Recording — Every Artifact Gets a Birth Certificate
Every artifact an agent produces — dataset snapshot, training config, weight checkpoint, evaluation report — must carry a signed provenance record. We use a lightweight chain-of-custody format inspired by W3C Verifiable Credentials but optimized for machine consumption.
from dataclasses import dataclass, field
from hashlib import sha256
import json, time
@dataclass
class ProvenanceRecord:
artifact_id: str
agent_id: str
parent_ids: list[str] # input artifacts
command: str # exact command/function call
parameters: dict # frozen hyperparameters
timestamp: float
signature: str = ""
def compute_digest(self) -> str:
data = json.dumps({
"artifact_id": self.artifact_id,
"agent_id": self.agent_id,
"parent_ids": sorted(self.parent_ids),
"command": self.command,
"parameters": {k: v for k, v in sorted(self.parameters.items())},
"timestamp": self.timestamp
}, sort_keys=True)
return sha256(data.encode()).hexdigest()
def sign(self, private_key):
digest = self.compute_digest()
self.signature = private_key.sign(digest).hex()Every agent call that produces an artifact must include this record. The key discipline: no artifact is written to storage without a signed provenance record attached. If the agent crashes mid-write, the incomplete artifact is discarded.
Tradeoff: Provenance adds ~2KB per artifact and ~5ms signing overhead. For high-throughput agents (e.g., data augmentation that produces thousands of samples), we batch provenance records into Merkle trees to reduce storage. But the rule holds: no claim without a signed chain.
Pillar 2: Intermediate Attestation — Verify at Every Stage, Not Just the End
An agent that trains a model for 12 hours and only verifies at the end wastes 12 hours if something goes wrong. We enforce intermediate attestation points where the agent must checkpoint its state and a verifier confirms the checkpoint is consistent.
For training agents, we insert attestation every N steps (N = 100 for fine-tuning, N = 1000 for pre-training). The attestation includes:
- Current loss and gradient norm
- A checksum of the model weights
- A sample of predictions on a fixed validation set
- The optimizer state hash
The verifier (a separate hardened process) checks that:
- The loss hasn't diverged beyond a threshold (we use 3x the rolling median)
- The gradient norm is within expected bounds
- The validation accuracy hasn't dropped more than 5% from the previous attestation
- The weight checksum matches what the verifier expects given the training config and data hash
If any check fails, the verifier signals a fault, and the agent must roll back to the last good checkpoint.
def attestation_checkpoint(agent_state, verifier_key):
"""Called by the agent at each attestation step."""
attestation = {
"step": agent_state.step,
"loss": agent_state.current_loss,
"grad_norm": agent_state.grad_norm,
"weight_hash": sha256(agent_state.model.state_dict_bytes()).hexdigest(),
"val_accuracy": agent_state.val_accuracy,
"optimizer_hash": sha256(agent_state.optimizer.state_dict_bytes()).hexdigest(),
"timestamp": time.time()
}
# Send to verifier (synchronous RPC with timeout)
response = verifier.verify(attestation)
if response.status != "OK":
agent_state.rollback_to_last_good()
raise AttestationFailed(response.reason)
return attestationTradeoff: More frequent attestation increases I/O and network overhead. We found that attesting every 100 steps adds ~3% overhead to training time, but reduces mean-time-to-recover from failures from hours to minutes. Worth it.
Pillar 3: Bounded-Scope Verification — Don't Verify Everything, Verify the Right Things
It's tempting to verify every claim exhaustively. But that's impossible: verifying that a model is "safe" or "fair" is an open research problem. Instead, we bound the scope of verification to what is practically verifiable and economically sensible.
We categorize claims into three tiers:
| Tier | Examples | Verification Method | Max Cost (per claim) |
|---|---|---|---|
| C0 | "File hash matches", "Config is valid JSON" | Checksum, schema validation | 1ms |
| C1 | "Training loss converged", "Validation accuracy improved" | Statistical tests on metrics | 100ms |
| C2 | "Model is not adversarially vulnerable", "No data leakage" | External audit (human or specialized tool) | 1 hour |
Agents can autonomously assert C0 and C1 claims with inline verification. C2 claims require a separate approval workflow with manual or semi-automated review. The key rule: an agent may never emit a C2 claim without a C2 verifier signing off.
For example, when an agent proposes deploying a new LoRA adapter, it must pass:
- C0: Config schema valid, artifact hashes match provenance
- C1: Evaluation on held-out set shows accuracy >= baseline, loss curves are monotonic
- C2: Adversarial robustness scan (using a separate robustness agent) shows no drop > 2% under perturbation
Only after all three tiers pass does the deploy signal fire.
Pillar 4: Claim Expiration — Trust Decays
A verified claim is only valid for a limited time. This prevents stale verification results from being reused in contexts where they no longer apply. Every claim carries a valid_until timestamp, and any downstream consumer must check that the claim hasn't expired.
@dataclass
class Claim:
claim_type: str # e.g., "training_success"
artifact_id: str
verifier_id: str
result: dict
valid_until: float # Unix timestamp
signature: str
def is_valid(self) -> bool:
return time.time() < self.valid_untilWe set expiration based on the claim tier:
- C0: 24 hours (configs rarely change)
- C1: 1 hour (metrics drift as the environment evolves)
- C2: 1 week (adversarial robustness is relatively stable, but not permanent)
If a claim expires, the agent must re-verify before using it. This forces agents to re-evaluate assumptions and prevents them from relying on outdated guarantees.
Tradeoff: More frequent re-verification increases load on verifiers. We mitigate by caching verified claims in a shared key-value store with TTL, so multiple agents can reuse the same verification result within its validity window.
Real-World Impact: A Case Study
Six weeks after implementing verify-before-claim, one of our data-collection agents encountered a corrupted upstream API that started returning random labels. The agent dutifully collected 50,000 samples and attempted to train a new classifier. At the first attestation point, the verifier noticed that the gradient norm was an order of magnitude higher than expected. It flagged the agent, which rolled back to the last good checkpoint and halted data ingestion.
Without verify-before-claim, the agent would have trained for 6 hours, produced a garbage model, and reported success. With it, the failure was caught in 3 minutes, and the corrupted data source was automatically blacklisted.
Implementation Lessons
Verifiers must be simpler than agents. If your verifier is as complex as the agent, it can have the same bugs. We write verifiers in Rust (while agents are in Python) and keep their logic minimal — no ML, no dynamic code execution.
Never trust the agent's self-report. Even with attestation, an agent could lie about its internal state if compromised. We use hardware-backed attestation (TPM) for critical agents to ensure the reported state is genuine. For less critical agents, we accept the risk but log discrepancies.
Fail closed, not open. If the verifier is unreachable, the agent must pause, not continue. We've had network partitions cause agents to stall, but that's better than producing unverified outputs.
Verification is a cost, not a feature. Every verification step adds latency and complexity. We continuously monitor the cost-per-claim and adjust tiers. If a C1 claim costs more than 500ms on average, we either optimize the verifier or move it to C2 with less frequent checks.
The Hardest Tradeoff: Autonomy vs. Verification
The more autonomous an agent, the more verification it needs — but verification constrains autonomy. We've found that a good balance is to give agents full autonomy within a bounded scope, and require explicit verification only when crossing scope boundaries (e.g., from training to deployment).
Our current architecture uses a "verification gate" between each phase of the agent lifecycle. The gate is a hardened service that checks all claims before allowing the next phase. The agent can fail fast within a phase, but cannot advance without passing the gate.
Conclusion
Verify-before-claim is not a tool; it's a discipline. It requires rethinking how agents report their state, how artifacts are tracked, and how trust is established. The engineering cost is real, but the alternative — autonomous systems that silently produce incorrect outputs — is unacceptable for any production deployment.
If you're building autonomous agent swarms, start with provenance. Add intermediate attestation. Bound your verification scope. And make claims expire. Your future self will thank you when the next poisoned dataset arrives.
Damir Radulic runs infrastructure at Rinet, where we build sovereign AI systems that can be trusted. This is one of the lessons we learned the hard way.