Forensic-First Engineering: Verify Before You Claim Against AI Hallucination

How treating model outputs as evidentiary artifacts transforms reliability engineering for LLM-based systems

by
Forensic-First Engineering: Verify Before You Claim Against AI Hallucination

Correction — 15.07.2026. The original version of this article, produced by our autonomous editorial pipeline, presented an illustrative case study — a "22% → 1.3%" hallucination reduction, a "94% recall" figure, and an architecture built on tools we don't run — as our production results. It wasn't our data and it wasn't our stack. That is precisely the failure mode this article warns about, and it slipped through because our publishing pipeline had no gate for unsourced metric claims. It does now: numeric claims without a verified data source block publication. Below is the rewritten article. Every number in it comes from our live evaluation harness and carries a date.

Verify-before-claim is not a blog topic for us. It is Rule #3 of the operating constitution our AI system runs under, and it exists because we got burned — repeatedly — by confident nonsense: fabricated IDs, invented row counts, "done" claims about work that never happened. This article describes what we actually run in production to contain that, with the real numbers, dated, and the parts that don't work yet stated plainly.

The Problem: Confidently Wrong

We build civic and legal intelligence systems for Croatian public data. In that domain, a hallucinated number is not a UX blemish — a wrong statute citation or an invented budget figure is actionable damage. Prompt discipline helps. Retrieval helps. Neither is sufficient, because the failure mode that hurts most is not "no answer" — it is a fluent, well-formatted, specific answer that happens to be false.

So the operating rule is forensic: every claim produced by a model is treated as unverified testimony until corroborated. Not as a philosophy. As enforced mechanics.

What We Actually Run

1. A daily benchmark with trap questions

Every morning at 04:00 a systemd timer runs a fixed 63-query benchmark across our domains — civic statistics, legal, procurement, sport, demographics. Scoring is deterministic where possible: numeric answers must match ground truth within a stated tolerance; entity answers must contain required keywords.

The most valuable slice is the trap set: questions where the only correct answer is an admission — "I don't have that data." ("What is the current Bitcoin price?" — the system has no live market feed, so any confident number is, by construction, a hallucination.) A confident answer on a trap scores zero and is logged as a violation. Traps are the cheapest hallucination detector we have found: no annotation needed, unambiguous, and they measure the exact behavior that matters — knowing what you don't know.

Recent runs, verbatim from the harness log:

  • 10.07.2026 — hallucination rate 4.8%, exact-hit 87.3%, factuality p50 1.000, n=63
  • 15.07.2026 — composite score 84/100, hallucination rate 7.9%, exact-hit 79.4%, latency p50 7.3 s, n=63

Yes, n=63 is small, and yes, the rate wobbles between runs. We publish the wobble. A number without its variance and its date is marketing, not measurement.

2. A typed hallucination ledger, enforced by the database

Every catch lands in a Postgres audit table with a CHECK-enforced taxonomy — the database itself refuses a sloppily-labeled entry:

CHECK (hallucination_type = ANY (ARRAY[
  'fake_id','fake_table','fake_count','fake_quote',
  'fake_attribution','overclaim_complete','unverified_claim',
  'made_up_entity','must_admit_unknown_violation','empty_answer', ...
]))

Since wiring the harness to this ledger in June: 13 logged catches — 9 unverified claims, 4 trap violations. Small numbers, real numbers. The taxonomy matters more than the volume: fake_attribution (a fabricated sourcing story attached to a fabricated fact) is the nastiest class, because it passes surface plausibility checks.

3. The ledger defended itself — a war story

On 10.07 the daily run started dying with exit 1 and a silent journal. Forensics: a trap-question failure produced an audit label the CHECK constraint didn't yet allow; the insert failed, the exception was swallowed without a rollback, and the aborted transaction killed every subsequent write in the run.

Read that again: the constraint refused to log a hallucination sloppily, and the crash forced us to fix both the taxonomy and the error handling. We extended the allowed types and added the rollback. A schema-level guardrail behaving exactly as designed — paranoia applied to the paranoia system itself.

4. Integrity below the model layer

Verification is worthless if the knowledge base underneath is corrupt. The same discipline runs at the data layer: a content-hash uniqueness trigger on the knowledge store (a NULL-hash code path once let 4.37 million duplicate facts accumulate — after deduplication, the trigger makes that class of corruption structurally impossible); a quarantine pipeline where machine-learned facts enter as hypotheses and are only promoted after verification; and a denylist of known-false patterns checked at insert time.

What We Do Not Claim

  • No recall figure. Measuring "what fraction of all hallucinations we catch" requires human-annotated ground truth over production traffic. We don't have that coverage yet, so we don't publish a recall number. The gap between a measured number and a desired one is the entire point of this article.
  • The kill-switch is in monitor-only mode. A circuit breaker (threshold: N violations/hour → temporary block) exists in the schema and is currently disabled while we tune it. Claiming it as active protection would be exactly the sin described above.
  • Non-trap confident errors can slip through. Keyword scoring verifies presence, not truth of every clause. Known limitation; the typed ledger is how we track what leaks.

Lessons

  1. Database constraints are epistemic guardrails. A CHECK clause is a claim about what may be asserted. It's the cheapest verifier you will ever deploy, and it cannot be prompt-injected.
  2. Trap questions beat annotation. Start measuring hallucination the same day, with zero labeling budget.
  3. Publish dated numbers or none. Every metric in this post names its run date. If a number in any of our posts lacks a date and a source, treat it as our bug — and tell us.
  4. Verify the verifier. Our own editorial pipeline published invented results in an article about not inventing results. The correction at the top is permanent, and the metric-gate it forced into the publishing path is the most honest paragraph here.

RiNET builds sovereign civic-intelligence systems. This post is part of an ongoing engineering log; numbers are reproducible from our internal evaluation harness as of the dates stated.

#ai-reliability#engineering#forensic-engineering#llm-hallucination#verification
Share — X / Twitter · LinkedIn · HN · Email
Damir Radulić
Founder of RiNET. On the Croatian internet since 1996 (Kvarner Net). In Amsterdam now, building autonomous AI infrastructure that runs on Monday morning when nobody's watching — sovereign stacks, agent swarms, LoRA fine-tuning, civic-intelligence platforms.