You Cannot Test Away Hallucination: Building a Statistical Guardrail Layer for Grounded Answers

Why traditional evals fail and how to deploy a real-time statistical guardrail for LLM outputs

by
You Cannot Test Away Hallucination: Building a Statistical Guardrail Layer for Grounded Answers

You Cannot Test Away Hallucination: Building a Statistical Guardrail Layer for Grounded Answers

Every team that deploys a large language model (LLM) into production eventually hits the same wall: the model confidently states something false. You add more test cases. You run human eval. You fine-tune. The next week, a new hallucination surfaces. This cycle is not a bug—it is a feature of language models. You cannot test away hallucination. The space of possible false statements is infinite; your test set is finite. The only viable approach is to build a runtime guardrail that statistically assesses whether an output is grounded in the provided context.

Why Evals Are Not Enough

Traditional evaluation pipelines rely on held-out test sets, often with manual annotation. You measure accuracy, F1, or BLEU on a few hundred examples. The model passes, you deploy. But production distributions shift. Users ask questions that are semantically novel. The model generalizes—sometimes incorrectly. A 98% accuracy on your test set means 2% of responses are potentially hallucinated. In a high-volume system, that 2% becomes a flood of bad answers.

Hallucination is not a classification error; it is a generation error. The model outputs tokens that are locally plausible but globally inconsistent with the provided context. No amount of static testing can catch every case because the model's internal representation is not a database—it is a probability distribution over tokens. You need a dynamic, per-output check.

The Statistical Guardrail Approach

Instead of trying to prevent hallucination, detect it after generation. The guardrail runs as a separate service that evaluates each LLM response against the source context before the answer reaches the user. It returns a groundedness score. Below a threshold, you fall back to a safe response (e.g., "I cannot find that information") or trigger a human review.

The core idea: a grounded answer should be statistically close to the context in embedding space, and its n-grams should overlap with the context more than a hallucinated answer would.

Two Metrics You Need

  1. Embedding Cosine Similarity: Embed the context and the response using a high-quality embedding model (e.g., BGE-M3 or E5-mistral-7b-instruct). Compute cosine similarity between the response embedding and the context embedding. If the response is grounded, it will lie near the context in the embedding space. A hallucination often drifts away. This metric captures semantic similarity but can be fooled by vague but unrelated answers.

  2. N-gram Overlap (ROUGE-L or BLEU): Compute the overlap of n-grams (unigrams, bigrams, trigrams) between the response and the context. Grounded answers tend to reuse entities, phrases, and facts from the context. Hallucinations introduce novel n-grams unsupported by the source. Use a weighted combination of ROUGE-L (recall-oriented) and BLEU (precision-oriented) to get a coverage and precision score.

Combine these two into a single score via a logistic regression or a simple weighted sum. Train the weights on a small set of human-annotated examples (grounded vs. hallucinated). You only need a few hundred examples to set a threshold.

Implementation Sketch

Assume you have a retrieval-augmented generation (RAG) pipeline. The context is the set of retrieved documents. The response is the LLM's answer.

import numpy as np
from sentence_transformers import SentenceTransformer
from rouge_score import rouge_scorer

# Load embedding model (e.g., BAAI/bge-m3)
embedder = SentenceTransformer("BAAI/bge-m3")

# Compute embedding similarity
def embedding_similarity(context, response):
    ctx_emb = embedder.encode(context, normalize_embeddings=True)
    resp_emb = embedder.encode(response, normalize_embeddings=True)
    return float(np.dot(ctx_emb, resp_emb))

# Compute ROUGE-L F1
scorer = rouge_scorer.RougeScorer(["rougeL"], use_stemmer=True)
def rouge_l_f1(context, response):
    scores = scorer.score(context, response)
    return scores["rougeL"].fmeasure

# Combined score (weights from training)
WEIGHT_EMB = 0.6
WEIGHT_ROUGE = 0.4
def groundedness_score(context, response):
    emb_score = embedding_similarity(context, response)
    rouge_score = rouge_l_f1(context, response)
    return WEIGHT_EMB * emb_score + WEIGHT_ROUGE * rouge_score

# Threshold determined from validation set
THRESHOLD = 0.7

def is_grounded(context, response):
    return groundedness_score(context, response) >= THRESHOLD

Deploy this as a microservice. The LLM response is only returned to the user if is_grounded returns True. Otherwise, return a fallback message or escalate.

Choosing the Right Threshold

Threshold selection is a business decision. Lower thresholds catch more hallucinations but increase false positives (rejecting good answers). Higher thresholds let more hallucinations through. Use a precision-recall curve on your validation set. Decide based on the cost of a hallucination vs. the cost of a false rejection. For medical or legal domains, you want near-zero hallucination rate, so accept more false positives. For a creative writing assistant, you might tolerate more hallucination.

Production Considerations

  • Latency: Embedding inference adds ~50-200ms depending on model size. Use a small, fast model like BGE-small or a distilled version. Cache context embeddings if the same context is reused. Run the guardrail on a separate GPU node or use CPU with ONNX Runtime.
  • Batching: If you have high throughput, batch guardrail calls. The embedding model benefits from batch processing.
  • Storage: Log every guardrail decision with score, context hash, and response. This data is gold for improving your threshold and training a more sophisticated classifier later.
  • Fallback strategy: When the guardrail rejects a response, consider re-querying the LLM with a stronger prompt like "Only answer if you are certain. If unsure, say 'I don't know'." Or route to a human-in-the-loop.

Beyond Simple Metrics

The two-metric approach is a starting point. You can extend it with:

  • Factual consistency models: Fine-tune a BERT-like model on the task of detecting entailment between context and response. Models like TrueTeacher or BARTScore can be used but are heavier.
  • Token-level log probabilities: The LLM's own confidence scores (e.g., average log probability of generated tokens) correlate with correctness. A low average log prob indicates the model is unsure. Combine this with the guardrail score.
  • Decomposition: Break the response into atomic claims (using an LLM) and check each claim against the context via a natural language inference (NLI) model. This is more thorough but expensive.

Case Study: Replacing SaaS Guardrails with On-Prem

At a previous engagement, we replaced a SaaS hallucination detection API with an on-prem guardrail using BGE-M3 embeddings and ROUGE-L. The SaaS solution cost $0.003 per API call and had a 95% recall at 90% precision. Our on-prem guardrail, running on a single NVIDIA T4, achieved 96% recall at 92% precision with a median latency of 80ms. We saved $12,000/month and eliminated data sovereignty concerns under the EU AI Act. The entire guardrail was deployed as a systemd service behind a FastAPI endpoint, consuming 4GB RAM and negligible CPU when idle.

The Bottom Line

Hallucination is not a testing problem. It is a runtime detection problem. By building a statistical guardrail layer that combines embedding similarity and n-gram overlap, you can catch a large fraction of hallucinations before they reach users. Start simple, collect production data, and iterate. You will never eliminate hallucination entirely, but you can make it a rare, controlled event.

No amount of testing will save you. A guardrail will.

#ai-infrastructure#evals#grounding#guardrails#hallucination#testing
Share — X / Twitter · LinkedIn · HN · Email
Damir Radulić
Founder of RiNET. On the Croatian internet since 1996 (Kvarner Net). In Amsterdam now, building autonomous AI infrastructure that runs on Monday morning when nobody's watching — sovereign stacks, agent swarms, LoRA fine-tuning, civic-intelligence platforms.