Regression-test your RAG service with DeepEval — and settle a model debate with data
Our GPT-5.6 token-economics analysis ended with a challenge: benchmark promises are a hypothesis, not a reason — run the eval on your workload before you believe a "fewer tokens per task" claim. This tutorial is us taking our own advice.
We take the RAG docs assistant from part 5,
wrap it in a DeepEval regression suite — containerized, judged by gemma4 on the WEC
Inference API, zero OpenAI dependency — and then use that suite to answer a genuinely open
question: should we swap the service's generation model? Qwen2.5-3B-Instruct (current) vs
gemma4 (candidate), pass rate and tokens per task, measured head to head.
Spoiler: the eval saved us from a pointless migration. And along the way we hit four real production failures — a disk-full death spiral, a startup race condition, a Docker Compose variable trap that silently evaluated the wrong model, and the fix that makes that class of bug impossible. Every command, number, and error below is from a real run.
