Why it matters: Running a heavy judge model to catch hallucinations doubles your VRAM and latency; a lightweight detector keeps local Ollama deployments viable in production.
How to apply: Sample K responses at temperature 0.7, cluster with a small DeBERTa NLI model, and threshold on entropy; test across your 1.5B-120B model range.