Judge Prompt Battlecards
Respond to common objections about LLM-as-judge RAG evaluation with calibrated, evidence-based language.
Noise objection The judge is noisy, so we should not use it. An engineer found inconsistent scores on borderline answers. Your line Agreed that noise matters. Let's compare it to human anchors and decide whether it is fit for monitoring, gating, or triage. Do not defend the judge as objective. Prove its reliability for a specific job. It validates the concern and turns it into a calibration plan. Human-only objection Only domain experts should score these answers. A legal lead is worried about automated decisions. Experts should define and adjudicate the hard cases. The judge can handle obvious regressions and…
Sign up free — one personalized lesson every day, matched to your role and goals.
Already have an account? Sign in