Write a Useful Eval Rubric
Build an eval rubric with task goal, observable criteria, edge cases, and thresholds.
A support copilot rubric says answers should be helpful, accurate, and friendly, but reviewers do not agree on what those words mean. Task goal -> observable criteria -> edge cases -> threshold The common trap is replacing vague product judgment with vague eval language. Helpful, high quality, and safe are aspirations until each is translated into visible behavior. Name the task Write the user task as: customer gets a billing answer they can act on without unnecessary escalation. Task wording keeps the rubric anchored in user success rather than model fluency. List failure modes Identify failures: wrong policy, missing account…
Sign up free — one personalized lesson every day, matched to your role and goals.
Already have an account? Sign in