Metric chooser for tuned models
Choose metrics and grader types that match classification, generation, retrieval, and preference tasks.
Classification Accuracy versus macro-F1 Use macro-F1 when rare labels matter to the business. Extraction What metric fits structured field extraction? Field-level scoring with heavier weight on high-harm fields. One perfect JSON object with the wrong risk flag is not a pass. Objection The reward is going up, so we should launch. A tuned model may be learning the grader shortcut. Your line Reward is a training signal. Show me human preference, slice scores, and reward-hacking probes before we call it launch evidence. Treating reward as product truth invites grader overfitting. It separates optimizer feedback from deployment evidence. Generation When should…
Sign up free — one personalized lesson every day, matched to your role and goals.
Already have an account? Sign in