Sort failures into the right intervention
Classify fine-tuning failures by root cause so the next intervention matches the problem.
Sort the failure by the first intervention it needs. Prompt/examples Retrieval/context Training data Grader/eval Policy decision sort-buckets common-trap-callout The model answers in paragraphs although the prompt never specified JSON output The model invents pricing because the current plan table is absent from context The rare cancellation edge case has only three examples in the dataset The LLM judge gives full credit to long answers with no citation Support and Legal disagree about whether refunds are allowed after 45 days The model follows examples well but uses the old tone for enterprise accounts The final answer is correct but the eval…
Sign up free — one personalized lesson every day, matched to your role and goals.
Already have an account? Sign in