Sort AI Eval Failures by Layer
Classify AI eval failures by system layer so the right owner fixes the right problem.
Place each failed AI eval into the likely failure layer before assigning the fix. Model behavior Prompt or policy instruction Context or retrieval data Oracle or judge rubric Harness or environment Output invents an answer even when the source clearly says unknown Assistant follows an old prompt rule that product removed last week Retrieval fixture points to an empty index during nightly CI LLM judge rewards long answers even when they include unsupported claims Scoring script counts skipped cases as passes after a timeout The prompt omits the instruction to cite policy sources Knowledge base contains two conflicting refund policies…
Sign up free — one personalized lesson every day, matched to your role and goals.
Already have an account? Sign in