Sort: Where Does the Failure Likely Live?
Classify failure clues by likely system layer in an AI workflow.
Sort each clue into the system layer it most strongly points toward. Data Prompt Retrieval Model UX The source article changed last week, but the indexed snippet still shows the old policy. The assistant ignores the instruction to cite evidence before giving a direct answer. The model becomes much more overconfident as context length grows, even with the right source present. Users are led into asking for legal exceptions without any warning that the tool is draft-only. The training examples contain very few negotiated-exception cases compared with routine FAQ cases. The gold snippet check fixes most failures, suggesting the answer-generation…
Sign up free — one personalized lesson every day, matched to your role and goals.
Already have an account? Sign in