Separate Model Failure from Harness Failure
Classify AI eval failures by failure layer before assigning a model fix.
The first question after a failed eval is not: how do we fix the model? The first question is: which layer failed? AI test systems include prompts, data fixtures, retrieval context, model versions, judge rubrics, scoring code, environment state, and reporting. Any one of those can create a red result. Layer classification prevents false work. A stale vector index looks like a reasoning failure. A brittle rubric looks like lower quality. A changed prompt variable looks like model drift. If the report does not show the layer, the team will argue from anecdotes. Use a simple field on every failure:…
Sign up free — one personalized lesson every day, matched to your role and goals.
Already have an account? Sign in