Describe the two core parts of a useful LLM eval: representative test data and explicit testing criteria.
The move: test the behavior you actually need, not the behavior that is easiest to score. The Test Data Your eval data should represent the inputs the system will see. Include routine cases, edge cases, and cases where the correct behavior is to escalate or refuse. For LLMOps, this usually means collecting real examples from tickets, chats, search queries, sales notes, or generated adversarial probes, then cleaning them into a stable test set. The Testing Criteria The criteria define what counts as correct. Sometimes that is exact match. Sometimes it is a rubric: grounded in source, no unsupported policy claim,…
Sign up free — one personalized lesson every day, matched to your role and goals.
Already have an account? Sign in