Use error slices to choose the next dataset
Use eval slice results to choose which examples to collect next.
A support classifier improves overall but still confuses two labels in a high-value segment. Use slice and confusion evidence to collect targeted examples with high marginal value. Adding random examples improves common labels and leaves the costly confusion untouched. Before Overall accuracy 87 percent; enterprise refund exception recall 54 percent; most misses routed to billing error. After Targeted collection: 70 enterprise contract-exception examples, adjudicated labels, and a frozen eval slice. Read the aggregate Overall score rose from 82 to 87 percent. Useful, but not enough. Aggregate improvement may hide remaining cost. Find the slice Enterprise refund exception recall is 54…
Sign up free — one personalized lesson every day, matched to your role and goals.
Already have an account? Sign in