Build a TF-IDF Baseline
Describe the steps of a TF-IDF baseline and interpret its first result.
You need a first classifier for 5,400 labeled support tickets before deciding whether a heavier model is justified. Clean split -> dummy floor -> TF-IDF model -> per-label error readout The common shortcut is to train a powerful model first, celebrate one aggregate score, and never learn whether simple lexical signal already solved most labels. Split cleanly Group related tickets by customer thread, then create train and validation sets. The split must estimate future behavior, not repeated-message memorization. Set dummy floor Train a DummyClassifier that predicts the most frequent label and record macro-F1. This tells you what a nearly useless…
Sign up free — one personalized lesson every day, matched to your role and goals.
Already have an account? Sign in