Identify dataset defects that make supervised fine-tuning learn the wrong pattern.
The training job is ready, but the data is not. Which items would damage a supervised fine-tuning dataset before training starts? Raw chat export pile Correct. Raw conversations are not automatically demonstrations. SFT needs the input plus the ideal output you want the model to imitate. Duplicate easy-case stack Correct. Many near-duplicates overweight a simple pattern and crowd out the edge cases that reveal whether the tune really improved behavior. Separate holdout eval tray Not an issue. A separate holdout set is healthy because it lets the team measure improvement on examples the model did not train on. Unresolved reviewer…
Sign up free — one personalized lesson every day, matched to your role and goals.
Already have an account? Sign in