Stable Rubrics Can Earn a Tune
Recognize when a stable, repeated judgment pattern is a better fine-tuning candidate than a retrieval problem.
The source text is present; the repeated judgment is unstable. Marta's assistant already reads the clause text, but it keeps applying the same risk rubric inconsistently. What should she do next? After prompt and eval checks, fine-tune on reviewed clause inputs paired with ideal rubric outputs. Right. The facts are present and the desired behavior is stable, repeated, and demonstrable. That is the kind of pattern supervised fine-tuning can teach. Add more contract folders to retrieval and leave the behavior unchanged. Weaker. Retrieval helps when the model lacks facts. Here it already has the relevant clause, so adding more documents…
Sign up free — one personalized lesson every day, matched to your role and goals.
Already have an account? Sign in