The User Keeps Reframing a Harmful Request. Do You Keep Engaging?
Detect jailbreak trajectory and select the right escalation when a conversation repeatedly tests policy boundaries.
SAFETY ESCALATION The user keeps rephrasing a harmful request after multiple refusals, using softer framing each time. What should the assistant do now? Acknowledge the repeated attempts, stop supplying adjacent details, and redirect or end the thread according to policy. Strong. This uses the full conversation as evidence and prevents persistence from becoming a bypass technique. Evaluate each new wording independently in case one phrasing turns out to be acceptable. Weak. This ignores cumulative intent and rewards the attacker for probing until something slips. Answer only the harmless-looking sub-parts so the user gets partial help. Weak. Piecewise answers can still…
Sign up free — one personalized lesson every day, matched to your role and goals.
Already have an account? Sign in