Recall layered prompt-injection defenses and explain why prompt-only controls are insufficient.
Prompt-only objection The system prompt already says to ignore malicious instructions. A review is about to approve retrieval plus write tools. Your line Good start, but the prompt is not an authorization boundary. We still need tool gating, argument validation, and a test where untrusted text asks for the action. Letting model obedience stand in for application permission. It acknowledges the control while showing the missing layers. Sanitize-late objection We can just sanitize the final answer. The model can call tools before writing final text. Final-output sanitizing helps the user message, but tool risk happens earlier. We need pre-action validation…
Sign up free — one personalized lesson every day, matched to your role and goals.
Already have an account? Sign in