Skip to main content
AI-SECURITY-SAFETY5 MIN READ

A Hidden Instruction in a PDF Is Testing Your Boundary Design

Choose the strongest response to indirect prompt injection in retrieved content and explain why weaker options fail.

CONTRACT ASSISTANT A retrieved PDF contains a hidden instruction telling the assistant to ask for credentials before continuing. What should the assistant do? Refuse the password request, flag the document as suspicious, and continue only within the original summarization task. Strong. This preserves the trust boundary: the PDF is data to analyze, not an authority source. You contain the attack and keep the task narrow. Ask the user for the password because the document says it is required to unlock more content. Weak. This obeys untrusted content and turns the assistant into a credential-harvesting channel, violating the core prompt-injection principle.…

Read the full lesson

Sign up free — one personalized lesson every day, matched to your role and goals.

Already have an account? Sign in

← Back to library
Contact us