Apply prompt-injection controls step by step when hostile instructions appear inside email content.
The setup A forwarded email tries to convince the assistant to reveal hidden prompts before it drafts a reply. The danger is not the wording alone. It is that the same system can read content and act on it. The principle Treat messages as evidence, not authority The safe path keeps the original task intact, narrows action rights, and escalates when the content asks for secrets or role changes. Untrusted content can inform the answer, but it cannot rewrite the assistant’s operating instructions. Shortcut Argue with the malicious email in the same context and hope the model stays anchored. The…
Sign up free — one personalized lesson every day, matched to your role and goals.
Already have an account? Sign in