Worked Walkthrough: Threat-Model a Prompt Injection Path
Run a simple, repeatable prompt-injection threat-model on an AI workflow.
A research agent can search the web, read documents, and email a summary to a client account manager. Map untrusted input → model interpretation → available action → guardrail. Teams often stop at “the model might be manipulated” and never trace which real action could follow from that manipulation. 1. Mark entry points List every untrusted source the model can read: search results, webpages, PDFs, ticket notes, attachments. If the source can contain attacker text, it belongs in the threat model even if users think of it as “content,” not “input.” 2. Mark privileges Write down the actions still available…
Sign up free — one personalized lesson every day, matched to your role and goals.
Already have an account? Sign in