Transformer Attention Pocket Guide
Recall the basic roles of query, key, and value in self-attention.
QKV What are the three roles in attention? What do attention weights do? Think weighted retrieval. They decide how much each value vector contributes to the current token's new representation. Weights come from query-key similarity, then are used to combine values. Self-attention versus recurrence Architecture contrast The transformer removed recurrence from the main sequence-transduction architecture. Attention means the model knows which words are important. A stakeholder asks for a simple explanation. Your line Attention weights show learned token-to-token weighting for this model and input; they are useful signals, not guaranteed human explanations. Treating attention weights as complete explanations. It gives…
Sign up free — one personalized lesson every day, matched to your role and goals.
Already have an account? Sign in