Calculate One Attention Head by Hand
Compute a simplified attention output for one target token.
Compute a simplified attention output for target token it using two source value vectors. Score compatibility, normalize to weights, then take the weighted sum of value vectors. The common trap is thinking attention copies the highest-scoring token. In most cases it blends value vectors according to normalized weights. Step 1: Start with weights Assume the query-key scores have already been normalized into weights: source A = 0.75, source B = 0.25. We skip raw dot products here to isolate the value-mixing mechanism. Step 2: Multiply values Source A value [2, 0] times 0.75 gives [1.5, 0]. Source B value [0,…
Sign up free — one personalized lesson every day, matched to your role and goals.
Already have an account? Sign in