Skip to main content
HOW-TRANSFORMERS-WORK5 MIN READ

Calculate One Attention Head by Hand

Compute a simplified attention output for one target token.

Compute a simplified attention output for target token it using two source value vectors. Score compatibility, normalize to weights, then take the weighted sum of value vectors. The common trap is thinking attention copies the highest-scoring token. In most cases it blends value vectors according to normalized weights. Step 1: Start with weights Assume the query-key scores have already been normalized into weights: source A = 0.75, source B = 0.25. We skip raw dot products here to isolate the value-mixing mechanism. Step 2: Multiply values Source A value [2, 0] times 0.75 gives [1.5, 0]. Source B value [0,…

Read the full lesson

Sign up free — one personalized lesson every day, matched to your role and goals.

Already have an account? Sign in

← Back to library
Contact us