Skip to main content
HOW-TRANSFORMERS-WORK5 MIN READ

Walk Through Multi-Head Attention

Explain how multi-head attention runs several learned attention projections and recombines them.

Explain multi-head attention without calling heads voters or mini-models. Separate the parallel subspaces: project, attend, concatenate, output-project. The common trap is saying each head votes on the final answer. Heads produce intermediate contextual vectors, not final text decisions. Step 1: Project per head The same token representations are sent through different learned Q, K, and V projections for each head. This creates separate compatibility and value spaces. Step 2: Attend in parallel Each head computes its own attention weights and value mixture. Different heads can route different patterns, although not every head has a clean human label. Step 3: Recombine…

Read the full lesson

Sign up free — one personalized lesson every day, matched to your role and goals.

Already have an account? Sign in

← Back to library
Contact us