Skip to main content
HOW-TRANSFORMERS-WORK5 MIN READ

Self-Attention Is a Context-Routing System

Explain self-attention as token-to-token routing of contextual information.

Self-attention is a learned routing system for context. The system view A transformer layer updates every token representation by letting it draw information from other token positions. In systems terms, attention defines weighted links and flows. The links are input-dependent: the same token can route context differently in different sentences. Query, key, and value Each token produces a query, a key, and a value. A query is matched against keys to produce scores. Those scores are normalized into attention weights. The weights combine value vectors into the token's next contextual representation. What attention is not Attention is not human attention…

Read the full lesson

Sign up free — one personalized lesson every day, matched to your role and goals.

Already have an account? Sign in

← Back to library
Contact us