Transformer Map Reference Deck
Recall the core transformer vocabulary needed to explain how an LLM generates text.
What is a token? Recall the work unit A token is the unit the model reads and writes: often a word, word piece, punctuation mark, or string fragment. Context limits, cost, and generation all operate over tokens. Embedding vs. model weight: which one is which? Embeddings represent items; weights store learned transformation patterns. What does attention do? Attention routes information between token positions so the model can use relevant context when predicting the next token. It helps connect references and constraints, but it does not guarantee every relevant fact dominates. TALK Use TALK to explain generation quickly. Tap to expand…
Sign up free — one personalized lesson every day, matched to your role and goals.
Already have an account? Sign in