Trace the inference loop a decoder-only transformer uses to generate text.
Inference loop The prompt is already in the context window. The model is about to produce the first visible token of its answer. Understanding the loop prevents overreading generation as a fully prewritten plan. Autoregressive generation Context -> logits -> decode -> append -> repeat The decoder produces a next-token distribution, not a complete essay in one shot. One-shot answer Misleading: it hides the iterative conditioning loop. Each token changes the conditions for the next token. Generation is iterative conditioning, not a hidden finished answer being copied out. 01 Score 02 Decode 03 Append Vocabulary distribution The model has processed…
Sign up free — one personalized lesson every day, matched to your role and goals.
Already have an account? Sign in