Explain why transformer self-attention made large-scale training more parallel than recurrent sequence models.
The transformer was a scaling move as much as an accuracy move. The old bottleneck Recurrent sequence models carry information through a chain. Token 200 depends on token 199, which depends on token 198, and so on. That makes the computation path long and harder to parallelize across positions. The transformer shift Self-attention lets every token position compare with other positions inside a layer. During training, much of this work becomes large matrix operations that parallel hardware handles well. That changed the speed of the experimental loop. The remaining trade-off Attention is not free. Comparing many token pairs can be…
Sign up free — one personalized lesson every day, matched to your role and goals.
Already have an account? Sign in