Attention Is All You Need:从 Attention 公式到 Transformer 结构与推理优化
前置文章已经说明了序列瓶颈、内积打分、Softmax 缩放和 GPU 内存墙。本文正式推导 Attention Is All You Need 中的 Scaled Dot-Product Attention:从 Q、K、V 的角色拆分开始,逐项解释 ,再把 mask、Multi-Head Attention、位置编码、KV Cache、GQA 和 FlashAttention 接到同一条线上。