When d_k is large, the dot products Q·Kᵀ grow in magnitude — their variance scales with d_k, pushing softmax into regions with very small gradients (saturation). Dividing by √d_k normalizes the variance back to ~1, keeping softmax in a stable gradient regime.
Scaled dot-product attention
Attention(Q,K,V)=softmax(dkQK⊤)V
Key insight: Concrete example: d_k = 64 → without scaling, dot products have std ≈ 8; after scaling by 1/√64 = 1/8, std ≈ 1.
02
Training & Optimization
03
Architecture Design
04
Training & Alignment
05
Inference & Deployment
20 questions across 5 categories. More coming soon.