Job Hunt

LeetCode resources and ML interview prep — for SWE and research roles.

✍️

Notes on the Job Hunt

Timing, information, connection & mindset

→

LeetCode / 力扣

Job Openings / 岗位

Career pages for major AI labs — US and China.

Community Posts

Only publicly verifiable, unexpired opportunities are listed.

+ Submit verified role

No currently verified community roles. Old April 2025 posts were archived instead of being shown as urgent.

ML Interview / 面经·八股

Core concepts with rigorous answers — for LLM researchers and practitioners.

01

Transformer & Attention

When d_k is large, the dot products Q·Kᵀ grow in magnitude — their variance scales with d_k, pushing softmax into regions with very small gradients (saturation). Dividing by √d_k normalizes the variance back to ~1, keeping softmax in a stable gradient regime.

Scaled dot-product attention
Attention(Q,K,V)=softmax ⁣(QK⊤dk)V\text{Attention}(Q,K,V) = \text{softmax}\!\left(\frac{QK^\top}{\sqrt{d_k}}\right)V

Key insight: Concrete example: d_k = 64 → without scaling, dot products have std ≈ 8; after scaling by 1/√64 = 1/8, std ≈ 1.

02

Training & Optimization

03

Architecture Design

04

Training & Alignment

05

Inference & Deployment

20 questions across 5 categories. More coming soon.