Posts tagged "interpretability"

18 posts

The Residual Stream: From AlexNet to DeepSeek-V4

29 min read

*A complete research dossier — the one idea, $\text{output} = \text{input} + F(\text{input})$, traced from computer vision in 2015 to the multi-lane residual streams of today's frontier models.* 📺

The Redundancy Trap: Why Single-Head Ablation Lies

20 min read

# The Redundancy Trap *I deleted the two attention heads with the largest positive direct effects in GPT-2 small. The model got **better**. Then I found the same failure across seven models — and in