Full-precision Llama-3.1 base models, run remotely via nnsight + NDIF (no local GPU). On this strict next-token metric larger models hold only modestly better at ≤20% removed, and all collapse by ~30–40% — yet free-generation fluency is far more forgiving (see the gallery below).
Source: metrics_scale.json · next-token top-1 / KL vs the intact model, 8 prompts, base models on NDIF. A centred middle band of the given fraction is set to identity. Directional (8 prompts, greedy), not a benchmark.