Posts tagged "llama"
4 posts
The Redundancy Trap: Why Single-Head Ablation Lies
# The Redundancy Trap *I deleted the two attention heads with the largest positive direct effects in GPT-2 small. The model got **better**. Then I found the same failure across seven models — and in
Where Does a Language Model Think? Finding and Removing the 'Workspace' Layers of Llama-3.1-8B
*A hands-on interpretability walkthrough. We watch concepts form layer-by-layer inside Llama-3.1-8B-Instruct, measure precisely which layers carry meaning, then **delete layers** — one at a time and
SafeConstellations: Mitigating Over-Refusals in LLMs Through Task-Aware Representation Steering
*Utsav Maskey · **Sumit Yadav** · Mark Dras · Usman Naseem* Accepted · ACL 2026 Main Conference · arXiv:2508.11290 Proceeding...
SafeConstellations: Mitigating Over-Refusals in LLMs Through Task-Aware Representation Steering (ACL 2026)
# SafeConstellations — ACL 2026 Main Conference A deep-dive companion blog for **SafeConstellations**, accepted at **ACL 2026 Main**. An inference-time method that reduces LLM over-refusal by up to