Posts tagged "ai-safety"
9 posts
SafeConstellations: Mitigating Over-Refusals in LLMs Through Task-Aware Representation Steering
*Utsav Maskey · **Sumit Yadav** · Mark Dras · Usman Naseem* Accepted · ACL 2026 Main Conference · arXiv:2508.11290 Proceeding...
Reading the Mind of Claude: Millions of Features Inside a Frontier AI
 In 2023, Anthropic learned to pull clean, single-meaning **features** out of a tiny one-layer model. The obviou
Circuit Tracing: How We Draw the Wiring Diagram of an AI's Thoughts
 In a companion post we toured *the biology of a language model* — watching Claude plan rh
The Biology of a Large Language Model: Tracing the Circuits Inside Claude
 We don't really *build* large language models — we **grow** them. We set up a trainin
Can an AI Look Inward? Emergent Introspection in Language Models
 Ask a language model what it's thinking and it will happily tell you — it describes its reasoning
Do AI Models Have Emotions? Inside Claude's Emotion Concepts
window.MathJax = { tex: { inlineMath: [['$','$'], ['\\(','\\)']], displayMath: [['$$','$$'], ['\\[','\\]']] }, svg: { fontCache: 'global' } }; ![A map of feeling — emotion concepts la
Zoom In: How to Read the Circuits Inside a Neural Network
window.MathJax = { tex: { inlineMath: [['$','$'], ['\\(','\\)']], displayMath: [['$$','$$'], ['\\[','\\]']] }, svg: { fontCache: 'global' } }; ![Curve detectors — a family of neurons
Mechanistic Interpretability & Sparse Autoencoders: Giving AI an X-Ray
window.MathJax = { tex: { inlineMath: [['$','$'], ['\\(','\\)']], displayMath: [['$$','$$'], ['\\[','\\]']] }, svg: { fontCache: 'global' } }; ![Dense vs. sparse activations — the SAE
Natural Language Autoencoders: Reading a Model's Mind in Plain English
window.MathJax = { tex: { inlineMath: [['$','$'], ['\\(','\\)']], displayMath: [['$$','$$'], ['\\[','\\]']] }, svg: { fontCache: 'global' } }; Inside a large language model, every "th