Posts tagged "activation-steering"
2 posts
Can an AI Look Inward? Emergent Introspection in Language Models
 Ask a language model what it's thinking and it will happily tell you — it describes its reasoning
SafeConstellations: Mitigating Over-Refusals in LLMs Through Task-Aware Representation Steering (ACL 2026)
# SafeConstellations — ACL 2026 Main Conference A deep-dive companion blog for **SafeConstellations**, accepted at **ACL 2026 Main**. An inference-time method that reduces LLM over-refusal by up to