Announcement_22
| Excited to share our new paper “Don’t Judge an LLM Only by Its Activations: Discovering Suppressed Safety Features via Counterfactual Activation Potential” – the revised version of our CAP paper, showing that safety-critical features can hide among a model’s inactive components, invisible to activation-focused interpretability, and that jailbreaks may work in part by suppressing them. Submitted to ICLR 2027, currently under review. arXiv | OpenReview |