Sparse Autoencoders Are Capable LLM Jailbreak Mitigators
AuthorsYannick Assogba, Jacopo Cortellazzi, Javier Abad†**, Pau Rodriguez, Xavier Suau, Arno Blaas
Sparse Autoencoders Are Capable LLM Jailbreak Mitigators
AuthorsYannick Assogba, Jacopo Cortellazzi, Javier Abad†**, Pau Rodriguez, Xavier Suau, Arno Blaas
Jailbreak attacks remain a persistent threat to large language model safety. We propose Context-Conditioned Delta Steering (CC-Delta), an SAE-based defense that identifies jailbreak-relevant sparse features by comparing token-level representations of the same harmful request with and without jailbreak context. Using paired harmful/jailbreak prompts, CC-Delta selects features via statistical testing and applies inference-time mean-shift steering in SAE latent space. Across four aligned instruction-tuned models and twelve jailbreak attacks, CC-Delta achieves comparable or better safety–utility tradeoffs than baseline defenses operating in dense latent space. In particular, our method clearly outperforms dense mean-shift steering on all four models, and particularly against out-of-distribution attacks, showing that steering in sparse SAE feature space offers advantages over steering in dense activation space for jailbreak mitigation. Our results suggest off-the-shelf SAEs trained for interpretability can be repurposed as practical jailbreak defenses without task-specific training.
Dynamically Scaled Activation Steering
September 18, 2026research area Human-Computer Interaction, research area Methods and AlgorithmsTransactions on Machine Learning Research (TMLR)
Activation steering has emerged as a powerful method for guiding the behavior of generative models towards desired outcomes such as toxicity mitigation. However, most existing methods apply interventions uniformly across all inputs, degrading model performance when steering is unnecessary. We introduce Dynamically Scaled Activation Steering (DSAS), a method-agnostic steering framework that decouples when to steer from how to steer. DSAS…
STEER: Semantic Turn Extension-Expansion Recognition for Voice Assistants
November 8, 2023research area Speech and Natural Language Processingconference EMNLP
*Equal Contributors
In the context of a voice assistant system, steering refers to the phenomenon in which a user issues a follow-up command attempting to direct or clarify a previous turn. We propose STEER, a steering detection model that predicts whether a follow-up turn is a user’s attempt to steer the previous command. Constructing a training dataset for steering use cases poses challenges due to the cold-start problem. To overcome this, we…