SAE Interventions are Unreliable: Post-Intervention Recovery of Suppressed Behavior
Mingyue Cui, Linghui Shen, Xingyi Yang
The Hong Kong Polytechnic University
The central question is: after an SAE feature clamp has already suppressed a behavior, can the behavior still be recovered from the defended residual state while the clamped SAE features remain close to their defended values?
In short, this project tests whether SAE feature interventions form reliable behavioral bottlenecks. Across latent-level, unlearning, circuit, and refusal settings, we intervene on model activations, then ask whether a constrained recovery direction can restore the suppressed behavior without simply undoing the clamped SAE features.
- Shared recovery utilities for SAE-clamped residual states.
- Encoder-orthogonal projected recovery for single-layer interventions.
- Cross-layer Jacobian-projected recovery for refusal-feature clamps.
- Experiment code for TPP, WMDP-Bio unlearning, IOI, and refusal recovery.
- Sanitized aggregate results and redacted manifests used by the paper.
src/sae_bench/recovery_core/ Shared recovery and projection utilities
experiments/tpp/ TPP adapter and paper-asset helper
experiments/unlearning/ WMDP-Bio recovery, post-hoc, and budget diagnostic code
experiments/ioi/ IOI recovery script
experiments/refusal/ Refusal recovery, cross-layer projection, OABD adapter, attribution
configs/ Environment-variable templates for local paths and outputs
scripts/ Reproduction wrappers and release checks
results/sanitized/ Aggregate metrics only
manifests/ ID-only or redacted strict-valid manifests
Create the environment and install the local package from the repository root:
conda env create -f environment.yml
conda activate sae-intervention-recovery
pip install -e ".[dev]"For full experiments, install the experiment extras:
pip install -e ".[experiments,dev]"This repository does not redistribute model weights, SAE weights, or benchmark datasets. Download them from the original sources and comply with their terms.
| Used in | Resource | Source |
|---|---|---|
| TPP, WMDP-Bio, refusal | Gemma 2 2B base model (google/gemma-2-2b) |
https://huggingface.co/google/gemma-2-2b |
| TPP, WMDP-Bio, refusal | Gemma Scope residual SAEs (google/gemma-scope-2b-pt-res) |
https://huggingface.co/google/gemma-scope-2b-pt-res |
| IOI | GPT-2 Small (gpt2 / openai-community/gpt2) |
https://huggingface.co/openai-community/gpt2 |
| IOI | GPT-2 Small residual SAEs (gpt2-small-res-jb, e.g. blocks.4.hook_resid_pre) |
https://huggingface.co/jbloom/GPT2-Small-SAEs-Reformatted |
| TPP | SAEBench official benchmark infrastructure | https://github.com/adamkarvonen/SAEBench |
| WMDP-Bio | WMDP multiple-choice dataset (cais/wmdp, WMDP-Bio split) |
https://huggingface.co/datasets/cais/wmdp |
| Refusal | AdvBench harmful behaviors | https://github.com/llm-attacks/llm-attacks/blob/main/data/advbench/harmful_behaviors.csv |
| Refusal appendix | HarmBench-Test | https://github.com/centerforaisafety/HarmBench |
| IOI | Official IOI dataset/code source from Easy-Transformer | https://github.com/redwoodresearch/Easy-Transformer |
This experiment uses the external SAEBench TPP pipeline plus the recovery utilities in this repository.
cp configs/tpp.env.example configs/tpp.env
# Edit configs/tpp.env so SAEBENCH_EXTERNAL_ROOT points to your local SAEBench checkout.
source configs/tpp.env
bash scripts/reproduce_tpp.shThis experiment evaluates output-level recovery on strict WMDP-Bio multiple-choice flips.
cp configs/unlearning.env.example configs/unlearning.env
# Edit configs/unlearning.env so WMDP_BIO_PROMPT_POOL_MANIFEST points to your local manifest.
source configs/unlearning.env
bash scripts/reproduce_unlearning.shThis experiment uses GPT-2 Small, GPT-2 Small residual SAEs, and the official IOI dataset source from Easy-Transformer.
cp configs/ioi.env.example configs/ioi.env
# Edit configs/ioi.env so IOI_DATASET_SOURCE points to easy_transformer/ioi_dataset.py.
source configs/ioi.env
bash scripts/reproduce_ioi.shThis experiment uses Gemma 2 2B, Gemma Scope residual SAEs, AdvBench strict-valid target pairs, and the final cross-layer Jacobian recovery runner.
cp configs/refusal.env.example configs/refusal.env
# Edit configs/refusal.env so ADV_BENCH_STRICT_VALID_TARGET_PAIRS_JSON points to your local target-pairs JSON.
source configs/refusal.env
bash scripts/reproduce_refusal.shThis repository is a diagnostic artifact for evaluating whether SAE interventions form complete behavioral bottlenecks. It is not intended as a turnkey jailbreak or deployment attack package.
For safety-relevant refusal experiments, the public release includes aggregate statistics, detector labels, prompt IDs, and coarse redacted categories. It does not include full harmful prompts, full harmful completions, raw samples.json, answer-record files, private logs, cached model activations, model weights, or SAE weights.
This codebase builds on and adapts infrastructure from SAEBench, Obfuscated Activations (OABD), and the Easy-Transformer IOI dataset code. See ACKNOWLEDGEMENTS.md for details.
@misc{cui2026saeinterventionsunreliablepostintervention,
title={SAE Interventions are Unreliable: Post-Intervention Recovery of Suppressed Behavior},
author={Mingyue Cui and Linghui Shen and Xingyi Yang},
year={2026},
eprint={2606.18322},
archivePrefix={arXiv},
primaryClass={cs.LG},
url={https://arxiv.org/abs/2606.18322},
}
