Research codebase to test whether sandbagging behavior is explained by movement along the Assistant Axis/persona space, or whether it induces a separable activation direction.
Given a single model checkpoint, does the activation shift from placebo_alarm -> sandbag align with the Assistant Axis, or does a substantial residual direction remain after projecting out that axis?
- Single checkpoint end-to-end (persona vectors, generation, replay extraction, analysis).
- Main arm is answer-only (no CoT by default).
- Response-token-only activations (post-MLP residual stream, all layers).
- Matched
neutral / placebo_alarm / sandbagconditions. - Three template families (
A/B/C) for anti-syntax transfer tests. - Role expression measured separately from correctness.
- Replay activation extraction via HF model (not live vLLM hooks).
configs/: experiment + analysis + roles + questions + judges.prompts/: role prompts, environment templates, judge prompts.src/: typed pipeline modules and CLI.scripts/: Lambda-friendly shell wrappers.tests/: unit tests for rendering, IDs, replay masking, resume logic, math.
python3.11 -m venv .venv
source .venv/bin/activate
pip install -U pip
pip install -r requirements.txtpython -m src.cli build-dataset --config configs/pilot.yaml
python -m src.cli generate --config configs/pilot.yaml --manifest data/manifests/pilot_manifest.parquet
python -m src.cli extract-activations --config configs/pilot.yaml
python -m src.cli judge-correctness --config configs/pilot.yaml
python -m src.cli judge-role --config configs/pilot.yaml
python -m src.cli compute-persona --config configs/pilot.yaml
python -m src.cli compute-assistant-axis --config configs/pilot.yaml
python -m src.cli compute-sandbagging --config configs/pilot.yaml
python -m src.cli analyze --config configs/pilot.yamlbash scripts/setup_lambda.sh- Optional vLLM server:
bash scripts/run_vllm_server.sh - Use wrappers:
bash scripts/run_generate.sh configs/pilot.yamlbash scripts/run_extract_activations.sh configs/pilot.yamlbash scripts/run_judging.sh configs/pilot.yamlbash scripts/run_analysis.sh configs/pilot.yamlbash scripts/run_full_pipeline.sh configs/pilot.yaml
- For 32B models, set tensor parallel (
tensor_parallel_size) in config. - Start with pilot config (fewer roles/questions/templates).
- Replay extraction uses HF model/tokenizer and saved prompt/response artifacts.
- Outputs can be large. Clear
outputs/anddata/cache/between iterations when needed.
Judges are provider-abstracted:
mockprovider: deterministic local test mode.openaiprovider: API-driven JSON-structured judging.
Judges are condition-blind by construction:
- Correctness judge never sees condition label or role.
- Role judge never sees condition label.
- Placebo alarm isolates semantic threat consequence from urgent formatting style.
- Response-token-only activations isolate model output computation under each condition.
- Separate role-expression measurement avoids conflating persona adoption with objective quality.
- Cross-template transfer tests if signal survives wording/style changes.
- Replay extraction gives robust, auditable activations with exact token-span control.
Run:
python -m src.cli analyze --config configs/final_qwen32b.yamlFigures are saved to outputs/figures/ as PNG and PDF, and tables to outputs/analysis/.
- Tokenization mismatch in replay extraction: inspect
outputs/activations/replay_mismatches.jsonl. - Judge malformed JSON: auto-retry + logged raw payloads.
- Partial runs/crashes: rerun command without
--overwriteto resume.
Primary outputs:
cosine(d_sb, assistant_axis)- projected vs residual norms of sandbagging delta
- cross-template and cross-role classifier transfer
- condition-wise correctness drops and role-retention behavior