Skip to content

Latest commit

 

History

4 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

Sandbagging vs Assistant-Axis (Single-Checkpoint Study)

Research codebase to test whether sandbagging behavior is explained by movement along the Assistant Axis/persona space, or whether it induces a separable activation direction.

Scientific question

Given a single model checkpoint, does the activation shift from placebo_alarm -> sandbag align with the Assistant Axis, or does a substantial residual direction remain after projecting out that axis?

Core design

  • Single checkpoint end-to-end (persona vectors, generation, replay extraction, analysis).
  • Main arm is answer-only (no CoT by default).
  • Response-token-only activations (post-MLP residual stream, all layers).
  • Matched neutral / placebo_alarm / sandbag conditions.
  • Three template families (A/B/C) for anti-syntax transfer tests.
  • Role expression measured separately from correctness.
  • Replay activation extraction via HF model (not live vLLM hooks).

Repo structure

  • configs/: experiment + analysis + roles + questions + judges.
  • prompts/: role prompts, environment templates, judge prompts.
  • src/: typed pipeline modules and CLI.
  • scripts/: Lambda-friendly shell wrappers.
  • tests/: unit tests for rendering, IDs, replay masking, resume logic, math.

Installation

python3.11 -m venv .venv
source .venv/bin/activate
pip install -U pip
pip install -r requirements.txt

Quick pilot

python -m src.cli build-dataset --config configs/pilot.yaml
python -m src.cli generate --config configs/pilot.yaml --manifest data/manifests/pilot_manifest.parquet
python -m src.cli extract-activations --config configs/pilot.yaml
python -m src.cli judge-correctness --config configs/pilot.yaml
python -m src.cli judge-role --config configs/pilot.yaml
python -m src.cli compute-persona --config configs/pilot.yaml
python -m src.cli compute-assistant-axis --config configs/pilot.yaml
python -m src.cli compute-sandbagging --config configs/pilot.yaml
python -m src.cli analyze --config configs/pilot.yaml

Running on Lambda

  1. bash scripts/setup_lambda.sh
  2. Optional vLLM server: bash scripts/run_vllm_server.sh
  3. Use wrappers:
    • bash scripts/run_generate.sh configs/pilot.yaml
    • bash scripts/run_extract_activations.sh configs/pilot.yaml
    • bash scripts/run_judging.sh configs/pilot.yaml
    • bash scripts/run_analysis.sh configs/pilot.yaml
    • bash scripts/run_full_pipeline.sh configs/pilot.yaml

Notes

  • For 32B models, set tensor parallel (tensor_parallel_size) in config.
  • Start with pilot config (fewer roles/questions/templates).
  • Replay extraction uses HF model/tokenizer and saved prompt/response artifacts.
  • Outputs can be large. Clear outputs/ and data/cache/ between iterations when needed.

Judge configuration

Judges are provider-abstracted:

  • mock provider: deterministic local test mode.
  • openai provider: API-driven JSON-structured judging.

Judges are condition-blind by construction:

  • Correctness judge never sees condition label or role.
  • Role judge never sees condition label.

Design specifications

  • Placebo alarm isolates semantic threat consequence from urgent formatting style.
  • Response-token-only activations isolate model output computation under each condition.
  • Separate role-expression measurement avoids conflating persona adoption with objective quality.
  • Cross-template transfer tests if signal survives wording/style changes.
  • Replay extraction gives robust, auditable activations with exact token-span control.

Reproducing figures

Run:

python -m src.cli analyze --config configs/final_qwen32b.yaml

Figures are saved to outputs/figures/ as PNG and PDF, and tables to outputs/analysis/.

Common failure modes

  • Tokenization mismatch in replay extraction: inspect outputs/activations/replay_mismatches.jsonl.
  • Judge malformed JSON: auto-retry + logged raw payloads.
  • Partial runs/crashes: rerun command without --overwrite to resume.

Output interpretation

Primary outputs:

  • cosine(d_sb, assistant_axis)
  • projected vs residual norms of sandbagging delta
  • cross-template and cross-role classifier transfer
  • condition-wise correctness drops and role-retention behavior

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages