A library for making RepE control vectors
-
Updated
Sep 24, 2025 - Jupyter Notebook
A library for making RepE control vectors
[ICLR 2025] General-purpose activation steering library
Steering vectors for transformer language models in Pytorch / Huggingface
A resource repository for representation engineering in large language models
KV Cache Steering for Controlling Frozen LLMs
[🏆 CHI26 Best Paper] CoBRA: Reproducible control of LLM agent behavior via classic social science experiments
Agents that talk through model internals — activations & J-space — instead of text. No words pass between them. The lab on top of brainscope + hidden-directions.
[EMNLP 2026 Main] Steering Geometry: Validating Human Value Geometry in LLM Steering Space.
Steering vectors with receipts: make one, catch one, deploy a calibrated one. pip install hidden-directions
Functional emotional architecture for LLMs — 42 systems, 1994 tests, 27 psychological theories. Emergent emotions via 7 ANIMA pillars: predictive processing, global workspace, autobiographical memory, ontogenic development, motivational drives, emotional discovery, computational phenomenology.
Lightweight representation engineering dataflow operations for agent developers.
Official code for "Activation Steering for Accent Adaptation in Speech Foundation Models" (Interspeech 2026). Parameter-free accent adaptation via mean-shift steering vectors — no weight updates, consistent WER reductions across 8 accents.
情绪旋钮:找到、读取并拨动大语言模型内部的情绪方向(不训练/不改权重)。Find, read and steer emotion directions in LLMs.
A preregistered hypothesis and experimental framework for testing task-induced appraisal-structured latent control representations in language models.
Located the number "five" inside Qwen2.5-32B and DeepSeek-V4-Flash (304B) and replaced it with "four" at 128 neurons/layer. The model still reads 5 but computes with it as 4, and in every test we ran it never noticed. Full records, reproducible on low-end hardware.
Reproducing Anthropic's Golden Gate Claude on an open model: extracting a concept direction from Qwen3-1.7B's residual stream with TransformerLens, then steering generation with a single vector addition — no fine-tuning, no gradients.
Pre-generation tool-call gating via linear probes on LLM hidden states. F1 ≈ 0.91–0.94 on BFCL v4, 14–22× faster than full generation. Cross-architecture transfer across Llama / Qwen / Phi / Mistral (3B–7B) with ≥96% retention.
Gemma 4 abliteration and mechanistic interpretability lab for refusal-direction extraction, activation steering, weight editing, and rigorous LLM evaluation.
Steer2Adapt: data-efficient inference-time LLM adaptation by composing steering vectors via Bayesian optimization over a semantic prior subspace.
To associate your repository with the representation-engineering topic, visit your repo's landing page and select "manage topics."