Steering vectors for transformer language models in Pytorch / Huggingface
-
Updated
Feb 21, 2025 - Python
Steering vectors for transformer language models in Pytorch / Huggingface
SOM network modified in order to control the latent space
[ICLR 2025] General-purpose activation steering library
Official code for "Activation Steering for Accent Adaptation in Speech Foundation Models" (Interspeech 2026). Parameter-free accent adaptation via mean-shift steering vectors — no weight updates, consistent WER reductions across 8 accents.
Qwen3-0.6B activation steering: style vectors, lens contamination eval, CPRR methodology
Evaluation framework of different methods for probing and steering LLMs activations to mitigate Chain-of-Thought Unfaithfulness. Research project by Giovanni M. Occhipinti (University of Bologna), Alessandro Abate e Nandi Schoots (University of Oxford).
Phase-aware LLM activation steering and linear probing. A memory-efficient, practical implementation of Representation Engineering (RepE) for safety research.
LatentBiopsy: Geometric Anomaly Detection for LLM Residual Streams.
Representation Rerouting for Agentic Safety: Defending LLM Agents against Prompt Injection via Circuit Breakers and Triplet Loss.
[🏆 CHI26 Best Paper] CoBRA: Reproducible control of LLM agent behavior via classic social science experiments
CRSM (Continuous Reasoning State Model): An asynchronous "System 2" architecture that implements Hierarchical State Sovereignty within a Mamba backbone. Unlike traditional search wrappers, CRSM uses Forward-Projected Planning and Sparse-Gated Injection to steer latent manifolds in real-time, decoupling strategic reasoning from token generation.
Mechanistic interpretability experiments on political control circuits, refusal behavior, concept steering, and late-decoder interactions in open LLMs.
Coordinate navigation for Transformer internal representations — zero training, any model
Functional emotional architecture for LLMs — 42 systems, 1994 tests, 27 psychological theories. Emergent emotions via 7 ANIMA pillars: predictive processing, global workspace, autobiographical memory, ontogenic development, motivational drives, emotional discovery, computational phenomenology.
Lightweight representation engineering dataflow operations for agent developers.
Training-time defense that redistributes LLM refusal via mean/covariance matching + KD, raising linear-ablation attack rank from K=1 to K≥16 (Llama-3.2-1B-Instruct)
Extending Anthropic's Persona Vectors to a vision-language model: affective image valence covertly shifts an extracted 'evil' direction in LLaVA-1.5-7B, invisibly to text-based harm monitors.
Steer2Adapt: data-efficient inference-time LLM adaptation by composing steering vectors via Bayesian optimization over a semantic prior subspace.
Independent reproduction & mechanistic study of K-Steering (Beyond Linear Steering, arXiv:2505.24535): when/why non-linear classifier-gradient steering beats the CAA linear baseline.
Local-first activation steering and representation-engineering lab with reproducible ActAdd demos, FastAPI workbench, React UI, and research-grounded limitations.
To associate your repository with the representation-engineering topic, visit your repo's landing page and select "manage topics."