[ICLR 2025] General-purpose activation steering library
-
Updated
Sep 18, 2025 - Python
[ICLR 2025] General-purpose activation steering library
Steering vectors for transformer language models in Pytorch / Huggingface
[🏆 CHI26 Best Paper] CoBRA: Reproducible control of LLM agent behavior via classic social science experiments
KV Cache Steering for Controlling Frozen LLMs
Agents that talk through model internals — activations & J-space — instead of text. No words pass between them. The lab on top of brainscope + hidden-directions.
Lightweight representation engineering dataflow operations for agent developers.
[EMNLP 2026 Main] Steering Geometry: Validating Human Value Geometry in LLM Steering Space.
Activation steering and trait monitoring for HuggingFace transformers
[Under Review] Not All Tokens Are Equally Useful for Steering: Robust Directions and Prefix Steering
Official code for "Activation Steering for Accent Adaptation in Speech Foundation Models" (Interspeech 2026). Parameter-free accent adaptation via mean-shift steering vectors — no weight updates, consistent WER reductions across 8 accents.
Steering vectors with receipts: make one, catch one, deploy a calibrated one. pip install hidden-directions
CRSM (Continuous Reasoning State Model): An asynchronous "System 2" architecture that implements Hierarchical State Sovereignty within a Mamba backbone. Unlike traditional search wrappers, CRSM uses Forward-Projected Planning and Sparse-Gated Injection to steer latent manifolds in real-time, decoupling strategic reasoning from token generation.
Turn a knob inside a small open model instead of writing a prompt. A reproducible RepE and CAA steering harness with an honest benchmark, including the failures.
Steer2Adapt: data-efficient inference-time LLM adaptation by composing steering vectors via Bayesian optimization over a semantic prior subspace.
Mechanistic interpretability experiments on political control circuits, refusal behavior, concept steering, and late-decoder interactions in open LLMs.
Phase-aware LLM activation steering and linear probing. A memory-efficient, practical implementation of Representation Engineering (RepE) for safety research.
The Anchor Paradox — stake, steerability, and the threshold of silicon life. Original theory and measurement framework by Simin Yuan.
Inspect, steer & monitor a real LLM on Apple Silicon — a browser-based SAE interpretability lab for Qwen Scope, powered by MLX
Functional emotional architecture for LLMs — 42 systems, 1994 tests, 27 psychological theories. Emergent emotions via 7 ANIMA pillars: predictive processing, global workspace, autobiographical memory, ontogenic development, motivational drives, emotional discovery, computational phenomenology.
Calibrated LLM observability toolkit: representation diagnostics, conformal risk calibration, verifier/control traces, and optional activation steering.
To associate your repository with the representation-engineering topic, visit your repo's landing page and select "manage topics."