A resource repository for representation engineering in large language models
-
Updated
Nov 14, 2024
A resource repository for representation engineering in large language models
Steering vectors for transformer language models in Pytorch / Huggingface
SOM network modified in order to control the latent space
[ICLR 2025] General-purpose activation steering library
A library for making RepE control vectors
Latent or Linguistic: Autonomous Behavior in LLMs. CS120 Final Project Submission
[🔥 ICLR 2026] - Misaligned Roles, Misplaced Images: Structural Input Perturbations Expose Multimodal Alignment Blind Spots
Official code for "Activation Steering for Accent Adaptation in Speech Foundation Models" (Interspeech 2026). Parameter-free accent adaptation via mean-shift steering vectors — no weight updates, consistent WER reductions across 8 accents.
Qwen3-0.6B activation steering: style vectors, lens contamination eval, CPRR methodology
Evaluation framework of different methods for probing and steering LLMs activations to mitigate Chain-of-Thought Unfaithfulness. Research project by Giovanni M. Occhipinti (University of Bologna), Alessandro Abate e Nandi Schoots (University of Oxford).
Phase-aware LLM activation steering and linear probing. A memory-efficient, practical implementation of Representation Engineering (RepE) for safety research.
LatentBiopsy: Geometric Anomaly Detection for LLM Residual Streams.
Representation Rerouting for Agentic Safety: Defending LLM Agents against Prompt Injection via Circuit Breakers and Triplet Loss.
[🏆 CHI26 Best Paper] CoBRA: Reproducible control of LLM agent behavior via classic social science experiments
CRSM (Continuous Reasoning State Model): An asynchronous "System 2" architecture that implements Hierarchical State Sovereignty within a Mamba backbone. Unlike traditional search wrappers, CRSM uses Forward-Projected Planning and Sparse-Gated Injection to steer latent manifolds in real-time, decoupling strategic reasoning from token generation.
Pre-generation tool-call gating via linear probes on LLM hidden states. F1 ≈ 0.91–0.94 on BFCL v4, 14–22× faster than full generation. Cross-architecture transfer across Llama / Qwen / Phi / Mistral (3B–7B) with ≥96% retention.
Mechanistic interpretability experiments on political control circuits, refusal behavior, concept steering, and late-decoder interactions in open LLMs.
Coordinate navigation for Transformer internal representations — zero training, any model
Functional emotional architecture for LLMs — 42 systems, 1994 tests, 27 psychological theories. Emergent emotions via 7 ANIMA pillars: predictive processing, global workspace, autobiographical memory, ontogenic development, motivational drives, emotional discovery, computational phenomenology.
To associate your repository with the representation-engineering topic, visit your repo's landing page and select "manage topics."