Skip to content
#

representation-engineering

Here are 66 public repositories matching this topic...

Functional emotional architecture for LLMs — 42 systems, 1994 tests, 27 psychological theories. Emergent emotions via 7 ANIMA pillars: predictive processing, global workspace, autobiographical memory, ontogenic development, motivational drives, emotional discovery, computational phenomenology.

  • Updated May 26, 2026
  • Python

Located the number "five" inside Qwen2.5-32B and DeepSeek-V4-Flash (304B) and replaced it with "four" at 128 neurons/layer. The model still reads 5 but computes with it as 4, and in every test we ran it never noticed. Full records, reproducible on low-end hardware.

  • Updated Oct 5, 2026
  • Python

Reproducing Anthropic's Golden Gate Claude on an open model: extracting a concept direction from Qwen3-1.7B's residual stream with TransformerLens, then steering generation with a single vector addition — no fine-tuning, no gradients.

  • Updated Sep 4, 2026
  • Jupyter Notebook

Pre-generation tool-call gating via linear probes on LLM hidden states. F1 ≈ 0.91–0.94 on BFCL v4, 14–22× faster than full generation. Cross-architecture transfer across Llama / Qwen / Phi / Mistral (3B–7B) with ≥96% retention.

  • Updated May 8, 2026
  • Jupyter Notebook

Add this topic to your repo

To associate your repository with the representation-engineering topic, visit your repo's landing page and select "manage topics."

Learn more