[EMNLP 2026 Main] Steering Geometry: Validating Human Value Geometry in LLM Steering Space.
-
Updated
Sep 9, 2026 - Python
[EMNLP 2026 Main] Steering Geometry: Validating Human Value Geometry in LLM Steering Space.
Four main takeaways: (1) LLMs are subject to pressure, they comply despite expressing distress; (2) LLMs are vulnerable to gradual boundary/value violations; (3) when LLMs refuse, they may ignore the response format requirements, so the query is retried; (4) we hypothesise there is a token pattern continuation attractor that might cause obedience.
A comprehensive toolkit for implementing, analyzing, and validating AI value alignment based on Anthropic's 'Values in the Wild' research.
A research substrate for developmental intelligence — where artificial organisms develop through experience, memory, plasticity, and environmental history.
The forge, distilled: an ontology of three weeks of alignment research — every direction tried, colored verified / falsified / open, each color backed by a named artifact. Products: justitia, proxylimen, fallacy-cutter. Full tree at tag forge-full-tree.
Modular, meta-reflective architecture for structuring deliberate, long-term, value-aligned improvement in autonomous agents and human-AI teams.
AI ethics framework built on Layer 0 Principle: ∀x, V(x) > 0. Combines philosophical depth with measurable implementation.
LoRA fine-tuning to internalize BOHDI virtues into model weights
TriEthix is a novel evaluation framework that systematically benchmarks frontier LLMs across three foundational ethical perspectives: virtue, deontology, and consequentialism in 3 steps: (Step-1) Moral Weights; (Step-2) Moral Consistency; and (Step-3) Moral Reasoning. TriEthix reveals robust moral profiles for AI Safety, Governance, and Welfare.
A unified framework: Collective Resonance → Strange Attractors → Value Alignment → Algorithmic Intentionality → Emergent Algorithmic Behavior
A minimal evolutionary agent experiment: can a hardwired value compass survive extreme pressure?
To associate your repository with the value-alignment topic, visit your repo's landing page and select "manage topics."