Provenance-first RAG that refuses to hallucinate.
-
Updated
Oct 7, 2026 - TypeScript
Provenance-first RAG that refuses to hallucinate.
Visual Abstention in Unified Multimodal Models: the Draw-or-Decline (DoD) benchmark and VisTA training for image editing models that decline infeasible requests.
Refusal gates for pi. Halt work that should not happen, and fail closed when they cannot tell.
Does DPO build a new safety representation, or amplify one the base model already had? Tracing a refusal direction across a Base to SFT to DPO chain.
We study whether categorical refusal tokens enable controllable and interpretable safety behavior in language models.
🔓 Ablate — directional ablation (abliteration) toolkit for open-source LLMs. Automatic censorship/refusal removal via residual-stream direction ablation, with KL-guided search, an LLM-judge harness, and one-call push to the Hub. pip install ablate-llm
Refusal-direction ablation on MoE language models, measured: what it costs in capability, which implementation choices break it, and whether community 'abliterated' uploads do anything. Scripts, prompt set and results; no weights.
Activation-level interventions that come with their controls — steer, ablate, patch, and SAE feature analysis for open-weight LLMs.
A probe suite that measures which conversation states an LLM cannot leave. Three arms, because two cannot tell obedience from token statistics; a null only counts when the design had the power to see the effect.
RefusalScope
Cross-architecture refusal-direction ablation study: Qwen 2.5 + Gemma 2/3/4. Mechanistic explanation for why Gemma 3 specifically admits single-layer jailbreaks.
An open reproduction of feature-level activation steering with the prompt set released, showing the capability tax that behavioural metrics miss
Reproducible refusal-rate evaluation harness for open-weight LLMs — adversarial-prompt benchmarks with byte-identical reruns.
How does this function's cost scale? Counted, not timed — and UNDETERMINED when no complexity class settles.
Every RAG eval reports faithfulness. None reports over-refusal beside it. An eval harness and policy engine for refusal correctness in regulated RAG: synthetic KYC/AML corpus, deterministic verifier, and a refusal policy you can commit, hash and audit.
Grounding, citation and refusal in a document generator: every claim carries the record behind it, and claims with no record are blocked before they ship. Demo data.
Bank and card statements into a reconciled ledger, with a plain-English refusal for the ones that do not add up: silently wrong money $62,832.40 to $0.00.
A document copilot that cites what it says and refuses when the evidence is not there: 33 questions, baseline 12/33 to harness 32/33, control 8/8 both ways.
中文企业公开报告 Hybrid RAG:Docling 解析 · Qdrant 稠密/稀疏检索 · 查询理解硬过滤 · 带引用生成与拒答 · 文档生命周期与评测看板
Public reference interfaces for proof-gated AI action, refusal, authority, and evidence boundaries.
To associate your repository with the refusal topic, visit your repo's landing page and select "manage topics."