Stanford NLP Python library for understanding and improving PyTorch models via interventions
-
Updated
Mar 6, 2026 - Python
Stanford NLP Python library for understanding and improving PyTorch models via interventions
Stop AI agents from doing things you didn't ask for.
Stanford NLP Python library for benchmarking the utility of LLM interpretability methods
Explainability of Deep Learning Models
🖼️ Enhance images effortlessly by adding or removing objects with the Qwen-Image-Edit-Object-Manipulator, ensuring realism and background detail.
Projet refait entièrement dans la v2 web
Implementation for the NeurIPS 2025 paper: An Analysis of Causal Effect Estimation using Outcome Invariant Data Augmentation
This project explores methods to detect and mitigate jailbreak behaviors in Large Language Models (LLMs). By analyzing activation patterns—particularly in deeper layers—we identify distinct differences between compliant and non-compliant responses to uncover a jailbreak "direction." Using this insight, we develop intervention strategies that modify
This is the Github repository for the preprint https://arxiv.org/abs/2505.19612
Causal inference of post-transcriptional regulation timelines from long-read sequencing in Arabidopsis thaliana
Validate and summarize human intervention and recovery segments in robot episodes.
Judea Pearl’s Causal Ladder, featuring Association, Intervention, and Counterfactual models.
Data package for the ON LiMiT feasibility study
To associate your repository with the intervention topic, visit your repo's landing page and select "manage topics."