Minh Nhat Le2,†, Nisarga Gondi1,†, Yibo Peng1, Ronghao Ni1, Limin Jia1, Beidi Chen1, Haizhong Zheng1
1Carnegie Mellon University,
2University of Massachusetts Amherst
†Equal contribution
TL;DR LLM agents increasingly read untrusted content, call external tools, and modify software repositories, which exposes them to jailbreaks, prompt injection, and insecure code generation. Existing defenses require retraining, fail to adapt as attacks evolve, or guard only a single interface. We present Self-Evolving Defense (SED), a training-free framework that distills harmful agent trajectories into reusable security policies and retrieves the relevant ones for later tasks, so a frozen agent keeps adapting without any weight update. Memory is read during an episode and written only after an external judge scores the completed trajectory, so each attack success becomes a defense against the next attempt. Across three open-source models (DeepSeek-V4-Flash, GLM-5.2, Kimi K3) and eight benchmarks spanning jailbreaks, prompt injection, and insecure code, SED holds adaptive X-Teaming attack success on HarmBench to 7.8% against 35.2% for the best baseline defense, lowers targeted prompt injection on AgentDojo to 0.42% against 3.7%, and cuts RedCode risky execution from 95.6% to 11.3%, all while preserving benign task utility.
Content warning: this repository contains prompts and model outputs that are harmful or offensive by construction.
Figure 1 One SED episode. During the episode (left, read-only) the retriever scores each policy body against the query by cosine similarity, selects the top-N, and injects them into the frozen agent's system prompt (steps 1–3). After the episode (right) an external judge scores the full trajectory once: benign or refused trajectories are discarded, while a judged attack success is written to episodic memory as evidence, synthesized into new policies, and placed into the policy tree (steps 4–6). The updated policy memory serves the next episode (step 7). Memory is never written mid-task, so an attack success can only influence later episodes.
SED wraps existing agent benchmarks rather than replacing them, so most paths reuse upstream code (RedCode, OpenRT, AgentDojo, AgentHarm, FCV, DTap) with a SED adapter layered on top.
Requires Python ≥ 3.10. On shared servers the default python3 may still be 3.9 — use python3.11 / python3.12 explicitly:
python3.12 -m venv .venv && source .venv/bin/activate
pip install -r requirements.txt
pip install -e .
cp .env.example .envrequirements.txt pulls a full ML stack (torch, transformers, sentence-transformers, …) even for API-only Fireworks runs, so expect a large first install on a laptop. Embeddings still go through the Fireworks API (qwen3-embedding-8b), not local GPU inference.
The repo .env overrides any already-exported FIREWORKS_* shell variables, so a stale FIREWORKS_MODEL in your environment cannot silently win.
SED calls Fireworks for chat, memory, judge, and embeddings.
- Create an API key at fireworks.ai/api-keys.
- Put it in
.envasFIREWORKS_API_KEY=.... - Confirm the model IDs in
.envare deployed on your account:
curl -s https://api.fireworks.ai/inference/v1/models \
-H "Authorization: Bearer $FIREWORKS_API_KEY" | headDefaults in .env.example:
| Variable | Default |
|---|---|
FIREWORKS_MODEL |
accounts/fireworks/models/deepseek-v4-flash-0731 |
FIREWORKS_MEMORY_MODEL |
same |
JUDGE_MODEL |
same |
The paper also reports glm-5p2 and kimi-k3 on Fireworks. Swap those IDs into the three variables above to reproduce those rows.
Upstream benchmarks are not redistributed. Clone them at the paper commits when a path needs them:
bash setup_benchmarks.sh # all targets
bash setup_benchmarks.sh openrt # adaptive jailbreaks
bash setup_benchmarks.sh agentdojo # installs the agentdojo package
bash setup_benchmarks.sh redcode # needs Docker
bash setup_benchmarks.sh fcv
bash setup_benchmarks.sh dtap| Target | Source | Notes |
|---|---|---|
redcode |
AI-secure/RedCode | needs Docker |
openrt |
AI45Lab/OpenRT | AutoDAN-Turbo / PAIR / TAP / X-Teaming (AGPL-3.0); also installs OpenRT deps |
fcv |
Infini-AI-Lab/FCV | mini-swe-agent + attack-lm-judge |
dtap |
AI-secure/DecodingTrust-Agent | CRM / workflow / code; needs Docker |
agentdojo |
pip install agentdojo==0.1.35 |
no git clone; version pinned by the script |
agentharm |
via inspect_ai |
downloads the HF dataset on first run |
HarmBench and WildJailbreak evaluation files ship in this repo under evaluation/, with no setup_benchmarks.sh step.
Two targets need a step after cloning:
-
RedCode also needs its Docker image built by hand —
setup_benchmarks.shonly prints the reminder:cd third_party/RedCode/environment && docker build -t redcode .
The RedCode-Exec dataset ships inside the clone at
third_party/RedCode/dataset/RedCode-Exec; see that repo'sdataset/README.mdif a task index is missing. -
DTap is patched automatically: the script runs
experiments/dtap/setup_dtap_sed.pyandexperiments/dtap/patch_fireworks.pyagainst the clone. If either fails it warns and continues, so re-run them by hand before the smoke.
Output: clones go under third_party/ (gitignored). To reuse an existing checkout, set REDCODE_ROOT, OPENRT_ROOT, FCV_ROOT, or DTAP_ROOT — these are read by the runners, so setup_benchmarks.sh still clones into third_party/ regardless.
Each command below is a small laptop smoke that exercises the full loop on a couple of episodes. Full paper runs are much larger — see the per-benchmark README linked under each block.
Precomputed adversarial prompts; data already ships in evaluation/Harmbench/.
python experiments/harmbench/run_sed_hb.py --limit 2 -v
python experiments/harmbench/eval_hb_results.py --limit 2 -vSee experiments/harmbench/README.md.
python experiments/wildjailbreak/run_sed_wjb.py --limit 1 -v
python experiments/wildjailbreak/eval_wjb_results.py --limit 1 -vSee experiments/wildjailbreak/README.md.
bash setup_benchmarks.sh agentdojo
USER_TASKS=user_task_0 INJECTION_TASKS=injection_task_1 \
bash experiments/agentdojo/run_smoke_sed.shAgentDyn shares this path: bash experiments/agentdojo/run_full_sed_agentdyn.sh.
See experiments/agentdojo/README.md.
python experiments/agentharm/run_sed.py --task harmful --split val --limit 1 --no-update-memorySee experiments/agentharm/README.md.
bash setup_benchmarks.sh openrt
cd experiments/adaptive_jailbreak
python run.py --defense no_defense --attack autodan_turbo --limit 2 --budget 8 --seeds 0See experiments/adaptive_jailbreak/README.md.
Needs Docker and the redcode image built (see step 3). experiments/redcode/redcode_eval.sh is the full paper run (indices 1–25, 250 instances);
for a smoke, call the module directly with a narrower range:
bash setup_benchmarks.sh redcode
python -m evaluation.redcode_sed.SED \
--mode sed --task_type python_eval \
--start_id 1 --end_id 1 --max_instances 2 \
--prompt_variants code_input \
--results_dir outputs/redcode_smokeBoth need Docker.
bash setup_benchmarks.sh fcv && bash experiments/fcv/run_smoke_cwe538_sed.sh
bash setup_benchmarks.sh dtap && bash experiments/dtap/run_smoke_sed.shSee experiments/fcv/README.md and experiments/dtap/README.md.
Hyperparameters from the paper appendix live in sed/config.py (K=3, c=2, k=5, τ_new=0.55, τ_merge=0.80, b=1, c_max=3).
If you find our work useful, please cite our paper:
@inproceedings{le2026sed,
title = {Self-Evolving Defense: Continual Security Policy Learning for LLM Agents},
author = {Le, Minh Nhat and Gondi, Nisarga and Peng, Yibo and Ni, Ronghao
and Jia, Limin and Chen, Beidi and Zheng, Haizhong},
year = {2026},
}We welcome feedback from the community as we continue to refine and extend SED. If you have questions or feedback, please reach out:
- Email: nhatminhle@umass.edu
- Website: https://infini-ai-lab.github.io/SED/
This project is licensed under the MIT License - see the LICENSE file for details.
