Skip to content
#

prompt-injection-defense

Here are 167 public repositories matching this topic...

A comprehensive reference for securing Large Language Models (LLMs). Covers OWASP GenAI Top-10 risks, prompt injection, adversarial attacks, real-world incidents, and practical defenses. Includes catalogs of red-teaming tools, guardrails, and mitigation strategies to help developers, researchers, and security teams deploy AI responsibly.

  • Updated Apr 3, 2026

PromptMe is an educational project that showcases security vulnerabilities in large language models (LLMs) and their web integrations. It includes 10 hands-on challenges inspired by the OWASP LLM Top 10, demonstrating how these vulnerabilities can be discovered and exploited in real-world scenarios.

  • Updated Sep 10, 2026
  • Python
agentlock

An adversarially benchmarked reference implementation for pre-action AI agent authorization. Provenance-based gating for LLM agent tool calls: deny-by-default permissions, parameter lineage, signed receipts, audit logging.

  • Updated Sep 18, 2026
  • Python

IntentFrame security plugin for Hermes Agent (Nous Research) — an external policy checkpoint that gates terminal, code, file, and cron tool calls before they run on your machine.

  • Updated Jun 28, 2026
  • Python

Personal finance agent that decides whether an expense is safe — pay in full, split, use installments, or wait — from a 90-day balance forecast over real transaction history, with a deterministic decision core and a narrow multimodal evidence layer for receipts and messages.

  • Updated Sep 15, 2026
  • Python

Open test set of .eml emails for checking how email security controls and AI mailbox assistants handle indirect prompt injection. It crosses three intents (system-prompt disclosure, data exfiltration, tool discovery) with three delivery methods (plaintext, hidden HTML, Base64), plus a benign control. All content is fictional and non-routable.

  • Updated Oct 2, 2026
  • Python

Detect and sanitize prompt injection attacks in Rails apps. Protects against direct injection (users hacking your LLMs via form inputs) and indirect injection (malicious prompts stored for other LLMs to scrape). ~70 detection patterns across 7 attack categories with configurable sensitivity levels. Now includes resource extraction detection pattern

  • Updated Feb 25, 2026
  • Ruby

Belay is an open-source, local-first security layer for AI coding agents (Claude Code, Codex, Cursor, OpenClaw, Hermes Agent and MCP) that blocks dangerous commands, secret leaks, and prompt injection at the tool-call boundary in under 100ms — no LLM in the decision path by default, no cloud, no phone-home.

  • Updated Sep 8, 2026
  • Rust

Add this topic to your repo

To associate your repository with the prompt-injection-defense topic, visit your repo's landing page and select "manage topics."

Learn more