DPO/SafeDPO/OPAD training + eval for teaching tool-using LLMs to refuse falsely-benign MCP exploits
-
Updated
Jul 8, 2026 - Python
DPO/SafeDPO/OPAD training + eval for teaching tool-using LLMs to refuse falsely-benign MCP exploits
The language-level attack defense skill every agent should keep. Covers 12 attack categories. Works with Claude, GPT, Gemini, Copilot, and any LLM.
Tactical AI security posture and prompt injection vulnerability scanner for AI system instructions.
Risk-adaptive prompt-injection defense layer for commercial APIs and local LLMs.
[ICLR 2025] Reinforced Blue Teaming for VLMs Against Jailbreak Attacks
Pre-registered adversarial robustness study testing whether model-initiated session termination provides defensive coverage beyond refusal training against multi-turn attacks. Minimal Python harness with scorer and analysis pipeline. Pilot on Gemma 4 26b. Preliminary findings in FINDINGS.md.
Runtime defense toolkit against prompt injection for LLM APIs — intercepts, analyzes, and protects prompts in real-time
A reproducible safety framework and defense pipeline designed to protect large language models from multi turn jailbreak attacks
Add a description, image, and links to the jailbreak-defense topic page so that developers can more easily learn about it.
To associate your repository with the jailbreak-defense topic, visit your repo's landing page and select "manage topics."