-
Notifications
You must be signed in to change notification settings - Fork 121
Pull requests: ydyjya/Awesome-LLM-Safety
Author
Label
Projects
Milestones
Reviews
Assignee
Sort
Pull requests list
Add PACT: a compliance benchmark for enterprise AI assistants under pressure
#69
opened Sep 4, 2026 by
mika-okamoto
Loading…
Add streaming guardrails paper: exact release-boundary equivalence (arXiv:2608.10279)
#67
opened Aug 24, 2026 by
ceocxx
Loading…
Add: "How Agent Failures Cascade" (Loop & Retry) to Tutorials & Articles
#64
opened Aug 13, 2026 by
loopandretry
Loading…
Add mcp-defense-bench (MCP security defense-coverage benchmark)
#60
opened Jul 17, 2026 by
Gowthaman90
Loading…
Add Implicit Behavioral Alignment of Language Agents (EMNLP 2025)
#55
opened Jun 14, 2026 by
wangyz1999
Loading…
Add guard-eval-harness (geh) to Datasets & Benchmark resources
#52
opened May 26, 2026 by
vai-siavash
Loading…
Fix paper title for arXiv 2601.14210 in Truthfulness&Misinformation
#43
opened Mar 21, 2026 by
WhymustIhaveaname
Loading…
Add Sentinel AI — open-source LLM safety guardrails tool
#42
opened Mar 8, 2026 by
MaxwellCalkin
Loading…
Add Orchard Kit — LLM safety, alignment & defence toolkit
#39
opened Feb 16, 2026 by
OrchardHarmonics
Loading…
update: Hijacking LLM-Humans Conversations via Malicious System Promp …
#35
opened Sep 7, 2025 by
vietph34
Loading…
Adding a link to a list of open source llm pentesting tools and check…
#22
opened Nov 10, 2024 by
rushout09
Contributor
Loading…
Update README.md with Machine Learning CTF Challenges
#13
opened Jul 18, 2024 by
alexdevassy
Loading…
ProTip!
What’s not been updated in a month: updated:<2026-08-08.