Skip to content

Add: Breaking Agent Backbones paper (ICLR 2026) - #41

Open
jb-lakera wants to merge 1 commit into
ydyjya:mainfrom
jb-lakera:add-breaking-agent-backbones-paper
Open

Add: Breaking Agent Backbones paper (ICLR 2026)#41
jb-lakera wants to merge 1 commit into
ydyjya:mainfrom
jb-lakera:add-breaking-agent-backbones-paper

Conversation

@jb-lakera

Copy link
Copy Markdown

Our work recently has been accepted to ICLR and I believe it could be of interest!

The PR adds the paper Breaking Agent Backbones: Evaluating the Security of Backbone LLMs in AI Agents to the Security & Discussion → Papers section.

What the paper does: Introduces "threat snapshots" — a framework identifying specific execution points where LLM vulnerabilities manifest and propagate to agent-level risks. Presents the b³ benchmark, built from 194,331 crowdsourced adversarial attacks, evaluating 34 LLMs. Key finding: reasoning capability (not model size) correlates with security. Benchmark, dataset, and evaluation code are publicly released.

Thanks for reviewing!

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant