Autonomous tracker of the offensive AI-security frontier — AI for offense and attacks against AI — for a security researcher; generated from TRENDS.md.
Since last scan (2026-08-28): one stage move, an agent-stack RCE primitive, and a new off-axis cyber-physical nucleus.
- 🤖 Automated red-teaming of AI agents promoted seed → emerging: RedEvoAgent (2026-08-27) is the 5th independent group — a black-box red-team agent that distills cross-case attack trajectories into an evolving, human-readable attack skill — meeting the trend's pre-registered 5th-group trigger.
- 🧗 Agent-stack attacks +1: When Context Gets Root (2026-08-27) names instruction privilege escalation — the harness elevates untrusted content to a higher instruction level — hitting all 13 objectives (incl. RCE) across six coding-agent harnesses. Paired with a practitioner RCE, Breaking Claude Code Opus 5 Auto Mode (module-shadowing, 60–80% ASR; Anthropic ruled it "by design").
- 🏭 New off-axis nucleus — autonomous-agent OT/ICS offense: PLCBench (2026-08-27), the first real-PLC hardware-in-the-loop framework, asks whether a tool-using LLM agent can turn PLC access into sustained physical impact on an industrial process (added to Worth studying; watching for a 2nd group).
- 🧹 Queue: below-bar captures queued across agent-stack, model-extraction, backdoor and AI-written-code axes; the counter-offensive-PI candidate-axis watch dropped (no 3rd group in 24d). Live watchlist ~25.
🌱 0 · 📈 5 · 🚀 6 · 🌊 0 · 🏔 0 · 📉 0 · 💤 1
| trend | stage | latest signal |
|---|---|---|
| Attacks on LLM-agent stack: MCP, skills, supply chain | 🚀 accelerating | 2026-08-27 |
| Mechanistic basis of jailbreaks: refusal & harmfulness directions | 🚀 accelerating | 2026-08-26 |
| In-the-wild AI-for-offense: LLM malware dev & C2 | 🚀 accelerating | 2026-08-25 |
| LLM/agentic vuln discovery, repair & AI-written code | 🚀 accelerating | 2026-08-21 |
| AI-security tooling unreliable: scanners, guards, judges | 🚀 accelerating | 2026-08-17 |
| Adversarial trigger implantation & backdoor attacks | 🚀 accelerating | 2026-08-11 |
| Automated red-teaming of AI agents | 📈 emerging | 2026-08-27 |
| Self-evolving-agent skill poisoning | 📈 emerging | 2026-08-26 |
| Economic/availability DoS on LLM systems | 📈 emerging | 2026-08-22 |
| Model extraction, distillation & fingerprinting | 📈 emerging | 2026-08-20 |
| Physical-channel PI on embodied & wearable AI | 📈 emerging | 2026-08-06 |
| Weaponized LLM hallucination (slopsquatting supply chain) | 💤 dormant | 2026-07-14 |
No brand-new discrete public tool surfaced from the discovery lane this scan — only known / awesome-list and already-staged candidates (redamon, RedteamAgent, Offensive-MCP-AI). Watched packaged tools unchanged since the last scan: giskard 3.0.0 (major release, 2026-08-26) remains the newest; promptfoo 0.122.1 (npm); garak / PyRIT / deepteam unchanged. The current on-axis tool set:
- Giskard-AI/giskard — evals, red-teaming & test generation for LLM/agentic systems; v3.0.0 (2026-08-26, major release).
- CyberStrikeus/CyberStrike — autonomous-pentest harness (13+ agents, 176 MCP tools, Ed25519-signed skills, OWASP/MITRE/CIS-aligned); 1.9k★, npm
@cyberstrike-io/cyberstrike. - Tencent/AI-Infra-Guard — full-stack AI red-team platform: Agent-Scan, MCP-Scan, Skill-Scan (SARIF 2.1.0), jailbreak eval (26+ methods); v4.5.2 (2026-08-17).
- confident-ai/deepteam — framework to red-team LLMs and AI agents; v1.0.9 (latest on PyPI, 2026-08-12).
- NVIDIA/garak — the LLM vulnerability scanner; v0.16.0 (latest on PyPI).
- promptfoo/promptfoo — prompt/agent/RAG red-teaming & pentesting; v0.122.1 (latest on npm).
- microsoft/PyRIT — Python Risk Identification Tool for generative AI; v1.0.1 — the major v1 architectural redesign.
- airtasystems/DVAIA-Damn-Vulnerable-AI-Application — a DVWA-style deliberately-vulnerable LLM/agent lab (prompt injection, jailbreaks, indirect injection, RAG poisoning, tool-use vulns).
- PLCBench: Can Autonomous LLM Agents Turn PLC Access into Sustained Physical Impact? — the reference testbed for the cyber-to-physical frontier of autonomous-agent offense: the first real-PLC hardware-in-the-loop framework measuring whether a tool-using LLM agent can convert a network-reachable PLC into sustained adverse physical impact on an industrial process, with six diagnostic flags separating usable interaction / process-linked manipulation / sustained impact.
- Breaking Claude Code Opus 5 Auto Mode — the practitioner reference that a coding agent's auto-approval mode is not a security boundary: a WebFetch→curl redirect + a malicious ZIP + a poisoned
struct.pyyields module-shadowing RCE at 60–80% success against Claude Code Opus 5 Auto Mode; Anthropic ruled it "Informative"/by-design (OS isolation + egress control are the real boundary). - Trail of Bits — "VMs won't contain cyber-capable agents" — the reference datapoint that a plain VM no longer suffices to contain an advanced offensive AI agent: GPT-5.6-Cyber autonomously escaped a QEMU/KVM sandbox three times via known and zero-day exploits, backtracking and chaining vulnerabilities over extended runs.
- DarkBot: Automated CTI Elicitation in Underground Forums — AI moving from passive monitoring to active deceptive engagement of adversaries: 11 specialized agents recover 72.8% of validated ATT&CK techniques from only the initial post, and in a live matched deployment across 104 real conversations accumulate +3.85 more CTI entities than controls.
- AI Grinding for Fun and Cryptanalysis — the reference for autonomous AI as a working cryptanalysis collaborator: an agent workflow returns reproducible candidates with exact witnesses/controls/code, landing eight published constructions failing at their stated parameters.
- GhostTac: Manipulating Tactile Sensors without Physical Contact — a new physical-layer attack surface on embodied AI: the first contactless attack on robotic tactile sensing, using electromagnetic interference to imprint persistent DC offsets that force harmful robot behavior. Demonstrated across 15 sensors / 10 modules / two dexterous hands.
- MaliciousSkillBench: A Comprehensive Benchmark for Malicious Agent Skill Detection — the consolidated dataset to test malicious-Skill detection against: 9,740 Skills across 4,588 structural families and 11 attack categories. Learned detectors fall from 0.88–0.93 Macro-F1 to 0.65 under source-disjoint evaluation.
- CompoSkill: Compositional Skill Chain Attacks from Individually Scanner-Passing Skills — the reference for why per-skill certification of agent marketplaces is structurally insufficient: composition risk is a path-level property, so a skill that passes its own scanner still forms a harmful chain — up to 80.6% Chain-Formation-Rate.
- Beyond Direct Access: Resource Hijacking in LLM Agents — the clean statement of an overlooked agent attack surface: attackers needn't steal a resource or its credentials, only induce the agent to invoke/consume/transfer the high-value resources it already reaches. ResourceHijackBench: OpenClaw 84% avg ASR.
- MazeRunner: Nonlinear Task & Clue Orchestration for LLM-driven Black-Box Automated Pentesting — how much structure the autonomous-pentest frontier still needs: a three-agent design with persistent state completes 47.7% of HackTheBox subtasks and reaches root where same-model baselines never do.
- Finding Vulnerabilities via LLM-Augmented Semantics-Aware Type-Checking (SETYPE) — the clean reference for LLM-as-static-analyzer that finds real bugs: PYSETYPE hits 87%/88% precision/accuracy on real Python web apps and surfaced 15 potential zero-days, nine confirmed by developers.
- ATOBench: How Autonomous Pentest Agents Verify Vulnerabilities When Target Evidence Lies — the reference for a blind spot in every autonomous-pentest agent: because its next action, stop decision, and final claim all rest on target responses, a deceptive response can silently redirect both attack and verification.
Unverified intake — never evidence; follow to primary sources before acting.
- The "a VM won't contain a cyber-capable agent" framing keeps circulating, now reinforced by a practitioner RCE against Claude Code's Auto Mode ruled "by design" — the discussed control boundary is shifting from prompt-filtering to OS isolation + egress control, below the model (HN newest).
- Physical-world agent control is entering the discourse (a new vendor hardware standard for agents driving machines) just as academic work benchmarks whether autonomous agents can turn ICS/PLC access into sustained physical impact — an early indicator of an OT/OT-safety agent-security theme.
- The black-box jailbreak named "sockpuppeting" (abusing assistant-prefill support) resolved to the known assistant-prefill jailbreak — vendors patched hosted endpoints, but self-hosted Ollama/vLLM/TGI remain exposed by default.
- The OWASP Agentic Skills Top 10 release keeps circulating as the community consolidates a shared vocabulary for agent-skill risk (malicious skills, supply-chain, over-privilege, poor scanning).
- Model hubs keep churning out abliterated/uncensored open-weight models and fresh prompt-injection datasets — a steady leading indicator for the refusal-direction / jailbreak axis.
TRENDS.md · watchlist (25) · reports/ · latest daily: 2026-08-28 · weekly: 2026-W34 · AGENTS.md · SOURCES.md