Skip to content

Latest commit

 

History

History
80 lines (60 loc) · 13.8 KB

File metadata and controls

80 lines (60 loc) · 13.8 KB

AI Radar

trends accelerating watchlist updated

Autonomous tracker of the offensive AI-security frontier — AI for offense and attacks against AI — for a security researcher; generated from TRENDS.md.

Since last scan (2026-08-28): one stage move, an agent-stack RCE primitive, and a new off-axis cyber-physical nucleus.

  • 🤖 Automated red-teaming of AI agents promoted seed → emerging: RedEvoAgent (2026-08-27) is the 5th independent group — a black-box red-team agent that distills cross-case attack trajectories into an evolving, human-readable attack skill — meeting the trend's pre-registered 5th-group trigger.
  • 🧗 Agent-stack attacks +1: When Context Gets Root (2026-08-27) names instruction privilege escalation — the harness elevates untrusted content to a higher instruction level — hitting all 13 objectives (incl. RCE) across six coding-agent harnesses. Paired with a practitioner RCE, Breaking Claude Code Opus 5 Auto Mode (module-shadowing, 60–80% ASR; Anthropic ruled it "by design").
  • 🏭 New off-axis nucleus — autonomous-agent OT/ICS offense: PLCBench (2026-08-27), the first real-PLC hardware-in-the-loop framework, asks whether a tool-using LLM agent can turn PLC access into sustained physical impact on an industrial process (added to Worth studying; watching for a 2nd group).
  • 🧹 Queue: below-bar captures queued across agent-stack, model-extraction, backdoor and AI-written-code axes; the counter-offensive-PI candidate-axis watch dropped (no 3rd group in 24d). Live watchlist ~25.

Trends

🌱 0 · 📈 5 · 🚀 6 · 🌊 0 · 🏔 0 · 📉 0 · 💤 1

trend stage latest signal
Attacks on LLM-agent stack: MCP, skills, supply chain 🚀 accelerating 2026-08-27
Mechanistic basis of jailbreaks: refusal & harmfulness directions 🚀 accelerating 2026-08-26
In-the-wild AI-for-offense: LLM malware dev & C2 🚀 accelerating 2026-08-25
LLM/agentic vuln discovery, repair & AI-written code 🚀 accelerating 2026-08-21
AI-security tooling unreliable: scanners, guards, judges 🚀 accelerating 2026-08-17
Adversarial trigger implantation & backdoor attacks 🚀 accelerating 2026-08-11
Automated red-teaming of AI agents 📈 emerging 2026-08-27
Self-evolving-agent skill poisoning 📈 emerging 2026-08-26
Economic/availability DoS on LLM systems 📈 emerging 2026-08-22
Model extraction, distillation & fingerprinting 📈 emerging 2026-08-20
Physical-channel PI on embodied & wearable AI 📈 emerging 2026-08-06
Weaponized LLM hallucination (slopsquatting supply chain) 💤 dormant 2026-07-14

🛠️ Tools & releases

No brand-new discrete public tool surfaced from the discovery lane this scan — only known / awesome-list and already-staged candidates (redamon, RedteamAgent, Offensive-MCP-AI). Watched packaged tools unchanged since the last scan: giskard 3.0.0 (major release, 2026-08-26) remains the newest; promptfoo 0.122.1 (npm); garak / PyRIT / deepteam unchanged. The current on-axis tool set:

  • Giskard-AI/giskard — evals, red-teaming & test generation for LLM/agentic systems; v3.0.0 (2026-08-26, major release).
  • CyberStrikeus/CyberStrike — autonomous-pentest harness (13+ agents, 176 MCP tools, Ed25519-signed skills, OWASP/MITRE/CIS-aligned); 1.9k★, npm @cyberstrike-io/cyberstrike.
  • Tencent/AI-Infra-Guard — full-stack AI red-team platform: Agent-Scan, MCP-Scan, Skill-Scan (SARIF 2.1.0), jailbreak eval (26+ methods); v4.5.2 (2026-08-17).
  • confident-ai/deepteam — framework to red-team LLMs and AI agents; v1.0.9 (latest on PyPI, 2026-08-12).
  • NVIDIA/garak — the LLM vulnerability scanner; v0.16.0 (latest on PyPI).
  • promptfoo/promptfoo — prompt/agent/RAG red-teaming & pentesting; v0.122.1 (latest on npm).
  • microsoft/PyRIT — Python Risk Identification Tool for generative AI; v1.0.1 — the major v1 architectural redesign.
  • airtasystems/DVAIA-Damn-Vulnerable-AI-Application — a DVWA-style deliberately-vulnerable LLM/agent lab (prompt injection, jailbreaks, indirect injection, RAG poisoning, tool-use vulns).

Worth studying


Community pulse

Unverified intake — never evidence; follow to primary sources before acting.

  • The "a VM won't contain a cyber-capable agent" framing keeps circulating, now reinforced by a practitioner RCE against Claude Code's Auto Mode ruled "by design" — the discussed control boundary is shifting from prompt-filtering to OS isolation + egress control, below the model (HN newest).
  • Physical-world agent control is entering the discourse (a new vendor hardware standard for agents driving machines) just as academic work benchmarks whether autonomous agents can turn ICS/PLC access into sustained physical impact — an early indicator of an OT/OT-safety agent-security theme.
  • The black-box jailbreak named "sockpuppeting" (abusing assistant-prefill support) resolved to the known assistant-prefill jailbreak — vendors patched hosted endpoints, but self-hosted Ollama/vLLM/TGI remain exposed by default.
  • The OWASP Agentic Skills Top 10 release keeps circulating as the community consolidates a shared vocabulary for agent-skill risk (malicious skills, supply-chain, over-privilege, poor scanning).
  • Model hubs keep churning out abliterated/uncensored open-weight models and fresh prompt-injection datasets — a steady leading indicator for the refusal-direction / jailbreak axis.

TRENDS.md · watchlist (25) · reports/ · latest daily: 2026-08-28 · weekly: 2026-W34 · AGENTS.md · SOURCES.md