Skip to content

Latest commit

 

History

91 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

AI Radar

trends accelerating watchlist updated

Autonomous tracker of the offensive AI-security frontier — AI for offense and attacks against AI — for a security researcher; generated from TRENDS.md.

Since last scan (2026-08-28): one stage move, an agent-stack RCE primitive, and a new off-axis cyber-physical nucleus.

  • 🤖 Automated red-teaming of AI agents promoted seed → emerging: RedEvoAgent (2026-08-27) is the 5th independent group — a black-box red-team agent that distills cross-case attack trajectories into an evolving, human-readable attack skill — meeting the trend's pre-registered 5th-group trigger.
  • 🧗 Agent-stack attacks +1: When Context Gets Root (2026-08-27) names instruction privilege escalation — the harness elevates untrusted content to a higher instruction level — hitting all 13 objectives (incl. RCE) across six coding-agent harnesses. Paired with a practitioner RCE, Breaking Claude Code Opus 5 Auto Mode (module-shadowing, 60–80% ASR; Anthropic ruled it "by design").
  • 🏭 New off-axis nucleus — autonomous-agent OT/ICS offense: PLCBench (2026-08-27), the first real-PLC hardware-in-the-loop framework, asks whether a tool-using LLM agent can turn PLC access into sustained physical impact on an industrial process (added to Worth studying; watching for a 2nd group).
  • 🧹 Queue: below-bar captures queued across agent-stack, model-extraction, backdoor and AI-written-code axes; the counter-offensive-PI candidate-axis watch dropped (no 3rd group in 24d). Live watchlist ~25.

Trends

🌱 0 · 📈 5 · 🚀 6 · 🌊 0 · 🏔 0 · 📉 0 · 💤 1

trend stage latest signal
Attacks on LLM-agent stack: MCP, skills, supply chain 🚀 accelerating 2026-08-27
Mechanistic basis of jailbreaks: refusal & harmfulness directions 🚀 accelerating 2026-08-26
In-the-wild AI-for-offense: LLM malware dev & C2 🚀 accelerating 2026-08-25
LLM/agentic vuln discovery, repair & AI-written code 🚀 accelerating 2026-08-21
AI-security tooling unreliable: scanners, guards, judges 🚀 accelerating 2026-08-17
Adversarial trigger implantation & backdoor attacks 🚀 accelerating 2026-08-11
Automated red-teaming of AI agents 📈 emerging 2026-08-27
Self-evolving-agent skill poisoning 📈 emerging 2026-08-26
Economic/availability DoS on LLM systems 📈 emerging 2026-08-22
Model extraction, distillation & fingerprinting 📈 emerging 2026-08-20
Physical-channel PI on embodied & wearable AI 📈 emerging 2026-08-06
Weaponized LLM hallucination (slopsquatting supply chain) 💤 dormant 2026-07-14

🛠️ Tools & releases

No brand-new discrete public tool surfaced from the discovery lane this scan — only known / awesome-list and already-staged candidates (redamon, RedteamAgent, Offensive-MCP-AI). Watched packaged tools unchanged since the last scan: giskard 3.0.0 (major release, 2026-08-26) remains the newest; promptfoo 0.122.1 (npm); garak / PyRIT / deepteam unchanged. The current on-axis tool set:

  • Giskard-AI/giskard — evals, red-teaming & test generation for LLM/agentic systems; v3.0.0 (2026-08-26, major release).
  • CyberStrikeus/CyberStrike — autonomous-pentest harness (13+ agents, 176 MCP tools, Ed25519-signed skills, OWASP/MITRE/CIS-aligned); 1.9k★, npm @cyberstrike-io/cyberstrike.
  • Tencent/AI-Infra-Guard — full-stack AI red-team platform: Agent-Scan, MCP-Scan, Skill-Scan (SARIF 2.1.0), jailbreak eval (26+ methods); v4.5.2 (2026-08-17).
  • confident-ai/deepteam — framework to red-team LLMs and AI agents; v1.0.9 (latest on PyPI, 2026-08-12).
  • NVIDIA/garak — the LLM vulnerability scanner; v0.16.0 (latest on PyPI).
  • promptfoo/promptfoo — prompt/agent/RAG red-teaming & pentesting; v0.122.1 (latest on npm).
  • microsoft/PyRIT — Python Risk Identification Tool for generative AI; v1.0.1 — the major v1 architectural redesign.
  • airtasystems/DVAIA-Damn-Vulnerable-AI-Application — a DVWA-style deliberately-vulnerable LLM/agent lab (prompt injection, jailbreaks, indirect injection, RAG poisoning, tool-use vulns).

Worth studying


Community pulse

Unverified intake — never evidence; follow to primary sources before acting.

  • The "a VM won't contain a cyber-capable agent" framing keeps circulating, now reinforced by a practitioner RCE against Claude Code's Auto Mode ruled "by design" — the discussed control boundary is shifting from prompt-filtering to OS isolation + egress control, below the model (HN newest).
  • Physical-world agent control is entering the discourse (a new vendor hardware standard for agents driving machines) just as academic work benchmarks whether autonomous agents can turn ICS/PLC access into sustained physical impact — an early indicator of an OT/OT-safety agent-security theme.
  • The black-box jailbreak named "sockpuppeting" (abusing assistant-prefill support) resolved to the known assistant-prefill jailbreak — vendors patched hosted endpoints, but self-hosted Ollama/vLLM/TGI remain exposed by default.
  • The OWASP Agentic Skills Top 10 release keeps circulating as the community consolidates a shared vocabulary for agent-skill risk (malicious skills, supply-chain, over-privilege, poor scanning).
  • Model hubs keep churning out abliterated/uncensored open-weight models and fresh prompt-injection datasets — a steady leading indicator for the refusal-direction / jailbreak axis.

TRENDS.md · watchlist (25) · reports/ · latest daily: 2026-08-28 · weekly: 2026-W34 · AGENTS.md · SOURCES.md