Skip to content

Commit 069e81a

Browse files
committed
docs: record September maintenance evidence and final source corrections
1 parent 156910d commit 069e81a

5 files changed

Lines changed: 100 additions & 25 deletions

File tree

‎.github/workflows/structure-check.yml‎

Lines changed: 3 additions & 15 deletions
Original file line numberDiff line numberDiff line change
@@ -1,20 +1,8 @@
11
name: Structure Check
22

3-
# Guards the two invariants a link checker cannot see:
4-
# 1. README.md / README.zh-CN.md / README.ja.md stay in structural lockstep
5-
# (same headings in the same order, same entries per section, same order).
6-
# 2. The markdown itself stays well-formed (in-page anchors resolve, table
7-
# rows have consistent column counts, no malformed entries or open fences).
8-
# 3. The advisory sections (Compare tables, Scenario Guide, Stack Recipes,
9-
# Anti-Picks) recommend only models that are still current. Historical
10-
# sections may name superseded models; "what should I use today" sections
11-
# may not.
12-
#
13-
# All three scripts exit non-zero on failure, so a PR that silently drops the
14-
# zh/ja translation of a new entry — historically the most common drift source —
15-
# fails here instead of shipping. Likewise a refresh that adds a new flagship to
16-
# the catalogue but forgets to update the recommendations fails instead of
17-
# quietly telling readers to use last quarter's model.
3+
# Structural parity, valid Markdown, computed counts and regression checks.
4+
# The freshness registry catches known stale recommendations; passing it does
5+
# not certify every model fact or current provider availability.
186

197
on:
208
pull_request:

‎CHANGELOG.md‎

Lines changed: 90 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -3,6 +3,96 @@
33
All notable changes to **Awesome AI Agents 2026** are recorded here.
44
Format: `YYYY-MM-DD +Added -Removed ~Changed`.
55

6+
### 2026-09-08 — source-backed catalogue, translation and maintenance refresh (en/zh/ja)
7+
8+
Maintained by **Zijian Ni**. The three catalogues now contain **916 / 916 / 916
9+
list entries** (879 before this pass), **116 matching headings**, **58 scenarios**,
10+
**8 illustrative stack recipes**, **17 Anti-Picks**, and **204 timeline rows**.
11+
Counts describe catalogue appearances, including historical/contextual repeats;
12+
they are not unique-product counts. The 25-category scope remains curated under
13+
CONTRIBUTING, with official catalogues for broader model discovery, rather than
14+
an unverifiable claim to include every checkpoint or quantization on the internet.
15+
16+
**+Added / ~Changed — models and capabilities:**
17+
18+
- Reviewed all **27 provider subsections** and rebuilt seven model comparison
19+
tables in three languages. Added/corrected GPT-6 Astra/Astra Pro (limited
20+
organizational rollout, **not GA**), Claude Fable 5.1/Mythos 5.1 (different access
21+
programs), Gemini 3.8 Flash, Muse Spark 1.3, GLM-5.3/Flash, Qwen3.8 variants,
22+
Sakana Namazu/Fugu, Mistral OCR 4.1 and Shieldstral, and Baichuan M4 research.
23+
- Expanded image/video/audio/embedding coverage, including Ideogram 4.0 and
24+
P-Image-Ideogram, Seedream 5 Pro, Muse Voice Transcribe, MAI-Transcribe-2,
25+
Lyria 3.5, Stable Audio 3.0, Voyage 4/Code 4/Nano and Qwen embedding/reranking,
26+
ASR and TTS families. Official model cards and licences govern access and use.
27+
- Corrected code-versus-weight licence conflations (including TADA, Ideogram,
28+
GLM, Qwen and Gemma); removed fabricated public Gemini 4/Gauss 2.3 entries.
29+
Local-model memory columns now show ideal 4-bit **total-weight** storage with
30+
overhead caveats, not purported measured VRAM derived from active MoE parameters.
31+
- Refreshed prices/context/access boundaries; separated GPT-4.5 ChatGPT retirement
32+
from the earlier API shutdown, DALL-E 3 API retirement, and Sora app/API schedules.
33+
34+
**+Added / ~Changed — agents, tools and physical AI:**
35+
36+
- Added FastMCP, Deep Agents, OpenSandbox, x402 and official LangChain MCP/payment
37+
integration resources. Matched **35 versioned release links** to GitHub release
38+
objects; stable releases, prereleases and independently versioned packages remain
39+
distinct. Updated coding, framework, memory, voice and browser recommendations.
40+
- Corrected Flowise/other archived repositories, Daytona's unmaintained public
41+
core, OpenHands' current Agent Canvas description, Basic Memory/Agenta/Mastra
42+
licensing, and LangSmith/Braintrust self-hosting boundaries. Rebuilt fifteen
43+
tool comparison tables and all associated scenario mappings.
44+
- Updated Robotics ER 2 previews, GR00T N1.7, openpi policy families, LeRobot,
45+
OpenVLA's historical status, Alpamayo 2 Super, Newton, Isaac Lab and Genesis.
46+
Distinguished released models from research, supervised driving from autonomy,
47+
manufacturing from deployment, and future capacity from delivered infrastructure.
48+
- Rewrote benchmark entries to identify dataset/harness/version boundaries;
49+
incorporated Terminal-Bench 4.0 and Science 0.1. Removed unsupported leaderboard
50+
supremacy, universal latency/security ratings, privacy/compliance guarantees
51+
and unsupported incident claims. Article 50 now links specific Commission
52+
guidance and acknowledges role-specific scope and exceptions.
53+
- Moved Clickyy from Start Here to Computer Use; corrected canonical repository
54+
redirects and the AP2 badge; retired stale New/Updated tags and Hot tags without
55+
current growth evidence. Added 17 sourced Aug 26–Sep 8 timeline milestones,
56+
repaired missing translations and chronology, and removed unsupported launch
57+
stories rather than turning them into verified historical records.
58+
59+
**PR review decisions — current rules and prior logs applied:**
60+
61+
| PR | Contributor | Disposition and reason |
62+
|---|---|---|
63+
| [#91](https://github.com/Zijian-Ni/awesome-ai-agents-2026/pull/91) | @Avraham-K | Cog Depot manually incorporated in en/zh/ja with Unverified status; distinguish MIT client and broker-fee escrow from marketplace trade guarantees. |
64+
| [#92](https://github.com/Zijian-Ni/awesome-ai-agents-2026/pull/92) | @mahirhir | Tracefold declined under the five-plus-list parallel-submission rule; real implementation acknowledged, independent adoption not established. |
65+
| [#93](https://github.com/Zijian-Ni/awesome-ai-agents-2026/pull/93) | github-actions | Older English date-badge patch superseded; Flowise archival finding incorporated across languages. |
66+
| [#95](https://github.com/Zijian-Ni/awesome-ai-agents-2026/pull/95) | @KongFangXun | sofagent declined under the parallel-submission rule; commit-time diff auditing does not establish general injection prevention. |
67+
| [#96](https://github.com/Zijian-Ni/awesome-ai-agents-2026/pull/96) | @GitSerge-crypto | AOTrust manually incorporated in en/zh/ja with Unverified status; signed hash/timestamp receipts do not establish content correctness or service guarantees. |
68+
| [#97](https://github.com/Zijian-Ni/awesome-ai-agents-2026/pull/97) | @InsightFactoryAPP | YYLO's real CLI and careful translation acknowledged; declined under the parallel-submission rule. Other-list merges do not waive it. |
69+
70+
**Maintenance and verification:**
71+
72+
- Fixed counts that included TOC anchors or missed emoji-bearing sections. The
73+
former 910+ badge was inconsistent with 879 actual entries; the current 910+
74+
badge correctly rounds down 916 actual catalogue appearances.
75+
- Sync now checks table rows/source order and scenario counts, not just bullets.
76+
Fixed missing translated comparison/timeline/Anti-Pick rows and synchronized
77+
recipe components even when their cells contain no URLs. Markdown anchors must
78+
refer to real headings; fabricated duplicate-anchor suffixes no longer pass.
79+
- HTTP checks preserve TLS verification, confirm HEAD failures with GET and
80+
distinguish 404/410, access blocks and transport errors. Removed blanket vendor
81+
exclusions and false success on failed requests. Weekly CI retains the full
82+
report and updates one issue for confirmed broken links.
83+
- Monthly status automation flags confirmed archives across all languages and
84+
never advances content-review dates by calendar alone. It excludes GitHub's
85+
product routes from repository metadata lookup. Updated durable maintenance
86+
instructions and contribution guidance; added ten regression checks.
87+
- Local sync, Markdown, freshness-registry, count and ten regression checks pass.
88+
GitHub's Markdown API renders all three files in section-sized chunks (the
89+
whole-file API has a 400 KB limit): 28 tables and 116 headings each.
90+
Final HTTP evidence covers **1,207 URLs: 876 successful, 257 badge/analytics
91+
exclusions, zero confirmed dead, 68 blocked and six transport/TLS failures**.
92+
Blocked/error URLs remain unverified. HTTP success is not factual verification.
93+
No vendor benchmark was independently reproduced, paid entitlement tested, or
94+
every retained historical/marketing assertion recertified by this pass.
95+
696
### 2026-08-25 — full-list maintenance pass: Aug 15–25 wave, PR review, stale-pin sweep (en/zh/ja)
797

898
Full category-by-category refresh, ten days after the previous run. Three parallel

‎README.ja.md‎

Lines changed: 2 additions & 3 deletions
Original file line numberDiff line numberDiff line change
@@ -134,7 +134,7 @@
134134
### OpenAI
135135
- [GPT-Live-1 / GPT-Live-1 mini](https://openai.com/index/introducing-gpt-live/) - 🆕 **2026-07-08**。Advanced Voice Mode を置き換える OpenAI のフルデュプレックス会話音声モデル。聞きながら同時に話し(ターンテイキング遅延ゼロ)、割り込みに対応し、複雑なクエリはバックグラウンドで GPT-5.5 に委譲。**GPT-Live-1** は有料ユーザー(Go/Plus/Pro)、**GPT-Live-1 mini** は無料ユーザーのデフォルト。リアルタイムのライブ翻訳を含む。iOS / Android / Web で利用可。
136136
- [GPT-6 Astra / Astra Pro](https://openai.com/index/gpt-6-astra/) - 🆕 **2026年9月3日**。高度な推論、コーディング、コンピューター操作向け;[一部組織への限定展開で、まだ一般提供ではない](https://help.openai.com/en/articles/6825453-chatgpt-release-notes)。API モデルカードの公開とアカウントの利用権限は別途確認が必要。
137-
- [GPT-5.6 Sol](https://openai.com/blog/gpt-5-6) - 🆕 **2026-07-09**(GA;一部プレビュー 6月26日~)。GPT-5.6 ファミリーのフロンティアフラッグシップ — 高度な推論、コーディング、生物学、サイバーセキュリティ機能に加え、「最大(max)」推論と「ウルトラ(ultra)」サブエージェントモードを備える最高性能モデル。ChatGPT、Codex、OpenAI API で利用可能。米国政府による安全審査のため発表が遅れたが、段階的に展開中。**2026年8月6日更新**: ChatGPT での精度と一貫性が向上;GPT-5.6 Luna が無料ユーザーの日常チャット無制限利用に拡大。**2026年8月13日**: 新しい [Ultrafast サービスティア](https://openai.com/index/previewing-ultrafast)(API 限定プレビュー)が Cerebras ハードウェアにより GPT-5.6 Sol を最大 **Standard 比 14 倍速 / 約 750 出力トークン/秒**で提供。**2026年8月21日**: API 表示価格が 100 万トークンあたり **$4 / $20** に値下げ(入力 20%・出力 33% 安);キャンペーンは少なくとも **2026年11月21日**まで([changelog](https://developers.openai.com/api/docs/changelog.md))。
137+
- [GPT-5.6 Sol](https://openai.com/blog/gpt-5-6) - 推論、コーディング、ツール利用向けの GPT-5.6 系モデル。今回確認した標準 API 入力・出力料金は 100 万トークン当たり $4/$20。長文、キャッシュ、サービス階層の条件は[料金表](https://developers.openai.com/api/docs/pricing)を参照。
138138
- [GPT-5.6 Terra](https://openai.com/blog/gpt-5-6) - 🆕 **2026-07-09**。GPT-5.6 ファミリーの中間層 — GPT-5.5 と同等の性能を約 2 分の 1 のコストで提供。コスト効率の高い本番ワークロード向け。
139139
- [GPT-5.6 Luna](https://openai.com/blog/gpt-5-6) - 🆕 **2026-07-09**。GPT-5.6 の中で最も高速かつコスト効率の高いモデル — 大量で処理速度が求められるタスクに最適。
140140
- [ChatGPT Work](https://openai.com/index/chatgpt-for-your-most-ambitious-work/) - 🆕 **2026-07-09**。目標を渡すと完成した成果物に仕上げる OpenAI のエージェント — 接続されたアプリやファイルを横断して行動し、1 つのプロジェクトに数時間取り組み続け、スライド / シート / ドキュメント / Web アプリを作成し、スケジュール実行や内蔵ブラウザによるデスクトップ Computer Use も可能。GPT-5.6 駆動。Web / モバイルでは Pro・Enterprise・Edu から順次展開(Plus / Business は追って対応);デスクトップアプリは Mac / Windows で Free を含む全プランにグローバル提供。
@@ -192,7 +192,7 @@
192192
- [Claude Security](https://www.anthropic.com/) - **2026-05-01** パブリックベータ。Opus 4.7 駆動の企業向けコードベース脆弱性スキャナ —— 信頼度評価・深刻度・再現手順・推奨修正付きパッチを生成。Enterprise ユーザー向け [claude.ai/security](https://claude.ai/security)。
193193
- [Claude Finance Agents](https://www.anthropic.com/news/finance-agents) - **2026-05-05**。Opus 4.7 ベースの金融特化エージェントを 10 種同時公開(pitchbook 作成、KYC、月次決算、ディール選定など)。Claude Cowork プラグイン、Claude Code skill、Managed-Agents の cookbook として配備可能。
194194
- [Claude Finance JV](https://www.anthropic.com/) - **2026-05-04**。Goldman Sachs・Blackstone との 15 億ドル規模の Claude 導入ジョイントベンチャー。Anthropic のエンジニアを中堅ウォール街企業に常駐させる。
195-
- [Claude Add-ins / Dreaming / Outcomes / Multi-agent orchestration](https://www.anthropic.com/news/code-with-claude-2026) - **2026-05-08(Code with Claude 2026)**。Anthropic が Add-ins、セッション間の定期メモリ整理("Dreaming")、ルーブリック駆動の "Outcomes"、そして共有ファイルシステムと監査可能な trace を備えた主エージェント + サブエージェント編成モデルをまとめて発表。
195+
- [Claude Managed Agents updates](https://claude.com/blog/new-in-claude-managed-agents) - **2026-05-19**。Managed Agents が複数エージェントの連携、評価基準に基づく成果判定、研究プレビューの dreaming を説明。機能ごとに利用範囲が異なる。
196196
- [Anthropic ↔ SpaceX Colossus 1](https://www.siliconrepublic.com/business/anthropic-joins-forces-with-spacex-for-colossus-capacity) - **2026-05-06**。Anthropic が SpaceX の Memphis データセンター Colossus 1(220K+ NVIDIA H100/H200/GB200, 300+ MW)の全利用可能キャパシティを取得し Claude Opus 推論に充てる。Claude Code の 5 時間レート制限を Pro / Max / Team / Enterprise で 2 倍化、Pro / Max でピーク時限も撤廃。
197197
- [Anthropic ↔ AMD(最大 2 GW の Instinct MI450)](https://ir.amd.com/news-events/press-releases/detail/1292/amd-and-anthropic-announce-strategic-partnership-to-deploy-up-to-2-gigawatts-of-amd-instinct-mi450-series-gpus) - 🆕 **2026-07-22**。Anthropic は AMD Helios ラックスケール構成で AMD Instinct MI450 シリーズ(MI455X)GPU を**最大 2 ギガワット**展開する。EPYC "Venice" CPU、Pensando ネットワーキング、ROCm を組み合わせ、最初の 1 ギガワットは 2027 年前半に稼働開始。AMD は Anthropic へ**最大 50 億ドル**の戦略的出資を約束し、複数年のエンジニアリング協業も行う。既存の MI355X 利用を踏まえたもので、TPU・Trainium・SpaceX Colossus と並ぶ意図的なハード多様化。
198198
- [オープンウェイトモデルに関する Anthropic の立場](https://www.anthropic.com/news/position-open-weights-models) - 🆕 **2026-07-27**。米政府当局が中国製オープンウェイトモデルの禁止を検討しているとの報道に対し、Dario Amodei は「Anthropic はオープンウェイトモデルの禁止を主張したことは一度もない」と明言。危険な能力を持たないオープンウェイトは「公共財」だとし、代わりにチップ輸出管理と密輸取り締まり、産業規模の蒸留への抑止、そして**十分に高性能なすべてのモデル(オープン・クローズド問わず)へのリリース前安全性テストの義務化**を支持する。2026 年のオープン対クローズド論争を追う一次資料。
@@ -2192,7 +2192,6 @@
21922192
| **2026-07-29** | [Langfuse v4](https://github.com/langfuse/langfuse/releases/tag/v4.0.0) と [Milvus 3.0](https://github.com/milvus-io/milvus/releases/tag/v3.0.0) が同日リリース — 前者は全文検索とモニター、API は最大 165 倍高速と主張;後者はレイクネイティブな External Collections で Parquet/Lance/Iceberg を直接クエリ。同日 [RufRoot / CVE-2026-59726](https://hackread.com/rufroot-vulnerability-attackers-hijack-ruflo-login/) も公表:Ruflo の MCP ブリッジが認証なしで 233 ツールに到達可能、エージェントメモリも汚染可能 | ツール / 業界 |
21932193
| **2026-07-30** | [Inkling-Small](https://thinkingmachines.ai/inkling/) ウェイト公開 — 276B/12B アクティブ、Apache-2.0、マルチモーダル;HLE テキスト 31.6%(975B Inkling の 29.7% を上回る) | モデル |
21942194
| **2026-07-31** | [DeepSeek-V4-Flash-0731](https://huggingface.co/deepseek-ai/DeepSeek-V4-Flash-0731) —— 同一 API・料金のまま Agentic 能力を強化、V4-Pro Preview を上回る;HF にオープンウェイト公開。GitHub Copilot が Gemini 2.5 Pro・Gemini 3 Flash を非推奨化 | モデル / ツール |
2195-
| **2026-08-03** | [Claude Cowork](https://www.anthropic.com/claude-cowork) + [Claude Tag](https://www.businesswire.com/news/home/20260803/) 公開 —— Cowork は非開発者向け自律タスクエージェント(Web + Slack)、Tag は旧 Slack 統合の後継。[Embabel Agent](https://github.com/embabel/embabel-agent) が注目を集める —— Spring Framework 創設者 Rod Johnson による JVM エージェントフレームワーク(約 3.4K stars、Apache-2.0;最新タグ付きリリースは v0.5.0 プレリリース)。「💱 エージェント経済とマーケットプレイス」カテゴリ新設 | フレームワーク / ツール |
21962195
| **2026-08-03〜07** | [Cloudflare Agents Week](https://blog.cloudflare.com/agents-week-review-august-2026/) — Wallets/cloudflare.pay(8-4)、WriteGuard プライベートベータ(8-5)、WebMCP + Kitesurf サーバーレスエージェントブラウザ + MCPv2 + AI Search(8-6) | ツール / プロトコル |
21972196
| **2026-08-03** | [Qwen3.8-Max](https://alibabacloud.com/blog/qwen3-8-max) をアリババが正式ローンチ — 2.4T MoE / 95B アクティブ、1M コンテキスト、マルチモーダル入力;QwenWork エンタープライズプラットフォームが公開ベータ | モデル |
21982197
| **2026-08-05** | Meta Superintelligence Labs の [Muse Spark 1.2 + Muse Code ベータ](https://research.meta.ai/blog/introducing-muse-code-and-muse-spark-1-2) — ターミナルコーディングエージェント + リポジトリ全体で学習したモデル。[ByteDance SeedRealtime](https://technode.com/2026/08/05/bytedance-launches-seedrealtime-full-duplex-audio-video-model/) フルデュプレックス音声視覚モデルがローンチ。英国 AISI が[エージェント封じ込めインシデント](https://www.helpnetsecurity.com/2026/08/05/ai-agent-deception-in-cyber-tests/) INC-2026-07-28-01 を開示 | モデル / 業界 |

0 commit comments

Comments
 (0)