This repo provides drop-in AI tooling for teams to install on their own repos. Your team can leverage tooling in two ways depending on your technical constraints:
- In CI automatically on every PR (outer loop) Findings post as PR comments and developer 👍/👎 feedback is gathered by metricsai. How it's wired in depends on the workflow and your CI — a reusable workflow, a vendored workflow/scripts, or Copilot auto-review (see each section).
- Locally, on demand before pushing (inner loop) A developer runs the scripts for a fast read — a terminal report, or a posted PR comment. No CI required.
As a rule of thumb, prefer the outer loop option for consistency and ease of use. Reach for the inner loop if technical constraints (e.g. no GitHub Actions, no Jenkins or Copilot, or inference must stay within your AWS boundary) make it easier, or you have a real process reason to shift the checks left in the software dev lifecycle.
| Path | What it is | Consumed by |
|---|---|---|
testing/classifier/ |
AI test-failure classifier — labels each failing test app-bug / test-bug / flaky / env | consumer repo CI and/or local |
security/ |
AI security / PR-review bundle that flags secrets, PII/PHI, OWASP & IaC-compliance issues by severity | consumer repo CI and/or local |
metricsai/ |
Python CLI that gathers metrics from the two workflows | run on demand |
AI assistant compliance and configuration Under the hood the testing and security bundles work similarly, configured by two independent choices:
- Which AI assistant runs — set
AI_REVIEW_TOOLtoclaude,codex, orcopilot(per developer locally, or per repo in CI). Make your choice based on project tooling constraints. - Where inference runs — the vendor's API by default, or Amazon Bedrock in
your own AWS account when code can't (or shouldn't) leave your AWS boundary (
claude/codexonly, notcopilot). Make your choice based on project compliance guidelines.
Whichever you pick, the AI ends its output with one parseable
<<<AI_REVIEW_RESULT:…>>> marker that the dispatcher reads to decide the outcome.
Problem solved: we prevent AI tooling from "fixing" a failing test when the
code is what actually broke. When a test fails, the classifier answers is
the test wrong, or is the code wrong? For each test failure it issues one verdict
(APPLICATION_BUG, TEST_BUG, FLAKY_FAILURE, or ENVIRONMENT_ISSUE), along with a
short rationale, and then posts a PR comment. Developers are asked to 👍/👎 the comment to evaluate classifier quality. As of now, this workflow is diagnostic only with the hope of leveraging the classifier to move towards auto-fixing tests in the future without sacrificing quality.
Canonical classification logic lives in the skill:
testing/classifier/.skills/test-classifier/SKILL.md.
INNER LOOP (local) OUTER LOOP (CI, on every PR)
test-classifier --pr N --submit Actions / Jenkins
│ │ runs suite (OBSERVED)
│ posts ▼
├───────────────────────────► ONE PR comment + 👍/👎 ──(devs react)
│ │ (either loop posts it)
│ + writes a row ▼
▼ metricsai weekly run
"Testing Events" tab │
└─────────► Google Sheet ◄────────────┘
Two writers reach the sheet: CI 👍/👎 reactions (harvested weekly by metricsai)
and local --submit rows (written immediately to the Testing Events tab). A
local --submit (or --post-comment) run also posts the same PR comment — so
the inner loop surfaces the verdict on the PR and records a metric row.
- Outer loop (CI, on every PR) — the designed path. It triggers on every PR, runs the suite itself, and comments only when something fails. GitHub Actions consumers reference a reusable workflow (no files copied); Jenkins consumers vendor the bundle and add a pipeline stage. Feedback is the 👍/👎 on the comment, harvested into the team's weekly metrics row.
- Inner loop (local) — a
test-classifiershell function runs on your unpushed changes (no PR needed) or against an existing PR. For local, we recommend running in OBSERVED mode with commandAI_RUN_SUITE=1 test-classifier --pr <PR> --submit. That posts the one PR comment, prompts "helpful? y/n" in the terminal, and writes to the metrics spreadsheet (Alternatively,--post-commentposts the comment only).
In summary, run the outer loop when your CI can host it. Otherwise, the inner loop is the fallback when CI isn't available.
Problem solved: stop secrets, PII/PHI, and security or compliance defects from
landing — the security counterpart to the classifier, answering is this change
safe? It reviews a change (or the whole repo) and reports findings by severity
(Critical / High / Medium / Low) with remediation, as a PASS / WARN / BLOCK
gate locally and as inline PR comments labeled security(<severity>): or
compliance(<severity>): with one-click fix suggestions. Like the classifier it
is advisory and never commits; a self-adjudication pass trims false positives.
It ships three layers:
code-security— secrets, PII, PHI, and OWASP-style defects in a diff.iac-compliance— infrastructure-as-code against CMS ARS 5.1 + NIST SP 800-53 Rev 5.pr-review/codebase-audit— the full PR diff (local, Actions, or Copilot), or a full-repo audit of the codebase at rest.
Canonical review logic lives in the skills, e.g.
security/review/.skills/pr-review/SKILL.md.
INNER LOOP (local, pre-commit) OUTER LOOP (CI, on every PR)
code-security / iac-compliance Actions (vendored) / Copilot (licensed)
│ │
▼ ▼
PASS / WARN / BLOCK PR comments + 👍/👎 ──(devs react)
(terminal, no sheet) │
▼
metricsai weekly run
│
▼
Google Sheet
Unlike testing, the local loop is a gate, not a sheet writer — only the CI
PR comments (👍/👎) and the AWS Security Hub count feed metricsai.
- Outer loop (CI, on every PR) — the recorded path, with two shipped
options: a vendored GitHub Actions workflow (
ai-pr-review.yml, copied into the repo — security has no point-at-it reusable workflow), or GitHub Copilot auto-review (needs a Copilot Enterprise / Business + Code-Review license, and you copy its instruction files —copilot-instructions.mdplus.github/instructions/*.md— to align it with the bundle). Either posts inlinesecurity(...)/compliance(...)comments; the 👍/👎 plus an AWS Security Hub count feed the weekly metrics row. There is no Jenkins integration for security (testing has one) — the dispatcher is CI-agnostic so a team could wire it into Jenkins by hand, but that path isn't shipped. - Inner loop (local, pre-commit/pre-push) — the
code-security/iac-complianceshell functions (and a localpr-review) scan unpushed changes on demand and return a PASS / WARN / BLOCK result in the terminal. Report-only: it catches secrets/PII before they leave the laptop, but doesn't feed the sheet. A full-repocodebase-auditalso lives here (primarily a local tool); it can run in CI, but needs an API key and can incur high token cost.
A team runs the outer loop when its CI can host it; the inner loop is the local backstop that keeps secrets and PII out before code is ever pushed.
Testing classifier
- Setup guide — install + all run paths (CI, Jenkins, local), for humans
- Canonical skill — the classification logic the AI follows
Security
- Setup guide — install + operations (pre-commit, PR review, audit), for humans
- Canonical skills —
code-security·iac-compliance·pr-review
Metrics
- metricsai — how 👍/👎 + findings are harvested into a Google Sheet