Agentic evaluation plugin for the Fusion Framework CLI. It uses the GitHub Copilot SDK to create a conversational session in which an LLM autonomously drives agent-browser to navigate, interact with, and judge a running Fusion application against eval criteria written in plain markdown.
- You have a Fusion application (cookbook or production app) and want to verify it works end-to-end.
- You want an LLM to autonomously explore and validate acceptance criteria without writing Playwright scripts.
- You want structured, evidence-backed pass/fail verdicts with screenshots, snapshots, and computed style checks.
The plugin follows a single-session agentic workflow: the Copilot agent reads your eval markdown, decides what to check, drives the browser through tool calls, collects evidence, and produces a structured JSON verdict in one session.
flowchart LR
A[CLI] -->|starts| B[Dev Server]
A -->|reads| C[eval/*.md]
A -->|creates| D[Copilot Session]
D -->|system prompt +\neval markdown| E[LLM]
E <-->|tool calls| F[agent-browser]
F <-->|headless Chrome| B
E -->|JSON verdict| A
sequenceDiagram
participant CLI as ffc copilot app eval
participant Server as Dev Server
participant SDK as Copilot SDK
participant LLM as LLM Agent
participant AB as agent-browser
participant Browser as Headless Chrome
CLI->>CLI: Resolve eval files from app/eval/*.md
CLI->>Server: Start dev server (ffc app serve)
CLI->>CLI: Wait for server ready
CLI->>AB: Reset daemon (kill stale sessions)
loop For each eval file
CLI->>CLI: Create timestamped run directory
CLI->>SDK: Create session with browser tools
SDK->>LLM: System prompt + eval markdown
loop Agent autonomously evaluates
LLM->>AB: browser_navigate(url)
AB->>Browser: Open page
LLM->>AB: browser_wait(networkidle)
LLM->>AB: browser_snapshot()
AB-->>LLM: Accessibility tree
LLM->>AB: browser_get_styles(selector)
AB-->>LLM: Computed CSS values
LLM->>AB: browser_screenshot()
AB-->>LLM: Image data
LLM->>AB: browser_errors()
AB-->>LLM: Console errors
end
LLM-->>CLI: JSON verdict {pass, reasoning, steps}
CLI->>CLI: Save verdict + evidence artifacts
end
CLI->>AB: Reset daemon
CLI->>Server: Stop dev server
CLI->>CLI: Exit 0 (all pass) or 1 (any fail)
# Evaluate all eval files in a cookbook
ffc copilot app eval ./cookbooks/app-react
# Or use the workspace-root shortcut
pnpm eval:app ./cookbooks/app-react
# Run a specific eval
ffc copilot app eval ./cookbooks/app-react --eval smoke
# Use a specific LLM model
ffc copilot app eval ./cookbooks/app-react --model claude-sonnet-4
# Point at an already-running server
ffc copilot app eval . --url http://localhost:3000/apps/my-appIf the application requires authentication, run the login flow once to persist MSAL tokens in the Chrome profile:
ffc copilot app eval ./cookbooks/app-react --login
# Alias
ffc copilot app eval ./cookbooks/app-react --logonThis opens a headed browser for interactive login. After you authenticate, press Ctrl+C. Subsequent eval runs reuse the saved session automatically.
Eval files are plain markdown files placed in the app's eval/ directory. The plugin passes the raw markdown directly to the LLM as the user prompt. No special parsing is required.
my-app/
├── eval/
│ ├── smoke.md
│ └── accessibility.md
├── src/
└── package.json
Write eval files as you would a user story or test plan. The LLM reads the markdown and decides how to verify each criterion.
---
name: smoke-test
---
## User Story
As a user, I need to see the application load successfully so I know
the deployment is healthy.
## Acceptance Criteria
- must see a heading with the application name
- must not have any JavaScript errors in the console
- should load within 5 seconds
- should display navigation elementsBoth story-driven formats (user story + acceptance criteria) and instruction-driven formats (numbered steps + assertions) work. The LLM adapts to the structure you provide.
ffc copilot app eval <path> [options]
| Option | Description | Default |
|---|---|---|
<path> |
Path to the Fusion application directory | required |
--eval <name-or-path> |
Run a specific eval by name or file path | All files in eval/ |
--port <port> |
Port for the app dev server | 3000 |
--host <host> |
Host for the app dev server | 0.0.0.0 |
--url <url> |
Skip server start and use an already-running URL | none |
--verbose |
Show agent-browser commands and app server output |
false |
--login, --logon |
Open a headed browser for interactive MSAL login | false |
-m, --model <model> |
LLM model to use, for example claude-sonnet-4 |
SDK default |
-o, --output <dir> |
Output directory for run artifacts | .tmp/copilot/ |
Each eval run creates a timestamped directory containing evidence and metadata:
.tmp/copilot/smoke-test_143052/
├── copilot-log.jsonl # Tool call log (timestamp, tool name, arguments)
├── copilot-response.md # Full LLM response text
├── verdict.json # Structured pass/fail verdict
└── evidence/
├── snapshot.txt # Accessibility tree snapshot
├── errors.txt # JavaScript console errors
├── url.txt # Final page URL
├── screenshot-*.jpg # Visual evidence screenshots
├── styles-*.txt # Computed CSS style captures
└── eval-*.json # JavaScript evaluation results
The LLM can use 18 browser tools during an eval session:
| Tool | Description |
|---|---|
browser_navigate |
Open a URL in the browser |
browser_snapshot |
Capture an accessibility tree with element refs (@e1, @e2) |
browser_screenshot |
Take a screenshot and return the image to the model |
browser_get_styles |
Get computed CSS styles of an element |
browser_eval |
Evaluate a JavaScript expression in the page context |
browser_click |
Click an element by ref, CSS selector, or text |
browser_fill |
Fill a form field (clears existing content first) |
browser_type |
Type text character by character |
browser_press_key |
Press a keyboard key (Enter, Tab, Escape, etc.) |
browser_hover |
Hover over an element |
browser_select |
Select an option from a dropdown |
browser_scroll |
Scroll the page or scroll an element into view |
browser_wait |
Wait for load state, text, element, or timeout |
browser_find |
Semantic element lookup by role, text, or label |
browser_errors |
Get JavaScript console errors |
browser_get_url |
Get the current page URL |
browser_go_back |
Navigate back in browser history |
browser_reload |
Reload the current page |
graph TD
subgraph "@equinor/fusion-framework-cli-plugin-copilot"
IDX[index.ts — plugin entry point]
CMD[commands/app/command.ts — CLI command tree]
EVAL[commands/app/eval.ts — Copilot session runner]
LOGIN[commands/app/login.ts — MSAL login flow]
RESOLVE[eval-resolve.ts — eval file discovery]
TOOLS[commands/app/tools/ — 18 browser tool wrappers]
UTILS[utils/ — agent-browser, daemon, server helpers]
end
IDX -->|registers| CMD
CMD -->|--login| LOGIN
CMD -->|resolves| RESOLVE
CMD -->|runs| EVAL
EVAL -->|creates| TOOLS
EVAL -->|uses| UTILS
LOGIN -->|uses| UTILS
TOOLS -->|calls| UTILS
- Node.js >= 20
- pnpm as the workspace package manager
- agent-browser installed globally or available on
PATH - GitHub Copilot access in VS Code, because the SDK authenticates through the GitHub Copilot extension
Important
This package is deliberately excluded from the repo's pnpm workspace (see
CODEMAP.md). Its @github/copilot-sdk dependency pulls in
~400MB of platform-specific native CLI binaries via optionalDependencies, which
would otherwise bloat the shared root pnpm-lock.yaml and CI pnpm-store cache for
every package in the monorepo.
It has its own install and build, run from its own directory:
cd packages/cli-plugins/copilot
pnpm install
pnpm buildpnpm install here resolves against the repo's shared local pnpm content-addressable
store, so there's no extra disk cost if you've already installed the root workspace —
but it produces its own node_modules and is not tracked in the root lockfile. The
vscode-jsonrpc patch and the agent-browser/koffi build-script exclusions this
package needs are declared locally in its own package.json under "pnpm", not in the
root pnpm-workspace.yaml.
Note
fusion-cli.config.ts loads this plugin optionally (via Promise.allSettled), so
the rest of the repo's ffc CLI works whether or not this package has been
installed/built.
Important
This package is outside the repo's pnpm workspace, so it is not versioned or
published through Changesets or the root ci.yml release-pkg job. There is no
automated release workflow — publish it manually from its own directory.
-
Add a new entry at the top of
CHANGELOG.md, above the previous version, following the existing Changesets-style format:## <new-version> ### Patch Changes - Describe the change here.
-
Bump the version and publish:
cd packages/cli-plugins/copilot pnpm version <patch|minor|major> pnpm install # refresh its own lockfile-less node_modules after the bump pnpm build pnpm publish --access public
-
Commit
package.json,CHANGELOG.md, and any source changes together, then push the commit and tag yourself — nothing does this automatically.