Skip to content

Add whatbroke to AI Evaluation & Testing collection - #3102

Open
arthi-arumugam-git wants to merge 1 commit into
pingcap:mainfrom
arthi-arumugam-git:add-whatbroke-ai-eval
Open

Add whatbroke to AI Evaluation & Testing collection#3102
arthi-arumugam-git wants to merge 1 commit into
pingcap:mainfrom
arthi-arumugam-git:add-whatbroke-ai-eval

Conversation

@arthi-arumugam-git

Copy link
Copy Markdown

Adding whatbroke to the AI Evaluation & Testing collection.

It's a CLI that diffs an LLM agent's recorded behavior (tool calls, arguments, latency) between two runs — e.g. before/after a model swap or prompt change — and reports what changed at the tool-call level, with multi-sample rates to separate real regressions from baseline flakiness. MIT, on npm as whatbroke-cli, fully offline.

It sits in the same lane as promptfoo/deepeval but answers "what changed between these two versions" rather than scoring against a rubric.

Thanks for maintaining these collections!

@vercel

vercel Bot commented Jul 24, 2026

Copy link
Copy Markdown

@arthi-arumugam-git is attempting to deploy a commit to the pingcap Team on Vercel.

A member of the Team first needs to authorize it.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant