Skip to content

Repository files navigation

sdrf-skills

Turn Claude Code, Cursor, OpenAI Codex, Gemini CLI, or OpenCode into an expert proteomics SDRF annotator.

Claude Code Skill Cursor Codex Gemini CLI OpenCode License: MIT SDRF Spec Skills

Pick a dataset → the agent fetches PRIDE + paper → you review a validated SDRF.

Structured skills that give AI assistants expert-level capabilities for annotating, validating, improving, and reviewing proteomics metadata in the SDRF format. Instead of guessing at ontology terms or SDRF rules, the agent follows the methodology of experienced annotators using real tools (OLS, PRIDE, PubMed). The specification data (column definitions, templates) lives in a git submodule and is read at runtime, so the skills stay current as the spec evolves.

Available skills

Sixteen skills, most in the sdrf: namespace (the two review-gate skills use portable hyphenated names):

Skill What it does
/sdrf:setup Guided dependency install (parse_sdrf, techsdrf) — conda or pip
/sdrf:knowledge SDRF format, column rules, ontology mappings, reserved words; plain-language explanations; ontology term lookup
/sdrf:templates Template selection, layers, and selection rules
/sdrf:metascreen Shortlist PRIDE / MassIVE / ProteomeXchange studies → resumable TSV
/sdrf:autoresearch Autonomous retained-improvement loop over a dataset or dataset class
/sdrf:annotate Plan + full workflow: PXD → PRIDE + paper → draft SDRF → validate
/sdrf:validate Systematic validation against templates + OLS ontology checking
/sdrf:fix Auto-fix common errors (UNIMOD swaps, case, format, artifacts)
/sdrf:review Comprehensive quality review + 5-dimension quality score cross-referenced to paper + PRIDE
$sdrf-adversarial-review Fresh-context, evidence-first review with a hash-bound verdict
$sdrf-annotate-reviewed Annotation orchestrator with isolated review, repair, and re-review
/sdrf:convert Choose and configure analysis pipelines from SDRF
/sdrf:design Detect batch effects, confounders, replication issues
/sdrf:contribute Contribute an annotated SDRF back to sdrf-annotated-datasets via PR
/sdrf:techrefine Verify/refine technical metadata from raw files via techsdrf
/sdrf:cellline Translate Cellosaurus records into SDRF cell-line columns

Installation

# 1. Clone WITH submodules (the spec data is a submodule):
git clone --recurse-submodules https://github.com/bigbio/sdrf-skills
# already cloned without them?  git submodule update --init --recursive

# 2. Install the deterministic helper tools (conda recommended — includes thermorawfileparser):
conda env create -f environment.yml && conda activate sdrf-skills
# pip alternative (thermorawfileparser not on PyPI):
#   pip install -r requirements.txt && pip install git+https://github.com/bigbio/techsdrf.git

Update the bundled spec any time with git submodule update --remote --recursive.

Setup by AI platform

Claude Code (plugin)
cd sdrf-skills && claude --plugin-dir .   # loads skills from the working tree

Start from the repo root (skills reference spec/ by repo-root-relative path). Then run /sdrf:setup, and use /sdrf:annotate PXD###### or /sdrf:validate your_file.sdrf.tsv. Marketplace install is not available yet — see #27.

Cursor

Ensure .cursor/rules/sdrf-skills.mdc is in your project; then ask "Follow the sdrf setup workflow" (Cursor does not run Claude Code's SessionStart hook).

Codex / Gemini CLI / OpenCode
  • Codex — follow .codex/INSTALL.md to symlink skills/ and spec/ into your Codex skills path.
  • Gemini CLI — auto-loads GEMINI.md from the repo root.
  • OpenCode — follow .opencode/AGENTS.md to wire the skills in.

For full annotation, configure the OLS, PRIDE, PubMed, and bioRxiv MCP servers, and validate with parse_sdrf validate-sdrf.

Usage

/sdrf:annotate PXD045678     → fetch PRIDE + paper → select templates → draft SDRF with OLS-verified terms → validate
/sdrf:validate file.sdrf.tsv → template + ontology validation
/sdrf:fix file.sdrf.tsv      → repair UNIMOD swaps, case, formats, artifacts (with changelog)
/sdrf:contribute PXD045678   → open a PR to bigbio/sdrf-annotated-datasets

Python tools

The repo is skills-first: new user-facing workflows go in skills/. tools/ holds deterministic helpers a skill can call (TSV parsing, OLS client, hallucination detection, quality scoring, auto-fix, cell-line enrichment, MassIVE fallback, and the review gate). Run them via the unified CLI:

python -m tools check  file.sdrf.tsv          # hallucinated terms / UNIMOD swaps
python -m tools score  file.sdrf.tsv          # quality score (0-100, 5 dimensions)
python -m tools fix    file.sdrf.tsv -o out.tsv
python -m tools verify UNIMOD:1 --label Acetyl
python -m tools review-gate gate              # enforce independent-review receipts

Adversarial review gate. Changed SDRFs are identified by SHA-256; a passing receipt is valid only for that exact content, and any edit makes it pending again. Changed artifacts are discovered from git (measured against the merge base), so the gate works from Claude Code, another assistant, CI, or a plain shell: python3 tools/review_gate.py gate --cwd <repo-root> (exit 1 = review still needed). On Claude Code a Stop hook runs the same check; the hook is a convenience, not the enforcement.

How it works

The skills/ directory is platform-agnostic markdown; each platform needs only a thin shim (.claude-plugin/, .cursor/rules/, .codex/, GEMINI.md, .opencode/) to discover and load it. The MCP tools an agent needs already exist (OLS, PRIDE, PubMed) — what was missing was the expertise: which ontology to search per column, how to read a paper for SDRF metadata, the common errors and their fixes, and what "good" annotation looks like. Skills encode that as step-by-step workflows, and the spec/ submodule keeps the column/template data current with no SKILL.md changes.

Contributing

Add a skill by creating skills/your-skill/SKILL.md with YAML frontmatter, writing the workflow in markdown, and referencing spec/ for any specification data (never hardcode). For questions about the SDRF specification itself, open an issue in bigbio/proteomics-metadata-standard.

Contributors

Maintained by the BigBio team.

License

MIT

About

Agentic plugins to annotate SDRF using Claude; Codex; Cursor; Gemini

Topics

Resources

Stars

15 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages