An MCP server that keeps a live catalog of the RDF-Connect ecosystem
(processors, runners, orchestrators and example pipelines), harvested periodically from GitHub the same
way rdf-connect.github.io does (github.data.ts:
searching repos tagged rdfc-processor / rdfc-runner / rdfc-orchestrator / rdfc-pipeline /
rdf-connect, then parsing their .ttl files for SHACL-described components), and exposes it to coding
agents as MCP tools so they can discover and wire up existing processors instead of guessing.
- Runs standalone (no VitePress), on a schedule, with the result cached to disk so a restart doesn't need to re-harvest before answering.
- Extracts each processor/runner/orchestrator's SHACL parameter shape (required vs. optional properties, datatypes, classes) into a structured form instead of just linking to the repo.
- Best-effort resolves the installable package name (npm
package.json/ PyPIpyproject.toml) so generatedowl:importspaths are closer to correct. - Adds a
generate_pipeline_skeletontool that drafts a wired-uppipeline.ttl(imports, channels,consistsOf, processor stubs) from a list of chosen processors.
npm install
npm run buildOptionally set a GitHub token to avoid low unauthenticated rate limits (the harvester makes dozens of GitHub API calls per run):
export GITHUB_TOKEN=ghp_...The server supports two MCP transports, and picks one at startup based on whether MCP_HTTP_PORT is set:
- stdio (default) - one process per client, talks JSON-RPC over stdin/stdout. What
npm startand most local MCP client configs (Claude Code, Claude Desktop) use. - HTTP (set
MCP_HTTP_PORT) - the Streamable HTTP transport, with stateful sessions (oneMcpServerper session, tracked by themcp-session-idheader). This is what you want to deploy remotely - e.g. on a VPC, behind nginx or another reverse proxy that terminates TLS. Besides/mcp, the HTTP server also serves a plain-HTML status page at/(live catalog stats plus the client config snippet to copy) and a JSON health check at/healthz.
npm start{
"mcpServers": {
"rdfc": {
"command": "node",
"args": ["/absolute/path/to/rdfc-mcp/server/dist/index.js"],
"env": { "GITHUB_TOKEN": "ghp_..." }
}
}
}MCP_HTTP_PORT=3001 npm startThen open http://localhost:3001/ for the status page, or point an MCP client that supports remote/HTTP
servers at the URL:
{
"mcpServers": {
"rdfc": {
"url": "https://rdfc-mcp.internal.example.com/mcp"
}
}
}The HTTP server itself has no authentication or TLS - it's meant to sit behind a reverse proxy (nginx,
etc.) on the same host or a private network that handles both. Don't expose MCP_HTTP_PORT directly to
the public internet.
docker build -t rdfc-mcp-server .The image defaults to the HTTP transport on port 3001 (ENV MCP_HTTP_PORT=3001), meant to be reverse-proxied
by nginx (or similar) running alongside it:
docker run -d --restart unless-stopped \
-p 127.0.0.1:3001:3001 \
-v rdfc-mcp-cache:/data \
-e GITHUB_TOKEN=ghp_... \
rdfc-mcp-serverBinding to 127.0.0.1 keeps the unauthenticated HTTP port off the network entirely; point nginx's
proxy_pass at http://127.0.0.1:3001/mcp and let it handle TLS/public exposure. Or via the included
docker-compose.yml (reads GITHUB_TOKEN/etc. from your shell/.env), which does the same thing:
docker compose up -dA named volume at /data keeps the harvested catalog across container restarts/redeploys; without it,
every fresh container starts from an empty cache and serves nothing useful until the first background
harvest finishes. Prewarm it before first use with:
docker run --rm -v rdfc-mcp-cache:/data --entrypoint node rdfc-mcp-server dist/harvest-cli.jsIf you'd rather run the container as a plain stdio server for a single local client instead, unset
MCP_HTTP_PORT and run with -i (not -t, which would corrupt the JSON-RPC stream over a pty):
docker run -i --rm -v rdfc-mcp-cache:/data -e MCP_HTTP_PORT= -e GITHUB_TOKEN=ghp_... rdfc-mcp-server| Env var | Default | Purpose |
|---|---|---|
MCP_HTTP_PORT |
unset (3001 in the Docker image) |
Set to serve HTTP (/mcp, /, /healthz) on this port instead of stdio. |
GITHUB_TOKEN |
none | Auth token for GitHub API calls (raises rate limits). |
RDFC_MCP_REFRESH_INTERVAL_MS |
21600000 (6h) |
How often to re-harvest GitHub in the background. |
RDFC_MCP_CACHE_DIR |
~/.cache/rdfc-mcp (/data in the Docker image) |
Where catalog.json is cached between restarts. |
On startup the server immediately serves whatever is in the disk cache (if any) while a fresh harvest
runs in the background if the cache is missing or older than the refresh interval; it then re-harvests on
that interval for as long as the process stays alive. There is deliberately no agent-facing tool to force a
refresh (it's slow and burns GitHub rate limit) - as an operator, prewarm or force-refresh the cache with
npm run harvest, a one-off CLI that runs the same harvester.
Resources
rdfc://guide- short primer on RDF-Connect concepts and the recommended tool workflow.rdfc://authoring-guide- JS/TS only. For when nothing in the catalog fits and you have to write a new processor targetingrdfc:NodeRunner. The tools/catalog only index.ttlmetadata; this covers the runtime APIs that metadata can't: theProcessor<T>lifecycle (@rdfc/js-runner), how aprocessor.ttlSHACL shape becomes your constructor args (includingrdfl:PathLens, the mechanism behind "extract a value at a configured SHACL path" parameters), the SDS message convention, and a testing/scaffold checklist. RDF-Connect processors aren't JS/TS-only though - forrdfc:PyRunner/rdfc:JvmRunnertargets there's no dedicated guide yet; useget_componentonTemplateProcessorPy/TestProcessorinstead.rdfc://catalog- the full harvested catalog as JSON.rdfc://shapes- the aggregated SHACL shapes graph (Turtle) forrdfc:Pipelineplus every discovered processor.
Tools
list_components- list/search processors, runners, orchestrators, pipeline examples, and other tagged repos (filter bykind,language,query).get_component- full detail on one component: parameters, source repo, package name, raw repo turtle.get_pipeline_examples- realpipeline.ttlfiles found in the ecosystem, for reference.get_shapes_graph- the aggregated SHACL shapes graph.generate_pipeline_skeleton- draft apipeline.ttlwiring the given processors together.catalog_status- freshness/counts/last error, for checking whether it's safe to trust the data yet (read-only - seenpm run harvestabove to force a refresh).
Prompts
build_pipeline- seeds a conversation with the primer plus a goal, pointing the agent at the tool workflow above.write_processor- seeds a conversation with a goal and alanguage(jsdefault, orpython/java). Forjsit includes the full authoring guide; forpython/javait instead points at that runner's template processor repo, since no dedicated guide exists for those yet.
HTTP-only routes (when MCP_HTTP_PORT is set - not part of the MCP protocol itself)
GET /- plain HTML status page: live catalog stats, the MCP client config snippet, and a table of available tools.GET /healthz- JSON liveness/status check for load balancers and container orchestration.
- Harvesting hits the public GitHub Search + REST APIs; without
GITHUB_TOKENyou'll likely hit the 60 req/hour unauthenticated limit after one or two harvests. generate_pipeline_skeletonis a best-effort starting point (straight-line chaining by declared reader/writer parameters, TODO placeholders for other required params) - always review its output.- Component parameter paths are assumed to live under the
rdfc:namespace when generating skeleton triples; a processor using a custom vocabulary property will need manual correction. - The HTTP transport uses stateful sessions kept in memory (
Map<sessionId, transport>) - if you run multiple replicas behind a load balancer, it needs to be session-affine (sticky) onmcp-session-id, or every replica must share the same in-memory state (they currently don't). - Request bodies over 4MB on the HTTP transport are rejected (
413) to bound memory use per request.