Turn a GitHub URL into a SHACL-validated Open Pulse Ontology graph β repositories, the people who built them, the organizations behind them, and the papers they cite.
π In production at
- imagingplaza.epfl.ch β discovery portal for EPFL imaging software.
- openpulse.science β broader EPFL/Swiss open-science software graph.
ββββββββββββββββββββββββββββββββ
github.com/X β β /v2/extract β β JSON-LD graph
β classify β gather context β
β β root + fan-out agents β - schema:SoftwareSourceCode
β β reconcile + resolve ROR β - schema:Person
β β SHACL validate β - org:Organization
ββββββββββββββββββββββββββββββββ - org:Membership
- pulse:Contribution
- schema:ScholarlyArticle
A single GitHub URL β a typed graph you can SPARQL. The pipeline combines deterministic provider lookups (GitHub REST, ORCID, ROR, Infoscience, GIMIE), nine RAG indices, and optional LLM agents.
# 1. Setup
git clone https://github.com/Imaging-Plaza/git-metadata-extractor.git
cd git-metadata-extractor
just install-dev
cp .env.example .env # fill in API_TOKEN, GME_GITHUB_TOKEN + one LLM credential
# 2. Run
just serve-dev # starts on http://localhost:1234
# 3. Extract
curl "http://localhost:1234/v2/extract/github.com/Imaging-Plaza/git-metadata-extractor?output_format=jsonld"You'll get a JSON-LD @graph like:
{
"@context": "https://open-pulse.epfl.ch/ontology/v2.1.2.jsonld",
"@graph": [
{
"@id": "urn:pulse:Imaging-Plaza/git-metadata-extractor",
"@type": "schema:SoftwareSourceCode",
"schema:name": "git-metadata-extractor",
"schema:license": "https://spdx.org/licenses/MIT.html",
"schema:author": [
{ "@id": "urn:pulse:caviri" }
],
"pulse:githubRepositoryHandle": "Imaging-Plaza/git-metadata-extractor"
},
{
"@id": "urn:pulse:caviri",
"@type": "schema:Person",
"schema:name": "Carlos Vivar Rios",
"org:hasMembership": [
{ "@id": "urn:pulse:caviri__https://ror.org/02hdt9m26" }
]
},
{
"@id": "urn:pulse:caviri__https://ror.org/02hdt9m26",
"@type": "org:Membership",
"org:organization": { "@id": "https://ror.org/02hdt9m26" }
},
{
"@id": "https://ror.org/02hdt9m26",
"@type": "org:Organization",
"schema:name": "Swiss Data Science Center"
}
]
}That person β membership β organization chain was inferred by the ROR resolver stages β the pipeline reads _company: "@SwissDataScienceCenter" from GitHub, hits the ROR index, and materializes a proper org:Membership triple. See docs/v2-pipeline.md for the full stage walkthrough.
# Async extraction (returns a job id)
curl -X POST http://localhost:1234/v2/extract \
-H "Content-Type: application/json" \
-d '{"url": "https://github.com/epfl-llm/meditron-7b"}'
# {"job_id": "0193ab12-...", "status": "queued"}
curl http://localhost:1234/v2/jobs/0193ab12-...
# Extract a person profile (fans out to their owned repos)
curl "http://localhost:1234/v2/extract/github.com/caviri"
# Extract an organization
curl "http://localhost:1234/v2/extract/github.com/Imaging-Plaza"
# Surface the internal pipeline fields too (gme-internal: + publiccode: namespaces)
curl "http://localhost:1234/v2/extract/github.com/X?include_internal_fields=true"
# Switch runtime per request
curl "http://localhost:1234/v2/extract/github.com/X?agent_runtime=rule_based"Swagger UI: http://localhost:1234/docs
| Doc | What's in it |
|---|---|
| docs/v2-pipeline.md | Every stage, load-bearing assumptions, affiliation strategy, env flags. Start here. |
| docs/architecture/overview.md | Layers, request lifecycle, the three runtimes, hallucination guards, caches |
| docs/cross-repo-contract.md | The index-layer split: who writes, who reads, version pinning |
| docs/getting-started.md | Install + first run, the long version |
| docs/v2-api-reference.md | /v2/extract, /v2/jobs, /v2/graph endpoints |
| docs/rag-indices.md | Nine RAG indices + federated layer |
| docs/v2-rag-tools.md | Agent-side RAG tools wired into the pipeline |
| docs/migration-v1-to-v2.md | /v1 β /v2 endpoint mapping |
| .env.example | Every env var with defaults and notes |
Versioned doc site: https://imaging-plaza.github.io/git-metadata-extractor/
git_metadata_extractor/
app.py # FastAPI app
api/ # /v2 routes (extract, jobs, auto-ingest, system)
pipeline/stages/ # 25 sequential pipeline stages
agents/llm/ # LLM-backed entity agents + RAG tools
agents/rule_based/ # deterministic counterparts
providers/ # GitHub, gimie, ROR, ORCID, Infoscience clients + RAG readers
schema/ # JSON Schema + JSON-LD context + Pydantic models
validation/ # strict-schema + SHACL validators
experimental/ # pi terminal-agent PoC (not production)
# RAG indices moved to https://github.com/sdsc-ordes/open-pulse-sources β
# the read-side providers import it as the `open_pulse_sources` library;
# ingest/embed and the /v2/indices management API live in that repo/service.
tests/v2/ # default test target
docs/ # MkDocs site source
The index layer is consumed twice β as an imported library (read side) and
as the gme-sources service image (write side) β and both share the same
DuckDB stores and Qdrant collections. They must be the same release:
| git-metadata-extractor | open-pulse-sources library | gme-sources image |
Notes |
|---|---|---|---|
3.0.0 (this release) |
v0.1.2 |
ghcr.io/sdsc-ordes/open-pulse-sources:0.1.2 |
first split release |
< 3.0.0 |
β | β | monolith; index layer was in-tree |
v0.1.1 and earlier are not supported by 3.0.0: they return a raw 500
instead of a 503 when a credential is missing, which the extract-side error
handling reads as an unexpected failure rather than a degraded index.
The library version lives in exactly one place β the
open-pulse-sources @ git+β¦@<tag> entry in pyproject.toml β and every
install path (just install-dev, CI, the Docker image) inherits it from
there. tests/v2/test_open_pulse_sources_pin.py fails the build if the
gme-sources image tag drifts from it, or if either default slips back to a
mutable ref (main / latest).
Bumping the child version therefore means: edit the pin in pyproject.toml,
match the image tag in tools/deploy/docker-compose.yml, add a matrix row
here, and run the test. For local cross-repo work, just install-dev
re-installs a checkout found at ./open-pulse-sources or
../open-pulse-sources as editable, overriding the pin.
Everything is in .env β copy .env.example and fill in what you need. Required minimum:
API_TOKENβ bearer token guarding/v2/extractand/v2/jobs/{id}. Fails closed (unset β503on every protected route). Generate withpython -c "import secrets; print(secrets.token_urlsafe(32))".GME_GITHUB_TOKENβ required for any real GitHub call.GIMIE_API_URLβ points at thegme-gimie-apisidecar (e.g.http://gme-gimie-api:15400); required for repository extraction (thegimiepackage was removed from the image).- One LLM credential β
RCP_TOKEN(EPFL),OPENAI_API_KEY, orOPENROUTER_API_KEY.
Optional (only when you use the feature): INFOSCIENCE_TOKEN, SELENIUM_REMOTE_URL, HF_TOKEN, ZENODO_TOKEN, OPENALEX_MAILTO, EPFL_GRAPH_USERNAME / EPFL_GRAPH_PASSWORD.
All ~40 env vars (with defaults + per-feature explanations) are in .env.example.
just test # fast loop via testmon (recompiles only what changed)
just test-full # full deterministic run
just lint # ruff
just type-check # mypy
just ci # lint + type-check + coverageIndex-layer test suites live in the
open-pulse-sources repo
(just test there).
docker build -t git-metadata-extractor -f tools/image/Dockerfile .
docker run -it --rm --env-file .env -p 1234:1234 \
-v ./data:/app/data --name gme --network dev \
git-metadata-extractorSelenium (for the link-veracity stage) and Qdrant (for the RAG indices) wire up through the devcontainer compose file. See docs/getting-started.md for the full setup.
- Quentin Chappuis β EPFL Center for Imaging
- Robin Franken β SDSC
- Carlos Vivar Rios β SDSC / EPFL Center for Imaging
Built at SDSC and the EPFL Center for Imaging.