This file provides guidance to Claude Code (claude.ai/code) when working with code in this repository.
irw is a Python client for the Item Response Warehouse (IRW), a harmonized repository of item response data (survey responses, test scores) hosted on Redivis. It lets researchers discover, filter, download, and reformat psychometric datasets.
Project map: ARCHITECTURE.md in ben-domingue/irw — which repo owns
what, where the data lives, and which document is authoritative when two disagree.
# Development install
pip install -e .
# Dependencies: redivis, pandas, numpy (see pyproject.toml)No build step required. Uses modern pyproject.toml with setuptools. Python 3.9–3.13 supported.
Redivis authentication is handled by the redivis client library itself (interactive browser login on first use).
All public functions are exposed through src/irw/api.py, which acts as a facade over two internal layers:
api.py (public surface — all user-facing functions)
├── operations/ (business logic: fetch, filter, list_tables, info, filter_info)
└── utils/
├── long2resp.py (long→wide format conversion for response matrices)
├── table_helpers.py (metadata lookups, BibTeX extraction)
└── redivis/ (Redivis API integration: datasets, tables, metadata, item_text, cache)
config.py holds every Redivis reference the package has: MAIN_REFS (the
response-data warehouses), SIM_REF (simulations), COMP_REF (competitions),
META_REF plus META_TABLES (metadata, tags, bibliography, collections), and
ITEMTEXT_REFS (item text). MAIN_REFS and ITEMTEXT_REFS are lists because
Redivis caps a dataset at 1000 tables. The same public API works across the
data sources.
NOM_REF (irw_nominal) is the nominal-response source, added in #9 to match
both R configs. It is tagged like main but has no collections, so
collections() errors for it rather than returning an empty frame — the same
rule as Rpkg/R/redivis-config.R. Note the naming difference across the two
clients: R calls the main source "core", Python calls it "main".
filter(), describe_filter() and get_filters() take source= (#28), and
the metadata lookups under them read that source's tables from
SOURCE_META_TABLES. TAG_SOURCES and COLLECTION_SOURCES in config.py are
R's .irw_tag_sources and .irw_collection_sources; test membership in them,
never source == "main". A filter a source cannot answer raises -- tag filters
for sim/comp, collection outside main, anything but n_responses,
n_actors and license for comp -- because an empty result reads as "no
matches". Main's metadata cache keys stay bare ("tags"); other sources'
carry the source ("tags:nom").
Adding a main warehouse: append a (user, dataset_ref) tuple to MAIN_REFS in config.py only. See docs/DEVELOPERS.md for details and test commands.
- Caching (
utils/redivis/cache.py): Redivis metadata tables are cached in-process to avoid redundant network calls. Cache is per-session only. fetch()return type varies: returns a single DataFrame for a single table name, or adict[name → DataFrame]for a list or a filtered Series.- Warning suppression in
__init__.py: silences a known noisy Redivis warning about missing reference IDs. - Internal naming: functions in
operations/use a leading underscore convention — they are not part of the public API.
irw.list_tables()— discover available datasetsirw.filter(**kwargs)— narrow by metadata (construct, sample size, license, etc.)irw.fetch(name)— download table(s) as pandas DataFrames in long format (id,item,resp)irw.long2resp(df)— pivot to wide-format response matrix for analysisirw.itemtext(name)/irw.save_bibtex(name)— retrieve item text and citations
Tests live in tests/ at the repo root and run with pytest:
pip install pytest
python -m pytest tests/ -qThey are offline — Redivis is mocked — so no credentials are needed.
tests/test_main_refs.py and tests/test_itemtext_refs.py are the templates to
follow when adding a warehouse or an item-text shard.
Linting is ruff check (configured in pyproject.toml, run as a CI job).
There is deliberately no formatter -- reflowing every file would cost the
whole package's git blame for a style win, and that trade was declined
(#33). The rule set is ruff's non-formatting default (E4, E7, E9, F):
real problems, not style. Adding a rule means owning the churn it causes, so
weigh that before widening it.
Follow existing style otherwise: numpydoc-style docstrings, and internal
helpers prefixed with _.
Type annotations are used — config.py and operations/version.py are fully
annotated and every public signature in api.py is. One constraint on them:
the package supports Python 3.9, so PEP 604 unions (Exception | None) need
from __future__ import annotations at the top of the file. A missing one
broke import irw on 3.9; tests/test_metadata_refs.py now checks for it.