Skip to content

Test: D4D for CM4AI (file mode, pre-staged inputs) #163

Description

@realmarcin

@d4dassistant Please create a D4D for CM4AI using file mode.

Dataset name: CM4AI
Input mode: file
Input location: data/sheets_d4dassistant/inputs/CM4AI/

The following 6 source files are pre-staged in the repo:

  • 2024.05.21.589311v1.full.txt (CM4AI preprint)
  • RePORT ⟩ RePORTER - CM4AI.txt (NIH RePORTER record)
  • creativecommons_org_licenses-by-nc-sa_row15.txt (license terms)
  • dataverse_10.18130_V3_B35XWX_row16.txt (Dataverse record 1)
  • dataverse_10.18130_V3_F3TD5R_row19.txt (Dataverse record 2)
  • doi_row3.json (DOI metadata)

Please run the file-mode workflow described in .github/workflows/d4d_assistant_create.md:

  1. ./src/github/validate_prerequisites.sh --dataset CM4AI --mode file
  2. Synthesize across all six sources (the value of file mode is multi-document synthesis).
  3. Validate the generated YAML against the schema and run validate_d4d_completeness.py.
  4. Open a PR with the generated CM4AI_d4d.yaml, metadata YAML, and HTML preview.

Test purpose: This is a deliberate end-to-end test of the GitHub D4D assistant in file mode. We plan to diff the output against the curated baseline at data/d4d_concatenated/curated/CM4AI_curated.yaml and against the existing claudecode_agent output as quality references.

Metadata

Metadata

Assignees

No one assigned

    Type

    No type

    Projects

    No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions