Skip to content

[EN] Extract data from AGBU Nubar Library (Paris) #67

Description

@ivbeg

Goal

Scrape metadata and data from AGBU Nubar Library (Paris) (Library and Archive).

Data Source

Name: AGBU Nubar Library (Paris)
URL: http://bnulibrary.org
Type: Library and Archive

Tasks

  • Analysis: Analyze the detailed structure of the AGBU Nubar Library (Paris) website.
  • Extraction: Develop a scraper/parser to extract the following fields (if available):
    • Title / Name
    • Date / Period
    • Author / Creator
    • Description / Abstract
    • URL to original object
      • Manuscript ID / Shelfmark
    • Material (Parchment/Paper)
    • Dimensions
  • Processing: Clean the data and save it as a structured file (CSV, JSONL).
  • Publication: Push the code and data to a new public GitHub repository.

Context

Founded in 1928, the AGBU Nubar Library is one of the most important centers for Armenian studies in the world. It holds Europe's largest collection of Armenian books, periodicals, and archives, with over 500,000 archival documents and 10,000 photographs. Its main archival collections, including the Andonian Collection on the Genocide and the archives of the Armenian Patriarchate of Istanbul, are currently being digitized, representing a foundational resource for Armeniana.

Deliverables

  1. Code: Python script (or other language) used for extraction.
  2. Data: The extracted dataset in a standard format (UTF-8 encoded).
  3. Documentation: A simple README.md explaining how to use the script and describing the data columns.

Resources

Metadata

Metadata

Assignees

No one assigned

    Labels

    extractionTask that require data extraction (scraping) skillstopic-cultureTasks dedicatated Armenian culture, language and history

    Type

    No type

    Projects

    No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions