Skip to content

[EN] Extract data from Gallica - Bibliothèque nationale de France (BnF) #69

Description

@ivbeg

Goal

Scrape metadata and data from Gallica - Bibliothèque nationale de France (BnF) (National Digital Library (France)).

Data Source

Name: Gallica - Bibliothèque nationale de France (BnF)
URL: https://gallica.bnf.fr
Type: National Digital Library (France)

Tasks

  • Analysis: Analyze the detailed structure of the Gallica - Bibliothèque nationale de France (BnF) website.
  • Extraction: Develop a scraper/parser to extract the following fields (if available):
    • Title / Name
    • Date / Period
    • Author / Creator
    • Description / Abstract
    • URL to original object
      • Manuscript ID / Shelfmark
    • Material (Parchment/Paper)
    • Dimensions
  • Processing: Clean the data and save it as a structured file (CSV, JSONL).
  • Publication: Push the code and data to a new public GitHub repository.

Context

Gallica is the digital library of the Bibliothèque nationale de France (BnF) and its partners. It provides access to millions of documents, including a significant collection of Armenian materials.
The BnF's full Armenian collection is estimated to contain ~350 manuscripts and 70% of all books printed in Armenian in the 16th and 17th centuries.

Deliverables

  1. Code: Python script (or other language) used for extraction.
  2. Data: The extracted dataset in a standard format (UTF-8 encoded).
  3. Documentation: A simple README.md explaining how to use the script and describing the data columns.

Resources

Metadata

Metadata

Assignees

No one assigned

    Labels

    extractionTask that require data extraction (scraping) skillstopic-cultureTasks dedicatated Armenian culture, language and history

    Type

    No type

    Projects

    No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions