Skip to content

[EN] Extract data from BULAC - Bibliothèque universitaire des langues et civilisations #72

Description

@ivbeg

Goal

Scrape metadata and data from BULAC - Bibliothèque universitaire des langues et civilisations (University Research Library).

Data Source

Name: BULAC - Bibliothèque universitaire des langues et civilisations
URL: https://www.bulac.fr
Type: University Research Library

Tasks

  • Analysis: Analyze the detailed structure of the BULAC - Bibliothèque universitaire des langues et civilisations website.
  • Extraction: Develop a scraper/parser to extract the following fields (if available):
    • Title / Name
    • Date / Period
    • Author / Creator
    • Description / Abstract
    • URL to original object
      • Manuscript ID / Shelfmark
    • Material (Parchment/Paper)
    • Dimensions
  • Processing: Clean the data and save it as a structured file (CSV, JSONL).
  • Publication: Push the code and data to a new public GitHub repository.

Context

BULAC is a major research library for languages and civilizations, associated with INALCO. It serves as a central portal for Armenian studies in France, providing curated access to numerous digital resources, including Armenian Rare Books (1512-1920), the National Library of Armenia's digital collections, and various academic journals and databases.

Deliverables

  1. Code: Python script (or other language) used for extraction.
  2. Data: The extracted dataset in a standard format (UTF-8 encoded).
  3. Documentation: A simple README.md explaining how to use the script and describing the data columns.

Resources

Metadata

Metadata

Assignees

No one assigned

    Labels

    extractionTask that require data extraction (scraping) skillstopic-cultureTasks dedicatated Armenian culture, language and history

    Type

    No type

    Projects

    No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions