Skip to content

[EN] Extract data from Goskatalog - State Catalogue of the Museum Fund of the Russian Federation #75

Description

@ivbeg

Goal

Scrape metadata and data from Goskatalog - State Catalogue of the Museum Fund of the Russian Federation (National Museum Database).

Data Source

Name: Goskatalog - State Catalogue of the Museum Fund of the Russian Federation
URL: https://goskatalog.ru
Type: National Museum Database

Tasks

  • Analysis: Analyze the detailed structure of the Goskatalog - State Catalogue of the Museum Fund of the Russian Federation website.
  • Extraction: Develop a scraper/parser to extract the following fields (if available):
    • Title / Name
    • Date / Period
    • Author / Creator
    • Description / Abstract
    • URL to original object
      • Dimensions
    • Medium/Technique
    • Provenance
  • Processing: Clean the data and save it as a structured file (CSV, JSONL).
  • Publication: Push the code and data to a new public GitHub repository.

Context

Goskatalog is the central database for all state museum collections in Russia. It contains a vast number of Armenian cultural artifacts, including paintings by major artists like Ivan Ayvazovsky. The Open Data Armenia community has created a dedicated dataset of Armenian-related artwork from Goskatalog, making this data easily accessible.

Deliverables

  1. Code: Python script (or other language) used for extraction.
  2. Data: The extracted dataset in a standard format (UTF-8 encoded).
  3. Documentation: A simple README.md explaining how to use the script and describing the data columns.

Resources

Metadata

Metadata

Assignees

No one assigned

    Labels

    extractionTask that require data extraction (scraping) skillstopic-cultureTasks dedicatated Armenian culture, language and history

    Type

    No type

    Projects

    No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions