Skip to content

feat/archive-scrape: find and download items from the Blissymoblics Internet Archive - #12

Open
klown wants to merge 5 commits into
inclusive-design:mainfrom
klown:feat/archive-scrape
Open

feat/archive-scrape: find and download items from the Blissymoblics Internet Archive#12
klown wants to merge 5 commits into
inclusive-design:mainfrom
klown:feat/archive-scrape

Conversation

@klown

@klown klown commented Jun 9, 2023

Copy link
Copy Markdown
Contributor

Description

Add a utility script that searches and downloads data from the Blissymbolics Internet Archive based on the research documented in the Archive Data Extraction Feasibility document.

Steps to test

  1. Set up a python virtual environment and install the Internet Archive's python library
  2. Run the extraction script

More details are provided in the README.

Expected behavior:

The relevant PDF and JP2 zip files are downloaded to the local directory.

NOTE: The size of the full download is approximately 4 gigabytes.

klown added 5 commits June 9, 2023 17:34
- search strings are closer to the full title, where possible
- search strings document better the titles sought
Instead of letting the exception kill the script, record the error on
the terminal and continue to the next item to download.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant