Skip to content

Latest commit

 

History

History
329 lines (242 loc) · 31.1 KB

File metadata and controls

329 lines (242 loc) · 31.1 KB

🤝 Contributing to the LINDAT Translation Wrapper of the ATRIUM project

Welcome! Thank you for your interest in contributing. This repository 1 provides a robust workflow for translating archival XML records (specifically ALTO XML and AMCR metadata) into English and other target languages. It addresses common challenges in digital archives, such as safely translating highly nested XMLs without breaking tags, namespaces, or OAI-PMH envelopes.

This document describes the project's capabilities, development workflow, code conventions, and rules for contributors.

📦 Release History

Version Release Type Key Features & Fixes
v0.10.5 Re-vendored atrium_document.py (set_source() fills sub-keys absent from a partial first write, so §1a's deferred origin checks resolve). .coveragerc stops omitting service/* — where the backend-less /translate HTTP 500 lived — with the floor held at 81 against a re-measured 82.87%. scheduled-smoke.yml now states what it actually validates: it re-runs the same hermetic suite as the PR lane, so its signal is fresh-install dependency drift rather than integration coverage — the live-backend gap that let that 500 ship is still open — and -rs makes skips visible. Concurrency scoped by github.event_name; release.yml declares one. Bumps: httpx2, huggingface-hub. Pre-release
v0.10.4 doc_id is inherited, never re-derived. Input is never the original document: alto-postprocess hands it PAGE_ALTO/<doc>/<doc>-1.alto.xml, record_doc_id() now takes the key from the --document-json baseline and falls back to canonical_doc_id() only for a standalone run. Per-file outputs keep the per-file name while the record, the log's file column, the default record filename and the paradata key all take the document's id; service/api.py derives the record filename it returns the same way. Re-vendored atrium_document.py, tests/test_document_originators.py -> canonical shared set, and dependency bumps. Pre-release
v0.10.3 End-to-end GHA pipeline for JSON input-output by atrium_document standard is refined, and tested for the draft JSON schema design. Pre-release
v0.10.2 Major GHA update with references to @v1 on the hub repo. Pre-release
v0.10.1 GHA: Exercises the release path end to end: version guard via the shared check_version.py, post-publish Trivy digest scan + SARIF, and buildkit SBOM/provenance attestations. No functional changes to the tool itself. Pre-release
v0.10.0 Added draft of atrium_document integrated into pipeline of data processing here. Refreshed paradata template. Pre-release
v0.9.0 Updated and bumped dependencies versions, edited GHA release workflow. Standartazied API service to the OpenAPI rules - aligned with agent-skill branch service as well. Pre-release
v0.8.1 Added version reading and license tests (cross-repository template). Dependency versions updated. LLM review performed with Fable. Added new tests and fixed existing ones. Pre-release
v0.8.0 Added agent_dev_log directory with issue logs and their digests+plans. Updated the paradata-related scripts with the atrium-project template. Updated processing to per-page call instead of per-block calls for ALTO XML inputs. Pre-release
v0.7.0 Multi-option backend finalized in theory (not tested in practice), added tests, and Docker GHA was updated. Pre-release
v0.6.2 Multi-option backend translation model (draft) and Docker GH Actions alignment. Pre-release
v0.6.1 Next round LLM review edits and Docker GH Actions alignment. Pre-release
v0.6.0 Security: Hardened XML parsing with lxml _SECURE_PARSER to prevent XXE. Reliability: Raised TranslationError with exponential back-off instead of logging corrupt output strings. Pinned all dependencies. Performance: Added --fast-align CLI flag to reduce ALTO API calls and LINDAT_* env vars for rate limiting. Fixes: Fixed empty-line KeyErrors, whitespace reflow issues in metadata, and improved file-saving directories for URL downloads. Pre-release
v0.5.1 Docker wrapper and small dependencies swap - fasttext-wheel Pre-release
v0.5.0 Dual-pass ALTO reconstruction (block + line translation with similarity-based token alignment); NMT-safe vocabulary sentinels + number-agreement guard; per-run license resolution & paradata logging Pre-release
v0.4.1 Pytest added for main functionality (Tests added to the repository) Pre-release
v0.4.0 Vocabulary added and overall enhancement (Added examples of updated files) Pre-release
v0.3.0 AMCR samples added + documentation expanded + paradata (Added paradata of outputs logging) Pre-release
v0.2.1 Draft version of ALTO/AMCR XML inputs only (Draft of AMCR is ready, Example of ALTO is included, Wrapped up with citation, license, and contribution draft) Pre-release
v0.1.0 Broad inputs support (and AMCR XML) (No logging of ALTO lines translation, AMCR XML support draft is included, No strict input format narrowing done yet) Pre-release
v0.0.2 Various input formats and focus on ALTO (No AMCR XML paths config, Broad inputs optionation like txt, pdf, xml..., Working draft version) Pre-release

v0.5.0 — translation logic & structure preservation (detail). ALTO documents are no longer translated with a single per-block pass. Each TextBlock is now translated twice: once as a whole block (the high-quality tokens that are written back) and once line-by-line (anchors only). _align_tokens_to_lines then partitions the block tokens into one bucket per physical TextLine using a ±50 % sliding-window difflib.SequenceMatcher search against each line anchor, and the tokens are redistributed across each line's String CONTENT attributes (greedy 1-to-1, last String absorbs the remainder). This preserves the original spatial layout while improving translation fluency. The change is covered by tests/test_alignment.py (token conservation, bucket count, clean-signal splits, edge cases) and the existing dual-pass assertions in tests/test_utils.py.


🏗️ Project Contributions & Capabilities

This pipeline contributes 4 major capabilities to the data translation lifecycle, as detailed in the section of the main README 🧠 Logic Overview.

1. Dedicated Archival XML Processing

The pipeline allows archives to safely translate structured documents without altering their spatial coordinates or metadata schemas.

  • ALTO XML Handling: Specifically targets and translates only the CONTENT attributes within TextBlock and TextLine elements natively. Reconstruction uses a dual-pass block/line translation plus similarity-based token alignment so that the original String positions are preserved (see the main README → ALTO Dual-Pass Reconstruction).
  • XML Metadata Handling: Uses deep recursive namespace extraction to parse specific elements based on custom XPaths and safely replace the text content. Works with any well-formed XML (AMCR/OAI-PMH or custom schemas).

2. Multi-Mode Translation Execution

Archive managers can choose processing modes based on their specific document types and workflows:

Mode Best For... Key Feature
ALTO XML Mode Scanned document archives Dual-pass translation + token alignment redistributes translated words back into the exact spatial CONTENT attributes.
XML Metadata Mode Highly nested metadata Safely handles OAI-PMH envelopes and translates specific targeted XPath fields in any well-formed XML.
Batch & URL Ingestion Large-scale collections Scans entire directories or downloads/sanitizes XMLs directly from REST URLs.

3. Automated Language & Quality Controls

A core contribution of this project is minimizing manual preprocessing and providing immediate review tools:

  • Language Identification: Source text is automatically analyzed using FastText 2. If the confidence score is low (< 0.2), the system safely defaults to Czech (cs) to keep the pipeline moving. In ALTO mode, detection runs once per TextBlock so every line in a block shares a consistent source language.
  • Sentence-Aware Chunking: Long texts are split at the highest-priority boundary found in each window, tried in strict order — newline (\n) → sentence-terminal punctuation (. , ! , ? ) → clause-level punctuation (; , , ) → word boundary — before being sent to the translation API. Keeping whole sentences together preserves NMT context; the word boundary is a fallback and a hard cut is the last resort, so mid-word truncation never occurs. The same shared chunker (processors/chunking.py) feeds the UDPipe lemmatiser.
  • QA Logging: Automatically produces a supplementary CSV file (file, page_num, line_num, text_src, text_tgt) for easy line-by-line manual QA review.
  • Schema Validation: Optionally validates metadata outputs against an XSD schema to guarantee post-translation structural integrity.

4. Seamless API & Configuration Integration

The project includes streamlined interfaces for reproducible archival processing:

  • LINDAT Integration: Direct connection to the LINDAT/CLARIAH-CZ Translation Service API (v2) 3.
  • Standardized Configs: Support for config.txt to define default input paths, target languages, and XPath lists, ensuring consistency across different archival teams.

Future work: the LINDAT translation backend may eventually be supplemented or replaced by an alternative open-source / locally hosted NMT model. This is not in scope for the current contribution and is recorded here only as a planned direction.


🌿 Branches & Environments

Branch Environment Rule
test Staging Base for all development. Always branch from test.
master Stable / Integration Merged exclusively by a human reviewer. Do not open PRs directly into master.
test    ←  feature-<name>
test    ←  bugfix-<name>
master  ←  (humans only, after test stabilises)

🏷️ Branch Naming

Type Pattern Example
New feature feature-<name> feature-amcr-validation
Bug fix bugfix-<name> bugfix-chunk-truncation
Hotfix on master hotfix-<name> hotfix-api-timeout

🔁 Contributor Workflow

  1. Create an issue (or find an existing one) describing the problem or feature.
  2. Branch from test:
git checkout test
git pull origin test
git checkout -b feature-<name>
  1. Implement your changes observing the project's code conventions.
  2. Run the minimum tests (see the Testing section).
  3. Open a Pull Request targeting the test branch.

📋 Pull Request Format

Every PR must include:

  • Issue link: Closes #<number> or Refs #<number>
  • Motivation: why the change is needed
  • Description of change: what was changed and how
  • Testing: what was run, what passed, what could not be executed

Use a Draft PR if the work is not ready for review.

**Do not open PRs into master — merging into master is exclusively the maintainers' responsibility.

Note on issue tracking: Issues reference the commits and PRs that resolved them — not the other way around. Commit messages describe what changed; the issue is the place to record why and link the resulting commits together.


✏️ Commit Messages

Format:

[type] concise description of what changed

Allowed types:

Type When to use
add Added content (general)
edit Edited existing content (general)
remove Removed existing content (general)
fix Bug fix
refactor Refactoring without behaviour change
test Adding or updating tests
docs Documentation only
chore Build, dependencies, CI configuration
style Formatting, no logic change
perf Performance optimisation

🧪 Code Conventions & Testing

Code Conventions

  • Comments: informative but short, may be LLM-generated, added when function name does not explain its functionality in detail
  • Argument types: set default type (e.g., int, list) for function arguments
  • Console flags: when a new one added, provide help message for it
  • Config files: when set of variables changes it should be reflected in repository documentation
  • Generated code: always should be manually launched and checked for mistakes before pushing

Minimum checks before every commit

Always run basic validation locally before pushing:

# 1. Python compilation check
python -m compileall -q .

# 2. Pre-commit hooks (runs black, isort, flake8, etc.)
pre-commit run --all-files

Note

If specific scripts or extraction modules are updated, please run a smoke-test against the data_samples/ directory to verify extraction integrity.


Running the test suite

The repository ships a lightweight pytest harness that requires no ML models or GPU for standard unit tests. Heavy tests that do require models or network access are marked slow and are excluded from the default run.

pip install -r requirements-test.txt  # pytest>=8.0 and pytest-cov only
pytest -m "not slow" --tb=short                              # fast — use before every commit
pytest --tb=short                                            # full suite (requires model setup)
pytest -m "not slow" --cov=. --cov-report=term-missing      # with coverage

tests/test_paradata.py (ParadataLogger, _sanitise) is shared across all repos. Repo-specific modules and GPU-heavy tests are marked @pytest.mark.slow and skipped by default.

Test layout, per-repo targets, and fixture conventions
tests/
├── __init__.py              # empty
├── conftest.py              # shared fixtures (tmp_path wrappers, sample data loaders)
├── fixtures/                # small static test-data files committed to the repo
└── test_<module>.py         # repo-specific unit tests

Per-repo targets:

Repository Test file Primary targets
atrium-nlp-enrich test_keywords.py _extract_surface_text, _extract_lemmas, _extract_legacy, extract_keywords, _sort_csv_file
atrium-alto-postprocess test_text_util.py Density/ratio helpers, detectors, pre_filter_line, parse_line_splits, categorize_line (ppl passed directly, no GPU), compute_quality_score
atrium-alto-postprocess test_utils.py directory_scraper, dataframe_results (Top-1 and Top-N), collect_images
atrium-translator test_utils.py _resolve_namespaces, validate_xml_with_xsd, process_alto_xml (incl. dual-pass call counts & redistribution), process_metadata_xml (mock translator injected)
atrium-translator test_alignment.py _align_tokens_to_lines — token conservation, bucket-per-line count, clean-signal splits, empty-block / single-line / empty-anchor edge cases
atrium-translator test_translator.py _chunk_text, _restore_tags, _load_vocabulary, _translate_with_vocabulary, number-agreement guard, boundary-priority chunking
atrium-translator test_lemmatizer.py _parse_conllu (and _parse_conllu_with_features), shared _chunk_text delegation

Slow tests — any test loading a model checkpoint, calling an external API, or requiring a GPU must be decorated with @pytest.mark.slow. Document in the PR description which resource it requires and how to enable it locally.

Fixtures — small, self-contained files committed under tests/fixtures/. Tests must not read from data_samples/ directly. Add a minimal fixture file in the same commit as any test that needs new sample data.

We have transitioned from black/isort/flake8 to Ruff for all linting and formatting, matching the overarching ATRIUM standard.

  1. Linting: Run ruff check . locally before opening a pull request. The CI environment utilizes the shared ruff.toml template.
  2. Testing: Our target is full structural test coverage. Execute tests using:
    pytest -m "not slow" --cov=. --cov-report=term-missing

[!NOTE]: Network-dependent tests (e.g., LINDAT endpoint interactions) and heavy ML model downloads (e.g., FastText weights) are marked @slow.


📁 Repository Documentation Management

Each documentation file has one target audience and one responsibility. Rules are not repeated — cross-references are used instead.

File Audience Responsibility
README.md GitHub visitors Project overview, workflow stages, quick start
CONTRIBUTING.md Developers Code conventions, branches, PRs, testing
  • Do not duplicate rules: if a rule is defined in CONTRIBUTING.md, other files reference it rather than copying it.
  • When changing a rule: update the canonical source and verify that referencing files still point correctly.

📞 Contacts & Acknowledgements

For support or specific archival integration questions, contact lutsai.k@gmail.com responsible for this GitHub repository 1.

  • Developed by: UFAL 4
  • Funded by: ATRIUM 5
  • APIs & Models:
    • LINDAT/CLARIAH-CZ Translation Service 3
    • Facebook's FastText model 2

©️ 2026 UFAL & ATRIUM

Footnotes

  1. https://github.com/ARUP-CAS/atrium-translator 2

  2. https://huggingface.co/facebook/fasttext-language-identification 2

  3. https://lindat.mff.cuni.cz/services/translation/ 2

  4. https://ufal.mff.cuni.cz/home-page

  5. https://atrium-research.eu/