TIKA-4831: Add content-based detection and a parser for GeoGebra files (ggb, ggs, ggt) - #3044
Conversation
0575303 to
882b69a
Compare
There was a problem hiding this comment.
Pull request overview
Adds first-class GeoGebra support to Apache Tika by introducing content-based detection for GeoGebra zip containers and a dedicated parser that extracts both metadata and user-visible text (plus a representative thumbnail as an embedded resource).
Changes:
- Added a
GeoGebraDetector(zip-commons) to recognizeggb/ggs/ggtby zip entry names (including streaming detection). - Added a
GeoGebraParser+GeoGebraXMLHandler(miscoffice) to extract GeoGebra metadata/text and emit the thumbnail as an embedded document. - Updated mime registry, zip specializations, module dependencies, and tests to cover new media types and parsing behavior.
Reviewed changes
Copilot reviewed 11 out of 17 changed files in this pull request and generated 3 comments.
Show a summary per file
| File | Description |
|---|---|
| tika-parsers/tika-parsers-standard/tika-parsers-standard-modules/tika-parser-zip-commons/src/test/java/org/apache/tika/detect/zip/GeoGebraDetectionTest.java | Adds detection tests for GeoGebra zip-based formats. |
| tika-parsers/tika-parsers-standard/tika-parsers-standard-modules/tika-parser-zip-commons/src/main/resources/META-INF/services/org.apache.tika.detect.zip.ZipContainerDetector | Registers GeoGebraDetector via ServiceLoader. |
| tika-parsers/tika-parsers-standard/tika-parsers-standard-modules/tika-parser-zip-commons/src/main/java/org/apache/tika/detect/zip/GeoGebraDetector.java | Implements content-based container detection for GeoGebra zip formats. |
| tika-parsers/tika-parsers-standard/tika-parsers-standard-modules/tika-parser-pkg-module/src/main/java/org/apache/tika/parser/pkg/ZipParser.java | Adds GeoGebra media types to zip specialization set. |
| tika-parsers/tika-parsers-standard/tika-parsers-standard-modules/tika-parser-miscoffice-module/src/test/java/org/apache/tika/parser/geogebra/GeoGebraParserTest.java | Adds parser tests for metadata, slide ordering, and thumbnail embedding behavior. |
| tika-parsers/tika-parsers-standard/tika-parsers-standard-modules/tika-parser-miscoffice-module/src/main/java/org/apache/tika/parser/geogebra/GeoGebraXMLHandler.java | SAX handler for extracting metadata/text from GeoGebra XML + inline rich text JSON. |
| tika-parsers/tika-parsers-standard/tika-parsers-standard-modules/tika-parser-miscoffice-module/src/main/java/org/apache/tika/parser/geogebra/GeoGebraParser.java | New parser for ggb/ggs/ggt, including embedded thumbnail + embedded resources. |
| tika-parsers/tika-parsers-standard/tika-parsers-standard-modules/tika-parser-miscoffice-module/pom.xml | Adds jackson-databind dependency for parsing structure.json and rich-text runs. |
| tika-core/src/test/java/org/apache/tika/TikaDetectionTest.java | Extends extension-based detection coverage for .ggp and .ggs. |
| tika-core/src/main/resources/org/apache/tika/mime/tika-mimetypes.xml | Adds .ggp mime type and marks GeoGebra zip-based types as sub-class-of application/zip. |
| CHANGES.txt | Documents new GeoGebra detection/parsing support (TIKA-4831). |
💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.
882b69a to
9cc62ba
Compare
9cc62ba to
fa2b17f
Compare
Mark application/vnd.geogebra.file and .tool as sub-classes of application/zip so the filename hint survives magic detection, and add the (IANA-registered) application/vnd.geogebra.slides (*.ggs) and application/vnd.geogebra.pinboard (*.ggp) types. Add a GeoGebraDetector to the zip container detectors that recognizes the formats without a filename by their well-known entries: geogebra.xml (worksheet), structure.json plus _slideN/geogebra.xml (Notes/Slides) and geogebra_macro.xml (tool), in ZipFile and streaming mode. The detection tests parse nameless streams so they exercise the container detector rather than the globs. Keep ZipParser.ZIP_SPECIALIZATIONS in sync with the new registry entries.
Parse geogebra.xml and, when present, geogebra_macro.xml with the pooled, hardened SAX path: construction title/author/date and the app name/version/format/id become metadata, and the user-visible text (string-literal expressions, rich-text content runs, captions, macro names and help texts) is emitted as XHTML paragraphs. Notes/Slides files emit one div per slide in structure.json order and set xmpTPg:NPages. The representative rendering - geogebra_thumbnail.png at the root, or the first slide's thumbnail - is emitted as an embedded document marked embeddedResourceType=THUMBNAIL so unpack sidecars identify the preview image; other slides' thumbnails are redundant renderings and are skipped. Any other embedded file (e.g. inserted pictures) is emitted as an embedded document under its full zip entry name.
fa2b17f to
34997b9
Compare
A crafted _slide id with more digits than an int holds threw an uncaught NumberFormatException out of the sort comparator; compare the digit strings by length then lexicographically instead. The regression test builds a two-slide container with an oversized id.
There was a problem hiding this comment.
Pull request overview
Copilot reviewed 13 out of 19 changed files in this pull request and generated no new comments.
Suppressed comments (1)
Previously missed (1) — in code that hasn't changed since the last review.
tika-parsers/tika-parsers-standard/tika-parsers-standard-modules/tika-parser-zip-commons/src/main/java/org/apache/tika/detect/zip/GeoGebraDetector.java:71
- In streaming detection mode, this detector always returns null from streamingDetectUpdate and waits for streamingDetectFinal to decide. Once both structure.json and a _slideN/geogebra.xml have been seen, the result is unambiguously GeoGebra Slides (GGS) and cannot be overturned by later entries, so returning GGS early would let DefaultZipContainerDetector/StreamingZipContainerDetector stop scanning the rest of a large ZIP.
public MediaType streamingDetectUpdate(ZipArchiveEntry zae, InputStream zis,
StreamingDetectContext detectContext) {
Names names = detectContext.get(Names.class);
if (names == null) {
names = new Names();
detectContext.set(Names.class, names);
}
names.update(zae.getName());
return null;
|
Let me know what your agent thinks of my agent's input. |
…ventions Metadata keys Tika coined are kebab-cased (geogebra:app-name, app-version, format-version); toolName, id and date are verbatim attribute names and stay. The component is named geogebra-parser explicitly. Parser: document metadata comes from the primary XML only (the macro XML of a worksheet contributes tool names, its nested construction no longer overrides title/author), text expressions yield every string literal (dynamic texts like "Area = " + a), structure.json only orders the slides that exist as _slideN/geogebra.xml (missing or corrupt structure.json falls back to numeric order, leading zeros aside), a root geogebra.xml next to slides is parsed too, the thumbnail falls back to the first slide that has one, geogebra_javascript.js is emitted as MACRO, pictures as INLINE with their slide page, other files as ATTACHMENT, housekeeping names only match at the root and in slide directories, unreadable entries and malformed XML are recorded and skipped instead of aborting the parse, structure.json reads are bounded and content JSON is shape-checked before parsing. Detector: getEntry lookups instead of a central directory walk (entries are only enumerated when structure.json exists), registered last in the SPI file so it does not outrank the existing zip detectors. Streaming detection is exercised by TestContainerAwareDetector with the .ggb fixture. CHANGES notes the upgrade behaviour.
|
Thanks, that was a genuinely useful pass. I've gone through all of it. The two things worth calling out: the metadata keys are now kebab-cased ( On the parser, the real bugs were the macro XML overwriting the worksheet metadata and the text expressions dropping anything that wasn't a single literal. Both fixed; for the latter I checked GeoGebra's source and strings are written between plain quotes with no escaping at all, so the One thing I left alone: Streaming detection is now covered by TestContainerAwareDetector. The fixtures are still synthetic though; I'm trying to get hold of a real GeoGebra file with a suitable licence and will add it when I have one. CHANGES is rewritten along your suggestion. The smaller items (bounded JSON reads, thumbnail fallback, leading zeros in slide ids, the hygiene list) are all in, each with a test. |
…on any line ending Spool and reset the embedded stream around detection like OpenDocumentParser does; split content text runs on CRLF and CR as well; write test zip entries as UTF-8 text or raw bytes.
testGeoGebra_classic.ggb was written by GeoGebra Classic 5.4 (three text objects, one of them dynamic), testGeoGebra_notes.ggs by GeoGebra Notes 5.4 (two pages, thumbnails, an inserted picture). Both get parser tests and replace the synthetic worksheet in TestContainerAwareDetector.
|
Real fixtures are in now: a worksheet written by GeoGebra Classic 5.4 (three text objects, one of them the dynamic |
|
Will merge on green ci. |
|
Great, thanks! |
Adds detection and parsing support for the GeoGebra file formats.
Issue: https://issues.apache.org/jira/browse/TIKA-4831
Detection
application/vnd.geogebra.file(*.ggb) andapplication/vnd.geogebra.tool(*.ggt) are nowsub-class-of application/zip, so the filename hint survives magic detection instead of being discarded in favor of plainapplication/zipapplication/vnd.geogebra.slides(*.ggs, zip-based) andapplication/vnd.geogebra.pinboard(*.ggp, JSON-based)GeoGebraDetector(zip container detector, ZipFile and streaming mode) recognizes the formats without a filename by their well-known entries:geogebra.xml(worksheet),structure.json+_slideN/geogebra.xml(Notes/Slides),geogebra_macro.xml(tool). Since a worksheet with macros contains bothgeogebra.xmlandgeogebra_macro.xml, the decision is made after all entry names have been seenZipParser.ZIP_SPECIALIZATIONSis kept in sync with the new registry entriesParser
A new
GeoGebraParser(miscoffice module) forggb/ggs/ggt:dc:title/dc:creator/geogebra:date, plusgeogebra:appName,geogebra:appVersion,geogebra:formatVersion,geogebra:id; Notes/Slides setxmpTPg:NPagesstructure.jsonordergeogebra_thumbnail.pngat the root, or the first slide's thumbnail) is emitted as an embedded document markedembeddedResourceType=THUMBNAIL, following the existing convention in the OOXML, ODF and iWork parsers, so/unpack/allsidecars identify the preview image; other slides' thumbnails are redundant renderings and are skippedXML is parsed through
XMLReaderUtils.parseSAXlike the other parsers in the module;structure.jsonand rich-text runs use Jackson (newjackson-databinddependency in the miscoffice module, version managed by the existing BOM import).Testing
GeoGebraDetectionTest(zip-commons) andGeoGebraParserTest(miscoffice) with craftedggb/ggs/ggtfixtures; the ggs fixture deliberately orders_slide1before_slide0instructure.jsonto pin the ordering behaviortika-core,tika-parser-zip-commons,tika-parser-miscoffice-moduleandtika-parser-pkg-modulepass, plusCompositeZipContainerDetectorTestin the integration testsvnd.geogebra.slides, text extracted, thumbnail emitted with theTHUMBNAILmarker