A GUI to produce PDFs or DjVus from scanned documents.
Scantpaper is a Linux application (it needs GTK3, SANE, and a Python 3 interpreter, all of which are available on other Unix-like systems like MacOS or BSD as well). It is the Python rewrite (v3) of the popular gscan2pdf.
- Scan single- or double-sided, with automatic interleaving
- OCR with tesseract for searchable PDF/A, without the need for temporary files
- Edit with crop, rotate, threshold, unsharp mask, and unpaper clean-up
- Save as PDF, DjVu, TIFF, PS, TXT, hOCR, or image files
- Recover crashed sessions and restore them on the next start
- Quick Start
- Description
- Command-line Options
- Diagnostics
- Configuration
- Dependencies
- Download, Installation & Removal
- Support
- Reporting Bugs
- Translations
- FAQs
- Known Limitations
- See Also
- History
- Author
- Thanks To
- Contributing
- Donate
- License
Install scantpaper and its dependencies (see Download, Installation & Removal), then:
- Start the application with
scantpaper(orpython3 -m scantpaper.appfrom a source checkout, withPYTHONPATH=srcset). Add--debug|info|warn|error|fatalto enable logging at the required level. - Scan one or several pages with File → Scan.
- Select the pages and create a PDF with File → Save.
- To make the saved PDF searchable, enable OCR in the scan window or run Tools → OCR before saving.
scantpaper provides a GUI for scanning, editing, and saving documents as PDF, DjVu, TIFF, PS, TXT, hOCR, SDB (scantpaper session), or image files (PNG, JPEG, PNM, GIF), and can prepend or append to an existing PDF. It supports batch scanning, metadata, OCR, and various editing tools.
Scans are acquired with SANE and held in a session database while you edit
them. When saving, PDFs are produced with img2pdf and OCR'd with ocrmypdf
(which produces PDF/A out of the box); DjVu export uses djvulibre-bin, TIFF
export uses libtiff, and images are written with ImageMagick.
┌─────────┐ ┌─────────────────┐ ┌──────────────┐ ┌──────────────────┐
│ SANE │ │ SQLite session │ │ edit tools / │ │ img2pdf / │──▶ PDF (PDF/A)
│ scanner │──▶│ (pages in temp │──▶│ OCR │──▶│ ocrmypdf │
│ │ │ directory) │ │ (tesseract) │ ├──────────────────┤──▶ DjVu
└─────────┘ └─────────────────┘ └──────────────┘ │ djvulibre-bin │
├──────────────────┤──▶ TIFF
│ libtiff │
├──────────────────┤──▶ PNG, JPEG, PNM, GIF
│ imagemagick │
└──────────────────┘
Page numbers are always consecutive (1, 2, 3, …). Deleting a page renumbers the remainder automatically, and editing a page's number in the document table moves that page to the corresponding position.
-
Single-sided: Each scan is appended at the end.
-
Double-sided: Scan all facing pages first (front 1, front 2, …, front n); they are appended in order. When you flip the stack the ADF feeds them in reverse (back of page n, then back of page n-1, …, back of page 1). Each reverse page is inserted immediately after its matching front page, producing a fully interleaved result: front 1, back 1, front 2, back 2, …, front n, back n.
scan pass 1 (fronts): front 1, front 2, ..., front n flip stack scan pass 2 (backs): back n, back n-1, ..., back 1 interleaved result: front 1, back 1, front 2, back 2, ..., front n, back n -
Extended mode (insert before page N): Each new scan is inserted before the selected page, advancing the insertion point for the next scan.
Automatic document feeder (ADF) and duplex scans now import every side of the document. Some Brother scanners (e.g. the DS-740D) prefetch the reverse side of a sheet as soon as the front side finishes reading; cancelling the scan session between pages discarded that buffered side, so only the first page of a duplex job was imported. Scantpaper no longer cancels the session between pages when scanning from a feeder, which preserves the prefetched side. The session is still cancelled at the end of the batch, when the requested page count is reached, on error, or when you cancel the scan.
For multi-page flatbed batches the behaviour follows gscan2pdf semantics: the Force new scan job between pages preference (enabled via Edit → Preferences, only available when Allow batch scanning from flatbed is enabled) controls whether the session is cancelled between flatbed pages; the session is always cancelled at the end of the batch.
- Scan: Options for device, page count, source document, side to scan, and device-dependent options (page size, mode, resolution, batch-scan, etc.). Optionally OCR each page on scan. When applying a scan-options profile, the scanner driver may revert an option (some drivers couple or alias options, e.g. page size and scan area); such an option is then set at most twice and dropped for the rest of that apply, so the scan still proceeds instead of ending in a reload-recursion error.
- Save: Save selected/all pages in multiple formats. Supports metadata. The Title, Author, Subject, and Keywords fields offer autocompletion, suggesting values from imported documents and values you have entered before. When saving as PDF, the progress bar tracks each page as it is written and reports the PDF conversion step. Imported images are stored in a compact format (JPEG for scanned pages, the original bytes for imported JPEG/PNG files, lossless PNG for bilevel or transparent pages), and PDF saves embed stored JPEG images directly instead of re-encoding them.
- Thumbnails: Page thumbnails are generated with a decimate-then-LANCZOS downscale for a sharp preview in the thumbnail panel without slowing down bulk imports, and bilevel or palette pages (common in scanned/OCR PDFs) are anti-aliased so their previews stay legible.
- Email as PDF: Attach pages as PDF to a blank email (requires xdg-email).
- Print: Print selected/all pages.
- Undo, Redo: Undo or redo the last action.
- Cut, Copy, Paste: Cut, copy, or paste selected pages.
- Delete: Remove selected pages.
- Select: Select all, odd, even, inverted, blank, dark, or modified pages, or pages without (up-to-date) OCR.
- Properties: Edit image metadata.
- Preferences: Configure default behaviors and frontends.
- Pan, Select, Select & Pan tools:
- Pan: Use the left mouse button to drag the image or canvas to pan the view
- Select: Use the left mouse button to select a rectangular box
- Select & Pan: Use the left mouse button to select a rectangular box, and the middle mouse button to drag the image or canvas to pan the view
- In all of the above, the mouse wheel zooms in or out.
- Zoom: 100%, fit to window, in, and out.
- Rendering: pages are shown at high quality whenever they are static, so the full-page view is easy to assess; a faster rendering is used only while you are panning or zooming, for smooth interaction.
- Rotate: 90° clockwise, 180°, and 90° anticlockwise.
| Action | Shortcut |
|---|---|
| New | Ctrl+N |
| Open | Ctrl+O |
| Scan | Ctrl+G |
| Save | Ctrl+S |
| Email as PDF | Ctrl+E |
| Ctrl+P | |
| Quit | Ctrl+Q |
| Undo | Ctrl+Z |
| Redo | Ctrl+Shift+Z |
| Cut | Ctrl+X |
| Copy | Ctrl+C |
| Paste | Ctrl+V |
| Delete | Del |
| Select | |
| All | Ctrl+A |
| Odd | Ctrl+1 |
| Even | Ctrl+2 |
| Invert | Ctrl+I |
| Blank | Ctrl+B |
| Dark | Ctrl+D |
| Modified | Ctrl+M |
| View | |
| Zoom in | + |
| Zoom out | − |
| Rotate 90° clockwise | Ctrl+Shift+R |
| Rotate 180° | Ctrl+Shift+F |
| Rotate 90° anticlockwise | Ctrl+Shift+C |
| Help | Ctrl+H |
- Threshold: Binarize images. The value is an ink-strength cutoff: pixels that differ from the paper colour by more than the given percentage are rendered black, so coloured text and annotations (stamps, highlighters) on white paper are kept. The default is 20; values saved by older versions are migrated automatically.
- Brightness / Contrast: Adjust brightness and contrast.
- Negate: Invert colours.
- Unsharp mask: Sharpen images.
- Crop: Crop selected pages.
- Clean up: Use unpaper to clean up scans.
- Split: Split pages vertically or horizontally.
- OCR: Use tesseract to create a text layer for the selected pages. The text layer is embedded into saved PDFs, making them searchable, and can be viewed and edited in the text layer window. Page pixels are fed to tesseract in memory and the stored image is reused, so OCR is faster than in previous versions.
- User-defined: Run user-defined commands.
%i- input filename%o- output filename%r- resolution
scantpaper supports the following options:
-
--device <device> [...]Specifies the device(s) to use, instead of getting the list from the SANE API. Useful for remote scanners. May be repeated, or given multiple space-separated devices. -
--help
Displays help and exits. -
--log=<log-file>
Specifies a file to store logging messages. On exit, the log is compressed to<log-file>.xz. -
--debug,--info,--warn,--error,--fatal
Defines the log level. Defaults to--debugif a log file is specified, otherwise--warn. -
--import=<PDF|DjVu|images>
Imports the specified file(s). For multi-page documents, a window is displayed to select required pages. -
--import-all=<PDF|DjVu|images>
Imports all pages of the specified file(s). -
--locale=<directory>Sets the directory containing translated messages. See Translations. -
--version
Displays the program version and exits.
$ scantpaper --version
scantpaper X.Y.Z(Replace X.Y.Z with your installed version.)
$ scantpaper --help
usage: scantpaper [-h] [--device DEVICE [DEVICE ...]]
[--import IMPORT_FILES [IMPORT_FILES ...]]
[--import-all IMPORT_ALL [IMPORT_ALL ...]] [--locale LOCALE]
[--log LOG] [--version] [--debug] [--info] [--warn]
[--error] [--fatal]
A GUI to produce PDFs or DjVus from scanned documents
options:
-h, --help show this help message and exit
--device DEVICE [DEVICE ...]
--import IMPORT_FILES [IMPORT_FILES ...]
--import-all IMPORT_ALL [IMPORT_ALL ...]
--locale LOCALE
--log LOG
--version show program's version number and exit
--debug
--info
--warn
--error
--fatal
Please see /usr/share/doc/C/scantpaper/documentation.html for more detail# Import every page of a PDF, letting you edit before saving
scantpaper --import-all ~/scans/document.pdf
# Import a PDF, choosing the pages to import in a dialog
scantpaper --import ~/scans/document.pdf
# Use a remote scanner
scantpaper --device "net:scanner.example.com:6566"Scanning is handled with SANE. PDF conversion uses img2pdf and ocrmypdf. TIFF export uses libtiff.
Saved PDFs are marked as created by scantpaper: the document info Creator
field and the XMP creator tool begin with scantpaper v<version>, followed
by the OCR toolchain used, e.g.
scantpaper v3.0.16 / OCRmyPDF 16.13.0 / Tesseract OCR-hOCR 5.5.0. The
Producer field continues to name the library that wrote the file.
To diagnose errors, start scantpaper from the command line with logging enabled:
PYTHONPATH=src python3 -m scantpaper.app --debugscantpaper creates a config file at ~/.config/scantpaperrc. The directory can be changed by setting $XDG_CONFIG_HOME. Preferences are usually set via Edit → Preferences.
If the config file cannot be read as JSON (for example after a partial write
or a hand edit), scantpaper rescues as many recognised settings as it can and
resets the rest to their defaults. The unreadable original is kept as a backup
at ~/.config/scantpaperrc.old, and the start-up message tells you which
settings were kept. Wrongly typed values (e.g. a string where a number is
expected) are corrected when possible; each correction is written to the log
rather than shown in the start-up message. Only values that cannot be corrected
are reported. The original file is never silently overwritten: saving on exit
waits until the message has been acknowledged.
Scan options saved in the configuration file by gscan2pdf 2.x (whose
default-scan-options block has no frontend list) are recognised and applied
to the scan dialog on startup, and are kept when the file is written back.
All session data (pages, edits, OCR, annotations) is stored in an SQLite
database in a temporary directory named scantpaper-????????, created under
$TMPDIR (or /tmp) by default. You can change this location in
Edit → Preferences. On exit the session directory is cleaned up.
If scantpaper crashes, the session directory survives. On the next start you are asked whether to restore it via File → Open crashed session.
Package names below are the Debian package names. Equivalent packages for other distributions are given in the wheel file installation instructions.
- gir1.2-gdkpixbuf-2.0
- gir1.2-gtk-3.0
- imagemagick
- img2pdf
- libtiff-tools
- ocrmypdf
- poppler-utils
- python3-gi
- python3-gi-cairo
- python3-iso639
- python3-pil
- python3-sane
- python3-tesserocr
- djvulibre-bin
- qpdf (for PDF encryption)
- unpaper
- xdg-utils
- python3-pytest-mock
- python3-pytest-cov
- python3-pytest-timeout
- python3-pytest-xvfb
- Python 3.10 or later.
- The dependencies listed below. They are installed
automatically when installing from a wheel or with
uv, but must be installed manually when running from a tarball or the repository.
-
Debian
sidshould automatically have the latest version.sudo apt update sudo apt install scantpaper
-
Ubuntu users can use the PPA:
sudo apt-add-repository ppa:jeffreyratcliffe/ppa sudo apt update sudo apt install scantpaper
In either case to remove scantpaper afterwards:
sudo apt remove scantpaperDownload .whl from Github.
# Install the C-libraries that pip cannot handle:
# For Debian/Ubuntu
sudo apt update
sudo apt install libgirepository-2.0-dev libcairo2-dev pkg-config python3-dev gir1.2-glib-2.0
# For Fedora
sudo dnf install gobject-introspection-devel cairo-devel pkgconf-pkg-config python3-devel
# For Arch
sudo pacman -S gobject-introspection cairo pkgconf python
# For Homebrew
brew install pygobject3 gobject-introspection cairo pkg-config
# Possibly upgrade pip
python3 -m pip install --upgrade pip
# Install from the wheel file, automatically including python dependencies
pip install scantpaper-x.x.x-py3-none-any.whlIf you haven't already, you will then probably have to add ~/.local/bin to
your path in order to find the new executable, after which you can start it with:
scantpaperTo then remove it:
pip uninstall scantpaperTo install the runtime dependencies with uv:
uv syncor with the additional development dependencies:
uv sync --extra testAfter which you can start it with:
uv scantpaperDownload .tar.gz from Github.
After installing the dependencies listed above:
tar xvfz scantpaper-x.x.x.tar.gz
cd scantpaper-x.x.x
PYTHONPATH=src python3 -m scantpaper.appBrowse the code at Github. After installing the dependencies listed above:
git clone https://github.com/carygravel/scantpaper.git
cd scantpaper
PYTHONPATH=src python3 -m scantpaper.appIn either of the above two cases, just delete the source directory to remove it.
- Mailing lists:
- gscan2pdf-announce (announcements)
- gscan2pdf-help (general support)
- Please read the FAQs first.
- Report bugs preferably against the Debian package or Debian Bugs.
- Alternatively, use the Github issue tracker.
- Include the log file created by
scantpaper --log=logwith your report. On exit the log is compressed tolog.xz, so submit that file.
scantpaper is partly translated into several languages. Contribute via Launchpad Rosetta.
- Scanner option translations come from sane-backends. Contribute via the sane-devel mailing list or SANE project.
To test updated .po files:
python3 dev/compile_mo.py --src po --out src/scantpaper/locale --domain scantpaper
PYTHONPATH=src python3 -m scantpaper.app --log=log --locale=localeSet locale variables as needed (e.g., for Russian):
LC_ALL=ru_RU.utf8 LC_MESSAGES=ru_RU.utf8 LC_CTYPE=ru_RU.utf8 LANG=ru_RU.utf8 LANGUAGE=ru_RU.utf8 PYTHONPATH=src python3 -m scantpaper.app --log=log --locale=localeIt may not be supported by SANE or your scanner. If you see it in scanimage --help but not in scantpaper, send the output to the maintainer.
Enable "Allow batch scanning from flatbed" in Preferences. Some scanners require additional settings.
The required package may not be installed (e.g., xdg-email, unpaper, imagemagick).
Set "# Pages" to "1" and "Batch scan" to "No".
Only changelogs from official Ubuntu builds are shown.
If the scanner is remote and not found automatically, specify the device:
scantpaper --device <device>Use pdftotext or djvutxt to extract text. Many viewers support searching the embedded text layer.
Create or edit ~/.config/gtk-3.0/gtk.css:
.rubberband,
rubberband,
flowbox rubberband,
treeview.view rubberband,
.content-view rubberband,
.content-view .rubberband {
border: 1px solid #2a76c6;
background-color: rgba(42, 118, 198, 0.2);
}
#scantpaper-ocr-output {
color: black;
}"scant" (https://en.wiktionary.org/wiki/scant) in this sense means "short (of)", as I am trying to digitalise my paperwork, and I liked the play on "scan".
Saving more than approximately 250 uncompressed scanned pages (at 300 dpi, 8-bit grayscale) produces a PDF exceeding 2 GiB. At that size, three tools in the save pipeline overflow 32-bit file offsets and produce truncated or corrupt output:
| Component | Tested version | Overflow point | Symptom |
|---|---|---|---|
| img2pdf (pikepdf engine, linearization) | 0.6.2 | 32-bit xref offsets in linearized output | Truncated PDF; first pages readable, later pages missing |
| Ghostscript (PDF/A conversion via ocrmypdf) | 10.07.1 | 32-bit file access in gs interpreter | Ghostscript error or corrupt output |
| pikepdf / qpdf (xref-stream linearization, metadata save) | pikepdf 10.5.0, qpdf 12.4.0 | 32-bit offsets in xref streams | "unable to find /Root dictionary"; PDF unopenable |
Scantpaper now estimates the output size before conversion and refuses to save when it would exceed 2 GiB, showing an error message suggesting fewer pages.
When updating dependencies, re-test by saving ~250 high-resolution uncompressed pages (e.g., 7000×5000 px grayscale TIFFs) and verifying the output PDF opens correctly in a PDF viewer. Note that Ghostscript's 64-bit integer support (needed for >2 GiB files) is build-dependent — see the Ghostscript documentation on word size.
- gscan2pdf (the Perl predecessor of scantpaper)
- XSane
- Scan Tailor
I started writing gscan2pdf as a Perl & Gtk2 project in 2006.
Version 2 switched to Gtk3, but kept the basic software architecture.
This stored the pages as temporary files with hashed names, which had a couple
of major disadvantages:
- Difficult to support PDF/A directly
- It was impossible to create documents with more than a few hundred pages, as it ran out of open file handles.
- In the event of a crash, it was tedious to recreate the document from the image files.
- AFAIK, Perl's support for Gtk4 never extended beyond that provided by introspection.
Therefore I decided in 2022 to completely rewrite gscan2pdf in Python and
renamed it for v3 scantpaper. The rewrite:
- Supports PDF/A by using
ocrmypdfto write PDFs - Stores all session data in a single Sqlite database
- Should be simple to migrate to Gtk4
See also the changelog for detailed release notes.
Jeffrey Ratcliffe (jffry at posteo dot net)
- All contributors (patches, translations, bugs, feedback)
- The SANE project
- The authors of
img2pdfandocrmypdf, without which this would have been much harder.
See contributing.
Copyright © 2006–2026 Jeffrey Ratcliffe jffry@posteo.net
This program is free software: you can redistribute it and/or modify it under the terms of the GNU General Public License v3 as published by the Free Software Foundation.
This program is distributed in the hope that it will be useful, but WITHOUT ANY WARRANTY; without even the implied warranty of MERCHANTABILITY or FITNESS FOR A PARTICULAR PURPOSE. See the GNU General Public License for more details.

