Sign inSign up

ofthemachine/pdf-toolkit

By ofthemachine

•Updated 19 days ago

Image
0

833

ofthemachine/pdf-toolkit repository overview

⁠pdf-toolkit — lean alpine PDF toolkit

Purpose-built RCTX for PDF post-processing: split, merge, extract, rotate, linearize, encrypt/decrypt, compress, OCR, text/table extraction, and programmatic PDF generation. No LibreOffice weight — pure PDF-in / PDF-out workflows.

⁠When to use this vs ofthemachine/libreoffice

  • pdf-toolkit (~200 MB) — PDF post-processing only. Use when the input is already PDF and the output is PDF or extracted text.
  • libreoffice (~700 MB) — needed only when you must render docx / pptx / xlsx / odt to PDF, or do format conversion through an office engine.

⁠Key components

⁠Native CLIs
  • qpdf — split, merge, linearize, encrypt/decrypt, repair, structural manipulation.
  • poppler-utils — pdftotext, pdftoppm, pdfinfo, pdfunite, pdfseparate, pdftohtml.
  • ghostscript (gs) — compression, RGB→CMYK, PDF/A conversion, image extraction.
  • mupdf-tools — mutool for text + image extraction, page rendering.
⁠Python
  • pypdf — merge / split / rotate / metadata; pure-Python.
  • pikepdf — Python wrapper around qpdf; structural-level manipulation.
  • pdfplumber — table extraction, layout-aware text extraction.
  • pdfminer.six — low-level text + layout extraction.
  • reportlab + fpdf2 — programmatic PDF generation.
  • ocrmypdf — wraps tesseract; auto-skips pages that already have text.
  • lxml — for tagged-PDF / PDF/UA XML manipulation.

⁠Image size

~250 MB. Alpine-based, no LibreOffice. The biggest single dependency is Ghostscript (~80 MB); skip it in a derived image if you don't need compression / format-conversion.

⁠Quick examples

make pdf-toolkit
make pdf-toolkit R=1
make pdf-toolkit I=1

Page count:

import subprocess, re
out = subprocess.run(["pdfinfo", "/input.pdf"], capture_output=True, text=True, check=True)
pages = int(re.search(r"Pages:\s+(\d+)", out.stdout).group(1))
print(pages)

Merge PDFs:

pdfunite a.pdf b.pdf c.pdf bundle.pdf

OCR a scanned PDF:

ocrmypdf --language eng --deskew --clean input.pdf ocred.pdf

⁠Pairing with operon TOOLs

Wrap qpdf-backed page counting as a TOOL (avoiding the LibreOffice render cost when the input is already PDF):

operon run ofthemachine/pdf-toolkit pdf_page_count.py \
  --record-ref tooling/pdf/page_count
operon tool register tools/pdf/page_count \
  --from-code tooling/pdf/page_count
operon invoke tools/pdf/page_count --file <pdf-ref>:/input.pdf

Or a bundle-assembly TOOL:

operon invoke tools/pdf/merge_bundle \
  --file <a-ref>:/in/a.pdf \
  --file <b-ref>:/in/b.pdf \
  --file <c-ref>:/in/c.pdf

Tag summary

Content type

Image

Digest

sha256:bf41a75aa…

Size

240.5 MB

Last updated

19 days ago

docker pull ofthemachine/pdf-toolkit