Purpose-built RCTX for PDF post-processing: split, merge, extract, rotate, linearize, encrypt/decrypt, compress, OCR, text/table extraction, and programmatic PDF generation. No LibreOffice weight — pure PDF-in / PDF-out workflows.
ofthemachine/libreofficepdf-toolkit (~200 MB) — PDF post-processing only. Use when the
input is already PDF and the output is PDF or extracted text.libreoffice (~700 MB) — needed only when you must render docx /
pptx / xlsx / odt to PDF, or do format conversion through an office engine.pdftotext, pdftoppm, pdfinfo, pdfunite,
pdfseparate, pdftohtml.gs) — compression, RGB→CMYK, PDF/A conversion, image extraction.mutool for text + image extraction, page rendering.~250 MB. Alpine-based, no LibreOffice. The biggest single dependency is Ghostscript (~80 MB); skip it in a derived image if you don't need compression / format-conversion.
make pdf-toolkit
make pdf-toolkit R=1
make pdf-toolkit I=1
Page count:
import subprocess, re
out = subprocess.run(["pdfinfo", "/input.pdf"], capture_output=True, text=True, check=True)
pages = int(re.search(r"Pages:\s+(\d+)", out.stdout).group(1))
print(pages)
Merge PDFs:
pdfunite a.pdf b.pdf c.pdf bundle.pdf
OCR a scanned PDF:
ocrmypdf --language eng --deskew --clean input.pdf ocred.pdf
Wrap qpdf-backed page counting as a TOOL (avoiding the LibreOffice render cost when the input is already PDF):
operon run ofthemachine/pdf-toolkit pdf_page_count.py \
--record-ref tooling/pdf/page_count
operon tool register tools/pdf/page_count \
--from-code tooling/pdf/page_count
operon invoke tools/pdf/page_count --file <pdf-ref>:/input.pdf
Or a bundle-assembly TOOL:
operon invoke tools/pdf/merge_bundle \
--file <a-ref>:/in/a.pdf \
--file <b-ref>:/in/b.pdf \
--file <c-ref>:/in/c.pdf
Content type
Image
Digest
sha256:bf41a75aa…
Size
240.5 MB
Last updated
19 days ago
docker pull ofthemachine/pdf-toolkit