Purpose-built RCTX for office document workflows: docx/pptx/xlsx ↔ PDF conversion, Word-accurate page-count via rendered PDF, OOXML manipulation, and PDF post-processing (split/merge/extract/linearize).
ofthemachine/python3 already carries the Python office libraries
(python-docx, pypdf, etc.) but has no native renderer. For accurate page
counts, format conversion, or PDF rendering you need a real office engine.
LibreOffice is the de-facto open-source choice and runs headlessly under a
non-privileged user.
This RCTX was proposed (and validated) against a real use case: building an operon TOOL for cross-machine docx page-counting. A pure-Python structural heuristic gave estimates within ±1-2 pages of Word on typical multi-page docx inputs; a soffice-backed TOOL would deliver Word-exact counts portably.
| Tool | Use |
|---|---|
soffice | LibreOffice headless. --headless --convert-to pdf is the canonical invocation. Supports docx/doc/odt/rtf/html/pptx/odp/xlsx/ods → many output formats. |
pdftotext / pdftoppm / pdfinfo / pdfunite | Poppler utilities. pdfinfo gives accurate page count. |
qpdf | Linearization, encryption/decryption, structural manipulation, repair. |
ghostscript (gs) | PostScript/PDF Swiss-army knife: compression, RGB→CMYK, PDF/A-conversion, image extraction. |
python-docx / python-pptx / openpyxl | Direct OOXML manipulation. |
pypdf / pdfplumber / pdfminer.six | PDF text & layout extraction. |
lxml | Low-level OOXML XML edit. |
Curated minimal set for consistent pagination across rebuilds:
fonts-liberation + fonts-liberation2 — Arial/Times/Courier equivalents (metric-compatible)fonts-noto-core — broad Latin / common-script coveragefonts-dejavu-core — fallbackfc-cache -f is run at build time so soffice sees them immediately.
Skipped: fonts-noto-cjk (~200 MB; add in a derived image if needed).
~700 MB. Most of the weight is LibreOffice itself (~400 MB) and the curated fonts (~100 MB). Python libs are slim by comparison.
Build locally:
make libreoffice
make libreoffice R=1 # build + run smoke test
make libreoffice I=1 # interactive shell
Word-accurate page count in a fraglet:
import subprocess, re, os
r = subprocess.run(
["soffice", "--headless", "--convert-to", "pdf",
"--outdir", "/tmp", "/input.docx"],
capture_output=True, check=True, timeout=120,
)
pdf = "/tmp/" + os.path.splitext(os.path.basename("/input.docx"))[0] + ".pdf"
info = subprocess.run(["pdfinfo", pdf], capture_output=True, text=True, check=True)
pages = int(re.search(r"Pages:\s+(\d+)", info.stdout).group(1))
print({"pages": pages, "method": "soffice+pdfinfo"})
Register a TOOL anchored to CODE that runs in this RCTX:
# 1. Record the CODE under a ledger ref
operon run ofthemachine/libreoffice page_count_soffice.py \
--record-ref tooling/docx/page_count_soffice
# 2. Register a TOOL anchored to the CODE
operon tool register tools/docx/page_count_soffice \
--from-code tooling/docx/page_count_soffice
# 3. Invoke with a docx FILE mounted at /input.docx
operon invoke tools/docx/page_count_soffice \
--file <docx-ref>:/input.docx
The TOOL ref tools/docx/page_count stays as the stable invocation surface;
swapping the backing CODE/RCTX is transparent to consumers.
Content type
Image
Digest
sha256:cae768c28…
Size
334.2 MB
Last updated
19 days ago
docker pull ofthemachine/libreoffice