Web scraping failsafe service as well as api based on ocr when scraping by traditional methods fail
147
Internal scraping engine of the SnapNScrape platform. Runs a two-stage pipeline: fast HTTP scrape first, then automatic Playwright + Tesseract OCR fallback for sites that block bots or rely on JavaScript rendering.
It is NOT meant to be run or accessed in isolation. It must be paired with its gateway:
pratham1uk/snapnscrape-gateway:v1.0.0 ← mandatory redis/redis-stack-server:latest ← mandatory
Use the full compose stack — do not run this image alone.
docker compose up
Ships with Chromium, Tesseract OCR, English language pack, and all required system libraries for headless browser operation. This is intentional — no additional setup needed on any machine.
TESSERACT_CMD Path to Tesseract binary (default: /usr/bin/tesseract)
Content type
Image
Digest
sha256:cf962f44d…
Size
574.2 MB
Last updated
5 months ago
docker pull pratham1uk/snapnscrape-scraper:1.0.0