Sign inSign up

pratham1uk/snapnscrape-scraper

By pratham1uk

•Updated 5 months ago

Web scraping failsafe service as well as api based on ocr when scraping by traditional methods fail

Image
Web analytics
0

147

pratham1uk/snapnscrape-scraper repository overview

⁠SnapNScrape — Scraper Service

Internal scraping engine of the SnapNScrape platform. Runs a two-stage pipeline: fast HTTP scrape first, then automatic Playwright + Tesseract OCR fallback for sites that block bots or rely on JavaScript rendering.

⁠What this image does

  • Stage 1: Direct HTTP scrape via httpx + BeautifulSoup (fast path)
  • Stage 2: If blocked or JS-heavy, launches headless Chromium, renders the full page, screenshots it, and runs Tesseract OCR to extract text from the image (failsafe path)
  • Internal REST API on port 8001 — not meant to be accessed directly
  • Built with FastAPI + Playwright + Tesseract

⁠This image is the backend engine

It is NOT meant to be run or accessed in isolation. It must be paired with its gateway:

pratham1uk/snapnscrape-gateway:v1.0.0 ← mandatory redis/redis-stack-server:latest ← mandatory

⁠Quick Start

Use the full compose stack — do not run this image alone.

docker compose up

⁠Why this image is large

Ships with Chromium, Tesseract OCR, English language pack, and all required system libraries for headless browser operation. This is intentional — no additional setup needed on any machine.

⁠Environment Variables

TESSERACT_CMD Path to Tesseract binary (default: /usr/bin/tesseract)

⁠Current Version: v1.0.0

⁠Part of the NScrape Series by pratham1uk

Tag summary

Content type

Image

Digest

sha256:cf962f44d…

Size

574.2 MB

Last updated

5 months ago

docker pull pratham1uk/snapnscrape-scraper:1.0.0