Sign inSign up

html2rss/botasaurus-scrape-api

By html2rss

•Updated 3 days ago

Scraper API bridge between html2rss and Botasaurus for rendered-page fetching.

Image
Integration & delivery
0

9.6K

html2rss/botasaurus-scrape-api repository overview

⁠html2rss Botasaurus Scrape API

Docker image for the scrape backend used by html2rss.

⁠Purpose

This service fetches rendered HTML from target pages and returns both HTML and best-effort response metadata. It is designed to support html2rss feed generation workflows where plain HTTP fetch is not enough.

⁠API

  • GET /health
  • POST /scrape

⁠html2rss Support

Use this image as the scraping component behind html2rss deployments. It provides:

  • anti-bot-aware navigation strategy (auto: google_get -> google_get_bypass -> get)
  • request isolation per scrape (separate runtime/profile, cleanup in finally)
  • SSRF guardrails for localhost/private/link-local/multicast/reserved ranges
  • stable response fields needed by downstream html2rss processing

⁠Quick Start

docker run --rm -p 4010:4010 html2rss/botasaurus-scrape-api:latest

Health check:

curl -s http://localhost:4010/health

Simple scrape:

curl -s -X POST http://localhost:4010/scrape \
  -H 'Content-Type: application/json' \
  -d '{"url":"https://example.com"}'

Tag summary

Content type

Image

Digest

sha256:40523fdfc…

Size

385.6 MB

Last updated

3 days ago

docker pull html2rss/botasaurus-scrape-api