Sign inSign up

ofthemachine/headless-browser

By ofthemachine

•Updated 18 days ago

Image
0

1.0K

ofthemachine/headless-browser repository overview

⁠headless-browser — Chromium + Playwright + parsing libs

Purpose-built RCTX for fetching pages that pure-HTTP can't reach: JavaScript- rendered SPAs, Cloudflare-protected sites, and anywhere bot-protection defeats a urllib/requests User-Agent spoof.

⁠Why this exists

ofthemachine/python3 carries requests and httpx — great for plain HTTP. But many modern sites are JS-rendered single-page applications or guarded by Cloudflare / hCaptcha / DataDome. Pure-HTTP returns 403 or empty bodies. Real Chromium driven by Playwright (with stealth patches) gets past most of these.

⁠Key components

  • Chromium — installed via Playwright's bundled installer, pinned at image build time.
  • Playwright — sync + async Python API.
  • playwright-stealth — JS-injected patches for navigator.webdriver, plugins, permissions, MIME types, and other tells. Significantly reduces detection by Cloudflare, hCaptcha, PerimeterX, DataDome.
  • selectolax / lxml / beautifulsoup4 / html5lib — fast and forgiving HTML parsers for downstream extraction.
  • httpx / requests — for the cases that don't need a browser.

⁠Image size

~900 MB. The bulk is Chromium itself (~250 MB compressed) plus its system dependencies installed by playwright install --with-deps. We skip firefox + webkit deliberately to save ~500 MB; derive a heavier image if you need cross-browser parity.

⁠Quick examples

Build locally:

make headless-browser
make headless-browser R=1
make headless-browser I=1

Cloudflare-protected fetch:

from playwright.sync_api import sync_playwright
from playwright_stealth import stealth_sync

with sync_playwright() as p:
    browser = p.chromium.launch(headless=True)
    page = browser.new_page()
    stealth_sync(page)
    page.goto("https://example.com/some-protected-page",
              wait_until="networkidle", timeout=30_000)
    page.wait_for_selector("article, main, body", timeout=15_000)
    html = page.content()
    browser.close()
    print(len(html))

⁠Pairing with operon TOOLs

Wrap as a TOOL for cache-source-style cross-machine fetching:

# Record CODE
operon run ofthemachine/headless-browser fetch_with_browser.py \
  --record-ref tooling/web/fetch_with_browser

# Register TOOL
operon tool register tools/web/fetch_with_browser \
  --from-code tooling/web/fetch_with_browser \
  --param-name URL --param-name WAIT_SELECTOR

# Invoke
operon invoke tools/web/fetch_with_browser \
  -p URL=https://example.com/some-page -p WAIT_SELECTOR=article

That gives any agent / pipeline cross-machine, memoized, bot-evading fetching with no host-side browser install.

Tag summary

Content type

Image

Digest

sha256:8ac6f2f48…

Size

848.7 MB

Last updated

18 days ago

docker pull ofthemachine/headless-browser