Sign inSign up

datweazel/caddy-portal

By datweazel

•Updated 1 day ago

Automatic HTTPS via Caddy + Let's Encrypt, configured by a single DOMAINS env var

Image
0

566

datweazel/caddy-portal repository overview

⁠caddy-portal

Automatic HTTPS for any web application, configured by a single environment variable. Spiritual successor to https-portal⁠, rebuilt on Caddy⁠ instead of Nginx.

Caution

**AI Disclaimer** I built the first version of this rewrite with the help of Claude Code. I still made sure I understand the code and its behavior so this isn't a copy-paste AI slop rewrite. I use AI as a tool not as a replacement for know-how.

⁠Why

  • HTTP/3 (QUIC) out of the box — no flags, no rebuild
  • Smaller image (~140 MB vs. ~200 MB) — single static Caddy binary, no acme-tiny, no cron, no DH-param generation
  • Caddy renews certs in-process — no cron job to monitor, renewal failures land in the same log stream
  • Modern TLS defaults that keep up with best practice automatically — no cipher lists to maintain
  • Structured JSON logs ready for Loki / Elastic / Datadog
  • Drop-in for the majority — same DOMAINS descriptor syntax, same persistent volume path, same VIRTUAL_HOST auto-discovery
  • Atomic config reloads via Caddy's admin API — no "reload failed silently" drift
  • No restart cascades — backend containers can restart freely (Caddy re-resolves DNS per request), and new domains can be hot-added via dynamic-env or VIRTUAL_HOST without touching caddy-portal itself. See Changing configuration at runtime⁠

What's deliberately not preserved: CUSTOM_NGINX_* env vars, bind-mounts of Nginx templates, and a handful of Nginx-specific tuning knobs. See docs/migration.md⁠ for the full list.

⁠Quick start

services:
  caddy-portal:
    image: datweazel/caddy-portal:1
    ports:
      - "80:80"
      - "443:443"
      - "443:443/udp"        # HTTP/3
    environment:
      DOMAINS: "example.com -> http://app:80"
      STAGE: staging         # use 'production' once DNS + reachability are confirmed
    volumes:
      - caddy-portal-data:/var/lib/https-portal

  app:
    image: your/app

volumes:
  caddy-portal-data:

Bring it up with docker compose up -d. Caddy obtains a certificate from Let's Encrypt's staging environment, then any request to https://example.com is reverse-proxied to your app container.

Once you've verified everything works, swap STAGE: staging → STAGE: production to get real (browser-trusted) certificates.

⁠DOMAINS descriptor reference

DOMAINS is a comma-separated list of descriptors. Each descriptor follows:

[ips] user:pass@host:port -> protocol://upstream(s) #stage

Everything except host is optional. Examples:

DescriptorWhat it does
example.comStatic site, serves a welcome page until you mount your own content under /var/www/vhosts/example.com
example.com -> http://app:80Reverse-proxy to app:80 over plain HTTP
example.com -> https://backend:8443Reverse-proxy to an HTTPS backend
example.com:8443 -> http://appListen on :8443 instead of :443
example.com => https://www.example.com307 redirect (use REDIRECT_CODE=301 for permanent)
www.example.com => https://example.com, example.com -> http://appThe classic www → apex setup
user:[email protected] -> http://appHTTP Basic Auth in front of the proxy
[10.0.0.0/8] example.com -> http://appAllow only the listed IPs; everyone else gets 403
example.com -> http://a:80|b:80Round-robin load balance across two upstreams
dev.example.com -> http://app #localOverride the global stage for this one domain (here: self-signed via Caddy's internal CA)

Multi-line YAML works too:

DOMAINS: >
  example.com -> http://app:80,
  api.example.com -> http://api:3000,
  admin.example.com -> http://admin:4000 #staging

⁠Configuration

The minimum you need is DOMAINS. Everything else has sensible defaults.

⁠Stages
STAGEWhat it doesWhen to use
productionReal Let's Encrypt certificates, trusted by browsersAfter you've verified DNS + reachability with staging
staging (default)Let's Encrypt staging certs — not browser-trusted. No rate-limit riskInitial testing
localSelf-signed certs via Caddy's internal CALocal dev where you don't have real DNS

A trailing #stage on a descriptor overrides the global stage for one domain only: example.com -> http://app #staging always uses staging even when STAGE=production.

⁠Environment variables

The full list. Anything not listed here is either an https-portal Nginx-ism that's intentionally dropped (see docs/migration.md⁠) or unsupported.

⁠Core
VariableDefaultMeaning
DOMAINS(none)Comma-separated descriptors; see above
STAGEstagingproduction | staging | local
FORCE_RENEWfalsetrue forces all certs to renew on next start
RENEW_MARGIN_DAYS30Renew this many days before expiry (Caddy renews around 1/3 lifetime by default; this is best-effort mapped)
MIGRATE_FROM_NGINXfalseOne-shot import of certs from a steveltn/https-portal volume layout
⁠Certificate algorithm
VariableDefaultMeaning
NUMBITS2048RSA key size; 4096 for stronger keys
CERTIFICATE_ALGORITHMrsaSet to prime256v1 to use ECDSA P-256 keys
⁠ACME challenges
VariableDefaultMeaning
DISABLE_TLS_ALPN_CHALLENGEfalseSet to true to force HTTP-01 only. Use when caddy-portal sits behind a TLS-terminating reverse proxy (Cloudflare orange-cloud, hosting-provider frontend, etc.) — see Running behind a TLS-terminating proxy⁠
⁠Routing
VariableDefaultMeaning
REDIRECT_CODE307Code used by => redirects; valid: 301, 302, 307, 308
INDEX_FILESindex.htmlSpace-separated list of index files for static sites
HSTS_MAX_AGE(unset)Emit Strict-Transport-Security: max-age=<N>
CLIENT_MAX_BODY_SIZE(unlimited)e.g. 20MB
GZIPonSet to off to disable response compression (Caddy emits zstd + gzip when enabled)
⁠Proxy & timeouts
VariableDefaultMeaning
KEEPALIVE_TIMEOUT(Caddy default)Server-side idle timeout
PROXY_CONNECT_TIMEOUT(Caddy default)Backend dial timeout
PROXY_SEND_TIMEOUT(Caddy default)Backend write timeout
PROXY_READ_TIMEOUT(Caddy default)Backend read timeout

Bare numbers are treated as seconds; Go duration strings (30s, 2m) work too.

⁠Logging
VariableDefaultMeaning
ACCESS_LOGoffoff | stdout | stderr | default (= /var/log/caddy/access.log) | custom path
ERROR_LOGstderrSame vocabulary as ACCESS_LOG, default path /var/log/caddy/error.log
ERROR_LOG_LEVELERRORCaddy log level: DEBUG, INFO, WARN, ERROR

Logs come out as JSON. Pipe them straight into your log aggregator.

⁠Custom Caddy snippets

For anything caddy-portal doesn't natively expose, three escape hatches:

VariableSpliced into
CUSTOM_CADDY_GLOBAL_BLOCKThe global { ... } options block
CUSTOM_CADDY_SERVER_BLOCKEvery site block
CUSTOM_CADDY_<HOST>_BLOCKOnly the matching site (example.com → CUSTOM_CADDY_EXAMPLE_COM_BLOCK)

The content is raw Caddyfile syntax — see caddyserver.com/docs/caddyfile⁠.

Example: add a security headers preset to every site:

environment:
  CUSTOM_CADDY_SERVER_BLOCK: |
    header {
      Referrer-Policy "strict-origin-when-cross-origin"
      Permissions-Policy "geolocation=(), microphone=()"
      X-Content-Type-Options "nosniff"
    }

⁠Auto-discovery

Mount the Docker socket and any container with a VIRTUAL_HOST env var gets auto-routed:

services:
  caddy-portal:
    image: datweazel/caddy-portal:1
    ports: ["80:80", "443:443", "443:443/udp"]
    environment:
      STAGE: production
    volumes:
      - caddy-portal-data:/var/lib/https-portal
      - /var/run/docker.sock:/var/run/docker.sock:ro    # ← discovery hook

  wordpress:
    image: wordpress
    environment:
      VIRTUAL_HOST: "blog.example.com"
      VIRTUAL_PORT: "80"

caddy-portal watches Docker events. When the wordpress container starts, a domain descriptor is written, the Caddyfile is re-rendered, and Caddy reloads atomically. Stop the container and the route is removed on the next reload.

VIRTUAL_HOST accepts the full descriptor syntax — including IP allow-lists and basic auth: VIRTUAL_HOST: "[10.0.0.0/8] admin@s3cret:blog.example.com".

⁠Running behind a TLS-terminating proxy

If something sits in front of caddy-portal that terminates TLS itself — a CDN like Cloudflare in proxy mode (orange cloud), a hosting-provider's frontend, a managed load balancer — then the TLS-ALPN-01 ACME challenge cannot reach caddy-portal. The proxy answers the TLS handshake with its own cert and never forwards the acme-tls/1 ALPN extension to origin.

You'll see this in the logs as:

"msg":"challenge failed","challenge_type":"tls-alpn-01"
"detail":"Cannot negotiate ALPN protocol \"acme-tls/1\" for tls-alpn-01 challenge"

Caddy will fall back to its next issuer (typically ZeroSSL or Google Trust Services), and one of them probably issues a cert via HTTP-01 instead. So in the short term it just works — you get a valid cert from a different CA than you expected, and the logs are noisy with LE failures.

The clean fix is to stop attempting TLS-ALPN-01 at all and have Caddy go straight for HTTP-01. Set DISABLE_TLS_ALPN_CHALLENGE=true:

services:
  caddy-portal:
    image: datweazel/caddy-portal:1
    environment:
      DOMAINS: "example.com -> http://app"
      STAGE: production
      DISABLE_TLS_ALPN_CHALLENGE: "true"

HTTP-01 works through TLS-terminating proxies because the validator hits port 80 and the proxy forwards the /.well-known/acme-challenge/ path to origin — Caddy serves the response, the proxy passes it back, validation succeeds.

What you see in the browser is the proxy's cert, not Caddy's. With Cloudflare in orange-cloud mode, the cert chain visible to end users comes from Cloudflare's edge (often issued by Google Trust Services or a different CA, regardless of what caddy-portal does). caddy-portal's own cert is only used between Cloudflare and your origin — and only if Cloudflare's SSL/TLS mode is set to "Full" or "Full (strict)".

Side effect of this flag: emitting issuer acme { ... } per-site replaces Caddy's default issuer fallback chain (Let's Encrypt → ZeroSSL → Google Trust Services) with a single Let's Encrypt issuer. If you want to keep fallback issuers AND disable TLS-ALPN-01, write the full issuer list yourself via CUSTOM_CADDY_SERVER_BLOCK.

If you'd rather not run behind a proxy at all and want LE to work directly, switching the Cloudflare DNS record from "Proxied" (orange) to "DNS only" (grey) makes TLS-ALPN-01 work again. You lose Cloudflare's caching, DDoS protection, and origin-IP hiding in exchange.

⁠Changing configuration at runtime

Two situations where caddy-portal doesn't need a restart.

⁠Backend container restarts (new Docker IP)

Caddy resolves upstream hostnames at request time, not at config load. When a container behind reverse_proxy app:80 is stopped and a new instance takes its place — even with a different Docker network IP — the next request re-resolves app against Docker's embedded DNS and connects to the new IP.

This works out of the box, no DYNAMIC_UPSTREAM or RESOLVER configuration required. The legacy nginx setup needed careful tuning here and silently broke when misconfigured; that whole class of bug is gone.

Verified end-to-end: stop a backend container, claim its IP with a temporary container so the backend gets reassigned a different IP on next start — existing connections fail over to the new IP without caddy-portal intervention.

⁠Adding, removing, or changing domains

Two patterns, depending on whether the new backend is a Docker container under your control.

Pattern A — VIRTUAL_HOST on the new container (preferred when applicable):

Mount the Docker socket once (see Auto-discovery⁠), then any neighbour container with a VIRTUAL_HOST env var is picked up automatically:

services:
  new-app:
    image: your/new-app
    environment:
      VIRTUAL_HOST: "new.example.com -> http://new-app:80"

docker compose up -d new-app → docker-gen sees the new container → caddy-portal re-renders the Caddyfile → Caddy reloads atomically. The new site is live in about one second.

Stop the container and the route is removed on the next reload, same loop.

Pattern B — write the new DOMAINS to dynamic-env:

Use this when the new backend isn't a Docker container, or when you want to change other settings (redirects, IP allow-lists, basic auth) without touching neighbouring containers. Bind-mount the dynamic-env directory:

services:
  caddy-portal:
    # ...
    volumes:
      - caddy-portal-data:/var/lib/https-portal
      - ./caddy-portal-env:/var/lib/https-portal/dynamic-env

Then on the host, write the complete new DOMAINS list to a file named DOMAINS:

echo "example.com -> http://app:80, new.example.com -> http://newapp:80" \
  > ./caddy-portal-env/DOMAINS

caddy-portal sees the file change (debounced ~1s via fs.watch), re-reads the merged environment, re-renders the Caddyfile, and runs caddy reload. The new site is live without dropping existing connections.

Important: a file in dynamic-env/ replaces the corresponding env var, it doesn't append. To add a domain you must write the full list including any pre-existing ones. Same semantics apply to every other env var overlaid this way.

⁠Live-tuning any other env var

The same mechanism works for every documented env var. Drop a file named after the variable, its contents become the new value, caddy-portal reloads within a second:

echo "120"     > ./caddy-portal-env/KEEPALIVE_TIMEOUT
echo "50MB"    > ./caddy-portal-env/CLIENT_MAX_BODY_SIZE
echo "stdout"  > ./caddy-portal-env/ACCESS_LOG
echo "31536000" > ./caddy-portal-env/HSTS_MAX_AGE

Filenames must be UPPERCASE env-var-style identifiers (matching ^[A-Z][A-Z0-9_]*$). Lowercase or non-conforming filenames are ignored — useful if you want to leave a README.txt in the directory.

Settings that affect cert issuance (STAGE, CERTIFICATE_ALGORITHM, NUMBITS) can also be changed live, but renewal happens lazily — Caddy won't rotate existing certs until they near expiry unless you also set FORCE_RENEW=true in the same dynamic-env change.

⁠Volume layout

Everything that needs to survive container restarts lives under /var/lib/https-portal:

/var/lib/https-portal/
├── caddy/                          ← Caddy's own data dir
│   ├── certificates/<acme-ca>/<domain>/   issued certs + metadata
│   ├── locks/                              locking primitives during renewal
│   └── ocsp/                               cached OCSP staples
├── dynamic-env/                    ← write env-var overrides here for live reload
│   └── <ENV_VAR_NAME>              file contents become the env value
├── .migrated-from-nginx            ← migration marker
└── <domain>/<stage>/...            ← legacy https-portal layout, only present after migration

For details on writing into dynamic-env/ to drive live reloads, see Changing configuration at runtime⁠.

⁠Coming from steveltn/https-portal

See docs/migration.md⁠ for the full guide. Short version:

  1. Change the image to datweazel/caddy-portal:1
  2. Add MIGRATE_FROM_NGINX: "true" to the env (one-shot)
  3. Add "443:443/udp" to ports (optional, enables HTTP/3)
  4. docker compose up -d
  5. Remove the migration flag after the first successful start

If you rely on CUSTOM_NGINX_* or bind-mounted .conf.erb files, caddy-portal is not for you — stay on steveltn/https-portal.

⁠Troubleshooting

⁠My site returns a 502

The reverse-proxy backend is unreachable. Check that the upstream container name resolves inside the Docker network (docker compose exec caddy-portal wget -O- http://your-backend:80) and that the backend is actually listening.

⁠Cert obtained from Let's Encrypt staging but browser doesn't trust it

Expected on STAGE: staging. The staging environment exists exactly to test without hitting production rate limits. Switch to STAGE: production once the staging cert works.

⁠Logs spam Cannot negotiate ALPN protocol "acme-tls/1"

caddy-portal is behind a TLS-terminating proxy that intercepts the TLS-ALPN-01 challenge. Set DISABLE_TLS_ALPN_CHALLENGE=true — full explanation in Running behind a TLS-terminating proxy⁠.

⁠My cert was issued by Google Trust Services / ZeroSSL instead of Let's Encrypt

Caddy's default issuer fallback chain. When the first issuer (Let's Encrypt) fails, Caddy automatically tries the next one. Often correlated with the ALPN issue above — fixing TLS-ALPN-01 access usually brings LE back. See Running behind a TLS-terminating proxy⁠ for the typical root cause.

⁠"no Docker socket mounted; auto-discovery disabled"

You're using VIRTUAL_HOST but didn't mount the socket. Add:

volumes:
  - /var/run/docker.sock:/var/run/docker.sock:ro
⁠HTTP/3 doesn't seem to be used

Browsers only switch to HTTP/3 after the server advertises it (via the alt-svc header on the first HTTP/2 response) and the client decides to upgrade. Confirm:

curl -sI https://your-domain.example/ | grep -i alt-svc
# alt-svc: h3=":443"; ma=2592000

If alt-svc is missing, you probably didn't expose 443:443/udp — Caddy emits the header only when the UDP listener is actually bound.

⁠Backend no longer receives some request headers, or clients get 431

Newer Caddy releases (2.11.4 and later) harden request handling, which can surface as breakage after updating caddy-portal:

  • Header fields containing _ or . (e.g. X_Api_Key, X.Trace) are dropped before they reach the backend.
  • Request headers are limited to 16 KiB. Larger ones (giant cookies, oversized tokens) get 431 Request Header Fields Too Large over HTTP/1.1; over HTTP/2 the client sees a protocol error instead.
  • Stalled request bodies or responses are cut off after 1 minute without any progress. Pauses between writes (e.g. SSE) don't count.

Each of these can be relaxed via CUSTOM_CADDY_GLOBAL_BLOCK. Only list what you actually need:

environment:
  CUSTOM_CADDY_GLOBAL_BLOCK: |
    servers {
      expected_underscore_headers X_Api_Key
      expected_dot_headers X.Trace
      max_header_size 64KB
      timeouts {
        read_body_idle 5m
        write_idle 5m
      }
    }

If you also set KEEPALIVE_TIMEOUT, caddy-portal already emits a servers block, and Caddy rejects a second one: the container exits right after startup. In that case unset KEEPALIVE_TIMEOUT and add idle <duration> to the timeouts block above instead.

⁠Container exits right after startup without a clear error

Caddy 2.11.6 and later don't print config errors when caddy run fails (caddyserver/caddy#7962⁠). To work around this, caddy-portal runs caddy validate once after a failed start. Look for the line [caddy-portal] caddy exited with code 1; running 'caddy validate' to surface the error: in the log. The actual error follows it, e.g. unrecognized global option: ....

If it reports config is valid instead, the config is fine and Caddy failed for a different reason (e.g. a port already in use).

⁠[caddy-portal] CUSTOM_NGINX_* will be ignored

Expected. See docs/migration.md#not-supported⁠ for the migration story.

⁠Reload didn't pick up my change

caddy-portal watches /var/run/domains and /var/lib/https-portal/dynamic-env with a 1-second debounce. If you're editing env vars on the host's compose file, those don't propagate into the container at runtime — for that, see Changing configuration at runtime⁠ or restart the container.

⁠License

MIT. See LICENSE⁠ and NOTICE⁠.

This project preserves the public configuration surface of steveltn/https-portal⁠ but is implemented from scratch in TypeScript on top of Caddy⁠.

Tag summary

Content type

Image

Digest

sha256:31860ea35…

Size

55.8 MB

Last updated

1 day ago

docker pull datweazel/caddy-portal