A container monitoring and auto-healing service with a web dashboard, REST API, and Prometheus metrics. Docker Auto-Heal watches your containers and restarts the ones that fail or go unhealthy, with cooldowns, exponential backoff, and automatic quarantine for containers that keep failing.
Full documentation, including source, is at github.com/TommyE123/docker-autoheal.
Images are published to both Docker Hub and GitHub Container Registry (GHCR) — use whichever you prefer.
Note: GHCR packages are private by default when first published. If
docker pull ghcr.io/tommye123/docker-autohealfails with an access/authentication error, the package hasn't been switched to public yet — use the Docker Hub image below in the meantime, ordocker login ghcr.iowith a token that has read access.
Docker Hub:
docker run -d \
--name docker-autoheal \
-v /var/run/docker.sock:/var/run/docker.sock:ro \
-v ./data:/data \
-p 3131:3131 \
-p 9090:9090 \
--restart unless-stopped \
tommye123/docker-autoheal:latest
GHCR:
docker run -d \
--name docker-autoheal \
-v /var/run/docker.sock:/var/run/docker.sock:ro \
-v ./data:/data \
-p 3131:3131 \
-p 9090:9090 \
--restart unless-stopped \
ghcr.io/tommye123/docker-autoheal:latest
Web UI: http://localhost:3131
/data, editable through the web UI, the REST
API, or config.json directly — there are no environment-variable settingsservices:
autoheal:
image: tommye123/docker-autoheal:latest # or ghcr.io/tommye123/docker-autoheal:latest
container_name: docker-autoheal
restart: unless-stopped
volumes:
- /var/run/docker.sock:/var/run/docker.sock:ro
- ./data:/data
ports:
- "3131:3131" # Web UI
- "9090:9090" # Prometheus metrics
labels:
- "autoheal=false" # keeps the monitor unselected under the default label-based
# selection; see docs/user/labels.md if you enable include_all
webapp:
image: nginx:alpine
labels:
autoheal: "true" # monitor this container
healthcheck:
test: ["CMD", "wget", "--quiet", "--tries=1", "--spider", "http://localhost"]
interval: 30s
timeout: 10s
retries: 3
docker compose up -d
Any container labelled autoheal=true is picked up automatically — both containers
already running when Auto-Heal starts, and new ones as they start.
All settings are managed through the web UI at http://localhost:3131 (Configuration
tab) or the /api/config* REST endpoints — see the "Configuration" section of the
full documentation
for the complete field reference.
Configuration is automatically persisted to /data/config.json. Export/import as JSON is
available from the web UI.
# Prometheus metrics
curl http://localhost:9090/metrics
# Service health
curl http://localhost:3131/health
# Interactive API docs (Swagger UI)
http://localhost:3131/docs
docker logs -f docker-autoheal
/var/run/docker.sock)Container not being monitored?
autoheal=true label (unless "monitor all containers" is enabled).docker logs docker-autoheal (set log level to DEBUG from the
Configuration tab for more detail).Won't start? Verify the Docker socket is accessible and ports 3131/9090 are free.
Container quarantined? It exceeded the configured restart threshold. View it in the
Web UI, fix the underlying issue, and it will auto-unquarantine once healthy — or
unquarantine it manually from the UI or POST /api/containers/{id}/unquarantine.
Full troubleshooting guide: docs/user/troubleshooting.md
This service requires read-only Docker socket access, which is a significant privilege. Recommended for production:
:ro, as shown above)MIT — see the LICENSE file.
Content type
Image
Digest
sha256:eaf11d1cb…
Size
67.4 MB
Last updated
4 days ago
docker pull tommye123/docker-autoheal