Sign inSign up

classifyre/all-in-one

By classifyre

•Updated about 12 hours ago

Classifyre All-in-One Docker image for local scanning and rapid prototyping.

Buildkit cache
Image
Machine learning & AI
Data science
Web servers
0

50K+

classifyre/all-in-one repository overview

Classifyre

⁠Classifyre All-in-One

Open-source investigation platform in a single Docker image — connect a source, detect leaked secrets, PII and security risks, and work findings into cases.

Classifyre turns the data scattered across the systems you already run — file shares, Confluence, Jira, SharePoint, S3, Git repositories, databases — into investigations you can act on. Detectors surface the evidence (leaked secrets and API keys, PII, security misconfigurations), and you work the results like an analyst: standing inquiries, duplicate detection, cases and hypotheses, with an AI autopilot doing the legwork in between.

This classifyre/all-in-one image is the fastest way to try Classifyre: the web UI, the API, the background worker, the Python scan workers, and a PostgreSQL database with pgvector — all in one container. One docker run, no dependencies to provision, nothing to connect. Ideal for local scanning, evaluations, and rapid prototyping.

For a team, or for anything that has to survive a machine dying, use the Helm chart⁠ (classifyre/classifyre-core) instead — it runs the same code as separate, independently scalable workloads.

📚 Full documentation: https://docs.classifyre.com⁠ — start with Deployment⁠ · Sources⁠ · Detectors⁠ · Investigations⁠


⁠Quickstart

docker run -d --name classifyre \
  -p 3000:3000 \
  --shm-size=1g \
  -v classifyre-pgdata:/var/lib/postgresql/data \
  -v classifyre-data:/var/lib/classifyre \
  -v classifyre-uv-cache:/cache/uv \
  classifyre/all-in-one:latest

Then open http://localhost:3000⁠.

The first boot initialises the database, applies the migrations and creates a workspace called default. On a laptop that takes a few minutes; watch it with docker logs -f classifyre.

Nothing leaves your machine. There is no account, no signup, and no telemetry you have not opted into.

What to do next: add a source (Sources → New), turn on detectors, run a scan, then work the findings into cases. See How it works⁠ and Flow⁠.

⁠Requirements

  • Memory: 4 GB minimum, 8 GB recommended. The container refuses to start below 4 GB.
  • CPU: 2 cores minimum (scans are CPU-bound).
  • Disk: ~10 GB for the image and database, plus room for what you scan.
  • Shared memory: always pass --shm-size=1g (PostgreSQL parallel queries need it).
  • On macOS/Windows, raise Docker Desktop to 4+ GB under Settings → Resources → Memory.
  • Architectures: linux/amd64 and linux/arm64 (Apple Silicon included).

⁠What is inside

ProcessPortPurpose
Caddy3000 (the only published port)Routes to the UI and API
Web UI3100 (internal)Next.js app + bundled docs at /docs
API + worker8000 (internal)REST, WebSockets, job queues, embeddings, scan orchestration
PostgreSQL 18 + pgvector5432 (internal, not exposed)All application data

⁠Scanning a folder on your machine

docker run -d --name classifyre \
  -p 3000:3000 --shm-size=1g \
  -v classifyre-pgdata:/var/lib/postgresql/data \
  -v classifyre-data:/var/lib/classifyre \
  -v classifyre-uv-cache:/cache/uv \
  -v "$HOME/Documents/case-files:/data/case-files:ro" \
  classifyre/all-in-one:latest

Then in the UI: Sources → New → Mounted Folder, path /data/case-files.

⁠Configuration

Everything is an environment variable (-e NAME=value or --env-file), and everything has a working default.

⁠Core
VariableDefaultPurpose
PORT_HTTP3000Port inside the container (or remap with -p 8080:3000).
CLASSIFYRE_BOOTSTRAP_NAMESPACEdefaultSlug of the workspace created on first boot. Set to "" to create your own in the UI.
CLASSIFYRE_MASKED_CONFIG_KEYauto-generatedEncrypts stored source credentials and MCP tokens. Back it up (see below) — without it, restored backups keep credentials encrypted and unreadable.
DEMO_MODEfalseRead-only mode; blocks every write. For shared demo instances.
TELEMETRY_DISABLED / DO_NOT_TRACKunsetSet either to 1 to disable anonymous usage telemetry.
CORS_ORIGINsame-originOnly needed if you serve the UI from another host.
⁠Database

Leave these unset to use the PostgreSQL server inside the image.

VariableDefaultPurpose
DATABASE_URLbundled serverSet it and the bundled server never starts (must have pgvector; user needs CREATE SCHEMA / CREATE EXTENSION).
POSTGRES_DBclassifyreDatabase name (bundled server only).
PGDATA/var/lib/postgresql/dataData directory (bundled server only).
CLASSIFYRE_AUTO_MIGRATEtruefalse skips migrations at startup (only if you run them yourself).
⁠Object storage (S3, MinIO, R2, …)

Scan logs go to disk by default; set S3_BUCKET to store them in object storage instead.

VariableDefaultPurpose
S3_BUCKETunsetSetting this switches scan logs from disk to S3.
S3_ENDPOINTAWSRequired for MinIO, R2, Backblaze, Garage, any non-AWS provider.
S3_REGIONus-east-1Bucket region.
S3_ACCESS_KEY_ID / S3_SECRET_ACCESS_KEYunsetOmit both to use the ambient credential chain.
S3_FORCE_PATH_STYLEtrueRequired by MinIO and most self-hosted providers; false for AWS.
S3_LOG_PREFIXrunner-logs/Key prefix for scan logs.
⁠Embeddings

A small embedding model (Xenova/all-MiniLM-L6-v2, 384 dimensions) ships baked into the image and runs on CPU, so semantic search and duplicate detection work offline out of the box.

VariableDefaultPurpose
EMBEDDING_PROVIDERtransformers-jsopenai-compatible to use a remote service instead.
EMBEDDING_BASE_URL / EMBEDDING_API_KEYunsetRemote endpoint for openai-compatible.
EMBEDDING_MODELXenova/all-MiniLM-L6-v2Changing this re-embeds everything.
EMBEDDING_ALLOW_REMOTE_MODELSfalsetrue allows downloading a different local model at runtime.
EMBEDDING_BATCH_SIZEautotunedLower it if embedding starves your scans.

AI providers (LLM detectors, investigation autopilot) are not environment variables — they are per-workspace settings under Settings → AI providers in the UI, with API keys encrypted at rest.

⁠Resources

The container sizes itself to the memory/CPU it was given. Override only with reason.

VariablePurpose
CLASSIFYRE_NODE_HEAP_MBNode heap ceiling. Bigger is not better past ~2.5 GB.
PG_SHARED_BUFFERS, PG_WORK_MEM, PG_EFFECTIVE_CACHE_SIZE, PG_MAX_CONNECTIONSPostgreSQL memory tuning.
MAX_CONCURRENT_RUNNERSScans running at once.
CLASSIFYRE_MAX_POOL_WORKERSDetector worker processes per scan.
CLASSIFYRE_MIN_MEMORY_MBThe 4 GB startup floor. Lower at your own risk.

⁠Upgrades & backup

docker pull classifyre/all-in-one:latest
docker rm -f classifyre
# re-run the same docker run as before, with the same volumes

Data lives in the volumes, so it carries over; migrations run automatically. Pin a version tag to decide when to upgrade. Downgrades are not supported.

Back up two things: the database (docker exec classifyre pg_dump -U postgres classifyre > classifyre.sql) and the key (docker exec classifyre cat /var/lib/classifyre/masked-config.key).

⁠Moving to Kubernetes

Nothing here is a dead end: export a workspace from this container and import it into a cluster install — sources, findings, cases and lineage all move. The Helm chart⁠ (classifyre/classifyre-core) runs the same images this one is composed from (classifyre/api, classifyre/web, classifyre/cli).


Classifyre is open source: https://github.com/classifyre/classifyre⁠ · Docs: https://docs.classifyre.com⁠

Tag summary

Content type

Image

Digest

sha256:9c45486bd…

Size

1.2 GB

Last updated

about 12 hours ago

docker pull classifyre/all-in-one