Sign inSign up

n8500x/pii-service

By n8500x

•Updated 7 months ago

PII detection, masking, and anonymization HTTP service with a bundled nginx TLS front end (HTTPS var

Image
0

2.3K

n8500x/pii-service repository overview

⁠pii-service

PII detection, masking, and anonymization HTTP service with a bundled nginx TLS front end (HTTPS variant).

A self-contained, air-gap-friendly service that detects and redacts personally identifiable information (PII) in text, and optionally proxies masked prompts to a configurable LLM backend and unmasks the response. This image builds on top of the FastAPI app with nginx and a self-signed TLS certificate baked in for HTTPS termination. For a plain-HTTP build with no nginx/TLS, see n8500x/pii-service-simple⁠.

⁠Overview

The service exposes a FastAPI application (pii_service.main:app) that:

  • Detects PII using GLiNER (NER), spaCy en_core_web_lg, and Microsoft Presidio.
  • Masks / unmasks PII with configurable patterns and session-based mapping.
  • Reports PII entity counts and confidence.
  • LLM integration: masks a prompt, sends it to a configurable LLM backend, and unmasks the response.

Detected entity types: PERSON, LOCATION, ORGANIZATION, EMAIL_ADDRESS, PHONE_NUMBER, DATE_TIME, CREDIT_CARD, IBAN_CODE, IP_ADDRESS, URL, DOMAIN.

All models (GLiNER urchade/gliner_base, microsoft/deberta-v3-base, spaCy, NLTK punkt_tab, tldextract Public Suffix List) are downloaded at build time and baked into the image, and the container runs fully offline (HF_HUB_OFFLINE=1, TRANSFORMERS_OFFLINE=1).

⁠Quick start

docker pull n8500x/pii-service

# FastAPI app on 8080; nginx HTTPS (8443) and HTTP proxy (18080) are bundled
docker run -d --name pii-service \
  -p 8080:8080 \
  -p 8443:8443 \
  -p 18080:18080 \
  -e LLM_ENDPOINT_URL=http://host.docker.internal:8000/v1 \
  -e LLM_MODEL_NAME=gpt-4 \
  n8500x/pii-service

A self-signed TLS certificate (CN=localhost, valid 10 years) is generated at build time and baked into /etc/nginx/certs, so no certificate mount is required. To use your own certificate, mount over that path:

docker run -d -p 8443:8443 \
  -v /path/to/certs:/etc/nginx/certs \
  n8500x/pii-service

Health check:

curl http://localhost:8080/health

⁠API

Interactive docs are served at /docs (Swagger UI), /redoc, and /openapi.json.

MethodEndpointPurpose
GET/healthService health check
GET/configCurrent service configuration
POST/analyzeDetect PII entities in text
POST/maskMask PII entities
POST/unmaskRestore masked PII
POST/llmMask prompt → query LLM → unmask response
POST/pii-reportPII entity counts / confidence report
GET/mappingRetrieve mask mappings
POST/fuzzyFuzzy match PII
GET/pii-typesList supported PII types
GET/stateSystem state

Analyze example:

curl -X POST http://localhost:8080/analyze \
  -H "Content-Type: application/json" \
  -d '{"text": "My email is [email protected] and phone is +1234567890"}'

Mask example:

curl -X POST http://localhost:8080/mask \
  -H "Content-Type: application/json" \
  -d '{"text": "My email is [email protected]"}'
# -> {"masked_text": "My email is <EMAIL_ADDRESS>", ...}

LLM with PII protection:

curl -X POST http://localhost:8080/llm \
  -H "Content-Type: application/json" \
  -d '{"prompt": "Analyze this email: [email protected]", "mask_pii": true, "unmask_response": true}'

⁠Configuration

Configured via environment variables (a .env is created from .env.example at build time):

VariableDescription
LLM_ENDPOINT_URLLLM backend URL (OpenAI-compatible or generic chat)
LLM_MODEL_NAMEModel name (e.g. gpt-4, mistral)
LLM_API_KEYBearer token (empty for local backends)
LLM_TEMPERATURESampling temperature (default 0.7)
LLM_MAX_TOKENSMax response tokens (default 2048)
GLINER_MODEL_NAMEGLiNER model (gliner_base)
GLINER_THRESHOLDConfidence threshold (default 0.5; lower = more sensitive)
PRESIDIO_USE_GLINEREnable GLiNER recognizer (true/false)
PRESIDIO_USE_NLPEnable spaCy NLP recognizer (true/false)

Ports exposed:

  • 8080 — FastAPI app (plain HTTP, direct)
  • 8081 — custom app internal (no TLS)
  • 8443 — HTTPS PII service (nginx, self-signed cert)
  • 8444 — HTTPS custom app
  • 18080 — HTTP nginx proxy to the FastAPI app (targeted by an OpenShift Route)
  • 8501 — Streamlit (optional)

Volumes: mount /etc/nginx/certs to supply your own TLS certificate/key (optional).

⁠Build

docker build -f Dockerfile.https -t n8500x/pii-service .

The build downloads and bakes in all NLP models, runs a 13-step validation suite, and generates the self-signed TLS certificate.

⁠Notes

  • Offline by design — no HuggingFace Hub, PyPI, or Public Suffix List fetches at runtime; suitable for air-gapped deployments. The image is large (full ML/NLP stack: GLiNER, DeBERTa-v3, spaCy en_core_web_lg, Presidio).
  • Self-signed certificate — the baked-in cert (CN=localhost) will trigger browser/client TLS warnings; supply a trusted cert for production.
  • OpenShift-ready — nginx is patched to non-privileged ports (443→8443, 8443→8444), runs temp/PID files under /tmp, and directories are group-0 writable for arbitrary UIDs. In an OpenShift Route deployment, the Route terminates TLS and forwards plain HTTP to the app (port 8080) or the nginx HTTP proxy (18080).
  • Pydantic v1 — the app pins Pydantic v1.x (BaseSettings built in).
  • The container startup runs internal validation and a test suite before the service stays up; check docker logs if it exits early.

Tag summary

Content type

Image

Digest

sha256:c9b2a4403…

Size

14.2 GB

Last updated

7 months ago

docker pull n8500x/pii-service