Typed decisions over any state in one forward pass. CPU-only sidecar for Laya engine (Apache-2.0)
377
Typed decisions over any state in one forward pass — an HTTP sidecar for the Laya decision engine (Apache-2.0, 27k★).
~0.2–0.6 s per decision · ~170 calls/min · CPU-only · stateless · Apache-2.0
The service is stateless and policy-free: it only answers typed questions
(choice / score / noul) over any state you send. What to do with an answer —
auto-act, log, escalate — always lives in your caller.
400s)Built and maintained by CodeCora — pragmatic AI infrastructure.
docker run -d --name laya \
-p 8080:8080 \
-v laya-hf-cache:/data/hf \
--memory 4g \
codecoradev/laya-service:latest
Weights (~644 MB) download from Hugging Face on first start and are cached in the
volume. Watch readiness — status: "ok" means the model is loaded:
curl http://localhost:8080/healthz
Environment: LAYA_CHECKPOINT (default multilingual), LAYA_DEVICE (default cpu),
AUTO_THRESHOLD (default 0.85). Multi-arch: amd64 + arm64.
curl -X POST http://localhost:8080/api/v1/decide \
-H 'Content-Type: application/json' \
-d '{
"state": {"body": "I was charged twice, refund now or I cancel!"},
"questions": {
"intent": {"type": "choice", "instructions": "What does the customer want?",
"criteria": {"refund": "money back", "technical": "bug", "shipping": "delivery"}},
"urgency": {"type": "score", "instructions": "How urgent?",
"criteria": ["routine", "soon", "critical"]},
"escalate": {"type": "noul", "instructions": "Needs a human now?"}
}
}'
Response (captured live):
{
"model": "multilingual",
"device": "cpu",
"latency_ms": 564.9,
"auto_threshold": 0.85,
"answers": {
"intent": {"type": "choice", "choice": "refund",
"probabilities": {"refund": 0.9999, "technical": 0.0, "shipping": 0.0001},
"confidence": 0.999, "action": {"act_probability": 1.0}},
"urgency": {"type": "score", "score": 1.9583,
"legend": {"0": "routine", "1": "soon", "2": "critical"},
"probabilities": {"0": 0.0006, "1": 0.0274, "2": 0.972},
"confidence": 0.8811, "action": {"act_probability": 1.0}},
"escalate": {"type": "noul", "noul": 0.9775, "confidence": 0.9775,
"action": {"act_probability": 1.0}}
}
}
| Type | criteria | Answer |
|---|---|---|
choice | object with 2–20 options | {"choice": "<key>", "probabilities": {…}, "confidence": 0–1} |
score | array of 2–10 ordered levels | {"score": <float>, "legend": {…}, "confidence": 0–1} |
noul | optional {"false": …, "true": …} labels | {"noul": <float 0–1>, "confidence": 0–1} |
state accepts a plain string, an email/ticket-like object, or a JSON document —
passed to the engine untouched. Keep it short: context is 1024 tokens (a ~2.5k-token
state takes ~17 s vs ~0.3 s at normal sizes).
| Route | Method | Notes |
|---|---|---|
/healthz | GET | Liveness + readiness ("status": "loading" while weights load) — open |
/api/v1/decide | POST | Typed decisions (canonical) |
/v1/decide | POST | Legacy alias, same contract |
/docs, /redoc, /openapi.json | GET | FastAPI API browser |
Errors are Jev-style: {"error": "…", "detail": "…"}.
400 schema/unknown-model · 401 wrong bearer (if you add auth) · 500 prediction
failure · 503 model still loading.
| Metric | Value |
|---|---|
| Latency, CPU | 235–600 ms per call |
| Throughput | ~170 calls/min serial |
| RSS (multilingual, fp32) | ~2.1 GB |
| Cold start | container boot + weight load (first start adds the HF download) |
confidence is calibrated (RLCD) — treat it as meaningful. Sweet spots: CS/ticket
triage (intent + urgency + needs-human), event triage, spam/phishing gates, small
model routing. Not for: long-context reasoning, >20-way choices, content moderation.
| Tag | Meaning |
|---|---|
latest | most recent build |
develop | tracks the develop branch head |
2026-09-28d | dated, pin these for stable deployments |
Apache-2.0. Laya is Apache-2.0 by Convai Innovations — this service only wraps it.
Content type
Image
Digest
sha256:9b44ac6e7…
Size
336.9 MB
Last updated
1 day ago
docker pull codecoradev/laya-service