Sign inSign up

codecoradev/laya-service

By codecoradev

•Updated 1 day ago

Typed decisions over any state in one forward pass. CPU-only sidecar for Laya engine (Apache-2.0)

Image
0

377

codecoradev/laya-service repository overview

⁠laya-service

Typed decisions over any state in one forward pass — an HTTP sidecar for the Laya⁠ decision engine (Apache-2.0, 27k★).

~0.2–0.6 s per decision · ~170 calls/min · CPU-only · stateless · Apache-2.0

The service is stateless and policy-free: it only answers typed questions (choice / score / noul) over any state you send. What to do with an answer — auto-act, log, escalate — always lives in your caller.

  • Typed answers, not free text — request schemas are validated before the model is touched (fail-fast 400s)
  • Calibrated confidence (RLCD) — common policy: auto-act at ≥ 0.85, escalate below
  • Nothing shared ships inside the image — the service is fail-closed; you bring your own auth at the edge

Built and maintained by CodeCora⁠ — pragmatic AI infrastructure.

⁠Quick start (~30 s)

docker run -d --name laya \
  -p 8080:8080 \
  -v laya-hf-cache:/data/hf \
  --memory 4g \
  codecoradev/laya-service:latest

Weights (~644 MB) download from Hugging Face on first start and are cached in the volume. Watch readiness — status: "ok" means the model is loaded:

curl http://localhost:8080/healthz

Environment: LAYA_CHECKPOINT (default multilingual), LAYA_DEVICE (default cpu), AUTO_THRESHOLD (default 0.85). Multi-arch: amd64 + arm64.

⁠Ask a question

curl -X POST http://localhost:8080/api/v1/decide \
  -H 'Content-Type: application/json' \
  -d '{
    "state": {"body": "I was charged twice, refund now or I cancel!"},
    "questions": {
      "intent":   {"type": "choice", "instructions": "What does the customer want?",
                   "criteria": {"refund": "money back", "technical": "bug", "shipping": "delivery"}},
      "urgency":  {"type": "score",  "instructions": "How urgent?",
                   "criteria": ["routine", "soon", "critical"]},
      "escalate": {"type": "noul",   "instructions": "Needs a human now?"}
    }
  }'

Response (captured live):

{
  "model": "multilingual",
  "device": "cpu",
  "latency_ms": 564.9,
  "auto_threshold": 0.85,
  "answers": {
    "intent": {"type": "choice", "choice": "refund",
               "probabilities": {"refund": 0.9999, "technical": 0.0, "shipping": 0.0001},
               "confidence": 0.999, "action": {"act_probability": 1.0}},
    "urgency": {"type": "score", "score": 1.9583,
                "legend": {"0": "routine", "1": "soon", "2": "critical"},
                "probabilities": {"0": 0.0006, "1": 0.0274, "2": 0.972},
                "confidence": 0.8811, "action": {"act_probability": 1.0}},
    "escalate": {"type": "noul", "noul": 0.9775, "confidence": 0.9775,
                 "action": {"act_probability": 1.0}}
  }
}

⁠Question types

TypecriteriaAnswer
choiceobject with 2–20 options{"choice": "<key>", "probabilities": {…}, "confidence": 0–1}
scorearray of 2–10 ordered levels{"score": <float>, "legend": {…}, "confidence": 0–1}
nouloptional {"false": …, "true": …} labels{"noul": <float 0–1>, "confidence": 0–1}

state accepts a plain string, an email/ticket-like object, or a JSON document — passed to the engine untouched. Keep it short: context is 1024 tokens (a ~2.5k-token state takes ~17 s vs ~0.3 s at normal sizes).

⁠Endpoints & errors

RouteMethodNotes
/healthzGETLiveness + readiness ("status": "loading" while weights load) — open
/api/v1/decidePOSTTyped decisions (canonical)
/v1/decidePOSTLegacy alias, same contract
/docs, /redoc, /openapi.jsonGETFastAPI API browser

Errors are Jev-style: {"error": "…", "detail": "…"}. 400 schema/unknown-model · 401 wrong bearer (if you add auth) · 500 prediction failure · 503 model still loading.

⁠Measured performance (2026-09, real E2E)

MetricValue
Latency, CPU235–600 ms per call
Throughput~170 calls/min serial
RSS (multilingual, fp32)~2.1 GB
Cold startcontainer boot + weight load (first start adds the HF download)

confidence is calibrated (RLCD) — treat it as meaningful. Sweet spots: CS/ticket triage (intent + urgency + needs-human), event triage, spam/phishing gates, small model routing. Not for: long-context reasoning, >20-way choices, content moderation.

⁠Tags

TagMeaning
latestmost recent build
developtracks the develop branch head
2026-09-28ddated, pin these for stable deployments

Apache-2.0. Laya is Apache-2.0 by Convai Innovations — this service only wraps it.

Tag summary

Content type

Image

Digest

sha256:9b44ac6e7…

Size

336.9 MB

Last updated

1 day ago

docker pull codecoradev/laya-service