Skip to content

HTTP API

Install pip install 'gemmadecision[serve]', then start gemmadecision serve and use http://127.0.0.1:8700. The default runtime is CPU ONNX. The service exposes decision and ranking endpoints backed by one resident model and a bounded batching queue. It is not an OpenAI chat-completions endpoint.

Routes

Method Path Purpose API key checked?
POST /v1/systemone Answer named choice, yes/no, and rubric questions. Yes, when configured.
POST /v1/rank Rank supplied candidates. Yes, when configured.
GET /health Readiness, model/backend metadata, and queue counters. No.
GET /ready Same response and readiness check as /health. No.
GET /v1/models Served model ID and typed-decisions task. No.
GET /docs, /redoc, /openapi.json Generated interactive documentation and schema. No.

When GEMMADECISION_API_KEY is set, send Authorization: Bearer <your-key> on inference requests. The CLI requires a key for a non-loopback bind unless --allow-unauthenticated is explicitly supplied. Health and schema endpoints remain public.

POST /v1/systemone

{
  "state": "The customer reports a duplicate charge and requests help today.",
  "questions": {
    "team": {
      "type": "choice",
      "instructions": "Choose the support team.",
      "criteria": {
        "billing": "Help with payments and duplicate charges.",
        "technical": "Help with product errors."
      }
    },
    "urgent": {
      "type": "noul",
      "instructions": "The customer explicitly requests help today."
    },
    "detail": {
      "type": "score",
      "instructions": "How detailed is this request?",
      "criteria": ["Too little detail", "Problem identified", "Problem and evidence supplied"]
    }
  }
}
Request field Type Required Meaning
state JSON value Yes Text or structured material to evaluate.
questions Object Yes 1–32 named question objects; their type selects the schema.
model String No Served model ID or an accepted alias.

instructions is required in each question and accepts a JSON value. Choice criteria are a 2–64 entry mapping; score criteria are a 2–64 element ordered list. Noul criteria are optional descriptions under false/true, with no/yes accepted as input aliases. Do not provide both aliases for one outcome. The normalized Noul answer uses false and true.

Response field Meaning
answers Mapping with exactly the requested question names.
Choice answer type: "choice", choice, probabilities, confidence, scores.
Noul answer type: "noul", noul (weight of true), probabilities, confidence, scores.
Score answer type: "score", score (expected zero-based level), probabilities, confidence, scores.
usage input_tokens, output_tokens, total_tokens, and cached_pairs.
model Full served model ID including immutable revision.
temperature Fixed value 4.136820402388508.
probabilities_source softmax_ranking_scores_external_temperature.
latency_ms Server-side elapsed inference handling time, including queueing.

Choice returns the largest score's label. Noul returns a float from 0 to 1, not a JSON boolean. Score returns sum(i * probabilities[str(i)]), which can lie between levels. All wire confidence fields are the largest probability, including when the preferred Noul outcome is false.

POST /v1/rank

{
  "state": "I cannot sign in to my account.",
  "question": "Which team should handle this?",
  "candidates": {
    "account": "Help with passwords and account access.",
    "billing": "Help with payments and invoices."
  }
}
Request field Type Default
state JSON value; context is an accepted input alias. Required.
candidates Label-to-string mapping or list of strings; answers is an input alias. Required, 2–64 candidates.
question JSON value Empty string.
model String Pinned served model ID.

The response has ranked, a list ordered from highest to lowest raw score. Each item contains rank (1-based), candidate (label), text, score, and probability. List input receives string-index labels. Ties preserve input order. The response also carries the same usage and probability-provenance fields as /v1/systemone.

For example, save a request as request.json and submit it with:

curl --fail-with-body http://127.0.0.1:8700/v1/rank \
  -H 'Content-Type: application/json' \
  -H "Authorization: Bearer $GEMMADECISION_API_KEY" \
  --data-binary @request.json

Model identity

The default ID is rajan2k/GemmaDecision-270M@785d530221c990671f29976902540101bb9c7647. Accepted aliases are rajan2k/GemmaDecision-270M, GemmaDecision-270M, and gemmadecision. Supplying another model ID fails validation; it does not download a different model. /v1/models returns a models list, not an OpenAI data response envelope.

Status codes and operational limits

Status Meaning
200 Successful inference or ready health response.
401 Missing or incorrect bearer key on an inference route.
413 POST body exceeds 2 MiB (2,097,152 bytes).
422 Invalid JSON/schema, unsupported model, invalid candidates, or token/input limits exceeded.
500 Inference backend failure or unavailable inference worker.
503 Full inference queue, or failed readiness check. Queue-full responses include Retry-After: 1.
504 Request exceeded the configured server deadline. No automatic retry is performed.

Errors use a detail field; schema-validation errors may contain a list of field errors rather than a string. Unknown fields are rejected. Request-size checking applies to POST bodies before inference. It does not extend model token limits: rendered state/question is capped at 2,048 tokens and each candidate at 768, including special tokens. Across all questions, a request may contain at most 256 candidate pairs. Text is never silently truncated.

Default server timeout is 60 seconds, queue capacity is 128, and a scheduler batch contains up to 32 requests. A timed-out request may have already started model work; returning a timeout cannot undo a running tensor operation.

The health counters completed_requests and model_batches describe successful requests and scheduler batch calls. A scheduler call can contain several encoder forwards. Neither counter is a billed-token meter. More detail is in the CLI and limits references.