Skip to content

Serve decisions over HTTP

Run one model process and share it across your application. The server uses Granian for Rust HTTP serving and ONNX Runtime on CPU for inference by default.

Start the server

Install the package in the environment that will run the server:

python -m pip install 'gemmadecision[serve]'
gemmadecision serve

The address is http://127.0.0.1:8700. The first start downloads the pinned model, then loads it. Later starts reuse the downloaded files. Keep this terminal open and use another terminal for requests.

Choose hardware explicitly when needed:

gemmadecision serve --device cpu
python -m pip install 'gemmadecision[serve,torch]'
gemmadecision serve --device cuda

After installing the Torch extra, use --device mps on a supported Apple Silicon installation. The default --backend auto --device auto uses ONNX on CPU; CUDA and MPS are explicit choices. CUDA requires a compatible NVIDIA driver and PyTorch build. --backend torch selects native Torch directly.

For native PydanticAI clients, also install gemmadecision[pydantic-ai] in the client environment. Plain SDK clients are part of the base package.

Check readiness after loading:

curl http://127.0.0.1:8700/ready

Open http://127.0.0.1:8700/docs for interactive request schemas.

Choose an option with curl

POST /v1/systemone accepts a state and one or more named questions. A choice question maps the labels your application uses to descriptions the model sees.

curl http://127.0.0.1:8700/v1/systemone \
  -H 'Content-Type: application/json' \
  -d '{
    "state": "The same card payment appears twice on my statement.",
    "questions": {
      "team": {
        "type": "choice",
        "instructions": "Which support team should handle this request?",
        "criteria": {
          "billing": "Investigate charges, duplicate payments and refunds.",
          "technical": "Investigate app crashes and login failures."
        }
      }
    }
  }'

Read the selected label from answers.team.choice. That answer also includes raw scores, derived probabilities and confidence, the largest derived probability. These numbers need validation on your application's data.

For an ordered list of options, use POST /v1/rank:

curl http://127.0.0.1:8700/v1/rank \
  -H 'Content-Type: application/json' \
  -d '{
    "state": "The same card payment appears twice on my statement.",
    "question": "Which action best addresses this request?",
    "candidates": {
      "check_payment": "Investigate the transaction records for a duplicate charge.",
      "reset_password": "Help the customer reset their account password."
    }
  }'

The ranked array is sorted from highest to lowest score. Mapping keys appear as candidate; descriptions appear as text.

Call the server from Python

Reuse a client instead of constructing one for every request:

from gemmadecision import DecisionClient

with DecisionClient() as client:
    answer = client.decide(
        "The same payment appears twice.",
        candidates={
            "billing": "Charges, duplicate payments and refunds",
            "technical": "App crashes and login failures",
        },
        question="Which team should handle this request?",
    )
    print(answer.choice)
    print(answer.probabilities)

For an async application:

import asyncio
from gemmadecision import AsyncDecisionClient

async def main():
    async with AsyncDecisionClient() as client:
        result = await client.rank(
            "The app closes whenever I try to sign in.",
            candidates={
                "technical": "Investigate app crashes and sign-in problems.",
                "billing": "Investigate payments and invoices.",
            },
        )
        print(result.ranked[0].candidate)

asyncio.run(main())

Both clients accept base_url, api_key and timeout. They reuse HTTP connections and do not retry decisions automatically. For native PydanticAI clients, see the PydanticAI guide.

Connect from another machine

Set a key in the server shell and bind to the network interface:

export GEMMADECISION_API_KEY="$(python -c 'import secrets; print(secrets.token_urlsafe(32))')"
gemmadecision serve --host 0.0.0.0 --port 8700

Provide the same key to your client through its environment or secret manager. DecisionClient and AsyncDecisionClient read GEMMADECISION_API_KEY and GEMMADECISION_BASE_URL automatically. Set the base URL to the server's actual hostname or address; 0.0.0.0 is a bind address.

Requests made with curl include the key as a bearer header:

curl "$GEMMADECISION_BASE_URL/v1/rank" \
  -H "Authorization: Bearer $GEMMADECISION_API_KEY" \
  -H 'Content-Type: application/json' \
  -d '{"state":"I cannot sign in.","candidates":["Account access support","Payment support"]}'

Use a TLS reverse proxy for an internet-facing service. The CLI requires a key when binding outside loopback unless you deliberately pass --allow-unauthenticated. /health, /ready, /v1/models and the API schema remain readable without an inference key.

Endpoints and tuning

Endpoint Purpose
POST /v1/systemone Named choice, yes/no and rubric questions
POST /v1/rank Ordered candidates with raw scores and derived probabilities
GET /v1/models The pinned model served by this process
GET /health, GET /ready Readiness and queue information
GET /docs Interactive API documentation

Start with the default batching settings. --max-batch-size limits candidate pairs per ONNX/Torch forward, while --batch-requests limits requests grouped by the server scheduler. --max-batch-tokens bounds padded tokens in an ONNX/Torch batch. The score cache is off unless you set --cache-size.

See measured performance before choosing concurrency, deployment for Docker and Modal, and troubleshooting for failures and input limits.