Troubleshooting¶
Start with the installed version and runtime information:
doctor reports installed libraries without loading the model. Server logs
appear in the terminal running gemmadecision serve.
Upgrade a 0.1.0 notebook¶
Version 0.1.0 installed Torch and Transformers by default. In a notebook with
preinstalled vision/audio packages, upgrading Torch alone can leave binary
dependencies incompatible. An error such as operator torchvision::nms does
not exist, followed by a failure to import Gemma3TextModel, can come from
that dependency conflict even though this model only handles text.
Version 0.2.0 uses ONNX Runtime on CPU by default. A plain install no longer installs or imports Torch, Transformers, torchvision, or torchaudio.
- In a notebook cell, run
%pip install --upgrade gemmadecision. - Restart the notebook kernel so it stops using already imported 0.1.0 modules.
- Run the local call again:
from importlib.metadata import version
from gemmadecision import decide
print(version("gemmadecision")) # Check that this is 0.2.0 or later.
print(decide("I was charged twice", choices=["billing", "technical"]))
The new default avoids the Torch dependency chain; it does not repair unrelated
Torch applications in the same environment. If you set GEMMADECISION_MODEL_PATH
to an old native Torch directory, unset it to use the default ONNX download,
or replace it with a complete ONNX directory. A legacy native directory still
selects Torch for compatibility.
For explicit native Torch or GPU use, install gemmadecision[torch]. Use a
clean environment or match the installed Torch, torchvision, and torchaudio
builds. Vision/audio packages are unnecessary for GemmaDecision: if nothing
else in that environment uses the incompatible package, removing it is another
option. Restart the kernel after dependency changes. The Torch backend preserves
the original import error and adds repair guidance for recognized media failures.
The first call is taking longer¶
The first local call or server start downloads the pinned runtime model files and loads the inference runtime. Optional GPU setup may also initialize kernels. Warm inference measurements exclude that work. Historical 0.1.0 download and latency measurements describe its Torch files, not the 0.2.0 ONNX default.
Download explicitly to separate download progress from startup:
python -m pip install 'gemmadecision[serve]'
gemmadecision download --output ./model
gemmadecision serve --model-path ./model --offline
For an HTTP server, wait until GET /ready succeeds. For local Python use,
keep the process running and reuse the default engine, or construct one
DecisionEngine and reuse it. Starting a new Python process for every decision
reloads the model each time.
If model verification fails, keep the complete pinned download intact. Point
the server at a fresh gemmadecision download --output directory instead of
mixing a head, tokenizer or encoder from another model release.
HTTP errors¶
| Status | What happened | What to check |
|---|---|---|
401 |
The inference key is missing or incorrect | Match the server's key in the client or Authorization: Bearer ... header |
413 |
The POST body exceeds 2 MiB | Send a smaller state and fewer questions |
422 |
Invalid schema, duplicate/empty candidates or an input limit | Read the response's detail; shorten the input or fix the request |
503 |
The inference queue is full, or the readiness endpoint reports an unavailable worker | Reduce caller concurrency; inspect server logs and /ready |
504 |
The server's decision deadline elapsed | Reduce the workload or raise --request-timeout after checking latency |
500 |
Inference failed | Inspect the server-side exception and available device memory |
The Python clients raise httpx.HTTPStatusError for HTTP failures. A client
timeout can occur before a server timeout if the client's deadline is shorter.
For example, increase the client deadline with DecisionClient(timeout=120)
only when the workload needs it. The clients do not automatically retry an
inference request; a timed-out operation that has already started may still
finish on the server.
Input limits¶
The serving contract accepts:
- 2,048 tokens for the rendered state and question together.
- 768 tokens per candidate description.
- 2–64 distinct candidates per question.
- At most 32 questions and 256 total candidate pairs per request.
- A POST body of at most 2 MiB.
Token counts include tokenizer special tokens. The package refuses overlong
inputs instead of silently truncating them. A custom, smaller
--max-batch-tokens setting can also reject an otherwise valid joint input.
Increasing a queue size does not extend input limits.
Authentication succeeds on health but fails on inference¶
Health, readiness, model information and API documentation are readable without
an inference key. Protected calls to /v1/systemone and /v1/rank still need
the configured key. SDK clients read GEMMADECISION_API_KEY; curl needs the
bearer header. See network serving.
The command uses the wrong Python environment¶
Install and run through the same interpreter:
python -m pip install 'gemmadecision[serve]'
python -m gemmadecision.cli doctor
python -m gemmadecision.cli serve --device cpu
Check python -m pip show gemmadecision if a different shell or notebook cannot
import the package. In a notebook, install into the interpreter used by its
kernel. The package supports Python 3.11 and later; the selected tensor backend
must also provide compatible wheels for that Python version and platform.
The HTTP server requires the serve extra; PydanticAI requires pydantic-ai.
A plain pip install gemmadecision is sufficient for local CPU decisions.
For explicit CUDA use, install the torch extra and confirm that PyTorch can
see the GPU:
Use --backend onnx --device cpu for the CPU runtime, or
--backend torch --device cuda for native GPU execution. Merely attaching a
GPU does not change the default ONNX backend. The CPU Dockerfile uses ONNX;
the optional vLLM backend requires its own supported Linux/CUDA environment.
See deployment.
PydanticAI rejects my output type¶
The model selects among finite options. Start with a supported output type:
from typing import Literal
from pydantic import BaseModel, Field
from pydantic_ai import Agent
from gemmadecision import GemmaDecisionModel
class TicketRoute(BaseModel):
team: Literal["billing", "technical", "review"] = Field(
description="Which team should handle this ticket?"
)
urgent: bool = Field(description="Does the ticket need prompt attention?")
agent = Agent(GemmaDecisionModel.local(), output_type=TicketRoute)
result = agent.run_sync("The same payment appears twice.")
print(result.output)
Unrestricted str fields, arbitrary dictionaries and unbounded numeric outputs
cannot be filled by choosing among known candidates. Use finite Literal/Enum
options, booleans or a described rubric, and handle generated prose or free-text
extraction in a separate component. The PydanticAI guide
explains the supported schemas.
GemmaDecisionModel.local() loads the model in your Python process.
GemmaDecisionModel() connects to an HTTP server. If you see a connection error
while expecting local inference, check which constructor you used.
The answer is a valid label, but the wrong one¶
Schema validation establishes the output type, not decision accuracy. Make option descriptions concrete and distinct, include an explicit review option when appropriate, and evaluate on representative labeled examples. See choosing options for probabilities, thresholds and none-of-the-above behavior.