Limits and result semantics¶
GemmaDecision chooses among alternatives you provide. It does not generate prose, extract arbitrary strings, or decide what tools your application is authorized to execute. This page describes the 0.2.0 package contract.
Input limits¶
| Item | Limit |
|---|---|
| Candidate choices in one ranking or choice question | 2–64 |
| Levels in one score question | 2–64 |
| Named questions in one request | 1–32 |
| Candidate pairs across all questions in one request | 256 |
| Rendered state plus question instructions | 2,048 tokens |
| Each candidate description | 768 tokens |
| HTTP POST body | 2 MiB / 2,097,152 bytes |
| Default ONNX/Torch batch size | 32 joint inputs |
| Default ONNX/Torch batch budget | 8,192 padded tokens |
| Default server queue capacity | 128 waiting requests |
| Default server request deadline | 60 seconds |
Token limits include tokenizer special tokens. They apply to rendered input, so nested state and structured instructions count toward the same 2,048-token state/question budget. Every candidate is concatenated with that state/question and a fixed separator, then checked against the backend's complete-input limit. The Torch encoder's larger architectural context does not increase the published state or candidate limits. ONNX and vLLM use a 3,072-token complete-input limit by default.
Text is never silently truncated. Invalid local requests raise validation errors; HTTP requests return 422. Candidate descriptions must be distinct and nonempty after rendering. For example, a request with five 64-option questions exceeds the 256-pair limit even though each question is individually valid.
A single ONNX/Torch joint input must fit the configured max_batch_tokens, since
even a one-element batch must respect that budget. The scheduler can split a
larger group of valid inputs into multiple forwards. It does not treat the
batch-token setting as a larger context window.
Scores, probabilities, and confidence¶
The model produces one raw scalar per joint input. Larger scores rank ahead of smaller scores. They are uncalibrated ranking scores, and values from different requests should not be compared as a common confidence scale.
The API applies softmax across a question's candidates at a fixed temperature
of 4.136820402388508. Consequently:
- Probabilities sum to one within that candidate set.
- Adding or removing candidates changes the normalized probabilities.
- A large probability is not a universal estimate that the decision is correct.
- An application fallback cutoff should be selected on its own held-out data.
The wire confidence field is the maximum probability. Noul.noul is the true
outcome's probability, while Score.score is the expected zero-based rubric
level. A score of 1.4 on a three-level rubric is valid; it is not necessarily
the index of the most probable level. Exact raw-score ties preserve input order.
Batch shape and numeric kernels can introduce small floating-point differences.
Torch strict=True preserves the published singleton batch shape for
comparisons, but it does not promise bit-identical results across devices or
software versions. ONNX strict=True also scores singleton batches; it does
not reproduce Torch kernels or guarantee identical scores. See the historical
Torch/vLLM numerical agreement checks.
PydanticAI field constraints¶
GemmaDecisionModel uses PydanticAI's decision-model translation. These field
shapes are supported inside a Pydantic output model by the tested PydanticAI
2.51 integration. Simple types such as bool and Literal can also be used
directly as the agent's output_type:
| Output shape | Translation / constraint |
|---|---|
bool |
One yes/no question; default boolean threshold is 0.5. |
Literal["a", "b", ...] or string Enum |
A choice with at least 2 and at most 64 alternatives. |
Whole-number Literal / Enum |
Finite categorical choices; arbitrary unbounded int is not supported. |
Described levels 0, 1, …, N-1 |
A rubric only when every numeric level has a description in the JSON schema. Otherwise they are categorical choices. Maximum 64 levels. |
Annotated[float, Field(ge=0, le=1)] |
A yes/no probability returned as a float, without boolean thresholding. A positive upper bound can also scale the units, e.g. 100 for a percentage; stepped multiple_of constraints are unsupported. |
list[Literal["a", "b", ...]] or list of string Enum |
One yes/no question per option, returning selected options. At least 2 string options; no list-length constraints. |
dict[Literal["a", "b", ...], bool] |
One yes/no question per finite key, retaining every key. Arbitrary free-form dictionary keys/values are unsupported. |
| Pydantic model of supported fields | Fields become named questions; nested models expand to their supported leaf fields. |
| Optional finite categorical field | Adds a distinct “none of these” choice. Optional booleans, probability fields, and rubrics are not interchangeable with this form. |
A rubric returned through the native typed model is rounded to a valid level
(half values round upward). The wire ScoreAnswer.score retains the fractional
expectation. Native decision metadata records this difference.
Put a described numeric rubric in a Pydantic model field: a bare union passed
as output_type can instead be interpreted as separate output routes.
Expanded questions still share the package's 32-question and 256-pair limits. A list with many alternatives can consume multiple question slots; the limit is not simply the number of Pydantic fields.
Free-form str, unconstrained numeric output, list[str], arbitrary objects,
and generative explanations are not expressible as ordinary decision fields.
Use a Literal, Enum, or another supported decision shape instead. Give
fields descriptions and use agent instructions to state what is being asked;
the user prompt carries the material to classify.
PydanticAI can also route among finite output types or tools. A tool whose
arguments cannot be represented may cause a handoff; it does not turn this
model into a text generator. The optional decision_route_threshold applies
to route selection, not as a universal confidence check for a single typed
answer. Keep tool authorization and side-effect policy in the application.
Runtime and memory¶
Python 3.11 or newer is required. The normal install uses CPU ONNX Runtime
and the Hugging Face tokenizers library. Torch, Transformers, Granian/FastAPI,
and PydanticAI are optional: install the torch, serve, or pydantic-ai
extras as needed. vLLM is optional and pinned to 0.30.0 for the pooling path.
The ONNX graph contains the encoder, normalized last-token pooling, and scalar head. It is an export of the pinned source model, not a separately trained model. Artifact format alone does not establish speed or prediction parity.
Torch supports CPU, CUDA, and MPS. Its encoder uses BF16 on CUDA when supported, and FP32 otherwise; normalization and the scalar head are FP32. vLLM requires a BF16-capable CUDA GPU; its encoder is BF16 and its scalar head runs FP32 on CPU. The published CPU and GPU measurements are hardware-specific, not a memory or latency guarantee for another machine.
The server starts one model process. Its Rust HTTP layer does not replace the ONNX/Torch/vLLM model computation. Cache entries store scores for exact complete joint inputs, never independent state and candidate embeddings. The cache is off by default.
Usage and cancellation¶
usage.input_tokens counts actual encoded joint inputs, including repeated
state/question tokens for distinct candidates. Inputs reused through the score
cache or deduplicated within a grouped batch are not charged again in that
count. cached_pairs includes those reused pairs. Token accounting may therefore
be assigned to the first request that encodes a shared input in a batch.
output_tokens is zero because ranking does not generate tokens. total_tokens
equals input_tokens. These are serving-work counters, not a model-provider
billing tariff.
Neither the SDK nor the server automatically retries inference. An HTTP or local native-model timeout stops waiting for the answer; an already running tensor operation may still finish. Check the HTTP status reference and configure application retry/fallback behavior explicitly.