Skip to content

Lightweight CPU runtime: 0.2.0

Version 0.2.0 runs local decisions with ONNX Runtime on CPU, without installing Torch or Transformers. The same decide() and rank() calls work. HTTP serving, PydanticAI, native Torch, and vLLM are optional extras.

In the CPU validation below, two short choices took 110 ms median. ONNX also preserved the Torch model's top choice on all 700 development cases. These results concern the stated workloads and model conversion; they do not establish a universal speedup or accuracy on a new task.

Install or upgrade

python -m pip install --upgrade gemmadecision==0.2.0

In a notebook, run %pip install --upgrade gemmadecision==0.2.0 in a cell, then restart the runtime/kernel before running the code below. Installing the upgrade alone does not replace already imported 0.1.0 modules in the running Python process.

from gemmadecision import decide

print(decide("I was charged twice", choices=["billing", "technical"]))
Add this capability Install
Local CPU decisions pip install gemmadecision
Granian Rust HTTP server pip install 'gemmadecision[serve]'
Native PydanticAI agents pip install 'gemmadecision[pydantic-ai]'
Native Torch, CUDA, or Apple MPS pip install 'gemmadecision[torch]'
vLLM on supported Linux/CUDA pip install 'gemmadecision[vllm]'

Extras combine, for example gemmadecision[serve,pydantic-ai]. The default auto path uses CPU ONNX. Explicit device="cuda" or device="mps" selects Torch and requires its extra. See hardware settings and serving.

Why the notebook dependency failure happened

The 0.1.0 package required Torch 2.13 or newer. Installing it into a notebook with Torch 2.11 and a matching torchvision 0.26 could upgrade Torch while leaving the existing torchvision binary tied to 2.11. The resulting operator torchvision::nms does not exist error can surface as a failure to import Gemma3TextModel: Transformers' model-loading path also checks installed optional vision/audio libraries, even for this text model.

The 0.2.0 default avoids that dependency chain. It does not replace or import Torch, torchvision, or torchaudio. It does not repair other applications that already have incompatible Torch packages. Explicit native Torch users should keep that environment's Torch/vision/audio builds compatible; the notebook recovery guide also covers old local model-directory settings.

What the runtime artifact contains

The graph includes the tuned encoder, normalized last-token pooling, and the trained scalar head. It returns the same kind of raw ranking score. The API still applies the existing softmax temperature, 4.136820402388508, and retains the choice, yes/no, and ordered-score response semantics.

Weights that round-trip exactly are stored in FP16 or BF16 and restored to their original FP32 values before computation. Other values remain FP32. The selected format uses lossless storage conversion; it is not INT8 quantization or FP16 model computation. Different runtime kernels can still produce small floating-point differences, which are measured below.

The graph is 538,514,811 bytes. Its complete runtime files total 573,083,893 bytes: approximately 573 MB / 546.5 MiB, including the tokenizer and configuration. This is a one-time cached model download, separate from the Python dependencies. Later calls in the same process reuse the loaded model.

The compact file avoids the size of a naive FP32 ONNX export. The original checkpoint already used BF16 storage and was about the same size; this is not a 50% reduction against that checkpoint. FP32 inference still needs memory for restored weights, activations, and runtime workspaces. Download size is not a RAM requirement. The validation process peaked at about 2.72 GiB RSS, but it ran Torch and ONNX sequentially in one process, so that number is not an isolated ONNX memory measurement or minimum.

The package verifies the pinned export manifest and every listed inference file. The source weights remain model revision 785d530221c990671f29976902540101bb9c7647; conversion does not retrain them.

Public PyPI installation

After publication, a fresh Python 3.13.3 environment on Modal, with four CPU cores and 8 GiB RAM, installed gemmadecision==0.2.0 directly from public PyPI. The exact example above returned billing. No Torch, torchvision, torchaudio, or Transformers packages were installed or imported, and both decide() and rank() worked again in a new process using the model cache offline.

Measurement Result
Ordinary pip install, including dependency resolution 7.5920 s
Compressed package wheels, including dependencies 54,019,036 bytes (~54 MB)
Published GemmaDecision wheel within that total 47,631 bytes
First exact example, including model download and loading 6.669 s
Model loading in a new process from the cache, offline 4.090 s
Warm decide() median for two descriptive choices, 34/35 joint tokens 83.49 ms

The 573 MB model download is separate from the package archives. Install time excludes virtual-environment creation and model download. The warm median covers ten calls after three warmups, using the descriptive two-choice fixture recorded in the report rather than the short first-call example. These are functional and latency observations from one run, not accuracy measurements or a speed comparison with the separate hosts used below.

The public installation report records the public package URL, dependency versions, exact samples, and verified model pins. It retains the shared harness's candidate-prefixed measurement keys; installation_source: public_pypi and the artifact URL identify this run. The published wheel SHA256 is 3b72a2bd3f0831ef623a96d5f9580c7829cbc0bae48a7af93d937029b9afae1d. Only two temporary artifact-path fields were removed from the report copy, with its original SHA256 recorded in the provenance entry. All 18 runtime Python files in the published wheel were also verified byte-for-byte against the candidate wheel and committed source; see the wheel/source comparison.

Fresh installation of the candidate wheel

A separate smoke test installed the 0.2.0 candidate wheel into a fresh Python 3.13.3 environment on Modal, with four CPU cores and 8 GiB RAM. This checked the built distribution before publication; it was not an install of 0.2.0 from public PyPI.

Measurement Result
Ordinary pip install, including dependency resolution 10.4266 s
Compressed package wheels, including dependencies 54,018,997 bytes (~54 MB)
GemmaDecision candidate wheel within that total 47,592 bytes
First decide("I was charged twice", choices=["billing", "technical"]), including model download and loading 8.448 s
Model loading in a new process from the cache, offline 5.359 s
Warm decide() median for two descriptive choices, 34/35 joint tokens 142.75 ms

The model files add a separate 573 MB one-time download. Package archive sizes exclude pip metadata, network overhead, and extracted installation size. Installation time excludes virtual-environment creation and model download. The first decision returned billing; that single example checks that inference works and is not an accuracy evaluation.

Torch, torchvision, torchaudio, and Transformers were absent from the fresh installation, and the inference checks made no forbidden framework import attempts. Both decide() and rank() also worked in a second process with network access to the model disabled and the previously downloaded cache.

The warm measurement used ten timed calls after three warmups, with the model already loaded. Its prompt and host differ from the same-host backend comparison below, so the two tables should not be used to infer a speed change.

The candidate installation report records dependency versions, wheel hashes, every timing sample, and model pins. Only two remote artifact-path fields were omitted from the copied report; a provenance entry records the original report hash. The candidate wheel SHA256 was 0d96bcfb0e852caeb8673b0ed7b9afca0612a1213f6d13ce8a4a86edcbd45341. It downloaded the verified export at Hub revision 254ac03b1ec6c96af3b9e605bd9c6f5cc0156fae.

Two additional candidate-wheel checks passed. A notebook compatibility simulation kept existing Torch 2.11.0+cpu and torchvision 0.26.0+cpu unchanged and returned billing for the same user example without importing either framework, even with a sentinel that rejects torchvision imports. This simulated that dependency setup; it did not run inside hosted Google Colab. The optional integrations check installed [serve,pydantic-ai] and passed a local native PydanticAI agent, an actual Granian HTTP server, unauthorized-request rejection, an authenticated SDK decision, and a remote native PydanticAI agent, all using CPU ONNX without Torch. These are functional checks, not additional accuracy measurements.

Conversion fidelity on development data

The check used the existing, frozen 700-example development set, with 2,300 candidate pairs. All 700 examples fit the input limits; none were skipped. Their joint candidate inputs ranged from 25 to 141 tokens. No final test or calibration data was opened, and neither weights nor temperature were fitted during this validation.

Development task family Cases Same top choice as CPU Torch FP32
Intent routing 300 300
Evidence relation 300 300
Rule compliance 100 100
Total 700 700 / 700 (100%)

Across the candidate scores, maximum absolute error was 0.00017345 and mean absolute error was 0.00000812. Maximum absolute derived-probability error was 0.00000908.

This measures conversion fidelity on development data, not a new blind accuracy benchmark. Matching the reference does not make its decisions correct, nor guarantee agreement for every future input or near tie. Existing model evaluations remain in the model card.

CPU latency on the same host

Both backends ran sequentially on the same remote Linux host with four available CPU cores and four inference threads. The CPU model was reported as unknown. The environment used Python 3.13.3, Torch 2.14.0+cpu, Transformers 5.17.0, and ONNX Runtime 1.30.0.

Each measurement covers a complete DecisionEngine.rank() call, including validation, tokenization, scoring, and response construction. The model was already loaded, the score cache was off, and there were three warmups per workload. HTTP, download, and model initialization are excluded. Both used a 4,096-padded-token budget and a maximum batch size of 16.

Joint tokens per candidate Choices Torch FP32 median ONNX median Timed calls per backend
29 2 149.2 ms 110.0 ms 10
29 4 186.3 ms 211.5 ms 10
128 2 417.2 ms 321.2 ms 10
128 4 683.9 ms 780.7 ms 10
512 2 1,262.3 ms 1,137.5 ms 5
512 4 2,457.2 ms 2,387.0 ms 5

ONNX was slower for the four-choice 29-token and 128-token workloads in this run. These small repeated samples support an environment-specific comparison, not a general latency or throughput guarantee. The 110 ms result applies to two short choices, not every decision. The primary installation benefit is avoiding the Torch/Transformers dependency stack for CPU inference.

Evidence and reproduction

Raw aggregate validation report contains every timing sample, task counts, environment, script hash, and export manifest hash. The report is copied unchanged. Its onnx_revision: null records that this check used a verified prepublication artifact, before an immutable Hub export revision was assigned.

Validation report SHA256
bc4851eca09d1f5ee2c40906f0e0cb9b667f79eac6e9a3c84b6c635b995b7566

Validated export manifest SHA256
f541a308c7cc4c75b645d2a026bde4771c0877e7129147cb260c1b8e305be178

The source includes the exporter, lossless-storage conversion, and validation harness. These are remote validation tools, not steps needed for ordinary inference. The separate 0.1.0 CPU report remains historical and uses different measurements.