Skip to content

Measured serving performance

These historical measurements cover the 0.1.0 Torch and vLLM implementations. Version 0.2.0 uses ONNX Runtime on CPU by default. These results do not measure that default or establish its speed, memory use, or numerical agreement.

On one NVIDIA H100 80GB HBM3, scoring 32 short joint inputs took 27.59 ms with batching, compared with 753.36 ms in the frozen published reference and 758.77 ms in this package's strict singleton mode. That is 27.31× faster than the reference for this specific workload. Four short candidates took 25.41 ms, compared with 94.35 ms in the reference.

These are local backend timings, including tokenization and the scalar head, with explicit GPU synchronization. They exclude HTTP, input validation, and downloads. Each workload had three warmups and only five measured repeats; these results establish a useful initial measurement, not a production SLA or model-accuracy result. Full samples and synthetic inputs are in h100-backend.json.

Device and model

Measurements used the fully tuned GemmaDecision v0.4.0 encoder and scalar head, model revision 785d530221c990671f29976902540101bb9c7647, with verified file hashes. The CUDA encoder used BF16; last-token normalization and the learned scalar head used FP32. The environment was Python 3.11.12, Torch 2.14.0, Transformers 5.17.0, CUDA 13.0, with two Torch CPU threads and TF32 matrix multiplication disabled. Source was loaded directly from the recorded source snapshot; the absent installed package version in the JSON is intentional.

The batch limit was 32 inputs and 8,192 padded tokens. The cache was off. The reference scores one state/candidate pair at a time. This package's strict mode preserves that numerical batch shape. Its normal mode buckets inputs by length, pads on the right, scores bounded batches, and restores original order.

Backend latency

All values below are medians in milliseconds. A pair means one complete joint state/question/candidate input, not one HTTP request. Long workloads can require several forward passes because the padded-token budget still applies.

Workload Tokens per pair Frozen reference Package strict Package batched Speedup vs reference
Short, 1 pair 49 23.35 23.63 23.64 0.99×
Short, 4 pairs 49–50 94.35 94.24 25.41 3.71×
Short, 16 pairs 49–52 373.87 383.23 26.01 14.37×
Short, 32 pairs 49–55 753.36 758.77 27.59 27.31×
Long, 1 pair 721 25.26 25.29 25.20 1.00×
Long, 4 pairs 721–722 99.40 100.64 28.68 3.47×
Long, 16 pairs 721–724 398.57 400.24 62.71 6.36×
Long, 32 pairs 721–727 792.84 806.43 100.67 7.88×

For the 32 short pairs, batched throughput was 1,160 pairs/s. This is not HTTP requests/s. The backend reported peak Torch-allocator memory of 726,264,320 bytes allocated and 838,860,800 bytes reserved. Those numbers exclude driver and non-Torch allocations and do not establish a minimum GPU VRAM requirement for all possible inputs.

Numerical agreement

On 12 hand-authored cases containing 96 candidate pairs, strict mode matched the reference scores exactly. Batched mode selected the same top candidate in all 12 cases; maximum absolute score difference was 0.0000152588, and maximum probability difference was 0.000000455138 using the fixed external temperature 4.136820402388508. The smallest observed reference top-two margin was about 0.0136. This fixture check cannot rule out different results on closer ties, other kernels, devices, or model versions.

The inputs are unlabeled serving fixtures. Agreement with the reference is not decision accuracy, JevBench performance, or a new calibration result. The experiment did not train, quantize, or change the model.

Real Rust HTTP server

The second run started the actual CLI and Granian 2.8.3 Rust HTTP runtime, then used the Python SDK over loopback HTTP. It used the same H100 class, model, precision, batch limits, a 1 ms batching window, and no score cache. Each request had one choice question with four candidates and a distinct case ID, so identical input deduplication did not inflate throughput. After three warmup requests, 32 requests per concurrency level were measured:

Concurrent clients HTTP requests/s Candidate pairs/s Request p50 Request p95 Scheduler batches
1 26.99 107.96 25.63 ms 70.17 ms 32
8 93.60 374.41 81.29 ms 121.53 ms 8
32 85.59 342.37 361.92 ms 366.58 ms 2

Request latency starts after the client concurrency semaphore is acquired; it includes server queueing, validation, tokenization, inference, and HTTP, but excludes waiting for a client slot. Throughput uses the elapsed time to complete each 32-request group. Scheduler batches count grouped engine calls; each can contain several bounded encoder forwards. Concurrency 32 increased latency and did not improve throughput over 8 in this small experiment.

With only 32 requests per level, the percentiles are noisy and do not describe steady-state production tail latency. There is no cross-machine network in this measurement. See h100-http.json for all results. Rust handles HTTP connections; the measured transformer math remains in Torch/SDPA. These results do not isolate a Rust-versus-Python HTTP speedup.

Real API checks

The successful HTTP run also verified all three question types (choice, noul, and score), the ranking endpoint, a native PydanticAI model returning a typed Literal/bool result, missing-key rejection (401), and an overlong state rejection (422). Readiness took 12.75 s using an existing local model directory; the first inference still needed additional warmup.

After shutting down the HTTP model process, a fresh process verified the simple decide() and rank() APIs with the pinned local files and network access to Hugging Face disabled. The first decide() including model load took 9.88 s; one subsequent call took 23.35 ms. These single calls are smoke checks, not latency distributions or uncached download timings.

The same process then ran a real PydanticAI Agent(GemmaDecisionModel.local(), output_type=Literal["billing", "technical"]). It returned the valid Literal billing, reported local://gemmadecision, and reused the same process-local engine. The warm Agent call took 154.12 ms. No external LLM or HTTP server was used for that local Agent check. Correct output types and engine reuse were tested; application accuracy was not scored.

Reproduction and evidence

Run the GPU scripts on a CUDA host with the package installed and a complete pinned model directory:

python scripts/benchmark_backends.py --help
python scripts/benchmark_server.py --model-path MODEL_DIR --output server-results.json

The backend report includes every fixture, five timing samples per workload, per-case parity, and environment metadata. The corresponding backend source hashes and HTTP source hashes identify the exact code used. Provenance records original and published JSON hashes; the HTTP report normalizes interpreter/model-directory paths for portability without changing measurements.

Estimated function compute was $0.09613 for the backend run and $0.04331 for the successful HTTP/API run: $0.13943 combined, excluding startup, storage, and transfer. These are rate-based estimates, not a billing receipt; see backend compute and HTTP compute. The first combined run's HTTP startup failed on a Granian keyword mismatch; its compute record preserves that failed status. The later successful HTTP run used the corrected CLI. A regression test now constructs the real Granian object while replacing only its serve() method, so constructor errors are caught without loading weights.

The preceding measurements cover the Torch backend. They establish no CPU/Mac latency claims. The separately measured vLLM backend is reported below. Validate the intended deployment hardware and workload before choosing concurrency and batch limits.

Input and batching contract

State/question and candidate limits remain 2,048 and 768 tokens respectively, including special tokens. The encoder validates complete joint inputs and never truncates text. A single pair longer than max_batch_tokens is rejected instead of exceeding that configured batch budget. Candidate and state cannot be cached as independent embeddings because this model encodes their concatenation jointly. Increasing an architectural context limit is not equivalent to validating a new serving input contract.

vLLM validation on H100

The optional vLLM 0.30.0 backend passed real GPU and HTTP/API checks on an H100 80GB HBM3. Batched PyTorch was faster for every backend workload in this experiment. This result concerns the optional GPU backends; it does not compare them with the current CPU ONNX default. This comparison uses the frozen reference and both backends measured sequentially in the same run with Torch 2.13.0+cu130, Python 3.12.3, Transformers 5.17.0, CUDA 13.0 and two CPU threads. It does not compare vLLM against the earlier Torch 2.14.0 run.

The model revision and 96 synthetic pairs were unchanged. Prefix and score caches were disabled; vLLM used LAST pooling, BF16 CUDA inference, eager execution, a 3,072-token joint limit and a CPU FP32 normalization/head. The backend timings below include tokenization and the head, exclude HTTP, and use three warmups plus five measured repetitions. PyTorch's 8,192 padded-token budget can split long workloads into multiple forwards; vLLM schedules its own input batches. These are the tested default configurations, not a search for each runtime's best tuning.

Workload Frozen reference p50 Batched Torch p50 vLLM p50 vLLM p95
Short, 1 pair 14.94 ms 14.97 ms 20.57 ms 21.24 ms
Short, 4 pairs 59.85 ms 16.64 ms 43.03 ms 43.31 ms
Short, 16 pairs 236.79 ms 17.53 ms 49.59 ms 49.93 ms
Short, 32 pairs 474.80 ms 18.93 ms 56.85 ms 59.69 ms
Long, 1 pair 16.23 ms 16.76 ms 21.90 ms 22.74 ms
Long, 4 pairs 64.57 ms 19.58 ms 46.12 ms 54.69 ms
Long, 16 pairs 257.96 ms 51.74 ms 63.59 ms 63.89 ms
Long, 32 pairs 516.84 ms 91.53 ms 115.11 ms 119.24 ms

For 32 short pairs, throughput was 562.90 pairs/s with vLLM, versus 1,690.71 pairs/s with batched Torch in that run. vLLM backend initialization took 68.56 s and its first request another 39.91 ms. Initialization includes engine setup with local weights, not downloading the model. The measurement script cannot report vLLM worker memory through the parent process's Torch allocator, so those fields are explicitly unavailable. See vllm-h100-backend.json.

Numerical agreement, not accuracy

The same 12 cases / 96 candidate pairs gave 12/12 top-candidate agreement between vLLM and the frozen reference. Maximum absolute score difference was 0.3179318905, mean absolute score difference 0.0661979032, and maximum probability difference 0.0087380903 using the unchanged temperature 4.136820402388508. Thus the measured outputs are not numerically equivalent, even though these cases retained their top choice. Batched Torch in the same run had maximum score difference 0.0000152588 and also agreed on all 12 top choices; strict Torch matched the reference exactly.

This is an unlabeled serving fixture check, not accuracy, a JevBench rerun or evidence that every close decision is preserved. Use strict Torch when matching the singleton reference's numerical behavior matters.

Real vLLM HTTP serving

The actual gemmadecision serve --backend vllm command ran behind Granian 2.8.3. After three warmup requests, each concurrency group contained 32 distinct requests with four candidates each, with caching off and a 1 ms microbatch window:

Concurrent clients HTTP requests/s Candidate pairs/s Request p50 Request p95
1 21.12 84.49 45.65 ms 56.11 ms
8 70.11 280.45 113.98 ms 124.41 ms
32 111.76 447.05 276.59 ms 282.13 ms

These are loopback HTTP SDK measurements; client-semaphore waiting is excluded from latency, while server queueing is included. The small sample counts do not establish production tail latency. No claim is made that cross-run HTTP differences from the Torch 2.14.0 experiment are attributable to vLLM alone. Startup to readiness was 62.42 s with existing local model files.

The run verified choice/noul/score responses, the ranking route, native PydanticAI Literal/bool output through the HTTP client, missing-key rejection (401), and oversized-state rejection (422). After the vLLM server stopped, the script also checked the simple decide() / rank() and local PydanticAI APIs. Those trailing simple_api checks used the then-default Torch backend, as their recorded engine metadata shows; they are not vLLM measurements. See vllm-h100-http.json.

vLLM reproducibility and runtime notes

python scripts/benchmark_backends.py --model-path MODEL_DIR --backend vllm --output backend-results.json
python scripts/benchmark_server.py --model-path MODEL_DIR --backend vllm --output server-results.json

The exact source hashes, compute record and portable-file provenance are published. Backend and HTTP checks completed successfully in one 230.26-second function, with estimated function compute $0.27498, including the reference/Torch comparisons and the subsequent API checks. The estimate excludes startup, storage and transfer; it is not a billing receipt.

The validation used the official vLLM Docker image with a launcher-pinned NumPy 2.4.6. That override conflicted with requirements of the image's unused mistral-common and lmcache packages. The tested Gemma pooling/API path completed, but those extra components were not validated. An independent, clean Linux/Python 3.12 resolution of gemmadecision[vllm] succeeded with NumPy 2.3.5 and no lmcache dependency. Let a fresh environment resolve dependencies; do not reproduce the image's forced NumPy override. The exact fresh installation was not separately GPU-tested. Details are in runtime notes.

The released PyPI wheel subsequently passed separate fresh CPU and H100 installation checks, covering local decisions, local native PydanticAI and real Granian HTTP serving. Those short functional checks establish that the published package works; their single-call timings are not additional performance benchmarks.

CPU latency

The published wheel was also measured on four CPU cores: 124 ms median for two 30-token candidate inputs, 189 ms for four, and 397 ms for two 128-token inputs. See the CPU timing table and raw samples for longer inputs, sample p95, cold-start costs and measurement limits.