gemmadecision · Python package 0.2.0
Small decisions. A few lines of Python.¶
Choose a support queue, classify a request, score an ordered rubric, or return a typed decision. Run GemmaDecision-270M locally, directly from your application.
Start in Python Browse use cases
from gemmadecision import decide
team = decide(
"A payment appears twice on my statement.",
choices=["billing", "technical"],
)
print(team)
The first call downloads and loads the public model. Later calls reuse it in the same process. The default uses ONNX Runtime on CPU without installing Torch or Transformers. No hosted-model API key or server is required for local use.
-
Use it in an application
Return a label with
decide(), or inspect all candidates withrank(). -
Keep outputs typed
Install the
pydantic-aiextra forLiteral,Enum, boolean and finite Pydantic fields. -
Share one model over HTTP
Install the
serveextra, then startgemmadecision serveand connect clients. -
Choose a deployment
Start on CPU with ONNX Runtime. Add the Torch extra for CUDA or Apple MPS; vLLM is optional on Linux/CUDA.
Pick the output you need¶
| Your application needs | Start here |
|---|---|
| One label from a known set | Route a support ticket |
| A yes/no signal | Check urgency |
| A score on described levels | Score urgency |
| Several fields in one response | Typed triage |
| Ordered supplied options | Rank responses |
| A review path for uncertain decisions | Human review |
What has been verified¶
The historical 0.1.0 Torch release passed fresh Modal CPU and H100 checks for direct Python calls, local PydanticAI and real HTTP serving. Those checks do not establish 0.2.0 ONNX performance. On four CPU cores, two short 30-token candidate inputs took 124 ms median in a small repeated measurement. CPU timings and installation evidence include the scope and raw results.
The recipes demonstrate ways to integrate a closed-choice ranker. Accuracy in each new application needs its own evaluation; the released model's measured domains are banking-support topics and three-way evidence relations. It does not generate free-form text. Model scope and limits.