Skip to content

gemmadecision · Python package 0.2.0

Small decisions. A few lines of Python.

Choose a support queue, classify a request, score an ordered rubric, or return a typed decision. Run GemmaDecision-270M locally, directly from your application.

Start in Python Browse use cases

pip install gemmadecision
from gemmadecision import decide

team = decide(
    "A payment appears twice on my statement.",
    choices=["billing", "technical"],
)
print(team)

The first call downloads and loads the public model. Later calls reuse it in the same process. The default uses ONNX Runtime on CPU without installing Torch or Transformers. No hosted-model API key or server is required for local use.

  • Use it in an application

    Return a label with decide(), or inspect all candidates with rank().

    Python quickstart

  • Keep outputs typed

    Install the pydantic-ai extra for Literal, Enum, boolean and finite Pydantic fields.

    PydanticAI guide

  • Share one model over HTTP

    Install the serve extra, then start gemmadecision serve and connect clients.

    Serving guide

  • Choose a deployment

    Start on CPU with ONNX Runtime. Add the Torch extra for CUDA or Apple MPS; vLLM is optional on Linux/CUDA.

    Deployment guide

Pick the output you need

Your application needs Start here
One label from a known set Route a support ticket
A yes/no signal Check urgency
A score on described levels Score urgency
Several fields in one response Typed triage
Ordered supplied options Rank responses
A review path for uncertain decisions Human review

What has been verified

The historical 0.1.0 Torch release passed fresh Modal CPU and H100 checks for direct Python calls, local PydanticAI and real HTTP serving. Those checks do not establish 0.2.0 ONNX performance. On four CPU cores, two short 30-token candidate inputs took 124 ms median in a small repeated measurement. CPU timings and installation evidence include the scope and raw results.

The recipes demonstrate ways to integrate a closed-choice ranker. Accuracy in each new application needs its own evaluation; the released model's measured domains are banking-support topics and three-way evidence relations. It does not generate free-form text. Model scope and limits.

PyPI package · Model card and weights · GitHub source