Recipes¶
Give GemmaDecision the context, a question, and the possible answers. These
recipes show how to put that decision into an application using
gemmadecision==0.2.0.
The default uses ONNX Runtime on CPU. The typed triage recipe
also needs pip install 'gemmadecision[pydantic-ai]'.
Every example runs locally. There is no server to start or API key to create. The first use downloads the published model; later runs reuse the download. Keep an engine or agent alive to reuse its loaded weights across requests. Each page has a complete Python example that you can save and run.
| I want to… | Start here | Interface |
|---|---|---|
| Send a support ticket to a team | Ticket routing | decide() |
| Identify a customer's intent | Intent classification | Choice |
| Assign sentiment labels | Sentiment | Choice |
| Ask a yes/no question | Urgency check | Noul |
| Assess a request against ordered levels | Urgency rubric | Score |
| Compare a claim with supplied evidence | Evidence relation | Choice |
| Choose among answers I already have | Response ranking | rank() |
| Recommend a tool from an approved list | Tool selection | Choice |
| Return several typed decisions together | PydanticAI triage | Agent |
| Send uncertain decisions to a person | Human review | Choice + application rule |
These are illustrative workflows, not accuracy claims. The model's published evaluations cover particular intent-routing and evidence-relation datasets; they do not establish performance on every recipe or your application's data. See the model card for the evaluated tasks and limitations. Test representative examples from your own domain before relying on a workflow.
The model ranks supplied alternatives. It does not write replies, extract arbitrary text, or execute tools. Its probability fields come from normalized ranking scores; they are not guaranteed probabilities of correctness. Adding or removing alternatives can change those probabilities.
The examples use short inputs. Version 0.2.0 accepts 2–64 distinct candidates per question, at most 32 questions and 256 candidate pairs per call, with 2,048 tokens for the state plus question and 768 tokens for each candidate. Oversized inputs raise an error rather than being silently truncated.
For runnable examples in a repository checkout, explore the recipe launcher:
For an application calling a separate process or machine, keep the same question objects and use the HTTP client.