Skip to content

PydanticAI: typed decisions from a local model

Start with in-process inference. Install the model runtime and integration:

pip install 'gemmadecision[pydantic-ai]'

New to GemmaDecision? The quickstart covers installation and your first local decision. The cookbook adds practical examples for routing, ranking, and review workflows.

There is no server to start, and no external model account or paid API key:

from typing import Literal

from pydantic_ai import Agent

from gemmadecision import GemmaDecisionModel


def main() -> None:
    router = Agent(
        GemmaDecisionModel.local(),
        output_type=Literal['billing', 'account', 'technical'],
        instructions='Which support team should handle this ticket?',
    )
    result = router.run_sync('I was charged twice.')
    print(result.output)


if __name__ == '__main__':
    main()

The constructor does not download or load weights. The first request downloads the pinned model once and loads it; subsequent calls reuse it. This is the same process-local CPU ONNX engine used by gemmadecision.decide() and gemmadecision.rank(). PydanticAI is optional; the base package keeps this framework out of a plain local install. The local model also supports await router.run(...): loading and inference run in a worker thread, keeping the event loop responsive. Engine locks serialize tokenizer and model access. This path does not implement server microbatching.

For explicit hardware settings, create DecisionEngine.from_pretrained(...) and pass it to GemmaDecisionModel.local(engine=engine). The process or caller owns the engine: exiting the model's async context does not unload shared weights. A timeout stops waiting and skips inference if it has not started; an already running model operation may finish in the background.

Use an existing server

For a separate server, use the remote constructor. Start a local server in one terminal:

pip install 'gemmadecision[serve,pydantic-ai]'
gemmadecision serve --host 127.0.0.1 --port 8700

Once it is ready, run this script in another terminal. The serving guide covers configuration and authentication.

import asyncio
from typing import Literal

from pydantic_ai import Agent

from gemmadecision import GemmaDecisionModel


async def main() -> None:
    async with GemmaDecisionModel(base_url='http://127.0.0.1:8700') as model:
        router = Agent(
            model,
            output_type=Literal['billing', 'account', 'technical'],
            instructions='Which support team should handle this ticket?',
        )
        result = await router.run('I was charged twice.')
        print(result.output)
        print(result.response.provider_details['probabilities'])


if __name__ == '__main__':
    asyncio.run(main())

The output is one of your supplied labels. The integration subclasses PydanticAI's public DecisionModel interface, which converts finite output schemas into ranking questions. It does not wrap a ranker as a text generator.

Several decisions in one request

Each field becomes a question; questions are evaluated together in one local engine call, or one HTTP request with a remote server. Use field descriptions for precise questions and enum member descriptions when the option names alone are ambiguous.

from typing import Literal

from pydantic import BaseModel, Field
from pydantic_ai import Agent

from gemmadecision import GemmaDecisionModel


class Ticket(BaseModel):
    team: Literal['billing', 'account', 'technical'] = Field(
        description='Which team should handle this ticket?'
    )
    urgent: bool = Field(description='Does the customer need attention today?')


def main() -> None:
    agent = Agent(GemmaDecisionModel.local(), output_type=Ticket)
    result = agent.run_sync('My card was charged twice. I need a refund today.')
    print(result.output.team)
    print(result.output.urgent)


if __name__ == '__main__':
    main()

See typed ticket triage for a complete workflow with a review route.

Supported PydanticAI decision schemas include bools, finite Literal/Enum choices, nested models of decision fields, and described rubric levels. Free-form str fields, arbitrary dictionaries, and unbounded numeric outputs are unsupported. Each choice or rubric allows at most 64 options, with at most 32 questions and 256 candidate pairs per request; input limits still apply. A list of labels becomes one yes/no question per label and counts toward that question limit. PydanticAI validates unsupported output schemas before sending a request. Multiple questions are independent judgments, not joint constrained generation, so enforce cross-field business rules in your application.

Reuse clients and route work

Set GemmaDecisionModel(base_url='http://your-server:8700', api_key='...') for another server. Or pass client=existing_async_client to reuse its pool. An injected client belongs to its caller and is not closed by the model. Requests honor PydanticAI's timeout and extra_headers settings. extra_body is rejected to keep the decision protocol explicit. Local models accept a timeout but reject extra_headers, which has no meaning without an HTTP request.

PydanticAI also supports typed output functions, tools and FallbackModel composition with this native interface. A decision selects a finite route or fills its finite typed fields; a separate generative model can handle a route that requires writing text. If tools perform real actions, the application remains responsible for its normal authorization rules.

from typing import Literal

from pydantic_ai import Agent

from gemmadecision import GemmaDecisionModel


def route(team: Literal['billing', 'account', 'technical']) -> str:
    """Select the queue for this ticket.

    Args:
        team: Which team should handle this ticket?
    """
    return f'Put this ticket in the {team} queue'


def main() -> None:
    agent = Agent(GemmaDecisionModel.local(), output_type=route)
    result = agent.run_sync('I was charged twice.')
    print(result.output)


if __name__ == '__main__':
    main()

Here route only returns a string; it does not send a ticket to an external queue. The tool selection recipe shows how to recommend a bounded action for an application to handle.

Confidence and performance

provider_details includes per-field confidence/distributions and probabilities_source='softmax_scores'. These numbers are probabilities derived from ranking scores. They are not measured probabilities of correctness on your data. Choose thresholds using held-out examples from your application. Changing decision_boolean_threshold changes the yes/no decision cutoff; it does not retrain the model or turn its scores into calibrated confidence. The human review recipe shows an explicit review option and an optional application-selected cutoff.

The model produces the decision in one response. PydanticAI's streaming API works by emitting the complete answer; it does not make inference faster or provide earlier partial tokens. Reuse a warm local model or server for low latency. The Python client and PydanticAI perform orchestration; inference speed comes from the serving backend and hardware.

Compatibility and testing

The pydantic-ai extra installs pydantic-ai-slim>=2.51,<3; the integration was tested against 2.51.0. PydanticAI is imported only when the integration module is used. Tests exercise actual PydanticAI agents against deterministic fake HTTP clients and local inference backends, including lazy shared loading, typed results, one-call multi-field decisions, streaming, output-function routing, response validation, timeouts, errors and client ownership. Those tests verify integration behavior, not model accuracy.

Reference: PydanticAI decision model documentation.