Skip to main content

LLM (Voice Mode)

Patter supports two voice architectures: See Engines for engine-mode reference. This page focuses on the llm= selector in pipeline mode.

Pipeline mode

Compose the three stages independently. Each provider reads its credentials from the environment by default.
Tool calling works across every provider — each adapter normalizes its vendor-specific streaming format to Patter’s unified {type: "text" | "tool_call" | "done"} chunk protocol, so your tools are defined once and run everywhere.
llm= and on_message are mutually exclusive. Pass one or the other on serve() — passing both raises a clear error at serve() time. When engine= is set, llm= is ignored (with a one-time warning in the logs). If neither llm= nor on_message is passed and OPENAI_API_KEY is set, Patter auto-constructs the default OpenAI LLM loop — existing 0.5.0 code still works.

Supported LLM providers

All classes accept api_key: str | None = None and fall back to the listed env var when it is omitted.

OpenAILLM

OpenAI Chat Completions with streaming + tool calling. Default model "gpt-4o-mini". For other OpenAI-compatible endpoints use the dedicated wrappers (GroqLLM, CerebrasLLM) — they subclass OpenAILLMProvider with the right base_url.

AnthropicLLM

Anthropic Messages API with native streaming and tool_use blocks, normalised to Patter’s chunk protocol. Default model "claude-haiku-4-5-20251001", default max_tokens=1024 (Anthropic requires an explicit cap on every request). Prompt caching is enabled by default — cache_control: { type: "ephemeral" } is attached to the system prompt and the last tool block, which cuts time-to-first-token on long system prompts and large tool catalogs. Pass prompt_caching=False to disable.
Install: pip install 'getpatter[anthropic]'.

GroqLLM

Hardware-accelerated Llama inference via Groq’s OpenAI-compatible Chat Completions API at https://api.groq.com/openai/v1. Default model "llama-3.3-70b-versatile".
Install: pip install 'getpatter[groq]'.

CerebrasLLM

Cerebras Inference API (OpenAI-compatible) at https://api.cerebras.ai/v1. Default model "gpt-oss-120b" — production tier, ~3000 tok/sec on WSE-3, no deprecation date. Pass model="llama3.1-8b" (8B params, sub-100ms TTFT) for the smaller free-tier alternative. The 404 model_not_found error includes a recovery hint listing other valid IDs (qwen-3-235b-a22b-instruct-2507, llama-3.3-70b on paid tier). Supports forwarding all OpenAI-style sampling kwargs (response_format, parallel_tool_calls, tool_choice, seed, top_p, frequency_penalty, presence_penalty, stop) and optional msgpack + gzip payload compression (enabled by default) — see Cerebras payload optimization. Failures retry once with exponential backoff and honour x-ratelimit-reset-* advisory headers; terminal errors raise PatterError.
Install: pip install 'getpatter[cerebras]'.

GoogleLLM

Google Gemini via the google-genai SDK. Supports the Gemini Developer API (API key) and Vertex AI (GCP project + location). Default model "gemini-2.5-flash".
Install: pip install 'getpatter[google]'.

CustomLLM (any OpenAI-compatible endpoint)

The industry-standard “Custom LLM” pattern: point Patter’s pipeline at any endpoint that speaks the OpenAI Chat Completions protocol (SSE streaming, optional tool calls). One provider covers:
  • agent runtimesHermes and OpenClaw presets (HermesLLM, OpenClawLLM) subclass this same engine with the right defaults baked in; prefer them when they exist,
  • local inference gateways — Ollama, vLLM, LM Studio (keyless OK),
  • your own service implementing /chat/completions.
CustomLLM is the canonical name for the generic engine also exported as OpenAICompatibleLLM — both construct the same class. Barge-in cancellation (including the pre-first-token abort for slow agent runtimes), the long-turn filler (long_turn_message), the spoken error fallback (llm_error_message), and usage-based cost attribution all work unchanged. All OpenAI-style sampling kwargs are forwarded.

Custom LLM via on_message

For cases the five built-in providers don’t cover — multi-model routing, local llama.cpp, an internal gateway, caching layers — drop llm= and plug an async on_message callback instead:
on_message and llm= cannot be used together. Combining them raises a clear error at serve() time — pick one.

Advanced: building a custom LLM provider

Three primitives are exported from the package barrel for users who need to plug in a custom LLM or tool dispatcher:
  • LLMChunk — the streaming-output type yielded by every LLMProvider.stream(...) implementation. Carries either a partial text delta, a tool-call delta, or a stream-end marker.
  • DefaultToolExecutor — the default tool dispatcher used by LLMLoop. Constructs from a tools= list and resolves both Python handler= callables and webhook_url= HTTP tools. Override its hooks to swap in custom error handling, telemetry, or authentication.
  • OpenAILLMProvider — the parent class shared by OpenAILLM, GroqLLM, CerebrasLLM. Sampling kwargs (temperature, top_p, seed, tool_choice, response_format, …) live here and are forwarded by every subclass.
These are stable public symbols mirrored byte-for-byte by the TypeScript SDK.

What’s next

STT

STT providers for pipeline mode.

TTS

TTS providers for pipeline mode.

Tools

Function calling (works across every LLM).

Engines

Speech-to-speech engines (OpenAI Realtime, ElevenLabs ConvAI).