LLM (Voice Mode)
Patter supports two voice architectures:
See Engines for engine-mode reference. This page focuses on the
llm selector in pipeline mode.
Pipeline mode
Compose the three stages independently. Each provider reads its credentials from the environment by default.{ type: "text" | "tool_call" | "done" } chunk protocol, so your tools are defined once and run everywhere.
llm and onMessage are mutually exclusive. Pass one or the other on serve() — passing both raises a clear error at serve() time. When engine is set, llm is ignored (with a one-time warning in the logs). If neither llm nor onMessage is passed and OPENAI_API_KEY is set, Patter auto-constructs the default OpenAI LLM loop — existing 0.5.0 code still works.Supported LLM providers
All classes accept an options object with
apiKey?: string and fall back to the listed env var when it is omitted.
OpenAILLM
OpenAI Chat Completions with streaming + tool calling. Default model"gpt-4o-mini".
AnthropicLLM
Anthropic Messages API with native streaming andtool_use blocks, normalised to Patter’s chunk protocol. Default model "claude-haiku-4-5-20251001". Pass maxTokens to override the default token cap.
Prompt caching is enabled by default — cache_control: { type: "ephemeral" } is attached to the system prompt and the last tool block, which cuts time-to-first-token on long system prompts and large tool catalogs. Pass promptCaching: false to disable.
GroqLLM
Hardware-accelerated Llama inference via Groq’s OpenAI-compatible Chat Completions API athttps://api.groq.com/openai/v1. Default model "llama-3.3-70b-versatile".
CerebrasLLM
Cerebras Inference API (OpenAI-compatible) athttps://api.cerebras.ai/v1. Default model "gpt-oss-120b" — production tier, ~3000 tok/sec on WSE-3, no deprecation date. Pass model: "llama3.1-8b" for the smaller free-tier alternative. The 404 model_not_found error includes a recovery hint listing other valid IDs.
Supports forwarding OpenAI-style sampling kwargs (responseFormat, parallelToolCalls, toolChoice, seed, topP, frequencyPenalty, presencePenalty, stop) and gzip request-body compression (enabled by default) — see Cerebras payload optimization. Failures retry once with exponential backoff and honour x-ratelimit-reset-* advisory headers; terminal errors throw PatterError.
GoogleLLM
Google Gemini via the Developer API (streaming SSE). Default model"gemini-2.5-flash".
CustomLLM (any OpenAI-compatible endpoint)
The industry-standard “Custom LLM” pattern: point Patter’s pipeline at any endpoint that speaks the OpenAI Chat Completions protocol (SSE streaming, optional tool calls). One provider covers:- agent runtimes — Hermes and OpenClaw presets (
HermesLLM,OpenClawLLM) subclass this same engine with the right defaults baked in; prefer them when they exist, - local inference gateways — Ollama, vLLM, LM Studio (keyless OK),
- your own service implementing
/chat/completions.
CustomLLM is the canonical name for the generic engine also exported as OpenAICompatibleLLM — both construct the same class. Barge-in cancellation (including the pre-first-token abort for slow agent runtimes), the long-turn filler (longTurnMessage), the spoken error fallback (llmErrorMessage), and usage-based cost attribution all work unchanged. All OpenAI-style sampling options are forwarded.
Custom LLM via onMessage
For cases the five built-in providers don’t cover — multi-model routing, local inference, an internal gateway, caching layers — drop llm and plug an async onMessage callback instead:
Advanced: building a custom LLM provider
Three primitives are exported from the package barrel for users who need to plug in a custom LLM or tool dispatcher:LLMChunk— the streaming-output type yielded by everyLLMProvider.stream(...)implementation. Carries either a partial text delta, a tool-call delta, or a stream-end marker.DefaultToolExecutor— the default tool dispatcher used byLLMLoop. Constructs from atoolsarray and resolves both inlinehandlercallables andwebhookUrlHTTP tools. Override its hooks to swap in custom error handling, telemetry, or authentication.OpenAILLMProvider— the parent class shared byOpenAILLM,GroqLLM,CerebrasLLM. Sampling options (temperature,topP,seed,toolChoice,responseFormat, …) live here and are forwarded by every subclass.LLMLoop— the orchestration loop wiring anLLMProvider, aDefaultToolExecutor, and the streaming output back to TTS.
What’s next
STT
STT providers for pipeline mode.
TTS
TTS providers for pipeline mode.
Tools
Function calling (works across every LLM).
Engines
Speech-to-speech engines (OpenAI Realtime, ElevenLabs ConvAI).

