auto and an
explicit MODEL_PROVIDER can reach it, with the same retry policy, deadlines,
streaming and structured output as the built-in providers.
Everything a feature knows about generation is the two ports (TextModel,
StructuredModel in apps/api/src/shared/application/ports/). A provider is
anything that satisfies both; retries, timeouts and the error taxonomy are shared,
not per-provider. Adding Gemini followed the steps below, so trace that commit for
a complete worked example.
Two routes to the ports
- OpenAI-compatible (
/chat/completions) is the fast path. OpenCode Zen and Gemini both hang offOpenAiCompatibleClient, so most providers need only a thin injectable adapter plus configuration — steps 1–4. - A native route (like Anthropic’s
/messages) needs a full adapter written from scratch. The contract it must honour: text plus streaming via anAsyncGenerator, structured output re-validated against the zod schema, a hard deadline per call, and failures mapped to theModelProviderErrortaxonomy (llm_errorfor permanent faults,llm_unavailablewhen a retry could help). Useanthropic-model.adapter.tsas the template.
Notes
Selection is config-time; retries are per-provider.
auto picks once at boot,
and a provider that stays down yields 503 llm_unavailable after the retry budget
— never a silent switch to another provider mid-request.The key never surfaces. Failure logs keep a truncated response body and the
client-facing message names neither the key nor the URL.
Structured output is promised, not trusted. The provider is asked to constrain
itself with the zod JSON Schema, and the answer is re-validated anyway. A provider
that ignores
response_format produces the schema error, not corrupt data.Every provider shares the same deadline and retry policy. There is no
per-provider tuning surface on purpose: the budget lives in
LLM_RETRY_MAX_ATTEMPTS, LLM_RETRY_BASE_DELAY_MS, LLM_RETRY_MAX_DELAY_MS
and the *_TIMEOUT_MS variables.