LLM Providers and Connections
An LLM connection is one endpoint an organization runs analysis against. It carries the address, the credential, the auth style and the structured-output mechanism — which is what a self-hosted model needs and what a stored API key could never express.
A provider kind is a wire protocol, and therefore which client the server runs. There are four:
| Kind | What it talks to |
|---|---|
anthropic | The hosted Claude API |
openai | The hosted OpenAI API |
mistral | The hosted Mistral API |
openai_compatible | Anything speaking /v1/chat/completions — your own model server, an OpenAI-compatible provider, or a gateway |
The fourth is the one that matters. Every self-hosted runtime and a dozen hosted providers speak the OpenAI Chat Completions format, so supporting the format rather than each vendor means adding a provider is a base URL and a catalog row, not a code change.
Self-hosted endpoints
Section titled “Self-hosted endpoints”Point OpenTremor at a model that runs inside your own network. The source under review never leaves it.
| Runtime | Base URL | Structured output |
|---|---|---|
| vLLM | http://host:8000/v1 | json_schema — real grammar-constrained decoding (xgrammar). The best option. |
| NVIDIA NIM | http://host:8000/v1 | json_schema |
| Text Generation Inference | http://host:8080/v1 | json_schema, sometimes json_object |
| Ollama | http://host:11434/v1 | json_object |
llama.cpp (llama-server) | http://host:8080/v1 | json_object, often prompt_only |
| LM Studio / LocalAI | http://host:1234/v1 | json_object |
Set auth_style: none for a server with no authentication — that is normal for an internal
endpoint behind network policy, and OpenTremor does not require a placeholder key for one.
OpenAI-compatible hosted providers
Section titled “OpenAI-compatible hosted providers”The same kind, a different address. Each needs its own catalog rows
(POST /admin/models with llm_backend: "openai_compatible"), because only you know what you
are paying per token.
| Provider | Base URL |
|---|---|
| DeepSeek | https://api.deepseek.com/v1 |
| Groq | https://api.groq.com/openai/v1 |
| Together AI | https://api.together.xyz/v1 |
| Fireworks AI | https://api.fireworks.ai/inference/v1 |
| xAI (Grok) | https://api.x.ai/v1 |
| Cerebras | https://api.cerebras.ai/v1 |
| OpenRouter | https://openrouter.ai/api/v1 |
| Nebius / Hugging Face Inference | vendor /v1 |
Gateways — LiteLLM, Portkey, Cloudflare AI Gateway — present the same interface. If you want failover, load balancing or spend caps across several providers, run one of those and point a single connection at it. OpenTremor deliberately does not reimplement that.
Structured output — the setting to get right
Section titled “Structured output — the setting to get right”Every analysis depends on the model returning a valid UnitAnalysis object. The hosted kinds each
have exactly one mechanism and use it. An openai_compatible endpoint could have any of four, so
it is a per-connection setting:
| Mode | Enforcement |
|---|---|
json_schema | The server rejects a response that doesn’t match. Strongest. |
tool_call | Forced function call. |
json_object | The server enforces JSON; the schema comes from the prompt. |
prompt_only | No server enforcement at all. |
The default for a new custom connection is prompt_only — an endpoint we have never seen has
told us nothing about what it enforces. Every non-tool_call mode is backed by a repair pass: a
response that fails schema validation is shown its own error and asked once more before the unit
fails. Use the connection probe to find out what your endpoint actually supports rather than
guessing:
curl -X POST "$API/orgs/$ORG/llm-connections/prod-vllm/test?model=qwen3-coder-30b" \ -H "Authorization: Bearer $TOKEN"{"reachable": true, "latency_ms": 812, "structured_output": true, "usage_reported": true, "error": null}reachable: true with structured_output: false is the interesting failure — the box is up and
the mode is wrong.
Who decides which addresses are reachable
Section titled “Who decides which addresses are reachable”An org admin creates connections. An operator decides where a custom address may point, via
llm.custom_endpoints — file/env only, never editable from the admin API, for the same reason
llm.credential_encryption_key isn’t: an admin who can widen it from inside the product can widen
it to the cloud instance metadata endpoint.
llm: custom_endpoints: enabled: false # cloud default — no custom base_urls at all allow_private_networks: false allowed_hosts: [] # exact hosts, or "*.suffix". Empty = any address the rules allowA self-hosted deployment ships enabled: true, allow_private_networks: true, because reaching
http://10.0.0.5:8000 is the entire point. A multi-tenant cloud instance ships the defaults
above, so signing up cannot turn the server into a probe of the operator’s network.
Two rules hold in both postures, and no setting turns them off:
- Link-local is always blocked (
169.254.0.0/16,fe80::/10). That is where cloud instance metadata lives. Allowing private networks means “my datacentre”, not “my instance’s IAM role”. - The resolved address is checked, not the hostname — otherwise any name with an A record
pointing at
127.0.0.1defeats the policy. Every address a name resolves to must pass, and the check runs again each time a connection is used, not only when it is saved.
The same policy governs the needs_review webhook URL, which is customer-supplied for the same
reasons.
Cost accounting
Section titled “Cost accounting”Prices come from the models catalog (Platform Admin → Models), which is keyed by provider kind and model name, not by connection. Two consequences:
- Rows for
openai_compatibleare not seeded. A price fordeepseek-chatwould also be applied to a self-hosted endpoint serving a model by that name, where the real per-token cost is zero. Add them yourself, with prices you choose. - A self-hosted connection catalogued at
0.00never trips an organization’s monthly USD budget. That is by design — there is no per-token cost to bound. Analysis-count entitlements (demo_analyses_used) are unaffected.
Some servers return no usage block at all. OpenTremor records that as unreported rather than
as zero tokens, and the probe reports it as usage_reported: false.
Compatibility with stored credentials
Section titled “Compatibility with stored credentials”POST /orgs/{org_id}/llm-credentials still works and still means the same thing. A credential is
stored as a connection whose id is its provider kind — which is what a hosted credential
always was. Nothing needs migrating, and llm_backend: "anthropic" in an analyze request keeps
resolving exactly as before.
New requests should send connection_id instead; exactly one of the two is required.