Skip to content

LLM Providers and Connections

An LLM connection is one endpoint an organization runs analysis against. It carries the address, the credential, the auth style and the structured-output mechanism — which is what a self-hosted model needs and what a stored API key could never express.

A provider kind is a wire protocol, and therefore which client the server runs. There are four:

KindWhat it talks to
anthropicThe hosted Claude API
openaiThe hosted OpenAI API
mistralThe hosted Mistral API
openai_compatibleAnything speaking /v1/chat/completions — your own model server, an OpenAI-compatible provider, or a gateway

The fourth is the one that matters. Every self-hosted runtime and a dozen hosted providers speak the OpenAI Chat Completions format, so supporting the format rather than each vendor means adding a provider is a base URL and a catalog row, not a code change.

Point OpenTremor at a model that runs inside your own network. The source under review never leaves it.

RuntimeBase URLStructured output
vLLMhttp://host:8000/v1json_schema — real grammar-constrained decoding (xgrammar). The best option.
NVIDIA NIMhttp://host:8000/v1json_schema
Text Generation Inferencehttp://host:8080/v1json_schema, sometimes json_object
Ollamahttp://host:11434/v1json_object
llama.cpp (llama-server)http://host:8080/v1json_object, often prompt_only
LM Studio / LocalAIhttp://host:1234/v1json_object

Set auth_style: none for a server with no authentication — that is normal for an internal endpoint behind network policy, and OpenTremor does not require a placeholder key for one.

The same kind, a different address. Each needs its own catalog rows (POST /admin/models with llm_backend: "openai_compatible"), because only you know what you are paying per token.

ProviderBase URL
DeepSeekhttps://api.deepseek.com/v1
Groqhttps://api.groq.com/openai/v1
Together AIhttps://api.together.xyz/v1
Fireworks AIhttps://api.fireworks.ai/inference/v1
xAI (Grok)https://api.x.ai/v1
Cerebrashttps://api.cerebras.ai/v1
OpenRouterhttps://openrouter.ai/api/v1
Nebius / Hugging Face Inferencevendor /v1

Gateways — LiteLLM, Portkey, Cloudflare AI Gateway — present the same interface. If you want failover, load balancing or spend caps across several providers, run one of those and point a single connection at it. OpenTremor deliberately does not reimplement that.

Structured output — the setting to get right

Section titled “Structured output — the setting to get right”

Every analysis depends on the model returning a valid UnitAnalysis object. The hosted kinds each have exactly one mechanism and use it. An openai_compatible endpoint could have any of four, so it is a per-connection setting:

ModeEnforcement
json_schemaThe server rejects a response that doesn’t match. Strongest.
tool_callForced function call.
json_objectThe server enforces JSON; the schema comes from the prompt.
prompt_onlyNo server enforcement at all.

The default for a new custom connection is prompt_only — an endpoint we have never seen has told us nothing about what it enforces. Every non-tool_call mode is backed by a repair pass: a response that fails schema validation is shown its own error and asked once more before the unit fails. Use the connection probe to find out what your endpoint actually supports rather than guessing:

Terminal window
curl -X POST "$API/orgs/$ORG/llm-connections/prod-vllm/test?model=qwen3-coder-30b" \
-H "Authorization: Bearer $TOKEN"
{"reachable": true, "latency_ms": 812, "structured_output": true,
"usage_reported": true, "error": null}

reachable: true with structured_output: false is the interesting failure — the box is up and the mode is wrong.

An org admin creates connections. An operator decides where a custom address may point, via llm.custom_endpoints — file/env only, never editable from the admin API, for the same reason llm.credential_encryption_key isn’t: an admin who can widen it from inside the product can widen it to the cloud instance metadata endpoint.

llm:
custom_endpoints:
enabled: false # cloud default — no custom base_urls at all
allow_private_networks: false
allowed_hosts: [] # exact hosts, or "*.suffix". Empty = any address the rules allow

A self-hosted deployment ships enabled: true, allow_private_networks: true, because reaching http://10.0.0.5:8000 is the entire point. A multi-tenant cloud instance ships the defaults above, so signing up cannot turn the server into a probe of the operator’s network.

Two rules hold in both postures, and no setting turns them off:

  • Link-local is always blocked (169.254.0.0/16, fe80::/10). That is where cloud instance metadata lives. Allowing private networks means “my datacentre”, not “my instance’s IAM role”.
  • The resolved address is checked, not the hostname — otherwise any name with an A record pointing at 127.0.0.1 defeats the policy. Every address a name resolves to must pass, and the check runs again each time a connection is used, not only when it is saved.

The same policy governs the needs_review webhook URL, which is customer-supplied for the same reasons.

Prices come from the models catalog (Platform Admin → Models), which is keyed by provider kind and model name, not by connection. Two consequences:

  • Rows for openai_compatible are not seeded. A price for deepseek-chat would also be applied to a self-hosted endpoint serving a model by that name, where the real per-token cost is zero. Add them yourself, with prices you choose.
  • A self-hosted connection catalogued at 0.00 never trips an organization’s monthly USD budget. That is by design — there is no per-token cost to bound. Analysis-count entitlements (demo_analyses_used) are unaffected.

Some servers return no usage block at all. OpenTremor records that as unreported rather than as zero tokens, and the probe reports it as usage_reported: false.

POST /orgs/{org_id}/llm-credentials still works and still means the same thing. A credential is stored as a connection whose id is its provider kind — which is what a hosted credential always was. Nothing needs migrating, and llm_backend: "anthropic" in an analyze request keeps resolving exactly as before.

New requests should send connection_id instead; exactly one of the two is required.