Skip to content

Self-Hosted LLM Endpoint

The reason to do this is data residency: an organization that cannot send its infrastructure code to a third-party API can still run the analysis, against a model it operates. This page is the end-to-end setup. For the reference on every field, see LLM Providers and Connections.

llm.custom_endpoints is operator-owned and file/env only — an org admin cannot widen it from the dashboard. On a self-hosted instance:

llm:
custom_endpoints:
enabled: true
allow_private_networks: true
# Optional, and worth setting: restrict which hosts are reachable at all.
allowed_hosts: ["*.internal.corp"]

Without enabled: true, creating a connection with a base_url returns a 400 naming this setting. Link-local addresses stay blocked regardless — see the reference page.

vLLM is the one to reach for: it is the only widely-deployed option with real grammar-constrained decoding, which is what makes structured output a guarantee rather than a hope.

Terminal window
vllm serve Qwen/Qwen3-Coder-30B-A3B-Instruct \
--host 0.0.0.0 --port 8000 \
--guided-decoding-backend xgrammar

Ollama works for evaluation and needs nothing but ollama serve; its endpoint is http://host:11434/v1 and its JSON enforcement is weaker, so expect json_object rather than json_schema.

Dashboard → LLM Connections → Add, or:

Terminal window
curl -X POST "$API/orgs/$ORG/llm-connections" \
-H "Authorization: Bearer $TOKEN" -H "Content-Type: application/json" \
-d '{
"connection_id": "prod-vllm",
"display_name": "Production vLLM (eu-west)",
"provider_kind": "openai_compatible",
"base_url": "http://llm.internal.corp:8000/v1",
"auth_style": "none",
"structured_output_mode": "json_schema"
}'

auth_style: "none" is correct for a server behind network policy with no authentication. For an endpoint served under a private CA, pass the bundle as tls_ca_pem — it is stored encrypted and never returned.

Terminal window
curl -X POST "$API/orgs/$ORG/llm-connections/prod-vllm/test?model=Qwen/Qwen3-Coder-30B-A3B-Instruct" \
-H "Authorization: Bearer $TOKEN"

One real analysis call against a two-line fixture. It exercises the whole path — address, auth, model name, structured output — so a pass means the connection works, not merely that something answered. Do this now rather than discovering a typo forty units into a hundred-unit job.

The allowed-models catalog gates every analysis call, and openai_compatible rows are not seeded (a seeded price would be wrong for your hardware). As a platform admin:

Terminal window
curl -X POST "$API/admin/models" \
-H "Authorization: Bearer $TOKEN" -H "Content-Type: application/json" \
-d '{"llm_backend": "openai_compatible", "model": "Qwen/Qwen3-Coder-30B-A3B-Instruct",
"input_price_per_1m_usd": 0, "output_price_per_1m_usd": 0}'

Zero is the honest price for hardware you already own — with the caveat that a zero-cost connection never trips an organization’s monthly USD budget. If you want internal chargeback, enter your own notional rate instead.

Terminal window
curl -X POST "$API/terraform-code-change/$NAMESPACE/analyze" \
-H "Authorization: Bearer $TOKEN" -H "Content-Type: application/json" \
-d '{"connection_id": "prod-vllm",
"model": "Qwen/Qwen3-Coder-30B-A3B-Instruct",
"raw_input": "'"$(cat change.diff)"'"}'

The GitHub integration takes the same connection_id, so pull-request checks run against your endpoint too — which is where an on-prem-only organization needs it most.

A job that fails with did not complete is usually the endpoint, not OpenTremor: those failures are treated as transient and retried within analysis.llm_max_attempts. A job that fails with a truncation error wants a larger analysis.llm_max_output_tokens — small models are more verbose, not less.