Self-Hosted LLM Endpoint
The reason to do this is data residency: an organization that cannot send its infrastructure code to a third-party API can still run the analysis, against a model it operates. This page is the end-to-end setup. For the reference on every field, see LLM Providers and Connections.
1. Let the deployment reach your network
Section titled “1. Let the deployment reach your network”llm.custom_endpoints is operator-owned and file/env only — an org admin cannot widen it from the
dashboard. On a self-hosted instance:
llm: custom_endpoints: enabled: true allow_private_networks: true # Optional, and worth setting: restrict which hosts are reachable at all. allowed_hosts: ["*.internal.corp"]Without enabled: true, creating a connection with a base_url returns a 400 naming this
setting. Link-local addresses stay blocked regardless — see the reference page.
2. Run a model server
Section titled “2. Run a model server”vLLM is the one to reach for: it is the only widely-deployed option with real grammar-constrained decoding, which is what makes structured output a guarantee rather than a hope.
vllm serve Qwen/Qwen3-Coder-30B-A3B-Instruct \ --host 0.0.0.0 --port 8000 \ --guided-decoding-backend xgrammarOllama works for evaluation and needs nothing but ollama serve; its endpoint is
http://host:11434/v1 and its JSON enforcement is weaker, so expect json_object rather than
json_schema.
3. Create the connection
Section titled “3. Create the connection”Dashboard → LLM Connections → Add, or:
curl -X POST "$API/orgs/$ORG/llm-connections" \ -H "Authorization: Bearer $TOKEN" -H "Content-Type: application/json" \ -d '{ "connection_id": "prod-vllm", "display_name": "Production vLLM (eu-west)", "provider_kind": "openai_compatible", "base_url": "http://llm.internal.corp:8000/v1", "auth_style": "none", "structured_output_mode": "json_schema" }'auth_style: "none" is correct for a server behind network policy with no authentication.
For an endpoint served under a private CA, pass the bundle as tls_ca_pem — it is stored
encrypted and never returned.
4. Probe it before trusting it
Section titled “4. Probe it before trusting it”curl -X POST "$API/orgs/$ORG/llm-connections/prod-vllm/test?model=Qwen/Qwen3-Coder-30B-A3B-Instruct" \ -H "Authorization: Bearer $TOKEN"One real analysis call against a two-line fixture. It exercises the whole path — address, auth, model name, structured output — so a pass means the connection works, not merely that something answered. Do this now rather than discovering a typo forty units into a hundred-unit job.
5. Catalogue the model
Section titled “5. Catalogue the model”The allowed-models catalog gates every analysis call, and openai_compatible rows are not seeded
(a seeded price would be wrong for your hardware). As a platform admin:
curl -X POST "$API/admin/models" \ -H "Authorization: Bearer $TOKEN" -H "Content-Type: application/json" \ -d '{"llm_backend": "openai_compatible", "model": "Qwen/Qwen3-Coder-30B-A3B-Instruct", "input_price_per_1m_usd": 0, "output_price_per_1m_usd": 0}'Zero is the honest price for hardware you already own — with the caveat that a zero-cost connection never trips an organization’s monthly USD budget. If you want internal chargeback, enter your own notional rate instead.
6. Use it
Section titled “6. Use it”curl -X POST "$API/terraform-code-change/$NAMESPACE/analyze" \ -H "Authorization: Bearer $TOKEN" -H "Content-Type: application/json" \ -d '{"connection_id": "prod-vllm", "model": "Qwen/Qwen3-Coder-30B-A3B-Instruct", "raw_input": "'"$(cat change.diff)"'"}'The GitHub integration takes the same connection_id, so pull-request checks run against your
endpoint too — which is where an on-prem-only organization needs it most.
What to expect
Section titled “What to expect”A job that fails with did not complete is usually the endpoint, not OpenTremor: those failures
are treated as transient and retried within analysis.llm_max_attempts. A job that fails with a
truncation error wants a larger analysis.llm_max_output_tokens — small models are more verbose,
not less.