Skip to content

Helm Chart (Kubernetes)

The Helm chart deploys the OpenTremor runtime on any Kubernetes cluster. It covers four of the workloads a real deployment runs, plus a development MongoDB:

WorkloadValues flagWhat it is
corealways onThe API — Deployment, Service, config Secret, ServiceAccount, optional Ingress/HPA/PDB/PVC
dashboarddashboard.enabledThe Next.js UI, on its own Deployment+Service
docsdocs.enabledThis documentation site, served as static files by nginx
task-schedulertaskScheduler.enabledThe “cron” core doesn’t have — see Task Scheduler
mongodbmongodb.enabledA single unauthenticated MongoDB pod, for development only

The three optional workloads are opt-in only because each needs an image published to your own registry — not because they are optional in any other sense.


  • Kubernetes 1.26+
  • Helm 3.12+
  • A reachable MongoDB instance (in-cluster or external)

The chart never invents a credential for you, and refuses to render rather than deploy something insecure. Prepare these before the first install.

1. config.auth.apiKey — the bootstrap credential

Section titled “1. config.auth.apiKey — the bootstrap credential”

The one credential that has no environment-variable override anywhere in core. It is the bootstrap superadmin key — usable with no database round-trip, always superadmin, by design the credential of last resort — so it can only come from the config file itself. That is why the chart renders the core config into a Secret rather than a ConfigMap.

Terminal window
helm upgrade --install opentremor helm/opentremor \
--set config.auth.apiKey="$(openssl rand -base64 32)" \
...

Leave it empty and the chart fails with an explanation instead of installing. This is deliberate: an OpenTremor with no api_key does not become less secure by degrees, it resolves every request to the anonymous owner of the default org — a multi-tenant system with its tenancy switched off — and no env var can fix it after the fact. To accept that knowingly (a throwaway single-user cluster), set config.auth.disableAuthentication: true.

The task scheduler authenticates with this same value, read from the same Secret rather than a copy that can drift — the endpoints it calls require is_superadmin, which an org-scoped service account from POST /auth/keys can never be.

Never stored in a ConfigMap. In production, create the Secret out of band and point the chart at it; the URI is injected as MONGODB_URI, which core’s config loader gives precedence over the file:

Terminal window
kubectl create secret generic opentremor-mongo-secret \
--namespace opentremor \
--from-literal=mongodb-uri="mongodb://opentremor:<password>@mongo:27017/opentremor?authSource=opentremor"
existingSecret:
name: opentremor-mongo-secret
key: mongodb-uri

For development, mongodb.enabled: true deploys a bundled instance and the URI is derived from its Service — config.storage.mongodb.uri can stay empty. With neither, the chart fails to render rather than starting a pod that cannot reach a database.

JWT_SECRET, LLM_CREDENTIAL_KEY and TWO_FACTOR_SECRET_KEY are read from env vars, so they get first-class plumbing through one Secret:

Terminal window
kubectl create secret generic opentremor-app-secrets \
--namespace opentremor \
--from-literal=jwt-secret="$(openssl rand -base64 32)" \
--from-literal=llm-credential-key="$(openssl rand -base64 32)" \
--from-literal=two-factor-secret-key="$(openssl rand -base64 32)"
appSecrets:
existingSecret: opentremor-app-secrets
keys:
jwtSecret: jwt-secret
llmCredentialKey: llm-credential-key
twoFactorSecretKey: two-factor-secret-key
defaultAdminPassword: "" # optional — DEFAULT_ADMIN_PASSWORD
Terminal window
kubectl create secret docker-registry your-registry-pull-secret \
--namespace opentremor \
--docker-server=<your-registry> \
--docker-username=<username> \
--docker-password=<password-or-token>

To keep the API key out of Helm’s hands entirely — SOPS, External Secrets, a Vault injector — render the whole core config yourself and point the chart at it. The chart then renders no config Secret of its own:

existingConfigSecret:
name: opentremor-core-config
key: config.yaml

Set taskScheduler.existingApiKeySecret alongside it, so the scheduler still gets a credential.


config: in values.yaml carries only the sections core reads from its config file: deployment, storage, auth, cors.

Everything else you might expect to find there is managed live through the admin API and reset to its hardcoded default at every boot, so a value in a config file — chart-rendered or not — is silently discarded: logging, server, session, registration, github, retention, jobs, org_offboarding. mcp and telemetry go further still: both are read once, synchronously, inside create_app() before any stored override could reach them, so they are fixed for the life of the process and not editable anywhere.

Set the live ones once the pod is up:

Terminal window
curl -X PATCH https://opentremor.example.com/admin/settings \
-H "X-API-Key: <admin-key>" -H "Content-Type: application/json" \
-d '{"server_public_base_url": "https://opentremor.example.com"}'

server_public_base_url makes report links point at the real domain instead of the pod-internal hostname request.base_url resolves to. If your ingress controller forwards X-Forwarded-Proto and X-Forwarded-Host, the server picks those up and this can be skipped.

With the dashboard deployed, set server_dashboard_base_url too — invite links and the GitHub App manifest redirect point at dashboard pages, and without it they silently drop the /dashboard prefix:

Terminal window
-d '{"server_dashboard_base_url": "https://opentremor.example.com/dashboard"}'

Everything is served from one host, path-routed — same-origin, no separate subdomains, no extra reverse-proxy layer. Each paths[] entry names its backend with service, which accepts three shorthands (api, dashboard, docs) resolved to this release’s real Service names, so nothing has to hardcode a release name. Any other value is used as a literal Service name. Omitting service means api.

ingress:
enabled: true
className: nginx
annotations:
nginx.ingress.kubernetes.io/proxy-read-timeout: "3600" # SSE keep-alive
hosts:
- host: opentremor.example.com
paths:
- path: /
pathType: Prefix
- path: /dashboard
pathType: Prefix
service: dashboard
- path: /docs
pathType: Prefix
service: docs
tls:
- secretName: opentremor-tls
hosts:
- opentremor.example.com

Routing to a workload you haven’t enabled fails at render time rather than producing an Ingress that points at a Service which doesn’t exist.


dashboard:
enabled: true
basePath: /dashboard
replicaCount: 1
image:
repository: <your-registry>/opentremor-dashboard
tag: "a1b2c3d"
service:
port: 3000

Server-side (Server Component) fetches reach the API over the cluster network via API_BASE, which the chart sets for you. Browser-side calls are same-origin and resolve through the ingress.

docs:
enabled: true
basePath: /docs
image:
repository: <your-registry>/opentremor-docs
tag: "a1b2c3d"

The docs image is stock nginx serving a static build, which writes its pid and cache to paths a read-only root filesystem forbids and expects its master process to start as root. It therefore has its own docs.podSecurityContext/docs.securityContext rather than the API’s hardened pair.

That pair drops every capability and then adds exactly three back:

capabilities:
drop: [ALL]
add: [CHOWN, SETGID, SETUID]

Without them the pod cannot start at all. nginx’s entrypoint chowns its cache directories before the master forks, so with no CHOWN it exits 1 on nginx: [emerg] chown("/var/cache/nginx/ client_temp", 101) failed (1: Operation not permitted) and CrashLoopBackOffs; the master then starts its workers as the unprivileged nginx user, which needs SETGID and SETUID. The other capabilities stay dropped and allowPrivilegeEscalation stays false.

Rebuilding the image on nginxinc/nginx-unprivileged is what removes the need for all three, and would let this go back to a bare drop: [ALL].

taskScheduler:
enabled: true
image:
repository: <your-registry>/opentremor-task-scheduler
tag: "a1b2c3d"
jobs:
- name: retention_sweep
method: POST
path: /admin/retention/run
heartbeat_minutes: 15
- name: org_purge
method: POST
path: /admin/organizations/purge-pending
heartbeat_minutes: 60
- name: jobs_sweep
method: POST
path: /admin/jobs/sweep
heartbeat_minutes: 15

The heartbeats are deliberately short and dumb; the real schedule lives on core, which self-gates or naturally no-ops when nothing is due. Adding a job as core grows more scheduled admin actions is a values change, not a code change. core.base_url is derived from the release, and the credential comes from the Secret core reads.

Single replica, strategy: Recreate — a second one would double every heartbeat for no benefit. The Service is not routed through the ingress; reach its status endpoint directly:

Terminal window
kubectl --namespace opentremor port-forward svc/opentremor-task-scheduler 8080
curl http://127.0.0.1:8080/jobs

Rotating config.auth.apiKey changes the Secret, but an env var sourced from a secretKeyRef does not restart a running pod — follow a rotation with kubectl rollout restart deployment/opentremor-task-scheduler.

mongodb:
enabled: true
persistence:
enabled: true
size: 8Gi

A single hand-written Deployment and Service — not a sub-chart — with no authentication, no replica set and no backups, so helm install produces something that boots. Without persistence.enabled, every restart of that pod loses the whole database. Use a real managed MongoDB via existingSecret for anything else.

It runs mongo:7.0, not 8, and that is deliberate. Every MongoDB 8.x build refuses to start on Linux 6.19 and newer, exiting immediately with MongoDB cannot start: Linux kernel versions 6.19 and newer has a known incompatibility with this version of MongoDB (SERVER-121912). A container shares the host’s kernel, so no image or securityContext works around it — on a current kernel the pod simply CrashLoopBackOffs. If you point existingSecret at your own MongoDB, this constraint is yours to check against the kernel your nodes run.


Once the pod is up:

Terminal window
curl -X POST https://opentremor.example.com/admin/storage/prepare-mongodb \
-H "X-API-Key: <admin-key>"

Surfaced in the UI as Platform Admin → Settings → Prepare a MongoDB instance. It is idempotent and covers the full current collection/index set. See Platform Admin — Preparing MongoDB.


replicaCount: 2
autoscaling:
enabled: true
minReplicas: 2
maxReplicas: 5
targetCPUUtilizationPercentage: 70
targetMemoryUtilizationPercentage: 75
podDisruptionBudget:
enabled: true
minAvailable: 1 # set this or maxUnavailable, never both

Any replica count above one requires appSecrets.existingSecret — see secret 3 above.


Core exposes /metrics. With the Prometheus Operator CRDs installed, the chart can register a ServiceMonitor pointing at it — the same series OpenTremor Monitoring’s Grafana dashboard reads:

metrics:
serviceMonitor:
enabled: true
interval: 30s
labels:
release: kube-prometheus-stack # match your Prometheus's serviceMonitorSelector

resources:
requests: { cpu: 100m, memory: 256Mi }
limits: { cpu: 500m, memory: 512Mi }
podSecurityContext:
runAsNonRoot: true
runAsUser: 1000
fsGroup: 1000
securityContext:
allowPrivilegeEscalation: false
readOnlyRootFilesystem: true
capabilities:
drop: [ALL]

The API and task-scheduler pods run with a read-only root filesystem, so HOME and XDG_CACHE_HOME are pointed at an emptyDir mounted on /tmp — the images’ Poetry-based entrypoints need a writable cache directory and the pod’s uid has no home of its own.


A typical pipeline builds and deploys on PR (dev) and merge to main (prod):

EventTriggerEnvironment
PR opened / updatedpull_request → maindev
Merge to mainpush → mainprod
EventTag formatExample
Pull requestpr-<number>-<sha7>pr-42-a1b2c3d
Merge to main<sha7>a1b2c3d

Four images are built, one per workload: core, dashboard, docs, task-scheduler. The dashboard and docs images take the base-path build args described above, so a pipeline that serves them under a prefix must pass the same prefix at build time that the ingress routes at deploy time.

Secrets (repo-level):

SecretDescription
REGISTRY_URLRegistry hostname, e.g. registry.company.com
REGISTRY_USERNAMERegistry user
REGISTRY_TOKENRegistry password / token

Variable (repo-level):

VariableDescription
AWS_REGIONAWS region, e.g. eu-west-1

GitHub Environments (Settings → Environments → New environment):

Create dev and prod. In each, add:

SecretDescription
AWS_ROLE_ARNIAM role ARN (OIDC)
EKS_CLUSTER_NAMEEKS cluster name
OPENTREMOR_API_KEYThe bootstrap credential passed to --set config.auth.apiKey

Add a required reviewer protection rule on prod to gate production deploys.

Keyless authentication via GitHub OIDC. In your AWS account:

  1. Create an OIDC identity provider: token.actions.githubusercontent.com
  2. Create two IAM roles (one per environment) with this trust policy:
{
"Effect": "Allow",
"Principal": { "Federated": "<oidc-provider-arn>" },
"Action": "sts:AssumeRoleWithWebIdentity",
"Condition": {
"StringLike": {
"token.actions.githubusercontent.com:sub": "repo:<org>/opentremor-deployments:environment:<dev|prod>"
}
}
}

Each role needs eks:DescribeCluster and the permissions to run helm upgrade.


values-dev.yaml and values-prod.yaml are complete environments. values-eks.yaml is an additional layer — ALB ingress annotations, including the load-balancer idle timeout the MCP endpoint’s SSE streams need — and is stacked on one of the other two, never used alone:

Terminal window
helm upgrade --install opentremor helm/opentremor \
-f helm/opentremor/values-dev.yaml \
-f helm/opentremor/values-eks.yaml \
--set config.auth.apiKey="$(cat ./api-key)" \
--namespace opentremor --create-namespace

Terminal window
helm upgrade --install opentremor \
helm/opentremor/ \
--namespace opentremor \
--create-namespace \
--set image.repository="<your-registry>/opentremor-core" \
--set image.tag="a1b2c3d" \
--set config.auth.apiKey="$(cat ./api-key)" \
--values helm/opentremor/values-prod.yaml \
--atomic \
--timeout 5m

Terminal window
helm template opentremor helm/opentremor/ \
--values helm/opentremor/values-prod.yaml \
--set config.auth.apiKey=placeholder \
--namespace opentremor
helm diff upgrade opentremor helm/opentremor/ \
--values helm/opentremor/values-prod.yaml \
--namespace opentremor

The chart reports every misconfiguration it finds in one pass rather than one per run, so a failed helm template lists everything that needs fixing.


Terminal window
helm uninstall opentremor --namespace opentremor
# PVCs are not deleted automatically
kubectl delete pvc -l app.kubernetes.io/instance=opentremor -n opentremor