Helm Chart (Kubernetes)
The Helm chart deploys the OpenTremor runtime on any Kubernetes cluster. It covers four of the workloads a real deployment runs, plus a development MongoDB:
| Workload | Values flag | What it is |
|---|---|---|
| core | always on | The API — Deployment, Service, config Secret, ServiceAccount, optional Ingress/HPA/PDB/PVC |
| dashboard | dashboard.enabled | The Next.js UI, on its own Deployment+Service |
| docs | docs.enabled | This documentation site, served as static files by nginx |
| task-scheduler | taskScheduler.enabled | The “cron” core doesn’t have — see Task Scheduler |
| mongodb | mongodb.enabled | A single unauthenticated MongoDB pod, for development only |
The three optional workloads are opt-in only because each needs an image published to your own registry — not because they are optional in any other sense.
Prerequisites
Section titled “Prerequisites”- Kubernetes 1.26+
- Helm 3.12+
- A reachable MongoDB instance (in-cluster or external)
The four secrets
Section titled “The four secrets”The chart never invents a credential for you, and refuses to render rather than deploy something insecure. Prepare these before the first install.
1. config.auth.apiKey — the bootstrap credential
Section titled “1. config.auth.apiKey — the bootstrap credential”The one credential that has no environment-variable override anywhere in core. It is the bootstrap superadmin key — usable with no database round-trip, always superadmin, by design the credential of last resort — so it can only come from the config file itself. That is why the chart renders the core config into a Secret rather than a ConfigMap.
helm upgrade --install opentremor helm/opentremor \ --set config.auth.apiKey="$(openssl rand -base64 32)" \ ...Leave it empty and the chart fails with an explanation instead of installing. This is deliberate:
an OpenTremor with no api_key does not become less secure by degrees, it resolves every
request to the anonymous owner of the default org — a multi-tenant system with its tenancy
switched off — and no env var can fix it after the fact. To accept that knowingly (a throwaway
single-user cluster), set config.auth.disableAuthentication: true.
The task scheduler authenticates with this same value, read from the same Secret rather than a
copy that can drift — the endpoints it calls require is_superadmin, which an org-scoped service
account from POST /auth/keys can never be.
2. The MongoDB URI
Section titled “2. The MongoDB URI”Never stored in a ConfigMap. In production, create the Secret out of band and point the chart at
it; the URI is injected as MONGODB_URI, which core’s config loader gives precedence over the
file:
kubectl create secret generic opentremor-mongo-secret \ --namespace opentremor \ --from-literal=mongodb-uri="mongodb://opentremor:<password>@mongo:27017/opentremor?authSource=opentremor"existingSecret: name: opentremor-mongo-secret key: mongodb-uriFor development, mongodb.enabled: true deploys a bundled instance and the URI is derived from
its Service — config.storage.mongodb.uri can stay empty. With neither, the chart fails to render
rather than starting a pod that cannot reach a database.
3. The signing and encryption keys
Section titled “3. The signing and encryption keys”JWT_SECRET, LLM_CREDENTIAL_KEY and TWO_FACTOR_SECRET_KEY are read from env vars, so they get
first-class plumbing through one Secret:
kubectl create secret generic opentremor-app-secrets \ --namespace opentremor \ --from-literal=jwt-secret="$(openssl rand -base64 32)" \ --from-literal=llm-credential-key="$(openssl rand -base64 32)" \ --from-literal=two-factor-secret-key="$(openssl rand -base64 32)"appSecrets: existingSecret: opentremor-app-secrets keys: jwtSecret: jwt-secret llmCredentialKey: llm-credential-key twoFactorSecretKey: two-factor-secret-key defaultAdminPassword: "" # optional — DEFAULT_ADMIN_PASSWORD4. Registry credentials
Section titled “4. Registry credentials”kubectl create secret docker-registry your-registry-pull-secret \ --namespace opentremor \ --docker-server=<your-registry> \ --docker-username=<username> \ --docker-password=<password-or-token>Bringing your own config file
Section titled “Bringing your own config file”To keep the API key out of Helm’s hands entirely — SOPS, External Secrets, a Vault injector — render the whole core config yourself and point the chart at it. The chart then renders no config Secret of its own:
existingConfigSecret: name: opentremor-core-config key: config.yamlSet taskScheduler.existingApiKeySecret alongside it, so the scheduler still gets a credential.
What the chart can and cannot configure
Section titled “What the chart can and cannot configure”config: in values.yaml carries only the sections core reads from its config file:
deployment, storage, auth, cors.
Everything else you might expect to find there is managed live through the admin API and reset to
its hardcoded default at every boot, so a value in a config file — chart-rendered or not — is
silently discarded: logging, server, session, registration, github, retention, jobs,
org_offboarding. mcp and telemetry go further still: both are read once, synchronously,
inside create_app() before any stored override could reach them, so they are fixed for the life
of the process and not editable anywhere.
Set the live ones once the pod is up:
curl -X PATCH https://opentremor.example.com/admin/settings \ -H "X-API-Key: <admin-key>" -H "Content-Type: application/json" \ -d '{"server_public_base_url": "https://opentremor.example.com"}'server_public_base_url makes report links point at the real domain instead of the pod-internal
hostname request.base_url resolves to. If your ingress controller forwards X-Forwarded-Proto
and X-Forwarded-Host, the server picks those up and this can be skipped.
With the dashboard deployed, set server_dashboard_base_url too — invite links and the GitHub App
manifest redirect point at dashboard pages, and without it they silently drop the /dashboard
prefix:
-d '{"server_dashboard_base_url": "https://opentremor.example.com/dashboard"}'Ingress routing
Section titled “Ingress routing”Everything is served from one host, path-routed — same-origin, no separate subdomains, no extra
reverse-proxy layer. Each paths[] entry names its backend with service, which accepts three
shorthands (api, dashboard, docs) resolved to this release’s real Service names, so nothing
has to hardcode a release name. Any other value is used as a literal Service name. Omitting
service means api.
ingress: enabled: true className: nginx annotations: nginx.ingress.kubernetes.io/proxy-read-timeout: "3600" # SSE keep-alive hosts: - host: opentremor.example.com paths: - path: / pathType: Prefix - path: /dashboard pathType: Prefix service: dashboard - path: /docs pathType: Prefix service: docs tls: - secretName: opentremor-tls hosts: - opentremor.example.comRouting to a workload you haven’t enabled fails at render time rather than producing an Ingress that points at a Service which doesn’t exist.
Workloads
Section titled “Workloads”Dashboard
Section titled “Dashboard”dashboard: enabled: true basePath: /dashboard replicaCount: 1 image: repository: <your-registry>/opentremor-dashboard tag: "a1b2c3d" service: port: 3000Server-side (Server Component) fetches reach the API over the cluster network via API_BASE, which
the chart sets for you. Browser-side calls are same-origin and resolve through the ingress.
docs: enabled: true basePath: /docs image: repository: <your-registry>/opentremor-docs tag: "a1b2c3d"The docs image is stock nginx serving a static build, which writes its pid and cache to paths a
read-only root filesystem forbids and expects its master process to start as root. It therefore has
its own docs.podSecurityContext/docs.securityContext rather than the API’s hardened pair.
That pair drops every capability and then adds exactly three back:
capabilities: drop: [ALL] add: [CHOWN, SETGID, SETUID]Without them the pod cannot start at all. nginx’s entrypoint chowns its cache directories before
the master forks, so with no CHOWN it exits 1 on nginx: [emerg] chown("/var/cache/nginx/ client_temp", 101) failed (1: Operation not permitted) and CrashLoopBackOffs; the master then
starts its workers as the unprivileged nginx user, which needs SETGID and SETUID. The other
capabilities stay dropped and allowPrivilegeEscalation stays false.
Rebuilding the image on nginxinc/nginx-unprivileged is what removes the need for all three, and
would let this go back to a bare drop: [ALL].
Task scheduler
Section titled “Task scheduler”taskScheduler: enabled: true image: repository: <your-registry>/opentremor-task-scheduler tag: "a1b2c3d" jobs: - name: retention_sweep method: POST path: /admin/retention/run heartbeat_minutes: 15 - name: org_purge method: POST path: /admin/organizations/purge-pending heartbeat_minutes: 60 - name: jobs_sweep method: POST path: /admin/jobs/sweep heartbeat_minutes: 15The heartbeats are deliberately short and dumb; the real schedule lives on core, which self-gates
or naturally no-ops when nothing is due. Adding a job as core grows more scheduled admin actions is
a values change, not a code change. core.base_url is derived from the release, and the credential
comes from the Secret core reads.
Single replica, strategy: Recreate — a second one would double every heartbeat for no benefit.
The Service is not routed through the ingress; reach its status endpoint directly:
kubectl --namespace opentremor port-forward svc/opentremor-task-scheduler 8080curl http://127.0.0.1:8080/jobsRotating config.auth.apiKey changes the Secret, but an env var sourced from a secretKeyRef does
not restart a running pod — follow a rotation with
kubectl rollout restart deployment/opentremor-task-scheduler.
Bundled MongoDB
Section titled “Bundled MongoDB”mongodb: enabled: true persistence: enabled: true size: 8GiA single hand-written Deployment and Service — not a sub-chart — with no authentication, no replica
set and no backups, so helm install produces something that boots. Without
persistence.enabled, every restart of that pod loses the whole database. Use a real managed
MongoDB via existingSecret for anything else.
It runs mongo:7.0, not 8, and that is deliberate. Every MongoDB 8.x build refuses to start on
Linux 6.19 and newer, exiting immediately with MongoDB cannot start: Linux kernel versions 6.19 and newer has a known incompatibility with this version of MongoDB
(SERVER-121912). A container shares the host’s
kernel, so no image or securityContext works around it — on a current kernel the pod simply
CrashLoopBackOffs. If you point existingSecret at your own MongoDB, this constraint is yours to
check against the kernel your nodes run.
Initialise MongoDB indexes
Section titled “Initialise MongoDB indexes”Once the pod is up:
curl -X POST https://opentremor.example.com/admin/storage/prepare-mongodb \ -H "X-API-Key: <admin-key>"Surfaced in the UI as Platform Admin → Settings → Prepare a MongoDB instance. It is idempotent and covers the full current collection/index set. See Platform Admin — Preparing MongoDB.
Scaling and availability
Section titled “Scaling and availability”replicaCount: 2
autoscaling: enabled: true minReplicas: 2 maxReplicas: 5 targetCPUUtilizationPercentage: 70 targetMemoryUtilizationPercentage: 75
podDisruptionBudget: enabled: true minAvailable: 1 # set this or maxUnavailable, never bothAny replica count above one requires appSecrets.existingSecret — see secret 3 above.
Metrics
Section titled “Metrics”Core exposes /metrics. With the Prometheus Operator CRDs installed, the chart can register a
ServiceMonitor pointing at it — the same series
OpenTremor Monitoring’s Grafana dashboard reads:
metrics: serviceMonitor: enabled: true interval: 30s labels: release: kube-prometheus-stack # match your Prometheus's serviceMonitorSelectorResource limits and security contexts
Section titled “Resource limits and security contexts”resources: requests: { cpu: 100m, memory: 256Mi } limits: { cpu: 500m, memory: 512Mi }
podSecurityContext: runAsNonRoot: true runAsUser: 1000 fsGroup: 1000
securityContext: allowPrivilegeEscalation: false readOnlyRootFilesystem: true capabilities: drop: [ALL]The API and task-scheduler pods run with a read-only root filesystem, so HOME and
XDG_CACHE_HOME are pointed at an emptyDir mounted on /tmp — the images’ Poetry-based
entrypoints need a writable cache directory and the pod’s uid has no home of its own.
CI/CD pipeline
Section titled “CI/CD pipeline”A typical pipeline builds and deploys on PR (dev) and merge to main (prod):
| Event | Trigger | Environment |
|---|---|---|
| PR opened / updated | pull_request → main | dev |
| Merge to main | push → main | prod |
Image tagging
Section titled “Image tagging”| Event | Tag format | Example |
|---|---|---|
| Pull request | pr-<number>-<sha7> | pr-42-a1b2c3d |
| Merge to main | <sha7> | a1b2c3d |
Four images are built, one per workload: core, dashboard, docs, task-scheduler. The dashboard and docs images take the base-path build args described above, so a pipeline that serves them under a prefix must pass the same prefix at build time that the ingress routes at deploy time.
Required CI configuration
Section titled “Required CI configuration”Secrets (repo-level):
| Secret | Description |
|---|---|
REGISTRY_URL | Registry hostname, e.g. registry.company.com |
REGISTRY_USERNAME | Registry user |
REGISTRY_TOKEN | Registry password / token |
Variable (repo-level):
| Variable | Description |
|---|---|
AWS_REGION | AWS region, e.g. eu-west-1 |
GitHub Environments (Settings → Environments → New environment):
Create dev and prod. In each, add:
| Secret | Description |
|---|---|
AWS_ROLE_ARN | IAM role ARN (OIDC) |
EKS_CLUSTER_NAME | EKS cluster name |
OPENTREMOR_API_KEY | The bootstrap credential passed to --set config.auth.apiKey |
Add a required reviewer protection rule on prod to gate production deploys.
AWS OIDC setup
Section titled “AWS OIDC setup”Keyless authentication via GitHub OIDC. In your AWS account:
- Create an OIDC identity provider:
token.actions.githubusercontent.com - Create two IAM roles (one per environment) with this trust policy:
{ "Effect": "Allow", "Principal": { "Federated": "<oidc-provider-arn>" }, "Action": "sts:AssumeRoleWithWebIdentity", "Condition": { "StringLike": { "token.actions.githubusercontent.com:sub": "repo:<org>/opentremor-deployments:environment:<dev|prod>" } }}Each role needs eks:DescribeCluster and the permissions to run helm upgrade.
Environment overlays
Section titled “Environment overlays”values-dev.yaml and values-prod.yaml are complete environments. values-eks.yaml is an
additional layer — ALB ingress annotations, including the load-balancer idle timeout the MCP
endpoint’s SSE streams need — and is stacked on one of the other two, never used alone:
helm upgrade --install opentremor helm/opentremor \ -f helm/opentremor/values-dev.yaml \ -f helm/opentremor/values-eks.yaml \ --set config.auth.apiKey="$(cat ./api-key)" \ --namespace opentremor --create-namespaceManual deploy
Section titled “Manual deploy”helm upgrade --install opentremor \ helm/opentremor/ \ --namespace opentremor \ --create-namespace \ --set image.repository="<your-registry>/opentremor-core" \ --set image.tag="a1b2c3d" \ --set config.auth.apiKey="$(cat ./api-key)" \ --values helm/opentremor/values-prod.yaml \ --atomic \ --timeout 5mDry-run / diff
Section titled “Dry-run / diff”helm template opentremor helm/opentremor/ \ --values helm/opentremor/values-prod.yaml \ --set config.auth.apiKey=placeholder \ --namespace opentremor
helm diff upgrade opentremor helm/opentremor/ \ --values helm/opentremor/values-prod.yaml \ --namespace opentremorThe chart reports every misconfiguration it finds in one pass rather than one per run, so a failed
helm template lists everything that needs fixing.
Uninstall
Section titled “Uninstall”helm uninstall opentremor --namespace opentremor
# PVCs are not deleted automaticallykubectl delete pvc -l app.kubernetes.io/instance=opentremor -n opentremor