Skip to content

Monitoring Pack (Grafana)

OpenTremor Core exposes a /metrics endpoint on its own — see Metrics for the catalogue. This package is the pre-built Grafana dashboard and Prometheus/Grafana provisioning stack that turns those metrics into something you look at rather than query by hand.


prometheus:
image: prom/prometheus:v3.13.0
volumes:
- ./prometheus/prometheus.yml:/etc/prometheus/prometheus.yml:ro
grafana:
image: grafana/grafana:13.1.0
environment:
GF_SECURITY_ADMIN_PASSWORD: admin
volumes:
- ./grafana/provisioning:/etc/grafana/provisioning:ro
- ./grafana/dashboards:/var/lib/grafana/dashboards:ro
ServiceURLCredentials
OpenTremor Corehttp://localhost:8000—
Prometheushttp://localhost:9090—
Grafanahttp://localhost:3000admin / admin

The Grafana dashboard is provisioned automatically — open Dashboards → OpenTremor.


Every metric carries an org_id label. The dashboard includes an Organization template variable (multi-select, defaults to “All”) at the top — every panel’s query is already scoped to it (org_id=~"$org_id"), so picking one or more orgs from the dropdown filters the entire dashboard at once, no per-panel editing needed. Requests that never resolve an org (health checks, public report links, the GitHub webhook, SSO callbacks before login) show up under _none.


opentremor-monitoring/
├── prometheus/
│ └── prometheus.yml Scrape config (target: opentremor-core:8000)
└── grafana/
├── provisioning/
│ ├── datasources/prometheus.yml Auto-registers Prometheus datasource
│ └── dashboards/dashboard.yml Points Grafana at the dashboards folder
└── dashboards/
└── opentremor.json Pre-built dashboard (import manually if needed)

In Grafana → Connections → Data sources → Add → Prometheus. Set the URL to where your Prometheus instance is reachable, e.g. http://prometheus:9090.

In Grafana → Dashboards → Import → Upload JSON file. Upload grafana/dashboards/opentremor.json. Select the Prometheus datasource created in step 1, then click Import.


SectionPanelPromQL highlight
HTTP TrafficRequest raterate(http_request_duration_seconds_count[2m])
HTTP TrafficLatency p50/p95/p99histogram_quantile(0.95, ...)
HTTP Traffic5xx error rateratio of status=~"5.." to total
HTTP TrafficRate by endpointgrouped by handler
HTTP TrafficRate by status codegrouped by status
Analysis PipelineIngest raterate(mcp_ingest_total[2m])
Analysis PipelineResources ingested/srate(mcp_ingested_resources_total[2m])
Analysis PipelineAnalysis submissions/srate(mcp_analysis_total[2m])
Analysis PipelineLLM response time (p50/p95/p99)histogram_quantile(0.95, rate(mcp_llm_analysis_duration_seconds_bucket[5m]))
FindingsBy severity (stacked)rate(mcp_findings_total[5m]) per severity
FindingsBy resource typerate(mcp_findings_total[5m]) per resource_type
TotalsStat panelsCumulative counters since last restart

Add a ServiceMonitor (if you use the Prometheus Operator) or annotate the pod for scraping:

values.yaml
podAnnotations:
prometheus.io/scrape: "true"
prometheus.io/port: "8000"
prometheus.io/path: /metrics

See Helm (Kubernetes) for the full values reference.