> ## Documentation Index
> Fetch the complete documentation index at: https://private-7c7dfe99-parallel-read-in-order-multi-part.mintlify.site/llms.txt
> Use this file to discover all available pages before exploring further.

# Set Up Prometheus Monitoring

Configure Prometheus to scrape ClickHouse metrics and visualize them with prebuilt Grafana dashboards using the kube-prometheus-stack and the ClickHouse Grafana mixin.

## Prerequisites

* A running ClickHouse Private deployment
* [Helm](https://helm.sh/) available on your workstation
* `kubectl` access to the target cluster
* The kube-prometheus-stack Helm chart and its container images available in your private registry (see [Airgap image preparation](#prepare-images-for-airgap) below)

## Steps

### 1. Install kube-prometheus-stack

Add the Prometheus community Helm repository to your local Helm client and pull the chart:

```sh theme={null}
helm repo add prometheus-community https://prometheus-community.github.io/helm-charts
helm repo update
helm pull prometheus-community/kube-prometheus-stack --version <version>
```

Then install it from your Helm repository:

```sh theme={null}
helm install kube-prometheus-stack \
  oci://<your-registry>/helm-charts/kube-prometheus-stack \
  --version <version> \
  --namespace monitoring \
  --create-namespace \
  --values kube-prometheus-stack-values.yaml
```

> **Note:** The kube-prometheus-stack bundles Prometheus, Alertmanager, Grafana, and the Prometheus Operator. Consult the [chart documentation](https://github.com/prometheus-community/helm-charts/tree/main/charts/kube-prometheus-stack) for the full list of images that must be mirrored to your private registry.

### 2. Enable PodMonitors in the ClickHouse Operator

The ClickHouse operator Helm chart includes PodMonitor definitions for ClickHouse Server and Keeper. Enable them in your operator Helm values:

```yaml theme={null}
metrics:
  enabled: true
  podmonitor:
    enabled: true
    scrape:
      clickhouseServer: true
      clickhouseKeeper: true
```

Apply the values by upgrading the operator release:

```sh theme={null}
helm upgrade clickhouse-operator \
  oci://<your-registry>/helm-charts/clickhouse-operator \
  --version <version> \
  --namespace clickhouse-operator-system \
  --values operator-values.yaml
```

This creates two PodMonitor resources:

| PodMonitor | Targets | Port |
| - | - | - |
| `clickhouse-server-metrics` | Pods labeled `app.kubernetes.io/name: clickhouse-server` | `prometheus` (8001, HTTP) or `prom-secure` (8004, HTTPS) |
| `clickhouse-keeper-metrics` | Pods labeled `app.kubernetes.io/name: clickhouse-keeper` | `prometheus` (8001, HTTP) or `prom-secure` (8004, HTTPS) |

Use the plain port for non-TLS deployments (`scrape.clickhouseServer`, `scrape.clickhouseKeeper`) and the secure port for TLS deployments (`scrape.clickhouseServerSecure`, `scrape.clickhouseKeeperSecure`). See [Expose Metrics over HTTPS](#expose-metrics-over-https-optional) for the full configuration.

Both PodMonitors carry the label `release: kube-prometheus-stack`, which matches the default `podMonitorSelector` of a kube-prometheus-stack Prometheus instance.

> **Note:** If your Prometheus instance uses a different release name, update the PodMonitor label selector accordingly by overriding the operator chart templates or configuring `prometheus.prometheusSpec.podMonitorSelector` in the kube-prometheus-stack values.

### 3. Verify Prometheus Targets

After deploying, confirm that Prometheus discovers and scrapes the ClickHouse targets:

```sh theme={null}
kubectl port-forward -n monitoring svc/kube-prometheus-stack-prometheus 9090:9090
```

Open `http://localhost:9090/targets` in a browser. Look for target groups named `podMonitor/clickhouse-operator-system/clickhouse-server-metrics` and `podMonitor/clickhouse-operator-system/clickhouse-keeper-metrics`. All targets should show a **State** of `UP`.

Run a test query to confirm metrics are flowing:

```promql theme={null}
-up{job="clickhouse-operator-system/clickhouse-server-metrics"}
```

> **Note:** The `job` label in Prometheus includes the namespace prefix: `clickhouse-operator-system/clickhouse-server-metrics` and `clickhouse-operator-system/clickhouse-keeper-metrics`. Use these full values in PromQL queries and Grafana dashboard filters.

### 4. Import the ClickHouse Grafana Mixin Dashboards

The [ClickHouse Grafana mixin](https://grafana.com/docs/grafana-cloud/monitor-infrastructure/integrations/integration-reference/integration-clickhouse/) provides prebuilt dashboards for ClickHouse server and keeper metrics. In an airgapped environment, import the dashboard JSON files manually.

#### Download dashboards

Download the mixin dashboard JSON files from the [Grafana integration page](https://grafana.com/docs/grafana-cloud/monitor-infrastructure/integrations/integration-reference/integration-clickhouse/) or export them from an existing Grafana instance.

#### Import via Grafana UI

1. Open Grafana (bundled with kube-prometheus-stack):
   ```sh theme={null}
   kubectl port-forward -n monitoring svc/kube-prometheus-stack-grafana 3000:80
   ```
2. Log in at `http://localhost:3000` (default credentials: `admin` / `prom-operator`).
3. Navigate to **Dashboards > Import**.
4. Upload each dashboard JSON file or paste its contents.
5. Select your Prometheus data source when prompted.

#### Import via ConfigMap

To manage dashboards as code, create a ConfigMap in the monitoring namespace with the Grafana sidecar label:

```yaml theme={null}
apiVersion: v1
kind: ConfigMap
metadata:
  name: clickhouse-grafana-dashboards
  namespace: monitoring
  labels:
    grafana_dashboard: "1"
data:
  clickhouse-overview.json: |
    { ... dashboard JSON ... }
```

The Grafana sidecar automatically picks up ConfigMaps with the `grafana_dashboard: "1"` label and loads the dashboards.

### 5. Configure Alert Rules (Optional)

Import the recommended ClickHouse alert rules into Prometheus by creating a `PrometheusRule` resource. See [Configure alerting](/cloud/clickhouse-private/how-to/configure-alerting) and [Metrics and alerts reference](/cloud/clickhouse-private/reference/metrics-and-alerts) for the full set of alert definitions.

Example:

```yaml theme={null}
apiVersion: monitoring.coreos.com/v1
kind: PrometheusRule
metadata:
  name: clickhouse-alerts
  namespace: monitoring
  labels:
    release: kube-prometheus-stack
spec:
  groups:
    - name: clickhouse-server
      rules:
        - alert: ClickhouseOperatorNotReconciling
          expr: avg(increase(last_cluster_reconcile[90m])) by (app) == 0
          for: 120m
          labels:
            severity: warning
          annotations:
            summary: "Operator has not reconciled {{ $labels.app }} in 2 hours"
        - alert: ClickHouseDataLoss
          expr: ClickHouse_CustomMetric_LostPartCount > 0
          labels:
            severity: critical
          annotations:
            summary: "Lost parts detected on {{ $labels.instance }}"
```

## Expose Metrics over HTTPS (Optional)

By default, Server and Keeper expose metrics on port `8001` over plain HTTP. Both can be configured to serve metrics exclusively over HTTPS on port `8004` (`prom-secure`). When TLS is required, the plain port `8001` binding is removed.

### Certificate Requirements

* A TLS certificate secret named `ch-<cluster-name>-cert-secret` in the cluster namespace, containing keys `tls.key`, `tls.crt`, and `chaincert.pem`

> For details on certificate generation, SAN requirements, and secret key names see [Configure Cert Manager for ClickHouse Certificates](/cloud/clickhouse-private/how-to/configure-cert-manager) or [Generate FIPS-Compliant Certificates for ClickHouse](/cloud/clickhouse-private/how-to/configure-fips-certificates).

### Expose Server Metrics over HTTPS

The Server pod exposes `prom-secure` (8004) only when `server.openSSL.required=true` is set in the cluster Helm values. This also disables all plain-text ports (8001, 8123, 9000), leaving only TLS listeners.

#### Enable

Upgrade the cluster Helm release with TLS required:

```sh theme={null}
helm upgrade <release-name> \
  oci://<your-registry>/helm-charts/onprem-clickhouse-cluster \
  --version <version> \
  --namespace <namespace> \
  --reuse-values \
  --set-json='server.openSSL.enabled=true' \
  --set-json='server.openSSL.required=true' \
  --set-json="server.openSSL.secret.name=\"ch-<cluster-name>-cert-secret\"" \
  --set-json='server.openSSL.secret.caKey="chaincert.pem"' \
  --set-json='server.openSSL.secret.certKey="tls.crt"' \
  --set-json='server.openSSL.secret.keyKey="tls.key"'
```

#### Update the PodMonitor

Enable the secure scrape endpoint in the operator Helm values. Disable `clickhouseServer` (8001) and enable `clickhouseServerSecure` (8004) instead:

```yaml theme={null}
metrics:
  enabled: true
  podmonitor:
    enabled: true
    scrape:
      clickhouseServer: false
      clickhouseServerSecure: true
    serverTLS:
      certSecret: clickhouse-tls-certs  # Secret copied to clickhouse-operator-system above
      caKey: chaincert.pem
      certKey: tls.crt
      keyKey: tls.key
      insecureSkipVerify: true  # Certs are issued for hostnames, not pod IPs
```

Apply by upgrading the operator release:

```sh theme={null}
helm upgrade clickhouse-operator \
  oci://<your-registry>/helm-charts/clickhouse-operator \
  --version <version> \
  --namespace clickhouse-operator-system \
  --values operator-values.yaml
```

### Expose Keeper Metrics over HTTPS

To expose Keeper metrics over HTTPS, Keeper must run in server-embedded mode. In this mode Keeper runs as a thread inside a `clickhouse-server` process and gains access to the TLS-capable HTTP handler stack, including `prom-secure` on port 8004.

> For a full explanation of server-embedded mode, image registry requirements, and what running Keeper this way implies for your deployment, see [Keeper-in-Server Mode](/cloud/clickhouse-private/explanation/keeper-in-server).

#### Enable

Upgrade the cluster Helm release to switch Keeper to the `clickhouse-server` image and enable server-embedded mode:

```sh theme={null}
helm upgrade <release-name> \
  oci://<your-registry>/helm-charts/onprem-clickhouse-cluster \
  --version <version> \
  --namespace <namespace> \
  --reuse-values \
  --set-json="keeper.image.repository=\"<your-registry>/clickhouse-server\"" \
  --set-json="keeper.image.tag=\"26.4.1.2596\"" \
  --set-json='keeper.featureFlags.runKeeperInServer=true' \
  --set-json='keeper.openSSL.required=true' \
  --set-json="keeper.openSSL.secret.name=\"ch-<cluster-name>-cert-secret\"" \
  --set-json='keeper.openSSL.secret.caKey="chaincert.pem"' \
  --set-json='keeper.openSSL.secret.certKey="tls.crt"' \
  --set-json='keeper.openSSL.secret.keyKey="tls.key"'
```

The operator rolls out Keeper pods one at a time.

> **Note:** Add `--set-json='keeper.featureFlags.disableNonSecureKeeperInServerPorts=true'` if you want to disable all non-TLS Keeper ports.

#### Update the PodMonitor

Once Keeper pods are running in server-embedded mode with TLS, enable the secure scrape endpoint in the operator Helm values.

```yaml theme={null}
metrics:
  enabled: true
  podmonitor:
    enabled: true
    scrape:
      clickhouseKeeper: false
      clickhouseKeeperSecure: true
    keeperTLS:
      certSecret: clickhouse-tls-certs
      caKey: chaincert.pem
      certKey: tls.crt
      keyKey: tls.key
      insecureSkipVerify: true
```

Apply by upgrading the operator release:

```sh theme={null}
helm upgrade clickhouse-operator \
  oci://<your-registry>/helm-charts/clickhouse-operator \
  --version <version> \
  --namespace clickhouse-operator-system \
  --values operator-values.yaml
```

The operator updates the `clickhouse-keeper-metrics` PodMonitor to scrape `prom-secure` (8004) over HTTPS using the provided certificate.

## Enable the Custom Metrics Handler (Optional)

ClickHouse Server exposes additional `ClickHouse_CustomMetric_*` metrics including lost part counts, replica read-only duration, data corruption errors, async insert memory, and more. These are derived from SQL queries against system tables and are not available on the standard Prometheus endpoint. The endpoint requires authentication.

Add an HTTP handler rule to the cluster Helm values.
:

```yaml theme={null}
server:
  config:
    http_handlers:
      rule:
      - url: /custom_metrics
        methods: GET
        handler:
          type: predefined_query_handler
          content_type: text/plain; charset=utf-8
          query: >-
            SELECT name, value, help, labels_serialized, type
            FROM merge(system, '^_custom_metrics_dictionary_.*')
            FORMAT Prometheus
```

The query scrapes only the tables — pre-aggregated, low-cardinality metrics such as `ClickHouse_CustomMetric_LostPartCount`, replica read-only duration, and memory usage.

Apply via helm upgrade:

```sh theme={null}
helm upgrade <release-name> \
  oci://<your-registry>/helm-charts/onprem-clickhouse-cluster \
  --version <version> \
  --namespace <namespace> \
  --reuse-values \
  --set-json='server.config={"http_handlers":{"rule":[{"url":"/custom_metrics","methods":"GET","handler":{"type":"predefined_query_handler","content_type":"text/plain; charset=utf-8","query":"SELECT name, value, help, labels_serialized, type FROM merge(system, '\''^_custom_metrics_dictionary_.*'\'') FORMAT Prometheus"}}]}}'
```

Then create a PodMonitor to scrape the endpoint. Use `scheme: http` and `targetPort: 8123` for non-TLS, or `scheme: https` and `targetPort: 8443` for TLS:

```yaml theme={null}
apiVersion: monitoring.coreos.com/v1
kind: PodMonitor
metadata:
  name: clickhouse-server-custom-metrics
  namespace: clickhouse-operator-system
  labels:
    release: kube-prometheus-stack
spec:
  namespaceSelector:
    any: true
  selector:
    matchLabels:
      app.kubernetes.io/name: clickhouse-server
  podMetricsEndpoints:
  - targetPort: 8443          # or 8123 for non-TLS
    path: /custom_metrics
    scheme: https             # or http for non-TLS
    basicAuth:
      username:
        name: <secret-name>
        key: username
      password:
        name: <secret-name>
        key: password
    tlsConfig:
      insecureSkipVerify: true  # omit for non-TLS
```

For the full list of custom metrics and their recommended alert thresholds, see [Metrics and alerts reference](/cloud/clickhouse-private/reference/metrics-and-alerts).

## Prepare Images for Airgap

The kube-prometheus-stack requires the following container images. Mirror them to your private registry before installation:

| Component | Image |
| - | - |
| Prometheus | `quay.io/prometheus/prometheus` |
| Alertmanager | `quay.io/prometheus/alertmanager` |
| Grafana | `docker.io/grafana/grafana` |
| Grafana sidecar | `quay.io/kiwigrid/k8s-sidecar` |
| Prometheus Operator | `quay.io/prometheus-operator/prometheus-operator` |
| kube-state-metrics | `registry.k8s.io/kube-state-metrics/kube-state-metrics` |
| node-exporter | `quay.io/prometheus/node-exporter` |

> **Note:** Exact image tags depend on the kube-prometheus-stack chart version you are deploying. Run `helm template` on the pulled chart to extract the exact image references for your version.
