Skip to main content
This guide covers how to configure ClickHouse and Keeper clusters using the operator.

ClickHouseCluster configuration

Basic configuration

Replicas and shards

  • Replicas: Number of ClickHouse instances per shard (for high availability)
  • Shards: Number of horizontal partitions (for scaling)
A cluster with replicas: 3 and shards: 2 will create 6 ClickHouse pods total.

Keeper integration

Every ClickHouse cluster needs a Keeper ensemble for coordination. Set exactly one of keeperClusterRef or externalKeeper; a cluster that sets both or neither is rejected. Point keeperClusterRef at a KeeperCluster this operator manages:
When keeperClusterRef.namespace is set, the operator must watch both namespaces. If WATCH_NAMESPACE is configured, include the ClickHouse and Keeper namespaces in that list.

Externally managed Keeper

Use externalKeeper to run ClickHouse against a Keeper ensemble the operator does not manage — one owned by another operator, running in a different Kubernetes cluster, on virtual machines or bare metal, or shared between several systems:
port defaults to 9181. host is a DNS name or an IP address; enclose IPv6 addresses in square brackets. Set tls: Enabled to connect over TLS, and give port the value the ensemble listens on for secure connections:
ClickHouse verifies the Keeper certificates against the system trust store, plus any caBundle you configure. See ClickHouse-Keeper communication over TLS. When several ClickHouse clusters share one ensemble, give each its own root znode through settings.extraConfig, so their replication metadata, SQL-defined users and named collections don’t collide. Create the znode before ClickHouse starts, and don’t change it afterwards:
Don’t put zookeeper.root in settings.extraReloadableConfig: a reload would switch the root of a running cluster and leave its replicated tables read-only. If the root znode has ACLs, they must let the cluster’s keeper-identity create child znodes. The operator does not change these servers, and it does not wait for them before it reconciles ClickHouse. Keeping the ensemble available is up to the user. Authentication works the same way as for an operator-managed Keeper: the operator reads the keeper-identity key from the cluster Secret and passes it to ClickHouse. See External Secret for the Secret layout.

KeeperCluster configuration

Leadership handover on termination

The Keeper container has a default preStop hook that asks the replica to hand off Raft leadership before it stops, and waits for the handover to complete before letting the container terminate. This runs on every graceful stop of the container, not just operator-initiated rollouts: a descheduler eviction, a node drain, or an autoscaler action all trigger the same hook, since it is a kubelet-level preStop handler rather than reconcile-loop logic. Without this, restarting the current leader leaves the ensemble without a leader until a new election completes, which can take several seconds under load and shows up to ClickHouse as Keeper connection loss and elevated query latency. The hook asks the replica to yield leadership (Keeper’s ydld four-letter-word command, sent over the http_control port) and then polls the replica’s status until it is no longer the leader, rather than firing the request and immediately letting the container stop — a fire-and-forget request can still lose the race with SIGTERM under write load, so waiting is what keeps the leaderless window bounded. The hook needs no configuration and degrades gracefully: on Keeper versions older than 26.3, which don’t yet serve the leadership commands used here, the request is a no-op and the container stops as it would without the hook. It runs a shell script, so it requires bash in the image; if you run a distroless or scratch Keeper image, override or remove it via containerTemplate.lifecycle — see API Reference.

Storage configuration

Configure persistent storage with dataVolumeClaimSpec, a standard Kubernetes PersistentVolumeClaimSpec. The operator turns it into a per-replica PersistentVolumeClaim mounted at the data path /var/lib/clickhouse:
The operator can modify an existing PVC only if the underlying StorageClass supports volume expansion.
Attaching extra disks in a multi-disk (JBOD) layout, running without a persistent volume, expanding capacity, custom storage policies, at-rest encryption, and the rules for what cannot change after creation are covered in the dedicated Storage and volumes guide.

Cluster domain

spec.clusterDomain sets the Kubernetes DNS suffix the operator uses when it builds the fully-qualified pod host names it writes into the ClickHouse server configuration. It defaults to cluster.local and exists on both ClickHouseCluster and KeeperCluster.
For ClickHouse replicas the operator creates a governing headless Service <cluster-name>-clickhouse-headless that serves client traffic. And per-replica services <cluster-name>-clickhouse-internal-<shard>-<index> that publish unready Pods for internal traffic and management requests from the operator. This can be used for replica recovery if it cannot become ready itself. Keeper nodes use pod names through their headless Service: <pod>.<headless-service>.<namespace>.svc.<clusterDomain>.
Only override this when your cluster’s kubelet runs with a --cluster-domain other than cluster.local. If the value does not match the real cluster domain, ClickHouse cannot resolve the Keeper and replica host names — coordination and Distributed queries fail with DNS resolution errors. Set the same value on the ClickHouseCluster and the KeeperCluster it references.

Pod configuration

Automatic topology spread and affinity

Distribute pods across availability zones:
Ensure your Kubernetes cluster has enough nodes in different zones to satisfy the spread constraints.
With a node autoscaler (Karpenter, Cluster Autoscaler) the spread constraint alone cannot provision a zone that has no nodes yet — the scheduler settles for the zones it can see. Set topologyMinDomains to the number of availability zones to force spreading:
topologyMinDomains requires topologyZoneKey. Pods stay Pending while fewer zones than requested are available, so keep it within the zones your cluster can provide.

Manual configuration

Arbitrary pod affinity/anti-affinity rules and topology spread constraints can be specified.

See API Reference for all supported Pod template options.

Pod disruption budgets

The operator creates a PodDisruptionBudget (PDB) for each cluster so that voluntary disruptions — node drains, rolling upgrades, autoscaler evictions — cannot take down enough pods to lose quorum or break availability. For ClickHouse clusters with more than one shard, one PDB is created per shard so a disruption in one shard cannot count against another.

Defaults

The operator picks safe defaults based on the cluster size so that a fresh apply already protects against accidental quorum loss. For a 3-shard ClickHouseCluster with replicas: 3, the operator creates three PDBs, one per shard, each with minAvailable: 1.

Overriding the defaults

Use spec.podDisruptionBudget to override either minAvailable or maxUnavailable (exactly one):
Or the maxUnavailable form, with a percentage:
Setting both minAvailable and maxUnavailable is rejected by the validating webhook. Pick one — Kubernetes itself does not allow both either.
You can also pass the unhealthyPodEvictionPolicy field through to the generated PDB — useful when you need to allow eviction of pods that are still in NotReady:

Policies

spec.podDisruptionBudget.policy lets you choose how aggressively the operator manages PDBs: Example — disable PDB management completely on a development cluster:
Example — keep your hand-crafted PDB next to the cluster and stop the operator from touching it:

Cluster-wide opt-out

PDB management can also be disabled cluster-wide via the operator’s ENABLE_PDB environment variable. With ENABLE_PDB=false, the operator skips the PDB reconcile step for every ClickHouseCluster and KeeperCluster regardless of their spec.podDisruptionBudget.policy, and does not watch PodDisruptionBudget resources at all. The operator’s ServiceAccount therefore does not need RBAC permissions on poddisruptionbudgets.policy/v1, which is useful when running the operator under a restricted ServiceAccount that intentionally omits those permissions.
This is intended for environments that ship their own disruption policies (e.g. through Gatekeeper / Kyverno) and want the operator out of the loop entirely.

Container configuration

Custom image

Use a specific ClickHouse image:

Container resources

Configure CPU and memory for ClickHouse containers:

Environment variables

Add custom environment variables:

Volume mounts

Add additional volume mounts:
It is allowed to specify multiple volume mounts to the same mountPath. Operator will create projected volume with all specified mounts.

Health probes

The operator configures liveness and readiness probes on every ClickHouse and Keeper container. No startup probe is set by default. ClickHouse and Keeper use the same default timing settings: ClickHouse opens the interserver listener before it loads tables and the client listeners after, so the liveness probe tolerates slow table loading, while a passing readiness probe means the replica serves queries. Override any probe through spec.containerTemplate on either resource. For example, customize the ClickHouse liveness probe:
A probe you set replaces the operator default entirely — it isn’t merged field by field. Specify the handler (httpGet, tcpSocket, or exec) and every timing field you care about; omitted fields fall back to the Kubernetes defaults (for example timeoutSeconds: 1, failureThreshold: 3), not to the operator values above.

Startup probe

Use startupProbe when a container needs more time to start than the liveness probe allows — with the defaults, about 60 seconds of initial delay plus 10 failed checks 5 seconds apart. Kubernetes holds back the liveness and readiness probes until the startup probe succeeds, so a slow start doesn’t trigger a restart loop. For example, give a Keeper replica that replays a large snapshot up to 10 minutes to start accepting connections:
After the startup probe succeeds, the regular liveness and readiness probes take over.

See API Reference for all supported Container template options.

TLS/SSL configuration

Configure secure endpoints

Pass a reference to a Kubernetes Secret containing TLS certificates to enable secure endpoints

SSL certificate secret format

It is expected that the Secret contains the server keypair:
  • tls.crt - PEM encoded server certificate
  • tls.key - PEM encoded private key
This format is compatible with cert-manager generated certificates.

ClickHouse-Keeper communication over TLS

If the KeeperCluster referenced by keeperClusterRef has TLS enabled, ClickHouseCluster would use secure connection to Keeper nodes automatically. With externalKeeper, set tls: Enabled instead, and set each node’s port to the ensemble’s secure client port (for example 9281). ClickHouseCluster verifies Keeper node certificates against the system trust store, plus any caBundle you configure. To trust a private CA (for example, a self-signed or internal CA), provide a custom CA bundle reference:

External Secret

By default the operator creates and owns a Secret containing the cluster’s internal credentials (interserver password, management password, keeper identity, cluster secret, named-collections key). The Secret is named after the cluster and lives in the cluster’s namespace. If you want to manage these credentials yourself — for example, sourcing them from HashiCorp Vault, AWS Secrets Manager, or External Secrets Operator — point the operator at a pre-existing Secret using spec.externalSecret:
The referenced Secret must reside in the same namespace as the ClickHouseCluster. The operator never deletes a Secret it did not create.

Required keys

The Secret must contain the following keys: A complete Secret looks like this:

Policy: Observe vs Manage

spec.externalSecret.policy controls how the operator handles missing required keys:
Even with policy: Manage the Secret must already exist in the namespace — the operator never creates the Secret itself, it only writes generated keys into an existing one. If the referenced Secret is missing, reconciliation is blocked with the ExternalSecretNotFound reason regardless of policy.
Pick Observe when an external system (Vault, ESO, sealed-secrets, GitOps) is the source of truth and you want the operator to fail loudly on misconfiguration. Pick Manage when you want self-sufficient bootstrapping but still want to retain ownership of the Secret object itself (for example, to back it up).

Status condition and troubleshooting

The operator exposes a ExternalSecretValid condition on ClickHouseCluster.status.conditions. Inspect it when reconciliation looks stuck:
Possible reasons: The operator requeues reconciliation while the Secret is invalid, so once you add the missing keys the next reconcile picks them up automatically — no need to bounce pods.
The set of required keys depends on the running ClickHouse version. named-collections-key is only validated once the operator’s version probe has detected ClickHouse 25.12 or newer. On older versions the key may be absent from the Secret. disk-encryption-key is required only when spec.settings.encryption is set.

Additional ports

The operator exposes a fixed set of ports on every ClickHouse Pod and its public headless Service: 8123 HTTP, 9000 native, 9009 interserver, 9001/9002 management, 9363 Prometheus metrics, and the TLS variants 8443/9440 when TLS is enabled. The interserver and management ports are additionally exposed through the per-replica internal Services so replicas and the operator can communicate before a replica becomes ready. To make ClickHouse listen on additional protocols — MySQL, PostgreSQL, gRPC, or any custom port — declare them in spec.additionalPorts:
The operator adds those ports to the Pod’s containerPorts and the public headless Service. The complete example lives at examples/custom_protocols.yaml.
additionalPorts only opens the ports on the Kubernetes side. It does not configure the ClickHouse server to listen on them. You also have to enable the matching protocol in spec.settings.extraConfig.protocols. Without that, the port is open on the Service but nothing inside the pod is answering.

End-to-end example: MySQL wire protocol

To expose ClickHouse over the MySQL wire protocol on port 9004:
After applying, verify from inside the cluster:

Field constraints

Reserved ports and names

The validating webhook rejects additionalPorts entries that would collide with ports the operator binds itself. All TLS-related ports are reserved unconditionally so that flipping spec.settings.tls.enabled later cannot break a previously valid cluster. The following names are also rejected — they are the operator’s internal protocol-type identifiers (not the human-readable aliases): A rejected request produces an error such as:

Version probe and upgrade channel

The operator does two independent things with cluster versions:
  1. Version reporting — for ClickHouseCluster, a Kubernetes Job runs the container image once to detect the running ClickHouse version; for KeeperCluster, the operator reads the server-reported version from running replicas. The detected version is recorded in .status.version and used by other reconciliation steps (e.g. the External Secret named-collections key is only required from ClickHouse 25.12).
  2. Upgrade channel — a periodic check against the public ClickHouse release feed (https://clickhouse.com/data/version_date.tsv). The operator reports whether a newer version is available via the VersionUpgraded status condition. It never upgrades the cluster on its own — the user is in control of the image tag.

Choosing a release channel

spec.upgradeChannel selects which set of upstream releases the operator compares against. Same field exists on both ClickHouseCluster and KeeperCluster.
Allowed values (validated by the CRD with the pattern ^(lts|stable|\d+\.\d+)?$): For production, pinning the channel to an explicit <major>.<minor> (e.g. 25.8) is generally preferred. It locks the cluster to the intended major release line and lets the operator surface a WrongReleaseChannel warning if any replica somehow drifts onto a different major — which matters especially when the image is referenced by a digest (@sha256:...) rather than by a human-readable tag. The empty default is fine for development clusters where major-version jumps are not a concern.

Status conditions

Two conditions surface the result of the probe and the upgrade check: Inspect them with:

Overriding the version probe Job

This applies to ClickHouseCluster only. KeeperCluster no longer runs a version-probe Job — its version is read directly from the running Keeper replicas — so spec.versionProbeTemplate is deprecated and has no effect there. The probe is implemented as a regular Kubernetes Job. If your cluster has admission policies that require specific Tolerations, node selectors, security contexts, or you want to limit how long completed probe Jobs linger, override the template via spec.versionProbeTemplate:
The container name version-probe is the operator’s default — the entry under containers: matches it by name, so the operator deep-merges the user-provided fields on top of the defaults.

Operator-wide controls

Two flags on the operator manager control the upgrade-check loop globally: Set --disable-version-update-checks=true in air-gapped environments or when egress to clickhouse.com is not allowed.

ClickHouse settings

Default user password

spec.settings.defaultUserPassword sets the password for the built-in default user. Provide the value from a key in a Secret (recommended) or a ConfigMap that you create, rather than inline in the CR:
Provide exactly one of secret or configMap, each with both name (the object) and key (the entry that holds the password).

Password types

passwordType tells ClickHouse how to interpret the value. It defaults to password (plaintext); the alternatives are hashed forms such as password_sha256_hex and password_double_sha1_hex. Prefer a hashed type so the plaintext is never stored. See the ClickHouse user settings for the full list.

Full example with a Secret

Create the Secret, then reference its key:
With passwordType: password, the in-pod clickhouse-client is configured with this password, which is handy for debugging.
For a hashed password, store the hash instead of the plaintext:

Using a ConfigMap

A ConfigMap works the same way, but its contents are not protected like a Secret. Use it only for non-sensitive or already-hashed values, such as a password_sha256_hex digest:
Do not put a plaintext password in a ConfigMap. Use a Secret for any plaintext (passwordType: password) value.

Custom users in configuration

Configure additional users in configuration files. Create a ConfigMap and Secret for user:
Add custom configuration to ClickHouseCluster:

Database sync

Enable automatic database synchronization for new replicas:
When enabled, the operator synchronizes Replicated and integration tables to new replicas. Every replica Pod carries the clickhouse.com/ReplicaInitialized readiness gate, so a new replica is published through the public headless Service only after the operator completes its initialization: the default database is converted to the Replicated engine and the schema is synchronized. Until then, the operator and the other replicas reach it through its internal Service. With enableDatabaseSync: false, the operator marks replicas initialized immediately, so readiness follows the container probes alone. The operator never drops a populated non-Replicated default database. Such a replica still starts serving client traffic after schema preparation, but the cluster reports SchemaInSync=False with reason DefaultDatabaseNotReplicated until you resolve that database yourself. Before a replica is removed on scale-down, the operator first unloads client traffic from it, waits until it stops being published, replicates its remaining data to the surviving replicas, and only then deletes it.

Server logging

Configure the ClickHouse server log through spec.settings.logger. Every field is optional with a safe default, so a cluster you never touch already logs at information to both the container console and a rotated file on disk.
The operator always keeps console logging on so that kubectl logs works, and layers file logging on top when logToFile is true. A cluster with the defaults renders this logger block:
The same spec.settings.logger block applies to a KeeperCluster; the operator writes its files under /var/log/clickhouse-keeper/ instead, and sets the Raft logger (raft_logs_level) to the same level.
Console logging stays on regardless of logToFile, so kubectl logs keeps working even when you disable file logging. Set jsonLogs: true when you ship logs to a structured log store that parses JSON.

System log tables

The operator enables the query_log, part_log, text_log, metric_log and asynchronous_metric_log system tables. Because they live on the same volume as your data, spec.settings.systemLogsTTLDays bounds them with a table TTL. The defaulting webhook sets it to 30 when you create a cluster without the field. Set 0 to disable the TTL explicitly.
The default applies only to new clusters. A cluster created before systemLogsTTLDays existed keeps unbounded system log tables until you set the field yourself.
The TTL takes effect on the next server restart. Override it per table, tune any other system log setting, or disable a table entirely with the @remove attribute — all through extraConfig:
A disabled table is not dropped: the existing data stays on disk until you drop the table yourself.
ClickHouse recreates a system table whose TTL changed and keeps the old data in a renamed table such as system.text_log_0. Drop those tables once you no longer need the history.

Custom configuration

Embedded extra configuration

Instead of mounting custom configuration files, you can directly specify additional ClickHouse configuration options. Add custom ClickHouse configuration using extraConfig:

Reloadable extra configuration

Any change to extraConfig restarts the ClickHouse pods, because the operator cannot know which settings apply at runtime. For settings ClickHouse reloads on the fly — named_collections is the common case — use extraReloadableConfig instead. The operator applies its changes through a configuration reload, without restarting pods:
Keep each top-level section in exactly one of extraConfig and extraReloadableConfig; the operator warns when both set the same section.
By using extraReloadableConfig you assert its settings are reload-safe. A restart-only setting placed here takes effect only at the next restart, without any error.

Embedded extra users configuration

You can also specify additional ClickHouse users configuration using extraUsersConfig. This is useful for defining users, profiles, quotas, and grants directly in the cluster specification.
The extraUsersConfig is stored in k8s ConfigMap object. Avoid plain text secrets there.

See documentation for all supported ClickHouse users configuration options.

Configuration example

Complete configuration example:

Periodic resync

Besides reacting to Kubernetes events, the operator re-reconciles every cluster on a fixed interval. This keeps status conditions fresh for state that no watch can observe — ClickHouse replica health, Keeper quorum — and retries remediation such as stuck-pod cleanup. The interval is controlled by the operator’s RESYNC_PERIOD environment variable (default 30s). It applies to KeeperCluster resources, where 0 disables the periodic resync. ClickHouseCluster resources poll ClickHouse system.warnings every 30 seconds, which already re-triggers reconciliation, so RESYNC_PERIOD currently has no effect on them.
The operator reconciles clusters concurrently; the number of parallel reconciles per controller is set by MAX_CONCURRENT_RECONCILES (default 4).

Pausing reconciliation

Set the clickhouse.com/pause-reconciliation: "true" annotation on a ClickHouseCluster or KeeperCluster to stop the operator from reconciling it. The operator reports ReconcileSucceeded=False with reason ReconciliationPaused on the paused resource and otherwise leaves it untouched — useful for manual intervention when the operator would fight your changes. While paused, the operator reports all state conditions (including Ready and Healthy) as Unknown with reason ReconciliationPaused — it does not observe the cluster. Remove the annotation to resume; the pending spec changes are applied on the next reconciliation.
Last modified on October 8, 2026