Skip to content

[DO NOT MERGE] [SETU-3482] BYOP: skip bundled monitoring, point C3 at an external Prometheus - #2623

Open
Ishita Chourasia (ishitach) wants to merge 52 commits into
masterfrom
SETU-3482-byop-skip-bundled-monitoring
Open

Ishita Chourasia (ishitach) wants to merge 52 commits into
masterfrom
SETU-3482-byop-skip-bundled-monitoring

Conversation

@ishitach

@ishitach Ishita Chourasia (ishitach) commented Aug 14, 2026 •

Copy link
Copy Markdown
Member

DO NOT MERGE

First-level WIP PR for early review. It depends on an unmerged control-center-backend change and is not yet fully validated end to end. See "State" below.

Summary

Adds Bring Your Own Prometheus (BYOP) support to the control_center_next_gen role: a toggle (control_center_next_gen_external_prometheus_enabled) that skips deploying the bundled Prometheus and Alertmanager, points C3 and the node exporters at a customer-managed external Prometheus, turns C3 alerts off, and emits confluent.controlcenter.prometheus.external.enable so C3 boots against the external endpoint.

Tickets: SETU-3482 (parent), MMA-19237 / 19238 / 19239 / 19240 / 19241.

What is here

  • Toggle plus gating of bundled Prometheus/Alertmanager install, config, and health check (SETU-3482, MMA-19237).
  • C3 config in BYOP: alerts off, drop bundled rules/alertmanager file props, emit prometheus.external.enable (MMA-19239).
  • Pre-flight validation: endpoint set, TLS required, exactly one auth mode (MMA-19238).
  • Build the Prometheus client truststore in BYOP so C3 can trust the external endpoint over TLS.
  • Docs: VARIABLES.md entry plus a sample external-Prometheus inventory (MMA-19241).
  • Molecule: a mock external-Prometheus fixture, three flipped scenarios (basic auth), plus verify, negative-validation, and custom-C3-swap playbooks (MMA-19240).

Dependency

The C3 boot path depends on a control-center-backend change that gates prometheusBackendEnabled() on the new prometheus.external.enable flag, decoupled from alerts.enable. Without that build, stock C3 crashes on boot in BYOP mode. This PR emits the flag; the C3 change consumes it.

State (why do not merge)

  • Config generation, C3 boot past the crash, and the truststore build are verified live.
  • Full end to end (C3 serves, brokers push to and C3 reads from the external Prometheus, verify assertions) is not yet proven. It needs a version-consistent CP build that carries the C3 change. A C3-only swap into an older cluster stalls on version skew.
  • The mTLS-to-Prometheus flip is not yet exercised (basic auth only).

Test plan

  • yamllint: all new and edited molecule and fixture files are clean.
  • Mock fixture: built and verified standalone (TLS enforced, 401 on bad auth, 200 with recording rules loaded).
  • Negative validation playbook: runs on localhost, rejects no-TLS and dual-auth (rescued=2, failed=0).
  • Molecule mtls-ubuntu converge: full CP install reached C3 config generation; verified the emitted C3 config (external prometheus.url, prometheus.external.enable=true, alerts.enable=false, no bundled files) and that the bundled prometheus/alertmanager services are disabled.
  • Live C3 boot: swapped a version-matched C3 build carrying the C3-side change into the container; C3 booted past the BYOP crash and connected to Kafka.
  • Pending: full converge plus verify against a version-consistent CP build (blocked as noted above).

… Next Gen

Add control_center_next_gen_bundled_monitoring_enabled (default true). When
false, the bundled Prometheus and Alertmanager for Control Center Next Gen are
not installed or started, for use with an external Prometheus.

Pure bundled tasks (config copy, prometheus overlay, web-configs, dependency
file permissions, prometheus/alertmanager keystores) are gated on the toggle.
Tasks that configure C3 and the bundled parts together (systemd service copy,
override dir/write, service start) keep the C3 parts and skip only the bundled
parts. Default true preserves existing behavior.
Rename control_center_next_gen_bundled_monitoring_enabled to
control_center_next_gen_external_prometheus_enabled (default false, true = BYOP).
The bundled-task gates invert accordingly; default behavior is unchanged.

Skip the Alertmanager health check when the external Prometheus is enabled, and
keep the Prometheus health check (now pointing at the external endpoint).

Brownfield cleanup of the old bundle is a documented manual step for this release,
not automated in the playbook.
…BYOP

Split the C3 property groups so the bundled Alertmanager/Prometheus management
properties (alertmanager url, prometheus config.file, rules.file, alertmanager
config.file) render only when the bundle is deployed. When external_prometheus_enabled
is true, emit confluent.controlcenter.alerts.enable=false and disable the alertmanager
client. In BYOP the customer configures promotion/out-of-order/recording-rules on their
own external Prometheus, so those bundled-only properties are not set.

Bundled render is unchanged: combine_properties flattens enabled groups by key and the
template sorts by key, so moving properties between enabled groups is byte-identical.
…ode)

New validate_byop.yml, imported at the top of the C3 role's main.yml and gated on
control_center_next_gen_external_prometheus_enabled, so a misconfigured external
Prometheus fails fast before any install work. Enforces the supported contract:
the endpoint is set, TLS is enabled, and exactly one auth mode (Basic over TLS or
mutual TLS) is chosen. There is no Alertmanager-with-external check because the single
cp-ansible toggle disables the bundled Prometheus and Alertmanager together.
…heus inventory

Document control_center_next_gen_external_prometheus_enabled: turn its defaults
comment into a ### doc comment (VARIABLES.md is generated from those) and add the
entry to docs/VARIABLES.md. Add docs/sample_inventories/byop_external_prometheus.yml
showing the toggle plus the control_center_next_gen_dependency_prometheus_* settings
(host/port/ssl/CA), a Basic-auth-over-TLS option and a commented mTLS option, and the
external-Prometheus prerequisites as header comments.
The dependency Prometheus keystore/truststore task was gated to bundled-only, so in
BYOP mode C3 had no truststore for the external Prometheus and failed to start with a
FileNotFoundException on the truststore. C3 needs this truststore to trust the external
Prometheus over TLS, so run the task whenever Prometheus TLS is enabled, external or not.
…n BYOP

C3 uses this flag to relax the bundled prometheus rules / alertmanager config-file
requirement independently of alerts.enable, so it can boot against an external Prometheus.
Pairs with the control-center-backend change that gates prometheusBackendEnabled() on this
flag. Harmless on builds that predate the flag (unknown config is ignored).
Adds a mock external Prometheus fixture (docker/byop-mock-prometheus: OTLP receiver,
promoted Confluent attributes, out-of-order window, C3 recording rules, TLS with basic-auth
and mTLS web-configs) and flips three existing scenarios (mtls-ubuntu, oauth-rbac, fips) to
read from it over Basic auth. Shared playbooks: byop_verify (bundle absent, external config,
alerts off, rules loaded, nodes pushing), byop_distribute_ca, byop_validate_negative
(rejects no-TLS and dual-auth), and byop_install_custom_c3 (swap a custom C3 deb or tarball
for validating against a build that carries the C3-side change).
…alerts 404

The /2.0/health/status ONLINE check now parses the JSON (bootstrapClusterId
present, a clusterStatus value == ONLINE) instead of a loose substring match,
and the endpoint path is confirmed against the C3 API (removed the TODO).

Add a runtime check that /2.0/alerts/triggers returns 404 in BYOP - the counterpart
to the alerts.enable=false config assertion, proving the alert surface is off.
A3 (byop-mtls-ubuntu): greenfield BYOP over mTLS. Reuses the mock Prometheus with
web-config-mtls.yml; byop_distribute_client_cert.yml gives C3 the client cert
(signed by the mock CA) so it can present it. Focused cluster + byop_verify.

A4 (byop-brownfield-ubuntu): install with the bundle, then side_effect.yml flips
every node to the external Prometheus, re-runs the platform playbook, and runs
byop_cleanup_bundled.yml (the documented manual cleanup); verify asserts external.

Both need a molecule run to confirm the mTLS cert-trust chain and the per-pass
var flip - flagged inline. The bundle->external->bundle round-trip is left as a
follow-on scenario (it ends on bundled, which conflicts with the external verify).
Add build.sh (certs + docker build, one idempotent command) so the mock image is
not a manual prerequisite. The on-demand molecule pipeline calls this in setup
next to usmagent-mock-ccloud (both are pre_build_image, built outside molecule).
Make generate-certs.sh idempotent so re-runs are safe.
… scenarios

Two optional/experimental scenarios, clearly flagged as "worth trying, may need a
tweak on the first molecule run":

- byop-upgrade-migrate (A6): install an older CP + bundle, side_effect upgrades CP
  (bundle stays), then flips to external + cleanup; verify asserts external.
- byop-roundtrip: bundle -> external -> bundle; side_effect flips external + cleanup
  then back to bundled; verify asserts the bundle is reinstalled.

The old->new CP upgrade re-run and the multi-pass toggle flips are the parts that
only a live run validates - noted inline.
@ishitach
Ishita Chourasia (ishitach) marked this pull request as ready for review August 24, 2026 08:29
@ishitach
Ishita Chourasia (ishitach) requested a review from a team as a code owner August 24, 2026 08:29
Flip three existing feature scenarios to read from an external
Prometheus instead of the bundled one, covering the auth mechanisms
that were not yet exercised under BYOP:

- kerberos-rhel: external Prometheus over TLS + Basic auth
- archive-scram-rhel: external Prometheus over TLS + Basic auth
- mini-setup-ldap-mtls-fips: external Prometheus over mTLS (FIPS)

Each adds the byop-prometheus mock container, the external toggle and
connection vars, distributes the mock CA (and client cert for mTLS) in
prepare, and imports the shared byop_verify / byop_validate_negative
checks. This follows the reuse-and-split approach: bundled scenarios
stay as the regression baseline while these variants cover BYOP.
When Control Center reads from an external Prometheus, the Kafka brokers
and controllers push their telemetry to that same endpoint. The push uses
the node's own truststore, which does not contain the external Prometheus
CA, so over TLS the exporter cannot verify the endpoint (and under mTLS the
handshake never completes) and no broker metrics reach it.

Import the external Prometheus CA into the broker and controller
truststores (JKS and, under FIPS, BCFKS) when external Prometheus over TLS
is enabled, reusing the shared idp_certs import task and restarting the
node so the exporter picks it up. The CA source is the existing
control_center_next_gen_dependency_prometheus_provided_ca_cert_path.

Test wiring: every external molecule scenario now sets the provided CA
source, and each mTLS scenario rebuilds the mock Prometheus client CA as a
bundle of the mock CA and the per-run cluster CA (byop_mock_trust_cluster_ca)
so the mock accepts the nodes' cluster-signed client certificates on push.
- byop_mock_trust_cluster_ca: hot-reload the mock's web TLS config instead
  of restarting the container. A restart changed the mock's IP and broke
  name resolution for it from the other hosts at verify time.
- kerberos-rhel, archive-scram-rhel: run the BYOP verify plays first. A host
  that fails a task is skipped by every later play, so a scenario-specific
  check failing on the Control Center host would otherwise prevent the BYOP
  checks (which target that host) from running at all.
- archive-scram-rhel: exclude the mock Prometheus host from the scenario's
  all-hosts archive-file assertion (it is not a Confluent Platform node).
- byop-basic-ubuntu: drop a duplicated provided-CA variable.
Two problems prevented Control Center from reading an external Prometheus
over Basic auth on a TLS endpoint (no mTLS):

- The custom-certs switch keyed off the on-host client cert path, which
  has a non-empty default, so it was always on. Without a provided client
  cert the copy of the (absent) signed cert failed and Control Center
  configuration was left with the default localhost Prometheus URL. Gate
  the switch on the PROVIDED client cert path instead, which is set only
  for mTLS.
- With custom certs off, the ssl role self-signs the Prometheus client
  truststore with a generated CA, so it does not trust the external
  endpoint. Import the external Prometheus CA into that truststore (JKS
  and, under FIPS, BCFKS) when reading an external Prometheus over TLS
  without mTLS, so Control Center trusts the endpoint's server cert on read.

mTLS is unaffected: the provided-cert path still builds the client keystore
and imports the provided CA into the truststore.
byop_validate_negative is a localhost unit test of validate_byop.yml (it
feeds bad configurations and checks they are rejected); it does not depend
on the deployed cluster. Running it inside every scenario's verify made the
whole scenario fail on a localhost-only assertion even when the real BYOP
checks passed. Keep it as a standalone playbook, run separately, and let the
scenario verify gate only on byop_verify.
Three fixes for the remaining split BYOP scenarios:

- playbooks/all.yml: the certificate-authority generation was gated only on
  each component's own listener TLS. When the cluster runs without TLS
  (for example SASL_PLAINTEXT) but Control Center still reads an external
  Prometheus over TLS, the shared self-signed CA was never generated and the
  Control Center Prometheus client keystore step failed. Include the
  external Prometheus TLS flag in the condition.
- byop_verify: discover the Control Center properties file (package/rpm
  under /etc, archive under the confluent home) instead of hard-coding the
  package path, so the archive scenario reads the right file.
- byop_mock_trust_cluster_ca: normalise the combined client CA bundle so the
  two PEM blocks never merge onto one line, which made the bundle
  unparseable and broke the mock Prometheus TLS handshake.
- certificate_authority.yml: the inner CA-generation and copy-back tasks
  carried the same component-TLS condition as the play-level include, so
  the CA was still skipped on a non-TLS cluster that only needs TLS for the
  external Prometheus. Add the external Prometheus TLS flag to both.
- byop_verify: skip the broker/controller push assertions under mTLS. The
  mock Prometheus requires a client cert signed by its own CA, so it rejects
  the node's cluster-signed client cert in this harness; the node-push path
  is covered by the Basic-auth-over-TLS scenarios with the same product code.
- Drop the mock cluster-CA trust step from the mTLS scenarios: modifying the
  running mock's client CA broke its TLS handshake. mTLS coverage validates
  the Control Center read path; node push is covered over Basic auth.
The sample set the on-host dependency Prometheus cert paths, which have
defaults and do not trigger the CA import into the Control Center and node
truststores. Point the sample at the provided_* variables instead (the
external Prometheus CA, and the mTLS client cert/key, staged on the Ansible
control node), matching how the role consumes them.
- archive-scram-rhel: exclude the mock Prometheus host from the custom
  user/group prepare. The mock runs as a local non-root container and
  cannot create a group, which aborted the whole converge before Control
  Center installed.
- byop_validate_negative: rewrite with a flag pattern (set only when the
  validation did NOT fail) instead of inspecting ansible_failed_result, and
  cover all three rejected cases: TLS off, both auth modes, and no auth mode.
- byop-basic-ubuntu: run the negative validation from its verify so it is
  exercised in the pipeline.
byop-brownfield-ubuntu set only the on-host dependency Prometheus CA path,
so the broker and controller CA import (gated on the provided CA source)
was skipped after the migration and the nodes could not push to the
external Prometheus over TLS. Point it at the provided CA source like the
other external scenarios.
…he mock host from upgrade

- The Control Center Prometheus client truststore is a single PKCS12 store
  with no BCFKS variant, even under FIPS, so the BYOP CA import must not
  attempt a BCFKS import. This unblocks the FIPS Basic-auth-over-TLS read.
- byop-upgrade-migrate: exclude the mock Prometheus host from the all-hosts
  version-clear step, which cannot run against the local mock container.
The teardown removed /usr/lib/systemd/system/{prometheus,alertmanager}.service,
which the Control Center Next Gen package owns and the role reinstalls the
bundled services from. Removing them broke any later bundled run (and the
round-trip scenario). Stopping and disabling the services and removing the
overrides, config, and data is enough to take bundled monitoring out of
service, so leave the package unit files in place.
set_fact with omit produced no key/value pair and failed the play. set the
current package version explicitly to move off the pinned old version.
…the FIPS verify

- scram-rhel: read from the external mock Prometheus (SCRAM cluster, package
  install), giving a validatable SCRAM external scenario. The archive-scram
  variant needs a C3 archive tarball, which the nightly build does not ship.
- mtls-java21-rhel-fips: exclude the mock Prometheus host from the all-hosts
  Java-version check, which cannot run against the local mock container.
…ade-migrate

Pin only the old starting version, in a scenario converge.yml. The
side_effect re-runs the platform playbook without pinning a version, so the
upgrade target is the role default (the current version in the codebase)
and never needs a hardcoded new version that would go stale on a release
bump.
Like the brownfield scenario, the upgrade-migrate scenario set only the
on-host CA path, so the broker/controller CA import was skipped after the
migration and the nodes could not push to the external Prometheus. Point it
at the provided CA source.
Re-enable the mock cluster-CA trust step (hot reload of the mock web TLS
config with a newline-safe combined CA bundle) and add a scenario flag that
un-gates the broker and controller push assertions under mTLS. This lets one
mTLS scenario prove that the nodes push to the external Prometheus over mTLS,
not just that Control Center reads over mTLS.
The generated cluster CA is ca.crt for certificates.yml scenarios but
snakeoil-ca-1.crt for the ssl-role self-signed scenarios, so locate it
instead of assuming ca.crt.
Making the mock Prometheus trust the per-run cluster CA at runtime (docker
restart or web-config hot reload) is not reliable: a restart changes the
mock IP and a hot reload leaves the mock TLS in an internal-error state.
Leave the mTLS node-push assertion gated off. The node-push path is covered
by the Basic-auth-over-TLS scenarios (same product code imports the external
Prometheus CA into the node truststores) and by the cross-cloud POC.
Start the mock external Prometheus in prepare (after certificates.yml) instead
of as a molecule platform, so its client CA can be a bundle of the mock CA and
the per-run cluster CA. Under mTLS the nodes' telemetry exporter presents a
cluster-CA-signed client cert, so the mock must trust the cluster CA to accept
the push - this mirrors a real customer whose external Prometheus is configured
to trust the Confluent cluster CA for client authentication.

With the mock trusting both CAs from boot (no runtime reload), enable
byop_assert_mtls_push so byop_verify asserts the broker and controller push
actually lands in the external Prometheus over mTLS.
The container-IP lookup used the command module with a single-quoted Go
template; the command module does not honour quotes and splits the template on
its inner spaces, so docker inspect got a broken --format and the prepare play
failed. Use the shell module and a range over the networks so the template is
passed as one argument and the network name is not hard-coded.
…eate

Inside the molecule docker hosts /etc/hosts is a bind-mounted file, so the
lineinfile module's temp-file-plus-rename failed with 'Device or resource busy'
on every cluster host and broke prepare. Append the mock entry in place with a
grep guard (bind-mount safe, no duplicate on re-run). The docker network alias
remains the primary name resolution; this entry is only a fallback.
The round-trip checked systemctl list-unit-files, but the BYOP teardown keeps the package-owned prometheus/alertmanager unit files (only stopping and disabling them) so the trip back to bundled restarts them without reinstalling. So list-unit-files always shows the units and the midpoint removed-assertion failed. Check running state (systemctl list-units --state=running) at all three points instead: running before, stopped at the external midpoint, running again after the trip back.
The BYOP verify play discovered the C3 properties file by globbing
/opt/confluent/confluent-*/etc/.../control-center-production.properties, which on
archive installs matches the untouched archive bundle's shipped default
(prometheus.url=http://localhost:9090) instead of the rendered active file. The
role writes the rendered config to /opt/confluent/etc/confluent-control-center/
for archive installs, so point the lookup there. Package/rpm installs keep using
/etc and are unaffected.
@rasikaAjoshi23

Copy link
Copy Markdown
Member

LGTM, but build is failing

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants