[DO NOT MERGE] [SETU-3482] BYOP: skip bundled monitoring, point C3 at an external Prometheus - #2623
Open
Ishita Chourasia (ishitach) wants to merge 52 commits into
Open
Ishita Chourasia (ishitach) wants to merge 52 commits into
Ishita Chourasia (ishitach) wants to merge 52 commits into
Conversation
… Next Gen Add control_center_next_gen_bundled_monitoring_enabled (default true). When false, the bundled Prometheus and Alertmanager for Control Center Next Gen are not installed or started, for use with an external Prometheus. Pure bundled tasks (config copy, prometheus overlay, web-configs, dependency file permissions, prometheus/alertmanager keystores) are gated on the toggle. Tasks that configure C3 and the bundled parts together (systemd service copy, override dir/write, service start) keep the C3 parts and skip only the bundled parts. Default true preserves existing behavior.
Rename control_center_next_gen_bundled_monitoring_enabled to control_center_next_gen_external_prometheus_enabled (default false, true = BYOP). The bundled-task gates invert accordingly; default behavior is unchanged. Skip the Alertmanager health check when the external Prometheus is enabled, and keep the Prometheus health check (now pointing at the external endpoint). Brownfield cleanup of the old bundle is a documented manual step for this release, not automated in the playbook.
…BYOP Split the C3 property groups so the bundled Alertmanager/Prometheus management properties (alertmanager url, prometheus config.file, rules.file, alertmanager config.file) render only when the bundle is deployed. When external_prometheus_enabled is true, emit confluent.controlcenter.alerts.enable=false and disable the alertmanager client. In BYOP the customer configures promotion/out-of-order/recording-rules on their own external Prometheus, so those bundled-only properties are not set. Bundled render is unchanged: combine_properties flattens enabled groups by key and the template sorts by key, so moving properties between enabled groups is byte-identical.
…ode) New validate_byop.yml, imported at the top of the C3 role's main.yml and gated on control_center_next_gen_external_prometheus_enabled, so a misconfigured external Prometheus fails fast before any install work. Enforces the supported contract: the endpoint is set, TLS is enabled, and exactly one auth mode (Basic over TLS or mutual TLS) is chosen. There is no Alertmanager-with-external check because the single cp-ansible toggle disables the bundled Prometheus and Alertmanager together.
…heus inventory Document control_center_next_gen_external_prometheus_enabled: turn its defaults comment into a ### doc comment (VARIABLES.md is generated from those) and add the entry to docs/VARIABLES.md. Add docs/sample_inventories/byop_external_prometheus.yml showing the toggle plus the control_center_next_gen_dependency_prometheus_* settings (host/port/ssl/CA), a Basic-auth-over-TLS option and a commented mTLS option, and the external-Prometheus prerequisites as header comments.
The dependency Prometheus keystore/truststore task was gated to bundled-only, so in BYOP mode C3 had no truststore for the external Prometheus and failed to start with a FileNotFoundException on the truststore. C3 needs this truststore to trust the external Prometheus over TLS, so run the task whenever Prometheus TLS is enabled, external or not.
…n BYOP C3 uses this flag to relax the bundled prometheus rules / alertmanager config-file requirement independently of alerts.enable, so it can boot against an external Prometheus. Pairs with the control-center-backend change that gates prometheusBackendEnabled() on this flag. Harmless on builds that predate the flag (unknown config is ignored).
Adds a mock external Prometheus fixture (docker/byop-mock-prometheus: OTLP receiver, promoted Confluent attributes, out-of-order window, C3 recording rules, TLS with basic-auth and mTLS web-configs) and flips three existing scenarios (mtls-ubuntu, oauth-rbac, fips) to read from it over Basic auth. Shared playbooks: byop_verify (bundle absent, external config, alerts off, rules loaded, nodes pushing), byop_distribute_ca, byop_validate_negative (rejects no-TLS and dual-auth), and byop_install_custom_c3 (swap a custom C3 deb or tarball for validating against a build that carries the C3-side change).
…alerts 404 The /2.0/health/status ONLINE check now parses the JSON (bootstrapClusterId present, a clusterStatus value == ONLINE) instead of a loose substring match, and the endpoint path is confirmed against the C3 API (removed the TODO). Add a runtime check that /2.0/alerts/triggers returns 404 in BYOP - the counterpart to the alerts.enable=false config assertion, proving the alert surface is off.
A3 (byop-mtls-ubuntu): greenfield BYOP over mTLS. Reuses the mock Prometheus with web-config-mtls.yml; byop_distribute_client_cert.yml gives C3 the client cert (signed by the mock CA) so it can present it. Focused cluster + byop_verify. A4 (byop-brownfield-ubuntu): install with the bundle, then side_effect.yml flips every node to the external Prometheus, re-runs the platform playbook, and runs byop_cleanup_bundled.yml (the documented manual cleanup); verify asserts external. Both need a molecule run to confirm the mTLS cert-trust chain and the per-pass var flip - flagged inline. The bundle->external->bundle round-trip is left as a follow-on scenario (it ends on bundled, which conflicts with the external verify).
Add build.sh (certs + docker build, one idempotent command) so the mock image is not a manual prerequisite. The on-demand molecule pipeline calls this in setup next to usmagent-mock-ccloud (both are pre_build_image, built outside molecule). Make generate-certs.sh idempotent so re-runs are safe.
… scenarios Two optional/experimental scenarios, clearly flagged as "worth trying, may need a tweak on the first molecule run": - byop-upgrade-migrate (A6): install an older CP + bundle, side_effect upgrades CP (bundle stays), then flips to external + cleanup; verify asserts external. - byop-roundtrip: bundle -> external -> bundle; side_effect flips external + cleanup then back to bundled; verify asserts the bundle is reinstalled. The old->new CP upgrade re-run and the multi-pass toggle flips are the parts that only a live run validates - noted inline.
…ad staged certs
Ishita Chourasia (ishitach)
marked this pull request as ready for review
August 24, 2026 08:29
…temp health-check debug
…default role vars)
…t from dest paths)
…TLS queries; revert debug
Flip three existing feature scenarios to read from an external Prometheus instead of the bundled one, covering the auth mechanisms that were not yet exercised under BYOP: - kerberos-rhel: external Prometheus over TLS + Basic auth - archive-scram-rhel: external Prometheus over TLS + Basic auth - mini-setup-ldap-mtls-fips: external Prometheus over mTLS (FIPS) Each adds the byop-prometheus mock container, the external toggle and connection vars, distributes the mock CA (and client cert for mTLS) in prepare, and imports the shared byop_verify / byop_validate_negative checks. This follows the reuse-and-split approach: bundled scenarios stay as the regression baseline while these variants cover BYOP.
When Control Center reads from an external Prometheus, the Kafka brokers and controllers push their telemetry to that same endpoint. The push uses the node's own truststore, which does not contain the external Prometheus CA, so over TLS the exporter cannot verify the endpoint (and under mTLS the handshake never completes) and no broker metrics reach it. Import the external Prometheus CA into the broker and controller truststores (JKS and, under FIPS, BCFKS) when external Prometheus over TLS is enabled, reusing the shared idp_certs import task and restarting the node so the exporter picks it up. The CA source is the existing control_center_next_gen_dependency_prometheus_provided_ca_cert_path. Test wiring: every external molecule scenario now sets the provided CA source, and each mTLS scenario rebuilds the mock Prometheus client CA as a bundle of the mock CA and the per-run cluster CA (byop_mock_trust_cluster_ca) so the mock accepts the nodes' cluster-signed client certificates on push.
- byop_mock_trust_cluster_ca: hot-reload the mock's web TLS config instead of restarting the container. A restart changed the mock's IP and broke name resolution for it from the other hosts at verify time. - kerberos-rhel, archive-scram-rhel: run the BYOP verify plays first. A host that fails a task is skipped by every later play, so a scenario-specific check failing on the Control Center host would otherwise prevent the BYOP checks (which target that host) from running at all. - archive-scram-rhel: exclude the mock Prometheus host from the scenario's all-hosts archive-file assertion (it is not a Confluent Platform node). - byop-basic-ubuntu: drop a duplicated provided-CA variable.
Two problems prevented Control Center from reading an external Prometheus over Basic auth on a TLS endpoint (no mTLS): - The custom-certs switch keyed off the on-host client cert path, which has a non-empty default, so it was always on. Without a provided client cert the copy of the (absent) signed cert failed and Control Center configuration was left with the default localhost Prometheus URL. Gate the switch on the PROVIDED client cert path instead, which is set only for mTLS. - With custom certs off, the ssl role self-signs the Prometheus client truststore with a generated CA, so it does not trust the external endpoint. Import the external Prometheus CA into that truststore (JKS and, under FIPS, BCFKS) when reading an external Prometheus over TLS without mTLS, so Control Center trusts the endpoint's server cert on read. mTLS is unaffected: the provided-cert path still builds the client keystore and imports the provided CA into the truststore.
byop_validate_negative is a localhost unit test of validate_byop.yml (it feeds bad configurations and checks they are rejected); it does not depend on the deployed cluster. Running it inside every scenario's verify made the whole scenario fail on a localhost-only assertion even when the real BYOP checks passed. Keep it as a standalone playbook, run separately, and let the scenario verify gate only on byop_verify.
Three fixes for the remaining split BYOP scenarios: - playbooks/all.yml: the certificate-authority generation was gated only on each component's own listener TLS. When the cluster runs without TLS (for example SASL_PLAINTEXT) but Control Center still reads an external Prometheus over TLS, the shared self-signed CA was never generated and the Control Center Prometheus client keystore step failed. Include the external Prometheus TLS flag in the condition. - byop_verify: discover the Control Center properties file (package/rpm under /etc, archive under the confluent home) instead of hard-coding the package path, so the archive scenario reads the right file. - byop_mock_trust_cluster_ca: normalise the combined client CA bundle so the two PEM blocks never merge onto one line, which made the bundle unparseable and broke the mock Prometheus TLS handshake.
- certificate_authority.yml: the inner CA-generation and copy-back tasks carried the same component-TLS condition as the play-level include, so the CA was still skipped on a non-TLS cluster that only needs TLS for the external Prometheus. Add the external Prometheus TLS flag to both. - byop_verify: skip the broker/controller push assertions under mTLS. The mock Prometheus requires a client cert signed by its own CA, so it rejects the node's cluster-signed client cert in this harness; the node-push path is covered by the Basic-auth-over-TLS scenarios with the same product code. - Drop the mock cluster-CA trust step from the mTLS scenarios: modifying the running mock's client CA broke its TLS handshake. mTLS coverage validates the Control Center read path; node push is covered over Basic auth.
The sample set the on-host dependency Prometheus cert paths, which have defaults and do not trigger the CA import into the Control Center and node truststores. Point the sample at the provided_* variables instead (the external Prometheus CA, and the mTLS client cert/key, staged on the Ansible control node), matching how the role consumes them.
- archive-scram-rhel: exclude the mock Prometheus host from the custom user/group prepare. The mock runs as a local non-root container and cannot create a group, which aborted the whole converge before Control Center installed. - byop_validate_negative: rewrite with a flag pattern (set only when the validation did NOT fail) instead of inspecting ansible_failed_result, and cover all three rejected cases: TLS off, both auth modes, and no auth mode. - byop-basic-ubuntu: run the negative validation from its verify so it is exercised in the pipeline.
byop-brownfield-ubuntu set only the on-host dependency Prometheus CA path, so the broker and controller CA import (gated on the provided CA source) was skipped after the migration and the nodes could not push to the external Prometheus over TLS. Point it at the provided CA source like the other external scenarios.
…he mock host from upgrade - The Control Center Prometheus client truststore is a single PKCS12 store with no BCFKS variant, even under FIPS, so the BYOP CA import must not attempt a BCFKS import. This unblocks the FIPS Basic-auth-over-TLS read. - byop-upgrade-migrate: exclude the mock Prometheus host from the all-hosts version-clear step, which cannot run against the local mock container.
The teardown removed /usr/lib/systemd/system/{prometheus,alertmanager}.service,
which the Control Center Next Gen package owns and the role reinstalls the
bundled services from. Removing them broke any later bundled run (and the
round-trip scenario). Stopping and disabling the services and removing the
overrides, config, and data is enough to take bundled monitoring out of
service, so leave the package unit files in place.
set_fact with omit produced no key/value pair and failed the play. set the current package version explicitly to move off the pinned old version.
…the FIPS verify - scram-rhel: read from the external mock Prometheus (SCRAM cluster, package install), giving a validatable SCRAM external scenario. The archive-scram variant needs a C3 archive tarball, which the nightly build does not ship. - mtls-java21-rhel-fips: exclude the mock Prometheus host from the all-hosts Java-version check, which cannot run against the local mock container.
…ade-migrate Pin only the old starting version, in a scenario converge.yml. The side_effect re-runs the platform playbook without pinning a version, so the upgrade target is the role default (the current version in the codebase) and never needs a hardcoded new version that would go stale on a release bump.
Like the brownfield scenario, the upgrade-migrate scenario set only the on-host CA path, so the broker/controller CA import was skipped after the migration and the nodes could not push to the external Prometheus. Point it at the provided CA source.
Re-enable the mock cluster-CA trust step (hot reload of the mock web TLS config with a newline-safe combined CA bundle) and add a scenario flag that un-gates the broker and controller push assertions under mTLS. This lets one mTLS scenario prove that the nodes push to the external Prometheus over mTLS, not just that Control Center reads over mTLS.
The generated cluster CA is ca.crt for certificates.yml scenarios but snakeoil-ca-1.crt for the ssl-role self-signed scenarios, so locate it instead of assuming ca.crt.
Making the mock Prometheus trust the per-run cluster CA at runtime (docker restart or web-config hot reload) is not reliable: a restart changes the mock IP and a hot reload leaves the mock TLS in an internal-error state. Leave the mTLS node-push assertion gated off. The node-push path is covered by the Basic-auth-over-TLS scenarios (same product code imports the external Prometheus CA into the node truststores) and by the cross-cloud POC.
Start the mock external Prometheus in prepare (after certificates.yml) instead of as a molecule platform, so its client CA can be a bundle of the mock CA and the per-run cluster CA. Under mTLS the nodes' telemetry exporter presents a cluster-CA-signed client cert, so the mock must trust the cluster CA to accept the push - this mirrors a real customer whose external Prometheus is configured to trust the Confluent cluster CA for client authentication. With the mock trusting both CAs from boot (no runtime reload), enable byop_assert_mtls_push so byop_verify asserts the broker and controller push actually lands in the external Prometheus over mTLS.
The container-IP lookup used the command module with a single-quoted Go template; the command module does not honour quotes and splits the template on its inner spaces, so docker inspect got a broken --format and the prepare play failed. Use the shell module and a range over the networks so the template is passed as one argument and the network name is not hard-coded.
…eate Inside the molecule docker hosts /etc/hosts is a bind-mounted file, so the lineinfile module's temp-file-plus-rename failed with 'Device or resource busy' on every cluster host and broke prepare. Append the mock entry in place with a grep guard (bind-mount safe, no duplicate on re-run). The docker network alias remains the primary name resolution; this entry is only a fallback.
The round-trip checked systemctl list-unit-files, but the BYOP teardown keeps the package-owned prometheus/alertmanager unit files (only stopping and disabling them) so the trip back to bundled restarts them without reinstalling. So list-unit-files always shows the units and the midpoint removed-assertion failed. Check running state (systemctl list-units --state=running) at all three points instead: running before, stopped at the external midpoint, running again after the trip back.
The BYOP verify play discovered the C3 properties file by globbing /opt/confluent/confluent-*/etc/.../control-center-production.properties, which on archive installs matches the untouched archive bundle's shipped default (prometheus.url=http://localhost:9090) instead of the rendered active file. The role writes the rendered config to /opt/confluent/etc/confluent-control-center/ for archive installs, so point the lookup there. Package/rpm installs keep using /etc and are unaffected.
This was referenced Aug 27, 2026
Rasika Joshi (rasikaAjoshi23)
approved these changes
Sep 1, 2026
Member
|
LGTM, but build is failing |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
DO NOT MERGE
First-level WIP PR for early review. It depends on an unmerged control-center-backend change and is not yet fully validated end to end. See "State" below.
Summary
Adds Bring Your Own Prometheus (BYOP) support to the
control_center_next_genrole: a toggle (control_center_next_gen_external_prometheus_enabled) that skips deploying the bundled Prometheus and Alertmanager, points C3 and the node exporters at a customer-managed external Prometheus, turns C3 alerts off, and emitsconfluent.controlcenter.prometheus.external.enableso C3 boots against the external endpoint.Tickets: SETU-3482 (parent), MMA-19237 / 19238 / 19239 / 19240 / 19241.
What is here
prometheus.external.enable(MMA-19239).Dependency
The C3 boot path depends on a control-center-backend change that gates
prometheusBackendEnabled()on the newprometheus.external.enableflag, decoupled fromalerts.enable. Without that build, stock C3 crashes on boot in BYOP mode. This PR emits the flag; the C3 change consumes it.State (why do not merge)
Test plan
mtls-ubuntuconverge: full CP install reached C3 config generation; verified the emitted C3 config (externalprometheus.url,prometheus.external.enable=true,alerts.enable=false, no bundled files) and that the bundled prometheus/alertmanager services are disabled.