diff --git a/CHANGELOG.md b/CHANGELOG.md index 9ff6724a..36087939 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -42,11 +42,15 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0 - **`components.klaus-gateway.versionRange` admits the 2.x line (`>=1.20.0 <3.0.0`)** (giantswarm/klaus-gateway#319). klaus-gateway 2.0.0 is the Slack-only gateway: the web and CLI channels, the OpenAI-compatible `/v1` front door and the Klaus instance path are removed there; the values keys this chart still forwards for them (`cli`, `lifecycle`, `upstream`, `agentgateway`, `routing.defaultTTL`, `a2a.saToken`) are accepted by the 2.x chart as no-ops, so the installations roll to it unchanged. Those keys leave this chart in a later change, once the 2.x line has rolled. - **The klaus-gateway ServiceMonitor carries the tenant label: `klausGateway.serviceMonitor`** (giantswarm/giantswarm#36711, giantswarm/klaus-gateway#316). The chart's monitor was on and carried the chart's own labels only, so Mimir routed its scrape to no tenant and every `klaus_gateway_*` series was lost on every installation (seen on gazelle, chart 1.19.1). `klausGateway.serviceMonitor.labels` sets `observability.giantswarm.io/tenant: giantswarm` and `klausGateway.serviceMonitor.enabled` is `auto`, resolved to the boolean the chart takes from the same answer as every other monitor of this chart. Pin first: `components.klaus-gateway.versionRange` is `>=1.20.0 <2.0.0` and `examples/customer-bom.yaml` pins `1.20.0`, the release that opens the labels key; an older chart's schema refuses it. `verify-klausgateway-otlp` asserts the resolution and the label; `verify-components-charts` names `1.20.0` as `UNRELEASED` until it exists. - **The mcp-kubernetes chart's own ServiceMonitor and its three Grafana boards are on** (giantswarm/giantswarm#36711). The chart ships a ServiceMonitor over the dedicated metrics port and the boards `administrator`, `security` and `cluster-operator`, and this chart turned none of them on, so an installation that runs the Kubernetes MCP server collected nothing from it. `mcp-kubernetes.mcpKubernetes.instrumentation.serviceMonitor.enabled` and `mcp-kubernetes.grafanaDashboards.enabled` follow the resolved `global.observability.metrics.serviceMonitor.enabled` (`auto | true | false`), the monitor carries `observability.giantswarm.io/tenant`, and the boards land in `Shared Org / Agent Platform` beside the platform's own. The chart's `prometheusRules` stay off: its three alerts link to runbook pages that do not exist yet. The floor `1.1.1` already carries every key, so no range moves. +- **Substrate's six metrics endpoints are collected: `substrate.metrics.podMonitor`** (giantswarm/giantswarm#36711). ate-api-server, atelet, atenet-router, atenet-egress, k8s-credential-provider and ate-controller each serve metrics, the fleet's collector discovers monitors and nothing else, and the chart shipped no monitor, so the whole ate registry was served and thrown away: actor crashes, lifecycle and scheduler durations, checkpoint and restore phases, actor CPU and memory, image-cache hits, worker pool desired against ready, router route duration and parking, and ate-controller's `controller_runtime_*` and `workqueue_*`. The block follows the resolved `global.observability.metrics.serviceMonitor.enabled` (`auto | true | false`) and carries `observability.giantswarm.io/tenant`. The floor `1.0.3` already carries the monitors (giantswarm/substrate#50, upstream kagent-dev/substrate#46, open). +- **The kagent controller's metrics are collected: `kagent.controller.metrics`, with the kagent chart's own ServiceMonitor** (giantswarm/giantswarm#36711). The controller served no endpoint before the line's `1.0.2` (`metricsserver.Options{BindAddress: "0"}`, which no variable changed, so the `METRICS_BIND_ADDRESS` this chart set had no effect); from `1.0.2` on it reads `METRICS_BIND_ADDRESS` and `METRICS_SECURE`, and the controller-runtime families and the line's `kagent_grpc_server_requests_total` / `..._request_duration_seconds` reach a scrape. `metrics.enabled: true`, plain HTTP on `:8080` (`secureServing: false`: the secure endpoint authenticates the scraper with a TokenReview and authorizes it with a SubjectAccessReview, which the fleet's collector does not do), `serviceMonitor.enabled: auto` following the resolved `global.observability.metrics.serviceMonitor.enabled`, the tenant label on the monitor. The floor `1.0.3` carries the keys. - **Every agent platform dashboard lands in one folder customers can reach: `Shared Org / Agent Platform`** (giantswarm/giantswarm#36711). The muster board (`Muster / MCP Gateway`) was loaded into the staff-only `Giant Swarm` organization, in a folder named after the component, so the people who run agents on the platform could not see it: `muster.observability.grafanaDashboard.folder` is `Agent Platform` and `.giantswarm.organization` is `Shared Org`. `Shared Org` is the organization every logged-in customer reaches as a Viewer and it carries the observability data of Giant Swarm managed components, which is where these boards belong; the `Giant Swarm` organization is staff-only. The boards that follow — the agentgateway gateway board, the platform's own overview — land in the same folder. - **The agentgateway packaging chart's own monitors and dashboard are on: `agentgateway.monitoring.enabled`** (giantswarm/giantswarm#36711, giantswarm/agentgateway#60). It follows the resolved `global.observability.metrics.serviceMonitor.enabled` like every other monitor of this chart (`auto | true | false`), at the fleet's 60s interval and with the `observability.giantswarm.io/tenant` label on both monitors. It turns on three objects the platform did not have: the **controller** ServiceMonitor, which nothing scraped, and with it the only view of the control plane that serves the data planes their config; upstream's proxy PodMonitor; and upstream's own Grafana board — Overview, Requests, LLM, **MCP tool calls**, Latency, **XDS**, **Runtime** — vendored in the packaging chart since 2.0.0 and never rendered. The MCP and runtime series it reads were already collected and had no consumer. The board is upstream's to maintain, so a version bump carries its fixes; the boards the platform writes itself (LLM cost per agent and per person) ship in this chart's `dashboards/` directory. ### Removed +- **`kagent.serviceMonitor` and the connectivity chart's kagent controller metrics Service and ServiceMonitor** (`templates/kagent/servicemonitor.yaml`; giantswarm/giantswarm#36711). Both were a workaround for a kagent chart that exposed only the API port; the line's chart ships the metrics Service and its ServiceMonitor since `1.0.2` (`kagent.controller.metrics`), and keeping ours would have scraped one endpoint twice. An installation that set `kagent.serviceMonitor.*` drops the key. + - **The Qwen3 small presets `qwen3-4b-instruct`, `qwen3-8b-fp8` and `qwen3-14b`** (giantswarm/agent-platform#591). They pinned 2025 checkpoints whose successors are smaller or stronger under the same licence and are replaced by the 24 GB line-up above (`qwen3-5-4b` for the 4B, `qwen3-5-9b-fp8` or `gemma-4-12b` for the 8B, `gpt-oss-20b` or `gemma-4-12b` for the 14B — which, at 28 GiB of BF16 weights, never fit a 24 GB card and was tuned for a 128 GB node). The upgrade removes their ConfigMaps; a model already served from one keeps serving (the `LLMInferenceService` is model-manager's object), and an installation that wants one back carries its file under `modelServing.presets` (UPGRADE.md). - **The presets no single card holds: `qwen3-5-27b`, `qwen3-coder-next` and `qwen3-5-35b-a3b`** (giantswarm/agent-platform#591): 54 GiB BF16, 80 GiB FP8 and 70 GiB BF16 of weights exceed the 48 GB GPU they would target; the presets above are their model-image successors of the same shape. The shipped set goes from twelve to eleven presets; `tests/verify-gpu-pool.py` leaves each side's own presets out of the golden comparison until the golden carries this change. diff --git a/Makefile.custom.mk b/Makefile.custom.mk index e0283103..b3fae985 100644 --- a/Makefile.custom.mk +++ b/Makefile.custom.mk @@ -277,21 +277,13 @@ verify-global: ## Assert the global.* contract behaviors (derived hostnames, gat if grep -q "$$pattern" /tmp/vg-mon.out; then echo "FAIL: monitor-gated render still contains $$pattern"; exit 1; fi; \ done @echo "ok: monitor gate" - @echo "--> the default render keeps the CNPG PodMonitor (fleet behavior) and renders NO kagent ServiceMonitor or metrics Service: the kagent line serves no /metrics (kagent.serviceMonitor.enabled: false)" - @helm template t $(CONNECTIVITY_DIR) $(VM) --set components.kagent.enabled=true --set postgres.enabled=true >/tmp/vg-mon-default.out 2>&1 || { cat /tmp/vg-mon-default.out; exit 1; } - @if grep -q 'kind: ServiceMonitor' /tmp/vg-mon-default.out; then echo "FAIL: the default render carries a kagent ServiceMonitor; the line's controller serves no /metrics and the monitor would sit at up=0"; exit 1; fi - @if grep -q 'kagent-controller-metrics' /tmp/vg-mon-default.out; then echo "FAIL: the default render carries the kagent controller metrics Service; nothing listens behind it on the line"; exit 1; fi - @grep -q 'enablePodMonitor: true' /tmp/vg-mon-default.out || { echo "FAIL: default render lost the CNPG PodMonitor"; exit 1; } - @grep -q 'helm.sh/resource-policy: keep' /tmp/vg-mon-default.out || { echo "FAIL: the CNPG Cluster lost helm.sh/resource-policy: keep"; exit 1; } - @echo "ok: default: no kagent monitor, CNPG PodMonitor + keep" - @echo "--> kagent.serviceMonitor.enabled=true (for when upstream serves metrics) renders the Service and the ServiceMonitor under the global gate" - @helm template t $(CONNECTIVITY_DIR) $(VM) --set components.kagent.enabled=true --set postgres.enabled=true --set kagent.serviceMonitor.enabled=true >/tmp/vg-mon-on.out 2>&1 || { cat /tmp/vg-mon-on.out; exit 1; } - @grep -q 'kind: ServiceMonitor' /tmp/vg-mon-on.out || { echo "FAIL: kagent.serviceMonitor.enabled=true renders no ServiceMonitor"; exit 1; } - @grep -q 'observability.giantswarm.io/tenant: giantswarm' /tmp/vg-mon-on.out || { echo "FAIL: the kagent ServiceMonitor lost the tenant label"; exit 1; } - @helm template t $(CONNECTIVITY_DIR) $(VM) --set components.kagent.enabled=true --set kagent.serviceMonitor.enabled=true --set global.observability.metrics.serviceMonitor.enabled=false 2>/dev/null | grep -q 'kind: ServiceMonitor' && { echo "FAIL: the global monitor gate no longer holds the kagent ServiceMonitor back"; exit 1; } || true - @echo "ok: the toggle under the global gate" - @echo "--> the kagent controller metrics Service selects kagent's own release instance (the pods' label), the ServiceMonitor this chart's Service" - @python3 -c 'import re,sys; docs=open("/tmp/vg-mon-on.out").read().split("\n---\n"); svc=[d for d in docs if "\nkind: Service\n" in d and re.search(r"^ name: t-kagent-controller-metrics$$", d, re.M)]; sys.exit("FAIL: the kagent controller metrics Service did not render") if len(svc)!=1 else None; sel=svc[0][svc[0].index(" selector:"):]; sys.exit("FAIL: the metrics Service does not select app.kubernetes.io/instance: kagent (the kagent release name the meta chart fixes):\n"+sel) if not re.search(r"^ app.kubernetes.io/instance: kagent$$", sel, re.M) else None; sys.exit("FAIL: the metrics Service selects this release (t) — under the meta chart that matches no pod (#305)") if re.search(r"^ app.kubernetes.io/instance: \"?t\"?$$", sel, re.M) else None; sm=[d for d in docs if "kind: ServiceMonitor" in d and "-kagent-controller\n" in d]; sys.exit("FAIL: the kagent ServiceMonitor did not render") if len(sm)!=1 else None; sys.exit("FAIL: the ServiceMonitor must select this chart\x27s Service (instance t)") if " app.kubernetes.io/instance: \"t\"" not in sm[0] else None; print("ok: metrics Service selects instance kagent; the ServiceMonitor selects this release\x27s Service")' + @echo "--> the default render keeps the CNPG PodMonitor (fleet behavior) and renders NO kagent ServiceMonitor or metrics Service: both are the kagent chart's own (controller.metrics)" + @helm template t $(CONNECTIVITY_DIR) $(VM) --set components.kagent.enabled=true --set postgres.enabled=true >/tmp/vg-mon-on.out 2>&1 || { cat /tmp/vg-mon-on.out; exit 1; } + @if grep -q 'kind: ServiceMonitor' /tmp/vg-mon-on.out; then echo "FAIL: this chart renders a ServiceMonitor; every monitor belongs to the component's own chart (the kagent controller's to controller.metrics.serviceMonitor, the agentgateway data plane's to the packaging chart)"; exit 1; fi + @if grep -q 'kagent-controller-metrics' /tmp/vg-mon-on.out; then echo "FAIL: this chart renders the kagent controller metrics Service; the kagent chart renders it under controller.metrics.enabled"; exit 1; fi + @grep -q 'enablePodMonitor: true' /tmp/vg-mon-on.out || { echo "FAIL: default render lost the CNPG PodMonitor"; exit 1; } + @grep -q 'helm.sh/resource-policy: keep' /tmp/vg-mon-on.out || { echo "FAIL: the CNPG Cluster lost helm.sh/resource-policy: keep"; exit 1; } + @echo "ok: no monitor of this chart's own, CNPG PodMonitor + keep" @echo "--> no kagent-targeting selector, Service name or hostname in this chart derives from .Release.Name (the standalone umbrella's one-release assumption)" @if grep -nE 'fullnameOverride \| default \.Release\.Name|fullnameOverride" \| default \(printf "%s-oauth2-proxy" \.Release\.Name' $(CONNECTIVITY_DIR)/templates/kagent/*.yaml; then echo "FAIL: a kagent template falls back to .Release.Name for a kagent-chart object; use agent-platform.kagent.fullname / agent-platform.kagent.releaseName"; exit 1; else echo "ok: kagent templates derive kagent names from the kagent helpers"; fi @echo "--> the CNPG CiliumNetworkPolicy renders only when postgres.enabled" diff --git a/helm/agent-platform-connectivity/README.md b/helm/agent-platform-connectivity/README.md index 672c67ec..a53b44d5 100644 --- a/helm/agent-platform-connectivity/README.md +++ b/helm/agent-platform-connectivity/README.md @@ -1089,12 +1089,8 @@ The kagent block is open in the schema, so the template refuses a key under `kag | kagent.controller.skillsInitImage.repository | string | `"kagent-skills-init"` | | | kagent.controller.auth.mode | string | `"trusted-proxy"` | | | kagent.controller.auth.userIdClaim | string | `"email"` | | -| kagent.controller.env[0].name | string | `"METRICS_BIND_ADDRESS"` | | -| kagent.controller.env[0].value | string | `":8080"` | | -| kagent.controller.env[1].name | string | `"METRICS_SECURE"` | | -| kagent.controller.env[1].value | string | `"false"` | | -| kagent.controller.env[2].name | string | `"OTEL_EXPORTER_OTLP_HEADERS"` | | -| kagent.controller.env[2].value | string | `"X-Scope-OrgID=giantswarm"` | | +| kagent.controller.env[0].name | string | `"OTEL_EXPORTER_OTLP_HEADERS"` | | +| kagent.controller.env[0].value | string | `"X-Scope-OrgID=giantswarm"` | | | kagent.controller.vpa.enabled | string | `"auto"` | | | kagent.controller.vpa.updateMode | string | `"InPlaceOrRecreate"` | | | kagent.controller.vpa.controlledValues | string | `"RequestsOnly"` | | @@ -1120,9 +1116,6 @@ The kagent block is open in the schema, so the template refuses a key under `kag | kagent.providers.anthropic.apiKeySecretRef | string | `"kagent-anthropic"` | | | kagent.providers.anthropic.apiKeySecretKey | string | `"ANTHROPIC_API_KEY"` | | | kagent.providers.anthropic.apiKey | string | `""` | | -| kagent.serviceMonitor.enabled | bool | `false` | | -| kagent.serviceMonitor.interval | string | `"60s"` | | -| kagent.serviceMonitor.labels."observability.giantswarm.io/tenant" | string | `"giantswarm"` | | | kagent.otel.tracing.enabled | string | `"auto"` | | | kagent.otel.tracing.exporter.otlp.endpoint | string | `"http://otlp-gateway.kube-system.svc:4317"` | | | kagent.otel.tracing.exporter.otlp.protocol | string | `"grpc"` | | diff --git a/helm/agent-platform-connectivity/templates/kagent/servicemonitor.yaml b/helm/agent-platform-connectivity/templates/kagent/servicemonitor.yaml deleted file mode 100644 index 6a2909d9..00000000 --- a/helm/agent-platform-connectivity/templates/kagent/servicemonitor.yaml +++ /dev/null @@ -1,70 +0,0 @@ -{{- /* Both gates apply: global.observability.metrics.serviceMonitor.enabled -switches every monitor object this chart renders (a vanilla cluster has no -monitoring.coreos.com CRDs); kagent.serviceMonitor.enabled keeps working -underneath as this object's own toggle. */ -}} -{{- if and (include "agent-platform.componentEnabled" (dict "root" $ "name" "kagent")) (include "agent-platform.serviceMonitor" .) .Values.kagent.serviceMonitor.enabled }} -{{- $kagentNs := .Values.kagent.namespaceOverride | default .Release.Namespace }} -{{- $releaseName := .Release.Name }} -{{- $monitor := .Values.global.observability.metrics.serviceMonitor }} -# Service exposing the kagent controller metrics port (:8080). -# The upstream controller-service.yaml only exposes port 8083 (API). We add a -# separate Service here so the ServiceMonitor below can target it without -# patching the kagent chart's single-port Service template. -# -# The Service is THIS chart's object (named and labelled after this release, -# which is what the ServiceMonitor's matchLabels select), but it selects the -# kagent chart's pods, which carry kagent's OWN release name in their instance -# label — `kagent`, the component key the meta chart fixes as the kagent release -# name (agent-platform.kagent.releaseName). Selecting this release's name here -# matched nothing under the meta chart: the Service had no endpoints and the -# controller was never scraped (giantswarm/agent-platform#305). ---- -apiVersion: v1 -kind: Service -metadata: - name: {{ $releaseName }}-kagent-controller-metrics - namespace: {{ $kagentNs }} - labels: - {{- include "labels.common" . | nindent 4 }} - app.kubernetes.io/component: kagent-metrics -spec: - type: ClusterIP - selector: - app.kubernetes.io/name: kagent - app.kubernetes.io/instance: {{ include "agent-platform.kagent.releaseName" . }} - app.kubernetes.io/component: controller - ports: - - name: metrics - port: 8080 - targetPort: 8080 - protocol: TCP ---- -# ServiceMonitor for the kagent controller's controller-runtime Prometheus metrics. -# Requires the Prometheus Operator CRD (monitoring.coreos.com/v1). -# The controller exposes metrics at :8080/metrics (HTTP, no TLS) when -# METRICS_BIND_ADDRESS=:8080 and METRICS_SECURE=false are set. -apiVersion: monitoring.coreos.com/v1 -kind: ServiceMonitor -metadata: - name: {{ include "name" . }}-kagent-controller - namespace: {{ $kagentNs }} - labels: - {{- include "labels.common" . | nindent 4 }} - {{- with ($monitor.labels | default .Values.kagent.serviceMonitor.labels) }} - {{- toYaml . | nindent 4 }} - {{- end }} -spec: - selector: - matchLabels: - app.kubernetes.io/name: {{ include "name" . | quote }} - app.kubernetes.io/instance: {{ $releaseName | quote }} - app.kubernetes.io/component: kagent-metrics - namespaceSelector: - matchNames: - - {{ $kagentNs }} - endpoints: - - port: metrics - path: /metrics - scheme: http - interval: {{ $monitor.interval | default .Values.kagent.serviceMonitor.interval | default "60s" }} -{{- end }} diff --git a/helm/agent-platform-connectivity/values.yaml b/helm/agent-platform-connectivity/values.yaml index 5b3853f6..bc1f8dcf 100644 --- a/helm/agent-platform-connectivity/values.yaml +++ b/helm/agent-platform-connectivity/values.yaml @@ -1173,21 +1173,12 @@ kagent: # @schema skipProperties: true; additionalProperties: true # stable human identity and matches muster's trusted-issuer subjectClaim, # so UI (oauth2-proxy) and A2A sessions key identically. userIdClaim: email - # Expose controller-runtime Prometheus metrics on :8080 (HTTP, no TLS). - # The default --metrics-bind-address is "0" (disabled) and the chart never - # sets it. We enable it here and wire a ServiceMonitor via extraObjects. - # metrics-secure=false avoids needing a cert for the in-cluster scrape. # X-Scope-OrgID routes OTLP traces and logs to the correct Loki/Tempo # tenant. The configmap template has no headers field so we inject via env. + # METRICS_BIND_ADDRESS and METRICS_SECURE are not set here: the kagent + # chart renders both from controller.metrics, and a second entry of the + # same name would render the variable twice. env: - - name: METRICS_BIND_ADDRESS - value: ":8080" - # false: HTTP scrape, no cert needed. The network policy on port 8083 - # already restricts access; HTTP on :8080 is acceptable for in-cluster - # scraping. Set to "true" + configure ServiceMonitor tlsConfig with a - # cert-manager-issued CA if HTTPS is required. - - name: METRICS_SECURE - value: "false" - name: OTEL_EXPORTER_OTLP_HEADERS value: "X-Scope-OrgID=giantswarm" # VerticalPodAutoscaler on the controller Deployment (templates/kagent/ @@ -1284,21 +1275,6 @@ kagent: # @schema skipProperties: true; additionalProperties: true apiKeySecretKey: ANTHROPIC_API_KEY apiKey: "" # Set per-cluster in secret-values; chart creates the Secret - # The controller metrics Service and ServiceMonitor (templates/kagent/ - # servicemonitor.yaml), under the global.observability.metrics.serviceMonitor - # gate (monitoring.coreos.com/v1). Off: the kagent line's controller serves - # no Prometheus /metrics, so the monitor would scrape a refused port and sit - # at up=0; nothing reads those metrics (the fleet's KagentControllerDown rule - # reads the Deployment's kube-state-metrics series). Back to true when - # upstream serves them. - serviceMonitor: - enabled: false - interval: 60s - # Extra labels on the ServiceMonitor metadata. The tenant label routes - # metrics to the correct Mimir instance in the GS Observability Platform. - labels: - observability.giantswarm.io/tenant: giantswarm - # OTel — traces and logs to the GS Observability Platform OTLP gateway. # X-Scope-OrgID header routes to the correct Loki/Tempo tenant. otel: diff --git a/helm/agent-platform/README.md b/helm/agent-platform/README.md index c39a67e0..ccd69b07 100644 --- a/helm/agent-platform/README.md +++ b/helm/agent-platform/README.md @@ -343,9 +343,8 @@ The map is merged into each component's own `nodeSelector` (`muster.nodeSelector | components.kagent.omitKeys[4] | string | `"modelConfigs"` | | | components.kagent.omitKeys[5] | string | `"oauth2ProxyIngress"` | | | components.kagent.omitKeys[6] | string | `"remoteMcpServers"` | | -| components.kagent.omitKeys[7] | string | `"serviceMonitor"` | | -| components.kagent.omitKeys[8] | string | `"uiRoute"` | | -| components.kagent.omitKeys[9] | string | `"substrateWorkerPool.podDisruptionBudget"` | | +| components.kagent.omitKeys[7] | string | `"uiRoute"` | | +| components.kagent.omitKeys[8] | string | `"substrateWorkerPool.podDisruptionBudget"` | | | components.kagent.omitEmptyKeys[0] | string | `"substrateWorkerPool.workerImage"` | | | components.kagent.omitEmptyKeys[1] | string | `"harness.image"` | | | components.kagent.enabled | bool | `false` | | @@ -761,7 +760,13 @@ The map is merged into each component's own `nodeSelector` (`muster.nodeSelector | kagent.controller.vpa.minAllowed.memory | string | `"128Mi"` | | | kagent.controller.vpa.maxAllowed.cpu | string | `"1900m"` | | | kagent.controller.vpa.maxAllowed.memory | string | `"1280Mi"` | | -| kagent.controller.metrics.enabled | bool | `false` | | +| kagent.controller.metrics.enabled | bool | `true` | | +| kagent.controller.metrics.bindAddress | string | `":8080"` | | +| kagent.controller.metrics.secureServing | bool | `false` | | +| kagent.controller.metrics.service.port | int | `8080` | | +| kagent.controller.metrics.serviceMonitor.enabled | string | `"auto"` | | +| kagent.controller.metrics.serviceMonitor.interval | string | `"60s"` | | +| kagent.controller.metrics.serviceMonitor.labels."observability.giantswarm.io/tenant" | string | `"giantswarm"` | | | kagent.controller.env[0].name | string | `"OTEL_EXPORTER_OTLP_HEADERS"` | | | kagent.controller.env[0].value | string | `"X-Scope-OrgID=giantswarm"` | | | kagent.ui.image.repository | string | `"giantswarm/kagent/ui"` | | @@ -797,9 +802,6 @@ The map is merged into each component's own `nodeSelector` (`muster.nodeSelector | kagent.providers.anthropic.apiKey | string | `""` | | | kagent.providers.anthropic.config.promptCaching | bool | `true` | | | kagent.providers.anthropic.config.cacheTTL | string | `"5m"` | | -| kagent.serviceMonitor.enabled | bool | `false` | | -| kagent.serviceMonitor.interval | string | `"60s"` | | -| kagent.serviceMonitor.labels."observability.giantswarm.io/tenant" | string | `"giantswarm"` | | | kagent.otel.tracing.enabled | string | `"auto"` | | | kagent.otel.tracing.exporter.otlp.endpoint | string | `"http://otlp-gateway.kube-system.svc:4317"` | | | kagent.otel.tracing.exporter.otlp.protocol | string | `"grpc"` | | @@ -1292,6 +1294,8 @@ The map is merged into each component's own `nodeSelector` (`muster.nodeSelector | kagent-crds.substrate.enabled | bool | `false` | | | substrate.createNamespace | bool | `false` | | | substrate.image.registry | string | `"gsoci.azurecr.io/giantswarm/substrate"` | | +| substrate.metrics.podMonitor.enabled | string | `"auto"` | | +| substrate.metrics.podMonitor.labels."observability.giantswarm.io/tenant" | string | `"giantswarm"` | | | substrate.postgres.enabled | string | `"auto"` | | | substrate.postgres.connectionString | string | `""` | | | substrate.postgres.schema | string | `"public"` | | diff --git a/helm/agent-platform/templates/_helpers.tpl b/helm/agent-platform/templates/_helpers.tpl index 596ddf83..7bd8e64d 100644 --- a/helm/agent-platform/templates/_helpers.tpl +++ b/helm/agent-platform/templates/_helpers.tpl @@ -1344,7 +1344,11 @@ answers, but only where the leaf is left at `auto`: own monitor; the chart takes a boolean), vm-manager.serviceMonitor.enabled, kserve-llmisvc-resources.kserve.llmisvc - .controller.serviceMonitor.enabled + .controller.serviceMonitor.enabled, + substrate.metrics.podMonitor.enabled (the six + workloads of the Substrate control plane), + kagent.controller.metrics.serviceMonitor.enabled + (the controller's own, from the line's 1.0.2) Two leaves have no `auto` form and are derived directly, off only: valkey.valkey.metrics.podMonitor.enabled — the valkey chart's own default is on; written false when monitors are off, left absent otherwise so the @@ -1405,6 +1409,8 @@ connectivity release both read the resolved value. */ -}} {{- include "agent-platform.shape.derive" (dict "values" $v "path" (list "agentgateway" "monitoring" "enabled") "value" $monitors) -}} {{- include "agent-platform.shape.derive" (dict "values" $v "path" (list "mcp-kubernetes" "mcpKubernetes" "instrumentation" "serviceMonitor" "enabled") "value" $monitors) -}} {{- include "agent-platform.shape.derive" (dict "values" $v "path" (list "mcp-kubernetes" "grafanaDashboards" "enabled") "value" $monitors) -}} +{{- include "agent-platform.shape.derive" (dict "values" $v "path" (list "substrate" "metrics" "podMonitor" "enabled") "value" $monitors) -}} +{{- include "agent-platform.shape.derive" (dict "values" $v "path" (list "kagent" "controller" "metrics" "serviceMonitor" "enabled") "value" $monitors) -}} {{- include "agent-platform.shape.derive" (dict "values" $v "path" (list "kagent" "otel" "tracing" "enabled") "value" $monitors) -}} {{- include "agent-platform.shape.derive" (dict "values" $v "path" (list "kagent" "otel" "logging" "enabled") "value" $monitors) -}} {{- include "agent-platform.shape.derive" (dict "values" $v "path" (list "vm-manager" "serviceMonitor" "enabled") "value" $monitors) -}} diff --git a/helm/agent-platform/values.yaml b/helm/agent-platform/values.yaml index 442a6a4f..421dd3e1 100644 --- a/helm/agent-platform/values.yaml +++ b/helm/agent-platform/values.yaml @@ -559,7 +559,6 @@ components: # @schema additionalProperties: true - modelConfigs - oauth2ProxyIngress - remoteMcpServers - - serviceMonitor - uiRoute # The worker PodDisruptionBudget is the connectivity chart's (the kagent # chart has no budget template for the WorkerPool; #472). @@ -2383,7 +2382,27 @@ kagent: # @schema skipProperties: true; additionalProperties: true # the same reason. Both go back on together when upstream serves metrics # (an upstream item under giantswarm/giantswarm#37742). metrics: - enabled: false + enabled: true + bindAddress: ":8080" + secureServing: false + service: + port: 8080 + # The chart's own ServiceMonitor over that Service, from 1.0.2 on. + # `auto` follows the resolved + # global.observability.metrics.serviceMonitor.enabled. The endpoint + # follows secureServing: port name, scheme, bearer token and TLS all + # derive from it. With secureServing on, the scrape is authorized only + # once -metrics-reader is bound to the collector's + # ServiceAccount (controller.metrics.serviceMonitor + # .prometheusServiceAccount), which is why this platform scrapes plain + # HTTP on a port closed to everything but its own namespaces. + serviceMonitor: + enabled: auto # @schema type: [boolean, string]; enum: [auto, true, false] + interval: 60s + labels: + # Mimir routes a scrape by this label; a monitor without it writes + # to no tenant of the GS observability platform. + observability.giantswarm.io/tenant: giantswarm # X-Scope-OrgID routes OTLP traces and logs to the correct Loki/Tempo # tenant. The configmap template has no headers field so we inject via env. # The meta chart drops the OTEL_EXPORTER_OTLP_HEADERS entry when both OTel @@ -2616,23 +2635,6 @@ kagent: # @schema skipProperties: true; additionalProperties: true promptCaching: true cacheTTL: "5m" - # The controller metrics Service and ServiceMonitor the connectivity chart - # renders under the global.observability.metrics.serviceMonitor gate (which - # detects monitoring.coreos.com/v1); this is the object's own toggle - # underneath that gate. Off: the kagent line serves no Prometheus /metrics - # (controller.metrics above), so the monitor would scrape a refused port - # and the target would sit at up=0 — and nothing reads those metrics: the - # fleet's KagentControllerDown rule reads the Deployment's - # kube-state-metrics series. Back to true together with controller.metrics - # when upstream serves them. - serviceMonitor: - enabled: false - interval: 60s - # Extra labels on the ServiceMonitor metadata. The tenant label routes - # metrics to the correct Mimir instance in the GS Observability Platform. - labels: - observability.giantswarm.io/tenant: giantswarm - # OTel — traces and logs to the GS Observability Platform OTLP gateway. # X-Scope-OrgID header routes to the correct Loki/Tempo tenant. `auto` (the # default) follows the resolved global.observability.metrics.serviceMonitor @@ -4638,6 +4640,20 @@ substrate: # @schema skipProperties: true; additionalProperties: true # follow it. The dev channel needs no registry override. image: registry: gsoci.azurecr.io/giantswarm/substrate + # One PodMonitor per workload that serves metrics: ate-api-server, atelet, + # atenet-router, atenet-egress, k8s-credential-provider and ate-controller, + # each pinned to the port name its own pod declares. The fleet's metrics + # collector discovers monitors and nothing else, so without these the whole + # ate registry is served and thrown away. `auto` follows the resolved + # global.observability.metrics.serviceMonitor.enabled; the chart's own + # default is off. + metrics: + podMonitor: + enabled: auto # @schema type: [boolean, string]; enum: [auto, true, false] + labels: + # Mimir routes a scrape by this label; a monitor without it writes to + # no tenant of the GS observability platform. + observability.giantswarm.io/tenant: giantswarm postgres: # Where the control plane keeps its state. `auto` (the default) follows # the platform: with the platform's CNPG Cluster on (postgres.enabled) the diff --git a/tests/verify-cluster-shape.py b/tests/verify-cluster-shape.py index d623283d..953c8995 100755 --- a/tests/verify-cluster-shape.py +++ b/tests/verify-cluster-shape.py @@ -261,10 +261,9 @@ def check_shape(meta: str, connectivity: str, ci: list[str], name: str, served: expect(got == yes(monitors), f"{where} kserve-llmisvc-resources values {'.'.join(KSERVE_MONITOR)} = {got!r}, want {yes(monitors)!r}") # The connectivity chart on its own, same served groups: the matching objects. - # kagent.serviceMonitor is off by default (the kagent line serves no - # /metrics), so the one ServiceMonitor the gate can produce is switched on - # here to see the gate resolve; the default is verify-global's. - rendered = render(connectivity, [*PARENT_REF, *ON, *apis(served), "--set", "kagent.serviceMonitor.enabled=true"]) + # The controller's ServiceMonitor is the kagent chart's own now, so this + # chart renders none of its own; the gate's own resolution is verify-global's. + rendered = render(connectivity, [*PARENT_REF, *ON, *apis(served)]) objects = kinds(rendered) cnp, netpol = objects.get("CiliumNetworkPolicy", 0), objects.get("NetworkPolicy", 0) expect(not (cnp and netpol), f"{where} connectivity renders both network-policy flavors") @@ -277,8 +276,11 @@ def check_shape(meta: str, connectivity: str, ci: list[str], name: str, served: # every shape; the chart's ClusterPolicy is the agent-sandbox one. expect(not re.search(r"kagent-declarative-pod-security|kagent-srt-settings|kagent\.dev/v1alpha2", rendered), f"{where} connectivity render carries a kagent v1alpha2 Agent mutation or object") - expect((objects.get("ServiceMonitor", 0) > 0) == monitors, - f"{where} connectivity ServiceMonitor count={objects.get('ServiceMonitor', 0)}, monitoring served={monitors}") + # This chart renders no monitor of its own any more: the agentgateway + # PodMonitor is the packaging chart's (#616) and the kagent controller's + # ServiceMonitor is the kagent chart's (controller.metrics.serviceMonitor). + expect(objects.get("ServiceMonitor", 0) == 0, + f"{where} connectivity renders {objects.get('ServiceMonitor', 0)} ServiceMonitor(s); every monitor belongs to the component's own chart") expect(objects.get("VerticalPodAutoscaler", 0) == (1 if vpa else 0), f"{where} connectivity VerticalPodAutoscaler count={objects.get('VerticalPodAutoscaler', 0)}, autoscaling served={vpa}") diff --git a/tests/verify-kagent-wiring.py b/tests/verify-kagent-wiring.py index 253ba7a4..7c9ff9a3 100755 --- a/tests/verify-kagent-wiring.py +++ b/tests/verify-kagent-wiring.py @@ -317,14 +317,14 @@ def check_retired_keys(values: dict[str, list[str]], conn_kagent: str) -> None: fail(f"the 0.10 wrapper's bundled example agents are still forwarded: {', '.join(left)}; the line ships none") controller = "\n".join(values.get("controller", [])) if "METRICS_" in controller: - fail("METRICS_* env forwarded to the controller; the line serves no Prometheus /metrics") - if not re.search(r"^metrics:\nenabled: false$", controller, re.M): - fail("kagent.controller.metrics.enabled is not false; the upstream chart's knob renders a Service to a port the line's controller never serves") + fail("METRICS_* env forwarded to the controller; the chart renders both variables from controller.metrics, and a second entry of the same name renders the variable twice") + if not re.search(r"^metrics:\n(?:.*\n)*?enabled: true$", controller, re.M): + fail("kagent.controller.metrics.enabled is not true; the chart's knob is what renders the metrics Service, the METRICS_* env and the ServiceMonitor over them") if "skillsInitImage" in controller: fail("kagent.controller.skillsInitImage forwarded; the line has no skills-init image") - if not re.search(r"^ serviceMonitor:\n enabled: false$", conn_kagent, re.M): - fail("kagent.serviceMonitor.enabled is not false in the values the connectivity chart receives; the monitor would scrape a refused port") - print("ok: no example agent, no METRICS_* env, controller metrics and the ServiceMonitor off") + if re.search(r"^ serviceMonitor:$", conn_kagent, re.M): + fail("kagent.serviceMonitor is still forwarded to the connectivity chart; the controller's monitor is the kagent chart's own (controller.metrics.serviceMonitor)") + print("ok: no example agent, no METRICS_* env, the controller's metrics and its own ServiceMonitor on") def check_substrate_pins(docs) -> None: diff --git a/tests/verify-target.py b/tests/verify-target.py index dcdbed48..dd74dea7 100755 --- a/tests/verify-target.py +++ b/tests/verify-target.py @@ -195,6 +195,40 @@ def strip(render: str) -> str: # names a failed turn's class; GOLDEN_REF admits 2.0.0. Written on BOTH sides so # the range compares equal; dropped once GOLDEN_REF carries the floor. KLAUS_GATEWAY_FLOOR_HOLD = ["--set", "components.klaus-gateway.versionRange=>=3.3.0 <4.0.0"] +# giantswarm/giantswarm#36711: this tree turns the substrate chart's PodMonitors +# on, which GOLDEN_REF's defaults leave off and whose block it does not carry at +# all. The whole block is written on BOTH sides so the forwarded values compare +# equal; dropped once GOLDEN_REF carries it. +SUBSTRATE_PODMONITOR_HOLD = [ + "--set", "substrate.metrics.podMonitor.enabled=false", + "--set", "substrate.metrics.podMonitor.labels.observability\\.giantswarm\\.io/tenant=giantswarm", +] +# giantswarm/giantswarm#36711: this tree turns the kagent controller's metrics +# and the chart's own ServiceMonitor on, which GOLDEN_REF leaves off (the line +# served no endpoint before 1.0.2, and this chart rendered the monitor itself). +# Held equal on BOTH sides; dropped once GOLDEN_REF carries it. +KAGENT_SERVICEMONITOR_HOLD = [ + "--set", "kagent.controller.metrics.enabled=false", + "--set", "kagent.controller.metrics.bindAddress=:8080", + "--set", "kagent.controller.metrics.secureServing=false", + "--set", "kagent.controller.metrics.service.port=8080", + "--set", "kagent.controller.metrics.serviceMonitor.enabled=false", +] +# The same change retires kagent.serviceMonitor, this chart's own key for the +# monitor it no longer renders. GOLDEN_REF still forwards the block, so it is +# dropped from BOTH sides by name; the other serviceMonitor blocks (muster's, +# oauth2-proxy's) go with it, which narrows the comparison until GOLDEN_REF +# carries the change. +SERVICEMONITOR_KEY = re.compile(r"^(\s+)serviceMonitor:\s*$") + + +def hold_retired_servicemonitor(here: str, there: str) -> tuple: + """The two renders without any forwarded serviceMonitor mapping.""" + def strip(render: str) -> str: + return drop_mapping(render, SERVICEMONITOR_KEY) + if (h := strip(here)) != here or strip(there) != there: + print("note: #36711 hold — the forwarded serviceMonitor blocks are left out of the golden comparison") + return h, strip(there) AGENTGATEWAY_IMAGES_HOLD = [ "--set", "agentgateway.controller.image.repository=giantswarm/agentgateway-upstream/controller", "--set", "agentgateway.controller.image.tag=2.0.0", @@ -601,11 +635,11 @@ def check_golden(meta: str, connectivity: str) -> None: # carries the klaus-gateway no-op key removal (#636) since 4.62.0. golden_only = {meta: [], connectivity: []} shapes = [ - ("meta default", meta, [*hold_608, *METRIC_LABELS_HOLD, *hold_iv, *MUSTER_DASHBOARD_HOLD, *AGENTGATEWAY_MONITORING_HOLD, *MCP_KUBERNETES_MONITORING_HOLD, *VM_MANAGER_MONITOR_HOLD, *KSERVE_MONITOR_HOLD, *KAGENT_CONTROLLER_RESOURCES_HOLD, *KLAUS_GATEWAY_FLOOR_HOLD]), - ("meta ci + engine off", meta, ["-f", f"{meta}/ci/ci-values.yaml", *ENGINE_OFF, *hold_608, *METRIC_LABELS_HOLD, *hold_iv, *MUSTER_DASHBOARD_HOLD, *AGENTGATEWAY_MONITORING_HOLD, *MCP_KUBERNETES_MONITORING_HOLD, *VM_MANAGER_MONITOR_HOLD, *KSERVE_MONITOR_HOLD, *KAGENT_CONTROLLER_RESOURCES_HOLD, *KLAUS_GATEWAY_FLOOR_HOLD]), - ("connectivity default", connectivity, [*VM, *METRIC_LABELS_HOLD, *hold_iv, *AGENTGATEWAY_IMAGES_HOLD, *KAGENT_VPA_CAP_HOLD]), - ("connectivity full", connectivity, [*CONN_FULL, *METRIC_LABELS_HOLD, *hold_iv, *AGENTGATEWAY_IMAGES_HOLD, *KAGENT_VPA_CAP_HOLD]), - ("connectivity backstage", connectivity, [*CONN_BACKSTAGE, *METRIC_LABELS_HOLD, *hold_iv, *AGENTGATEWAY_IMAGES_HOLD, *KAGENT_VPA_CAP_HOLD]), + ("meta default", meta, [*hold_608, *METRIC_LABELS_HOLD, *hold_iv, *MUSTER_DASHBOARD_HOLD, *AGENTGATEWAY_MONITORING_HOLD, *SUBSTRATE_PODMONITOR_HOLD, *KAGENT_SERVICEMONITOR_HOLD, *MCP_KUBERNETES_MONITORING_HOLD, *VM_MANAGER_MONITOR_HOLD, *KSERVE_MONITOR_HOLD, *KAGENT_CONTROLLER_RESOURCES_HOLD, *KLAUS_GATEWAY_FLOOR_HOLD]), + ("meta ci + engine off", meta, ["-f", f"{meta}/ci/ci-values.yaml", *ENGINE_OFF, *hold_608, *METRIC_LABELS_HOLD, *hold_iv, *MUSTER_DASHBOARD_HOLD, *AGENTGATEWAY_MONITORING_HOLD, *SUBSTRATE_PODMONITOR_HOLD, *KAGENT_SERVICEMONITOR_HOLD, *MCP_KUBERNETES_MONITORING_HOLD, *VM_MANAGER_MONITOR_HOLD, *KSERVE_MONITOR_HOLD, *KAGENT_CONTROLLER_RESOURCES_HOLD, *KLAUS_GATEWAY_FLOOR_HOLD]), + ("connectivity default", connectivity, [*VM, *METRIC_LABELS_HOLD, *hold_iv, *AGENTGATEWAY_IMAGES_HOLD, *KAGENT_VPA_CAP_HOLD, *KAGENT_SERVICEMONITOR_HOLD]), + ("connectivity full", connectivity, [*CONN_FULL, *METRIC_LABELS_HOLD, *hold_iv, *AGENTGATEWAY_IMAGES_HOLD, *KAGENT_VPA_CAP_HOLD, *KAGENT_SERVICEMONITOR_HOLD]), + ("connectivity backstage", connectivity, [*CONN_BACKSTAGE, *METRIC_LABELS_HOLD, *hold_iv, *AGENTGATEWAY_IMAGES_HOLD, *KAGENT_VPA_CAP_HOLD, *KAGENT_SERVICEMONITOR_HOLD]), ] for label, chart, flags in shapes: here = helm(chart, flags) @@ -615,6 +649,7 @@ def check_golden(meta: str, connectivity: str) -> None: here, there = hold_dashboards(here, there, chart == meta) here, there = hold_opus_5_5_price(here, there) here, there = hold_dataplane_buffer(here, there) + here, there = hold_retired_servicemonitor(here, there) if chart == meta: here, there = drop_new_roster_entries(here, there) here, there = hold_llmd_only(here, there)