Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
4 changes: 4 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -42,11 +42,15 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0
- **`components.klaus-gateway.versionRange` admits the 2.x line (`>=1.20.0 <3.0.0`)** (giantswarm/klaus-gateway#319). klaus-gateway 2.0.0 is the Slack-only gateway: the web and CLI channels, the OpenAI-compatible `/v1` front door and the Klaus instance path are removed there; the values keys this chart still forwards for them (`cli`, `lifecycle`, `upstream`, `agentgateway`, `routing.defaultTTL`, `a2a.saToken`) are accepted by the 2.x chart as no-ops, so the installations roll to it unchanged. Those keys leave this chart in a later change, once the 2.x line has rolled.
- **The klaus-gateway ServiceMonitor carries the tenant label: `klausGateway.serviceMonitor`** (giantswarm/giantswarm#36711, giantswarm/klaus-gateway#316). The chart's monitor was on and carried the chart's own labels only, so Mimir routed its scrape to no tenant and every `klaus_gateway_*` series was lost on every installation (seen on gazelle, chart 1.19.1). `klausGateway.serviceMonitor.labels` sets `observability.giantswarm.io/tenant: giantswarm` and `klausGateway.serviceMonitor.enabled` is `auto`, resolved to the boolean the chart takes from the same answer as every other monitor of this chart. Pin first: `components.klaus-gateway.versionRange` is `>=1.20.0 <2.0.0` and `examples/customer-bom.yaml` pins `1.20.0`, the release that opens the labels key; an older chart's schema refuses it. `verify-klausgateway-otlp` asserts the resolution and the label; `verify-components-charts` names `1.20.0` as `UNRELEASED` until it exists.
- **The mcp-kubernetes chart's own ServiceMonitor and its three Grafana boards are on** (giantswarm/giantswarm#36711). The chart ships a ServiceMonitor over the dedicated metrics port and the boards `administrator`, `security` and `cluster-operator`, and this chart turned none of them on, so an installation that runs the Kubernetes MCP server collected nothing from it. `mcp-kubernetes.mcpKubernetes.instrumentation.serviceMonitor.enabled` and `mcp-kubernetes.grafanaDashboards.enabled` follow the resolved `global.observability.metrics.serviceMonitor.enabled` (`auto | true | false`), the monitor carries `observability.giantswarm.io/tenant`, and the boards land in `Shared Org / Agent Platform` beside the platform's own. The chart's `prometheusRules` stay off: its three alerts link to runbook pages that do not exist yet. The floor `1.1.1` already carries every key, so no range moves.
- **Substrate's six metrics endpoints are collected: `substrate.metrics.podMonitor`** (giantswarm/giantswarm#36711). ate-api-server, atelet, atenet-router, atenet-egress, k8s-credential-provider and ate-controller each serve metrics, the fleet's collector discovers monitors and nothing else, and the chart shipped no monitor, so the whole ate registry was served and thrown away: actor crashes, lifecycle and scheduler durations, checkpoint and restore phases, actor CPU and memory, image-cache hits, worker pool desired against ready, router route duration and parking, and ate-controller's `controller_runtime_*` and `workqueue_*`. The block follows the resolved `global.observability.metrics.serviceMonitor.enabled` (`auto | true | false`) and carries `observability.giantswarm.io/tenant`. The floor `1.0.3` already carries the monitors (giantswarm/substrate#50, upstream kagent-dev/substrate#46, open).
- **The kagent controller's metrics are collected: `kagent.controller.metrics`, with the kagent chart's own ServiceMonitor** (giantswarm/giantswarm#36711). The controller served no endpoint before the line's `1.0.2` (`metricsserver.Options{BindAddress: "0"}`, which no variable changed, so the `METRICS_BIND_ADDRESS` this chart set had no effect); from `1.0.2` on it reads `METRICS_BIND_ADDRESS` and `METRICS_SECURE`, and the controller-runtime families and the line's `kagent_grpc_server_requests_total` / `..._request_duration_seconds` reach a scrape. `metrics.enabled: true`, plain HTTP on `:8080` (`secureServing: false`: the secure endpoint authenticates the scraper with a TokenReview and authorizes it with a SubjectAccessReview, which the fleet's collector does not do), `serviceMonitor.enabled: auto` following the resolved `global.observability.metrics.serviceMonitor.enabled`, the tenant label on the monitor. The floor `1.0.3` carries the keys.
- **Every agent platform dashboard lands in one folder customers can reach: `Shared Org / Agent Platform`** (giantswarm/giantswarm#36711). The muster board (`Muster / MCP Gateway`) was loaded into the staff-only `Giant Swarm` organization, in a folder named after the component, so the people who run agents on the platform could not see it: `muster.observability.grafanaDashboard.folder` is `Agent Platform` and `.giantswarm.organization` is `Shared Org`. `Shared Org` is the organization every logged-in customer reaches as a Viewer and it carries the observability data of Giant Swarm managed components, which is where these boards belong; the `Giant Swarm` organization is staff-only. The boards that follow — the agentgateway gateway board, the platform's own overview — land in the same folder.
- **The agentgateway packaging chart's own monitors and dashboard are on: `agentgateway.monitoring.enabled`** (giantswarm/giantswarm#36711, giantswarm/agentgateway#60). It follows the resolved `global.observability.metrics.serviceMonitor.enabled` like every other monitor of this chart (`auto | true | false`), at the fleet's 60s interval and with the `observability.giantswarm.io/tenant` label on both monitors. It turns on three objects the platform did not have: the **controller** ServiceMonitor, which nothing scraped, and with it the only view of the control plane that serves the data planes their config; upstream's proxy PodMonitor; and upstream's own Grafana board — Overview, Requests, LLM, **MCP tool calls**, Latency, **XDS**, **Runtime** — vendored in the packaging chart since 2.0.0 and never rendered. The MCP and runtime series it reads were already collected and had no consumer. The board is upstream's to maintain, so a version bump carries its fixes; the boards the platform writes itself (LLM cost per agent and per person) ship in this chart's `dashboards/` directory.

### Removed

- **`kagent.serviceMonitor` and the connectivity chart's kagent controller metrics Service and ServiceMonitor** (`templates/kagent/servicemonitor.yaml`; giantswarm/giantswarm#36711). Both were a workaround for a kagent chart that exposed only the API port; the line's chart ships the metrics Service and its ServiceMonitor since `1.0.2` (`kagent.controller.metrics`), and keeping ours would have scraped one endpoint twice. An installation that set `kagent.serviceMonitor.*` drops the key.

- **The Qwen3 small presets `qwen3-4b-instruct`, `qwen3-8b-fp8` and `qwen3-14b`** (giantswarm/agent-platform#591). They pinned 2025 checkpoints whose successors are smaller or stronger under the same licence and are replaced by the 24 GB line-up above (`qwen3-5-4b` for the 4B, `qwen3-5-9b-fp8` or `gemma-4-12b` for the 8B, `gpt-oss-20b` or `gemma-4-12b` for the 14B — which, at 28 GiB of BF16 weights, never fit a 24 GB card and was tuned for a 128 GB node). The upgrade removes their ConfigMaps; a model already served from one keeps serving (the `LLMInferenceService` is model-manager's object), and an installation that wants one back carries its file under `modelServing.presets` (UPGRADE.md).
- **The presets no single card holds: `qwen3-5-27b`, `qwen3-coder-next` and `qwen3-5-35b-a3b`** (giantswarm/agent-platform#591): 54 GiB BF16, 80 GiB FP8 and 70 GiB BF16 of weights exceed the 48 GB GPU they would target; the presets above are their model-image successors of the same shape. The shipped set goes from twelve to eleven presets; `tests/verify-gpu-pool.py` leaves each side's own presets out of the golden comparison until the golden carries this change.

Expand Down
22 changes: 7 additions & 15 deletions Makefile.custom.mk
Original file line number Diff line number Diff line change
Expand Up @@ -277,21 +277,13 @@ verify-global: ## Assert the global.* contract behaviors (derived hostnames, gat
if grep -q "$$pattern" /tmp/vg-mon.out; then echo "FAIL: monitor-gated render still contains $$pattern"; exit 1; fi; \
done
@echo "ok: monitor gate"
@echo "--> the default render keeps the CNPG PodMonitor (fleet behavior) and renders NO kagent ServiceMonitor or metrics Service: the kagent line serves no /metrics (kagent.serviceMonitor.enabled: false)"
@helm template t $(CONNECTIVITY_DIR) $(VM) --set components.kagent.enabled=true --set postgres.enabled=true >/tmp/vg-mon-default.out 2>&1 || { cat /tmp/vg-mon-default.out; exit 1; }
@if grep -q 'kind: ServiceMonitor' /tmp/vg-mon-default.out; then echo "FAIL: the default render carries a kagent ServiceMonitor; the line's controller serves no /metrics and the monitor would sit at up=0"; exit 1; fi
@if grep -q 'kagent-controller-metrics' /tmp/vg-mon-default.out; then echo "FAIL: the default render carries the kagent controller metrics Service; nothing listens behind it on the line"; exit 1; fi
@grep -q 'enablePodMonitor: true' /tmp/vg-mon-default.out || { echo "FAIL: default render lost the CNPG PodMonitor"; exit 1; }
@grep -q 'helm.sh/resource-policy: keep' /tmp/vg-mon-default.out || { echo "FAIL: the CNPG Cluster lost helm.sh/resource-policy: keep"; exit 1; }
@echo "ok: default: no kagent monitor, CNPG PodMonitor + keep"
@echo "--> kagent.serviceMonitor.enabled=true (for when upstream serves metrics) renders the Service and the ServiceMonitor under the global gate"
@helm template t $(CONNECTIVITY_DIR) $(VM) --set components.kagent.enabled=true --set postgres.enabled=true --set kagent.serviceMonitor.enabled=true >/tmp/vg-mon-on.out 2>&1 || { cat /tmp/vg-mon-on.out; exit 1; }
@grep -q 'kind: ServiceMonitor' /tmp/vg-mon-on.out || { echo "FAIL: kagent.serviceMonitor.enabled=true renders no ServiceMonitor"; exit 1; }
@grep -q 'observability.giantswarm.io/tenant: giantswarm' /tmp/vg-mon-on.out || { echo "FAIL: the kagent ServiceMonitor lost the tenant label"; exit 1; }
@helm template t $(CONNECTIVITY_DIR) $(VM) --set components.kagent.enabled=true --set kagent.serviceMonitor.enabled=true --set global.observability.metrics.serviceMonitor.enabled=false 2>/dev/null | grep -q 'kind: ServiceMonitor' && { echo "FAIL: the global monitor gate no longer holds the kagent ServiceMonitor back"; exit 1; } || true
@echo "ok: the toggle under the global gate"
@echo "--> the kagent controller metrics Service selects kagent's own release instance (the pods' label), the ServiceMonitor this chart's Service"
@python3 -c 'import re,sys; docs=open("/tmp/vg-mon-on.out").read().split("\n---\n"); svc=[d for d in docs if "\nkind: Service\n" in d and re.search(r"^ name: t-kagent-controller-metrics$$", d, re.M)]; sys.exit("FAIL: the kagent controller metrics Service did not render") if len(svc)!=1 else None; sel=svc[0][svc[0].index(" selector:"):]; sys.exit("FAIL: the metrics Service does not select app.kubernetes.io/instance: kagent (the kagent release name the meta chart fixes):\n"+sel) if not re.search(r"^ app.kubernetes.io/instance: kagent$$", sel, re.M) else None; sys.exit("FAIL: the metrics Service selects this release (t) — under the meta chart that matches no pod (#305)") if re.search(r"^ app.kubernetes.io/instance: \"?t\"?$$", sel, re.M) else None; sm=[d for d in docs if "kind: ServiceMonitor" in d and "-kagent-controller\n" in d]; sys.exit("FAIL: the kagent ServiceMonitor did not render") if len(sm)!=1 else None; sys.exit("FAIL: the ServiceMonitor must select this chart\x27s Service (instance t)") if " app.kubernetes.io/instance: \"t\"" not in sm[0] else None; print("ok: metrics Service selects instance kagent; the ServiceMonitor selects this release\x27s Service")'
@echo "--> the default render keeps the CNPG PodMonitor (fleet behavior) and renders NO kagent ServiceMonitor or metrics Service: both are the kagent chart's own (controller.metrics)"
@helm template t $(CONNECTIVITY_DIR) $(VM) --set components.kagent.enabled=true --set postgres.enabled=true >/tmp/vg-mon-on.out 2>&1 || { cat /tmp/vg-mon-on.out; exit 1; }
@if grep -q 'kind: ServiceMonitor' /tmp/vg-mon-on.out; then echo "FAIL: this chart renders a ServiceMonitor; every monitor belongs to the component's own chart (the kagent controller's to controller.metrics.serviceMonitor, the agentgateway data plane's to the packaging chart)"; exit 1; fi
@if grep -q 'kagent-controller-metrics' /tmp/vg-mon-on.out; then echo "FAIL: this chart renders the kagent controller metrics Service; the kagent chart renders it under controller.metrics.enabled"; exit 1; fi
@grep -q 'enablePodMonitor: true' /tmp/vg-mon-on.out || { echo "FAIL: default render lost the CNPG PodMonitor"; exit 1; }
@grep -q 'helm.sh/resource-policy: keep' /tmp/vg-mon-on.out || { echo "FAIL: the CNPG Cluster lost helm.sh/resource-policy: keep"; exit 1; }
@echo "ok: no monitor of this chart's own, CNPG PodMonitor + keep"
@echo "--> no kagent-targeting selector, Service name or hostname in this chart derives from .Release.Name (the standalone umbrella's one-release assumption)"
@if grep -nE 'fullnameOverride \| default \.Release\.Name|fullnameOverride" \| default \(printf "%s-oauth2-proxy" \.Release\.Name' $(CONNECTIVITY_DIR)/templates/kagent/*.yaml; then echo "FAIL: a kagent template falls back to .Release.Name for a kagent-chart object; use agent-platform.kagent.fullname / agent-platform.kagent.releaseName"; exit 1; else echo "ok: kagent templates derive kagent names from the kagent helpers"; fi
@echo "--> the CNPG CiliumNetworkPolicy renders only when postgres.enabled"
Expand Down
11 changes: 2 additions & 9 deletions helm/agent-platform-connectivity/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -1089,12 +1089,8 @@ The kagent block is open in the schema, so the template refuses a key under `kag
| kagent.controller.skillsInitImage.repository | string | `"kagent-skills-init"` | |
| kagent.controller.auth.mode | string | `"trusted-proxy"` | |
| kagent.controller.auth.userIdClaim | string | `"email"` | |
| kagent.controller.env[0].name | string | `"METRICS_BIND_ADDRESS"` | |
| kagent.controller.env[0].value | string | `":8080"` | |
| kagent.controller.env[1].name | string | `"METRICS_SECURE"` | |
| kagent.controller.env[1].value | string | `"false"` | |
| kagent.controller.env[2].name | string | `"OTEL_EXPORTER_OTLP_HEADERS"` | |
| kagent.controller.env[2].value | string | `"X-Scope-OrgID=giantswarm"` | |
| kagent.controller.env[0].name | string | `"OTEL_EXPORTER_OTLP_HEADERS"` | |
| kagent.controller.env[0].value | string | `"X-Scope-OrgID=giantswarm"` | |
| kagent.controller.vpa.enabled | string | `"auto"` | |
| kagent.controller.vpa.updateMode | string | `"InPlaceOrRecreate"` | |
| kagent.controller.vpa.controlledValues | string | `"RequestsOnly"` | |
Expand All @@ -1120,9 +1116,6 @@ The kagent block is open in the schema, so the template refuses a key under `kag
| kagent.providers.anthropic.apiKeySecretRef | string | `"kagent-anthropic"` | |
| kagent.providers.anthropic.apiKeySecretKey | string | `"ANTHROPIC_API_KEY"` | |
| kagent.providers.anthropic.apiKey | string | `""` | |
| kagent.serviceMonitor.enabled | bool | `false` | |
| kagent.serviceMonitor.interval | string | `"60s"` | |
| kagent.serviceMonitor.labels."observability.giantswarm.io/tenant" | string | `"giantswarm"` | |
| kagent.otel.tracing.enabled | string | `"auto"` | |
| kagent.otel.tracing.exporter.otlp.endpoint | string | `"http://otlp-gateway.kube-system.svc:4317"` | |
| kagent.otel.tracing.exporter.otlp.protocol | string | `"grpc"` | |
Expand Down

This file was deleted.

Loading
Loading