Skip to content

nighthawk-ingress: record router CPU and memory per stage - #41

Draft
ygao-g wants to merge 2 commits into
mainfrom
bench-resources
Draft

ygao-g wants to merge 2 commits into
mainfrom
bench-resources

Conversation

@ygao-g

@ygao-g ygao-g commented Oct 1, 2026

Copy link
Copy Markdown
Owner

Two commits add CPU and memory observability to the nighthawk-ingress router-capacity benchmark. A run today records how many rps the router held, but not which container saturated or what the actors and workers cost. Tracking PR; not for merge yet.

nighthawk-ingress: fetch per-stage CPU and memory for a finished run. fetch_run_resources.py reads a finished run's stats.jsonl and queries Managed Prometheus for actor CPU and memory and the router request rate, aligned to the run's stages. --system-metrics adds kubernetes.io/container/* where the cluster exports it. The GCP project and cluster label come from --project and --cluster, or PROJECT_ID and CLUSTER_NAME from the environment. The run's atelet pods are matched by the commit sha in the tag, or named with --atelet.

nighthawk-ingress: sample router CPU and memory during a run. runner.py --sample-resources polls kubelet cAdvisor from inside the runner pod for the router, runner and worker containers, and snapshots Envoy admin stats before and after the session. It writes resources.jsonl, resources-by-stage.json and envoy-stats.jsonl next to stats.jsonl. run-dev.sh --sample-resources passes the flag through. Off by default, with no change to the orchestrator or the Prow template. It needs resource-sampler-rbac.yaml applied by hand; nodes/proxy get also authorizes kubelet exec, so benchmark clusters only.

Validated on a dev GKE cluster: Envoy used 1.80 mean / 1.99 max of a 2-core pin in the 60 s testing stage at 10304 rps, 14% CFS-throttled. Known limit: cAdvisor refreshes every 10 to 20 s, so 10 s adjusting stages read low; the testing stage is reliable. 49 stdlib unittest tests, no new dependencies.

🤖 This PR was developed with AI assistance. I have reviewed and tested all changes.

ygao-g added 2 commits October 1, 2026 07:37
fetch_run_resources.py reads a run's stats.jsonl and queries Managed
Prometheus for the run's ate.actor.stats.* and router request rate,
picking the run's atelet and router pods by name because two installs
share one cluster label. It writes resources.jsonl and
resources-by-stage.json. --system-metrics adds the GKE container
metrics, which the benchmark cluster does not export yet.
Runs record no CPU or memory for the router, the runner, or the worker
pods, so a capacity result cannot say which container saturated. With
--sample-resources, the runner now polls kubelet cAdvisor for those
containers and snapshots the router's Envoy admin stats before and after
the session. Both outputs align to the stats.jsonl stages.

The flag is off by default. It needs the RBAC in
resource-sampler-rbac.yaml, which the orchestrator does not apply.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant