Repository navigation
Conversation
fetch_run_resources.py reads a run's stats.jsonl and queries Managed Prometheus for the run's ate.actor.stats.* and router request rate, picking the run's atelet and router pods by name because two installs share one cluster label. It writes resources.jsonl and resources-by-stage.json. --system-metrics adds the GKE container metrics, which the benchmark cluster does not export yet.
Runs record no CPU or memory for the router, the runner, or the worker pods, so a capacity result cannot say which container saturated. With --sample-resources, the runner now polls kubelet cAdvisor for those containers and snapshots the router's Envoy admin stats before and after the session. Both outputs align to the stats.jsonl stages. The flag is off by default. It needs the RBAC in resource-sampler-rbac.yaml, which the orchestrator does not apply.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Two commits add CPU and memory observability to the
nighthawk-ingressrouter-capacity benchmark. A run today records how many rps the router held, but not which container saturated or what the actors and workers cost. Tracking PR; not for merge yet.nighthawk-ingress: fetch per-stage CPU and memory for a finished run.
fetch_run_resources.pyreads a finished run'sstats.jsonland queries Managed Prometheus for actor CPU and memory and the router request rate, aligned to the run's stages.--system-metricsaddskubernetes.io/container/*where the cluster exports it. The GCP project and cluster label come from--projectand--cluster, orPROJECT_IDandCLUSTER_NAMEfrom the environment. The run's atelet pods are matched by the commit sha in the tag, or named with--atelet.nighthawk-ingress: sample router CPU and memory during a run.
runner.py --sample-resourcespolls kubelet cAdvisor from inside the runner pod for the router, runner and worker containers, and snapshots Envoy admin stats before and after the session. It writesresources.jsonl,resources-by-stage.jsonandenvoy-stats.jsonlnext tostats.jsonl.run-dev.sh --sample-resourcespasses the flag through. Off by default, with no change to the orchestrator or the Prow template. It needsresource-sampler-rbac.yamlapplied by hand;nodes/proxy getalso authorizes kubelet exec, so benchmark clusters only.Validated on a dev GKE cluster: Envoy used 1.80 mean / 1.99 max of a 2-core pin in the 60 s testing stage at 10304 rps, 14% CFS-throttled. Known limit: cAdvisor refreshes every 10 to 20 s, so 10 s adjusting stages read low; the testing stage is reliable. 49 stdlib
unittesttests, no new dependencies.🤖 This PR was developed with AI assistance. I have reviewed and tested all changes.