Skip to content

benchmarking: append the suspend/resume phase percentiles to stats.jsonl from the runner - #2148

Draft
Lucky Abolorunke (Oneimu) wants to merge 4 commits into
agent-substrate:mainfrom
Oneimu:phase-logs-runner
Draft

Lucky Abolorunke (Oneimu) wants to merge 4 commits into
agent-substrate:mainfrom
Oneimu:phase-logs-runner

Conversation

@Oneimu

Copy link
Copy Markdown
Collaborator

benchmarking: append the suspend/resume phase percentiles to stats.jsonl from the runner

Stacked on #2009 (base branch phase-logs-analysis until it merges).

Why

#2009 adds the tool that turns the Checkpoint/Restore timing breakdown records into percentiles, but nothing in the automated runs calls it: the orchestrator deletes the worker and ate-system pods as soon as a test's runner exits, so the records never reach stats.jsonl or anything that reads it. The runner is the one process that holds the run's identity, owns stats.jsonl, and is still alive while the pods are — so it does the collection itself, the way it already collects cluster facts.

What

  • locust/phase_breakdown.py (new): after locust finishes, lists the app=atelet pods in ate-system and the ate.dev/worker-pool pods in benchmark-workloads through the Kubernetes client, reads their logs since the run started (a restarted container's previous log too), parses the records with phase_report, and appends to stats.jsonl:
    • one row per layer / op / sandbox class / kind / scope / phase — metric: phase_<layer>_<op>_<class>_<kind>_<scope>_<phase>, measurements: those dimensions as fields plus count, p50_ms, p90_ms, p95_ms, max_ms, all strings like the runner's other rows;
    • one phase_breakdown_summary row: pods read and failed, record count, and the atelet checkpoint/restore record counts next to locust's SuspendActor/ResumeActor request counts. Fewer records than requests means a node log rotated during the run; the runner's log says so.
      Additive like the cluster facts: a failure here never costs the locust rows.
  • locust/runner.py: --phase-breakdown (default on, --no-phase-breakdown to skip), called after the cluster-facts block.
  • automation/manifests/runner-job.yaml.tmpl: benchmark-runner gains get/list on pods and get on pods/log in both namespaces.
  • locust/Dockerfile: copies phase_breakdown.py and analysis/phase_report.py into the image.
  • analysis/phase_report.py: parse_lines() and stats_rows() for the above (the CLI is unchanged).
  • Docs: analysis/README.md "Automated runs", the flag in benchmarking/README.md, a pointer in the runbook.

Testing

  • python3 -m unittest discover -s benchmarking/locust/unit_tests — 18 tests (4 new, mocked CoreV1Api: rows appended in shape, rotation warning, unreadable pod skipped, restarted container read twice, empty run still writes the summary, failed requests don't trigger the warning).
  • python3 -m unittest discover -s benchmarking/analysis — 22 tests.
  • Live: against a GKE cluster with 3 workers, after a 10-cycle glutton run, append_phase_breakdown read the six pods and wrote 46 phase rows plus the summary (12 atelet checkpoints vs 10 locust suspends + 2 shutdown suspends). The template's RBAC was applied and a log read impersonating benchmark-runner succeeded in both namespaces.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant