Skip to content

benchmarking/analysis: aggregate suspend/resume phase breakdown log r… - #2009

Open
Lucky Abolorunke (Oneimu) wants to merge 4 commits into
agent-substrate:mainfrom
Oneimu:phase-logs-analysis
Open

Lucky Abolorunke (Oneimu) wants to merge 4 commits into
agent-substrate:mainfrom
Oneimu:phase-logs-analysis

Conversation

@Oneimu

@Oneimu Lucky Abolorunke (Oneimu) commented Sep 30, 2026 •

Copy link
Copy Markdown
Collaborator

benchmarking/analysis: aggregate suspend/resume phase breakdown log records

Why

#1941 makes atelet and ateom-microvm write one structured Checkpoint timing breakdown / Restore timing breakdown record per operation per layer, so that SuspendActor / ResumeActor latency can be attributed to the phases inside it. A log line is not a benchmark result, though: something has to get those records off the cluster and turn them into percentiles. Today nothing in benchmarking/ reads node-side logs — the runner keeps its own logs.txt/traces.txt and the orchestrator tails the locust job. This PR adds that consumer, so a run can answer "where did the 6 s go?" from data rather than from one anecdote.

What

Everything lives under benchmarking/analysis/; no production code changes, and no new log output — this only reads what the binaries already write.

  • collect_logs.sh — runs kubectl logs for every atelet pod (ate-system) and every worker pod (benchmark-workloads) for the run's window and writes them to --dest, next to the locust artifacts. Two records per lifecycle operation is the entire volume it collects.
  • phase_report.py — parses the records from those dumps (tolerant of kubectl line prefixes and unrelated lines) and prints:
    • Phase percentiles per layer, operation, scope and phase (count, p50, p90, p95, max), with the atelet ateom_checkpoint / ateom_restore bucket split further by the ateom rows.
    • An unattributed row for the two checkpoint layers: the total minus what the logged phases account for, counting the three concurrent captures once. It is the time the instrumentation does not yet name.
    • Waterfalls of the slowest N operations, with the ateom record nested under the atelet phase it decomposes and the gap between the two (RPC and queueing between the layers).
    • Failed operations are excluded from the percentiles via the error.type marker on the record.
    • --csv writes phase_percentiles.csv for run-over-run comparison.
  • test_phase_report.py — 10 unittest tests covering parsing of both layers, nanosecond timestamps, prefixed lines, failure exclusion, the residual arithmetic and the waterfall join.
  • README.md — the record shapes, how to collect and aggregate, how to read each section and its concurrency caveats. benchmarking/README.md gains a short section pointing at it.
  • docs/dev/suspend-resume-phase-breakdown.md — an end-to-end runbook for a dev cluster: deploy the benchmark workloads on a micro-VM pool, drive glutton suspend/resume load with a sized working set, collect (kubectl or Cloud Logging), report, read the tables, and the failure modes that come up in practice. Linked from CONTRIBUTING.md beside the local micro-VM guide.

Scope and follow-ups

  • Standalone tool. Nothing invokes it yet. benchmarking/automation/orchestrator.py deletes the workload and ate-system pods right after each test, so the collection has to run before that teardown. The follow-up is to call collect_logs.sh and phase_report.py from run_test() before the teardown and write the phase percentiles into the run's --dest as stats.jsonl rows (the suite's existing result format: timestamp, tag, test_name, metric, measurements), so they travel with the rest of the run's artifacts.

Follow-up (#2148): the automated runs will consume this through the locust runner, which reads the pod logs itself before the orchestrator's teardown and appends the percentiles to stats.jsonl — that lands in a separate PR stacked on this one (branch phase-logs-runner; PR link to follow), so this PR stays the tool and its docs.

@bowei Bowei Du (bowei) added area/benchmarking area/observability kind/feature An enhancement / feature request or implementation labels Sep 30, 2026
Comment thread benchmarking/analysis/collect_logs.sh
Comment thread benchmarking/analysis/phase_report.py Outdated
Comment thread benchmarking/analysis/phase_report.py
Comment thread benchmarking/analysis/phase_report.py Outdated
Comment thread docs/dev/suspend-resume-phase-breakdown.md Outdated
Comment thread benchmarking/analysis/phase_report.py Outdated
Comment thread benchmarking/analysis/README.md
Comment thread benchmarking/analysis/README.md Outdated

@cheftako Walter Fender (cheftako) left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Looks good. 1 suggestion for resilience

Comment thread benchmarking/analysis/phase_report.py
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

area/benchmarking area/observability kind/feature An enhancement / feature request or implementation

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants