Skip to content

perf(coding/tui): freeze the streaming live blocks so a frame costs the tail - #1

Merged
rsbin1178 merged 4 commits into
mainfrom
feat/streaming-frame-projection
Oct 7, 2026
Merged

rsbin1178 merged 4 commits into
mainfrom
feat/streaming-frame-projection

Conversation

@rsbin1178

Copy link
Copy Markdown
Owner

Phase 1 of the TUI CPU work. It removes the dominant per-frame cost: the two live blocks
(the streaming answer draft and the streaming reasoning section) were re-rendered in
full on every frame, which is O(body) per frame and O(body^2) over one message.

Commits

  • 5b55f9c assemble a streaming live body from a frozen prefix
  • fdfda67 freeze the live Thinking rows and stop hashing them per frame
  • 8201df5 keep link references on one side of a frozen Markdown prefix

Measured

Real entry point tui.Run in a PTY (120x40), 627-message session, 1500 deltas at 150/s
with interleaved tool round trips, external 100 ms CPU sampling; HEAD abd4e1e vs this
branch, interleaved rounds so machine drift cannot favour one build.

workload child CPU interval peak alloc GC
reasoning-mix 37.82% -> 20.64% (-45%) 75.5 / 81.8 / 84.9 -> 44.6 / 34.8 / 37.4 2.94 -> 1.67 GB 247 -> 149
answer-heavy 54.70% -> 23.97% (-56%) 115.0 / 116.1 / 110.1 -> 45.5 / 46.7 / 46.3 4.66 -> 2.98 GB 336 -> 214

One frame of a growing body (-benchmem): Markdown 32 KB 22.5 ms -> 0.20 ms,
Thinking 32 KB 1.44 ms -> 38 us. Idle is unchanged (~2-3% of one core, no periodic spike).

Correctness

A prefix may be frozen only where rendering it alone reproduces the whole document's rows
byte for byte. Pinned by TestLiveMarkdownMatchesWholeRender and its fuzz target, by
TestLiveThinkingMatchesSettledSection and its fuzz target, and by a streaming corpus
that includes the link-reference cases (a usage before its definition, a definition whose
label spans lines, brackets in fenced code) plus mutation checks that restore the old rule.

Notes

No visible-output change, no frame-rate change, no fork change. The pinned Bubble Tea
fork's 60 Hz flush loop is the idle floor and is untouched.

A streaming answer draft is re-rendered in full on every frame: the Markdown
cache is keyed by content, so a growing body misses every time and pays a
whole-document glamour render. That is O(body) per frame and O(body^2) over one
message, and a CPU profile of the real TUI under a paced stream attributes 36%
of all samples to it (`renderManagedTranscript` -> `transcriptStore.buildTail`
-> `renderMarkdownBlockBody` -> `markdownRenderer.renderLive` ->
`glamour.TermRenderer.Render`).

`markdownRenderer.renderLive` now keeps the rendered rows of a frozen prefix and
renders only the newly frozen region plus the live tail. A boundary may be
frozen only where rendering the prefix alone reproduces the whole document's
rows byte for byte: a blank line outside a fenced code block whose preceding
block is a plain top-level paragraph, not followed by a definition-list marker,
in a prefix that defines no link reference and is valid UTF-8. The seam is
reconstructed by rendering the prefix's last block together with the tail and
dropping that block's own rows.

Measured on the real TUI entry point in a PTY, 3 runs each, 627-message session
with 1500 deltas at 150/s, before -> after:

  answer-heavy     54.1-54.6% -> 24.6-24.8% CPU (-54% rel), peak 103.8-115.4 -> 47.6-57.1
  reasoning-mix    37.8-37.9% -> 22.6-24.2% CPU (-39% rel), peak  77.7-86.5  -> 47.2-55.1
  allocations      4.44 GB -> 2.77 GB, GC 323 -> 204 (answer-heavy)

`BenchmarkLiveMarkdownFrame` on one growing live body: 32 KB 22.5 ms -> 0.20 ms
per frame.

Equivalence is pinned by a streaming test over a corpus of document shapes, by
invalidation tests for a replaced body and for a geometry change, and by a fuzz
target comparing the assembled render with a whole-document render.
… per frame

Three more pieces of per-frame work that grew with the streaming body, found by
re-profiling after the live-Markdown fix. On the real TUI entry point in a PTY,
3 interleaved rounds of each build, reasoning-mix / answer-heavy child CPU:

  HEAD              37.95% / 54.94%
  live-Markdown fix 23.08% / 25.49%
  this change       20.92% / 24.33%   (-44.9% / -55.7% vs HEAD)

- `renderThinkingBlock` now assembles the live thought from the rows frozen at
  its last line break and renders only the tail. The split is exact because
  `ansi.Wrap(body) == ansi.Wrap(body[:k]) + ansi.Wrap(body[k:])` when `body[:k]`
  ends on a line break, and because lipgloss pads every line of a block to its
  widest line. One frame of a growing thought: 193 us -> 6.4 us at 4 KB,
  1.44 ms -> 36.6 us at 32 KB, 5.68 ms -> 134 us at 128 KB.

- `sanitizeToolText` returns the caller's own string when it carries no control
  sequence, no control character and no surrounding space, which is what a
  streamed thought is. It was 52% of the live Thinking frame's allocation.

- `transcriptStore.updateRevision` no longer hashes the live record's text on
  every frame. Only a search reads that text, so the digest moved behind
  `fingerprint()`, which a search scan calls; the frame's revision keeps the
  geometry and the record heights. This frame was 9.7% of samples.

A profile after the change no longer shows `renderThinkingBlock` or
`updateRevision` among the application frames; what leads is scheduler handoff
and the renderer's PTY writes.

Each change is pinned by tests that fail without it: a streaming equivalence
test and an fuzz target for the live Thinking path, a whole-thought-per-frame
mutation that breaks the bounded-work assertion, a frozen-prefix invalidation
test, an allocation assertion for clean prose, and a revision test asserting the
frame does not read the live rows while a search fingerprint does.
…n prefix

goldmark resolves link references over the whole document, so a bracket on
either side of a freeze boundary can change the other side's rows. The guard was
a per-line `[...]:` pattern applied only in one direction, which left two holes:

- a usage before its definition froze into the prefix, so the frame rendered
  `[later]` literally while a whole render resolved it:
    "Intro.\n\nUse [later] first.\n\n[later]: /url\n\nTail.\n"
- a definition whose label spans lines is valid CommonMark but invisible to a
  per-line test, so the definition sat in the frozen prefix and a usage in the
  tail lost its scope:
    "A.\n\n[x\ny]: https://e.com \"t\"\n\nB.\n\nC uses [x y] here.\n"

The guard is now a single rule: a boundary is freezable only if the frozen
prefix carries no `[` at all. That covers both directions and the multiline
label, and makes the per-line definition pattern unnecessary. Fenced code is
exempt because its brackets are code; autolinks, linkified bare URLs and inline
links resolve without definitions and stay freezable. A body that carries a
bracket falls back to a whole render, which is correct but not incremental.

Both documents now render byte-identically to a whole render. Four corpus cases
(`forward ref`, `multi label`, `code bracket`, `shortcut refs`) are replayed by
the streaming equivalence test and the fuzz target, and a focused boundary test
pins the freeze point; all of them fail under a mutation that restores the old
per-line rule.
CI lints the pull request's own diff (`--new-from-rev=HEAD~`), and this branch was
never pushed, so its changes had never been linted. Two suppressions and four
mechanical fixes:

- `markdownFrozenRows` and `liveMarkdownBoundary` are boundary scanners whose
  branches are the conditions of one ordered pass, so they carry `//nolint:gocyclo`
  with the reason the repo uses for the same shape elsewhere.
- the row-count to hash-bucket conversion in `updateRevision` is `//nolint:gosec`:
  a row count is never negative.
- `markdownRowsAfter` and `markdownIndented` use integer ranges.
- `TestSanitizeToolTextLeavesCleanProseUntouched` is `//nolint:paralleltest`:
  AllocsPerRun measures the whole process and refuses a parallel test.
- the two `len` comparisons use `assert.Len`, and the whitespace wsl_v5 asks for
  before a `for`, a `return`, an assignment or a fuzz seed was added.

No behavior change: build, vet, the TUI tests, the race run and the PTY suite pass,
and `golangci-lint --new-from-rev=abd4e1e` reports no issues on the diff.
@rsbin1178
rsbin1178 merged commit fade106 into main Oct 7, 2026
4 checks passed
@rsbin1178
rsbin1178 deleted the feat/streaming-frame-projection branch October 8, 2026 01:43
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant