Repository navigation
perf(coding/tui): freeze the streaming live blocks so a frame costs the tail - #1
Merged
Merged
Conversation
A streaming answer draft is re-rendered in full on every frame: the Markdown cache is keyed by content, so a growing body misses every time and pays a whole-document glamour render. That is O(body) per frame and O(body^2) over one message, and a CPU profile of the real TUI under a paced stream attributes 36% of all samples to it (`renderManagedTranscript` -> `transcriptStore.buildTail` -> `renderMarkdownBlockBody` -> `markdownRenderer.renderLive` -> `glamour.TermRenderer.Render`). `markdownRenderer.renderLive` now keeps the rendered rows of a frozen prefix and renders only the newly frozen region plus the live tail. A boundary may be frozen only where rendering the prefix alone reproduces the whole document's rows byte for byte: a blank line outside a fenced code block whose preceding block is a plain top-level paragraph, not followed by a definition-list marker, in a prefix that defines no link reference and is valid UTF-8. The seam is reconstructed by rendering the prefix's last block together with the tail and dropping that block's own rows. Measured on the real TUI entry point in a PTY, 3 runs each, 627-message session with 1500 deltas at 150/s, before -> after: answer-heavy 54.1-54.6% -> 24.6-24.8% CPU (-54% rel), peak 103.8-115.4 -> 47.6-57.1 reasoning-mix 37.8-37.9% -> 22.6-24.2% CPU (-39% rel), peak 77.7-86.5 -> 47.2-55.1 allocations 4.44 GB -> 2.77 GB, GC 323 -> 204 (answer-heavy) `BenchmarkLiveMarkdownFrame` on one growing live body: 32 KB 22.5 ms -> 0.20 ms per frame. Equivalence is pinned by a streaming test over a corpus of document shapes, by invalidation tests for a replaced body and for a geometry change, and by a fuzz target comparing the assembled render with a whole-document render.
… per frame Three more pieces of per-frame work that grew with the streaming body, found by re-profiling after the live-Markdown fix. On the real TUI entry point in a PTY, 3 interleaved rounds of each build, reasoning-mix / answer-heavy child CPU: HEAD 37.95% / 54.94% live-Markdown fix 23.08% / 25.49% this change 20.92% / 24.33% (-44.9% / -55.7% vs HEAD) - `renderThinkingBlock` now assembles the live thought from the rows frozen at its last line break and renders only the tail. The split is exact because `ansi.Wrap(body) == ansi.Wrap(body[:k]) + ansi.Wrap(body[k:])` when `body[:k]` ends on a line break, and because lipgloss pads every line of a block to its widest line. One frame of a growing thought: 193 us -> 6.4 us at 4 KB, 1.44 ms -> 36.6 us at 32 KB, 5.68 ms -> 134 us at 128 KB. - `sanitizeToolText` returns the caller's own string when it carries no control sequence, no control character and no surrounding space, which is what a streamed thought is. It was 52% of the live Thinking frame's allocation. - `transcriptStore.updateRevision` no longer hashes the live record's text on every frame. Only a search reads that text, so the digest moved behind `fingerprint()`, which a search scan calls; the frame's revision keeps the geometry and the record heights. This frame was 9.7% of samples. A profile after the change no longer shows `renderThinkingBlock` or `updateRevision` among the application frames; what leads is scheduler handoff and the renderer's PTY writes. Each change is pinned by tests that fail without it: a streaming equivalence test and an fuzz target for the live Thinking path, a whole-thought-per-frame mutation that breaks the bounded-work assertion, a frozen-prefix invalidation test, an allocation assertion for clean prose, and a revision test asserting the frame does not read the live rows while a search fingerprint does.
…n prefix
goldmark resolves link references over the whole document, so a bracket on
either side of a freeze boundary can change the other side's rows. The guard was
a per-line `[...]:` pattern applied only in one direction, which left two holes:
- a usage before its definition froze into the prefix, so the frame rendered
`[later]` literally while a whole render resolved it:
"Intro.\n\nUse [later] first.\n\n[later]: /url\n\nTail.\n"
- a definition whose label spans lines is valid CommonMark but invisible to a
per-line test, so the definition sat in the frozen prefix and a usage in the
tail lost its scope:
"A.\n\n[x\ny]: https://e.com \"t\"\n\nB.\n\nC uses [x y] here.\n"
The guard is now a single rule: a boundary is freezable only if the frozen
prefix carries no `[` at all. That covers both directions and the multiline
label, and makes the per-line definition pattern unnecessary. Fenced code is
exempt because its brackets are code; autolinks, linkified bare URLs and inline
links resolve without definitions and stay freezable. A body that carries a
bracket falls back to a whole render, which is correct but not incremental.
Both documents now render byte-identically to a whole render. Four corpus cases
(`forward ref`, `multi label`, `code bracket`, `shortcut refs`) are replayed by
the streaming equivalence test and the fuzz target, and a focused boundary test
pins the freeze point; all of them fail under a mutation that restores the old
per-line rule.
CI lints the pull request's own diff (`--new-from-rev=HEAD~`), and this branch was never pushed, so its changes had never been linted. Two suppressions and four mechanical fixes: - `markdownFrozenRows` and `liveMarkdownBoundary` are boundary scanners whose branches are the conditions of one ordered pass, so they carry `//nolint:gocyclo` with the reason the repo uses for the same shape elsewhere. - the row-count to hash-bucket conversion in `updateRevision` is `//nolint:gosec`: a row count is never negative. - `markdownRowsAfter` and `markdownIndented` use integer ranges. - `TestSanitizeToolTextLeavesCleanProseUntouched` is `//nolint:paralleltest`: AllocsPerRun measures the whole process and refuses a parallel test. - the two `len` comparisons use `assert.Len`, and the whitespace wsl_v5 asks for before a `for`, a `return`, an assignment or a fuzz seed was added. No behavior change: build, vet, the TUI tests, the race run and the PTY suite pass, and `golangci-lint --new-from-rev=abd4e1e` reports no issues on the diff.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Phase 1 of the TUI CPU work. It removes the dominant per-frame cost: the two live blocks
(the streaming answer draft and the streaming reasoning section) were re-rendered in
full on every frame, which is O(body) per frame and O(body^2) over one message.
Commits
5b55f9cassemble a streaming live body from a frozen prefixfdfda67freeze the live Thinking rows and stop hashing them per frame8201df5keep link references on one side of a frozen Markdown prefixMeasured
Real entry point
tui.Runin a PTY (120x40), 627-message session, 1500 deltas at 150/swith interleaved tool round trips, external 100 ms CPU sampling; HEAD
abd4e1evs thisbranch, interleaved rounds so machine drift cannot favour one build.
One frame of a growing body (
-benchmem): Markdown 32 KB 22.5 ms -> 0.20 ms,Thinking 32 KB 1.44 ms -> 38 us. Idle is unchanged (~2-3% of one core, no periodic spike).
Correctness
A prefix may be frozen only where rendering it alone reproduces the whole document's rows
byte for byte. Pinned by
TestLiveMarkdownMatchesWholeRenderand its fuzz target, byTestLiveThinkingMatchesSettledSectionand its fuzz target, and by a streaming corpusthat includes the link-reference cases (a usage before its definition, a definition whose
label spans lines, brackets in fenced code) plus mutation checks that restore the old rule.
Notes
No visible-output change, no frame-rate change, no fork change. The pinned Bubble Tea
fork's 60 Hz flush loop is the idle floor and is untouched.