Dense span-assembly profiling
This run profiles the current dense render_snapshot() path on released Alacritty and tui-lipan. The reproducible patch does not change production renderable_content_lines. Compact history, conversion, and a specialized compact renderer stay out of this pass.
The question was whether the span builder could consume row structure more directly and skip work that dense and compact rows share. The answer is yes, but the shared work is not row traversal.
What this run asked
A released 253×64 snapshot is about 275-290 µs. Live display_iter is only about 28-49 µs of that. The rest was unattributed framework assembly. Subtractive Criterion stages on the released builder split that remainder:
dense snapshot
├── row traversal / cell acquisition
├── span boundary detection
├── attribute comparisons
├── text / combining-character handling
├── span construction / pushes
└── snapshot allocation / initializationTwo probes sit beside that chain. One walks dense rows by Line index instead of GridIterator. The other compares fg / bg / style flags and constructs Style only when a run actually changes. Unit tests require both to match the current snapshot text, spans, and hyperlinks.
Environment
Rozi cedd4a6 plus the documentation and experiment archive in this worktree, Rust 1.98.1, and released engine/framework versions 0.26.0/0.9.0. AMD Ryzen 7 5700X3D, 16 threads, CPU scaling reported at 89% of max MHz. Criterion used 20 samples, two-second target measurement, and one-second warm-up. Values below are mean estimates.
135 framework terminal unit tests passed, including the three stage checks. Component timings include benchmark overhead. They lined up closely enough here to treat as an attribution, not as a cycle-accurate partition.
Stage times
Same-day measurements. 253×64 live viewport, 5,000 history rows, offset 0. Times in µs except wrap_flags.
| Stage | plain | unicode | dense |
|---|---|---|---|
walk_cells | 28.0 | 28.0 | 27.9 |
walk_viewport | 32.4 | 32.5 | 32.3 |
walk_dense_rows | 9.1 | 9.1 | 8.9 |
map_styles | 95.6 | 95.5 | 95.8 |
detect_runs | 215.4 | 217.5 | 215.2 |
detect_runs_cell | 68.3 | 67.7 | 67.8 |
accumulate_text | 75.2 | 76.0 | 76.1 |
build_lines | 281.6 | 285.0 | 286.9 |
join_text | 283.5 | 288.6 | 283.7 |
wrap_flags | 0.185 | 0.185 | 0.188 |
snapshot | 288.8 | 293.7 | 287.7 |
lines_dense_rows | 297.9 | 294.1 | 296.8 |
lines_attr_keyed | 161.2 | 162.7 | 161.5 |
lines_dense_attr | 132.9 | 133.1 | 132.6 |
wrap_flags is nanoseconds. Join versus build_lines is in the noise on dense. The "dense" corpus fills every column with a unique letter. Those cells still share the default style, so the builder still emits one span per row.
Attribution
Plain snapshot, subtracting adjacent stages:
dense snapshot 288.8 µs
├── GridIterator walk 28.0
├── point_to_viewport 4.4
├── style_from_term_cell per cell 63.2
├── Style PartialEq per cell 119.8
├── text, hyperlinks, span pushes 66.2
└── join, wrap flags, cache init 7.2accumulate_text minus viewport walk is 42.8 µs, so most of the 66 µs after detect_runs is pushing characters. Span allocation, hyperlink scans, and flushes are the rest. Wrap flags are free. Snapshot bookkeeping after the line vec exists is a few microseconds.
That is 1.73% of a 16.7 ms 60 Hz frame for one pane. It is not a user-visible hitch. It is also not row storage. 183 µs of the 289 µs is constructing a 14-field Style and comparing it on every cell, including 200 trailing spaces that never change the run.
Comparing cell attributes without building Style is 68 µs, not 215 µs. The current equality path is the expensive one.
What actually moves the total
Dense row indexing makes a bare walk 3.5× faster (9 µs versus 28 µs). Rewriting the full builder to use that walk, while still mapping Style on every cell, does not win. lines_dense_rows is 298 µs against 282 µs for the production loop. Traversal was never the remainder.
Constructing Style only at run boundaries is. lines_attr_keyed is 161 µs, 43% under build_lines, and the tests say the snapshot matches. Combining that with dense rows reaches 133 µs. Once style work is gone, the 19 µs iterator tax shows up, which is when a parts() consumer would start to matter.
I would still not feed compact tails into span assembly yet. The dense builder was spending 120 µs deciding that 16,192 default cells are one run. Released Alacritty pays that too. Fix that boundary first. A compact prefix + repeated tail span is a later, smaller play, and it still has to wait on the snapshot storing every blank column as text.
Decision
Shared compact/live row
rejected
Dense viewport + compact history
strongly promising
Deferred/batch compaction
rejected
Bulk structural history API
validated
Snapshot/render traversal
substantially improved
Remaining blocker
short-line ingestion conversion cost
Span assembly
Style-per-cell mapping and PartialEq dominate
Next investigation
attr-keyed run detection on released tui-lipanDo not ship an engine fork. Conversion is still the storage blocker. The live 43 µs bulk-read gap from the previous day is small against a frame, and most snapshot time was never compact history. The next useful patch is in tui-lipan, on released dense rows, and it would help Rozi whether or not compact history lands.
Production dependencies and the sibling framework remain unchanged. Follow-up implementation and compact remasurement are in Attr-keyed terminal spans.