Skip to content

Dense span-assembly profiling ​

This run profiles the current dense render_snapshot() path on released Alacritty and tui-lipan. The reproducible patch does not change production renderable_content_lines. Compact history, conversion, and a specialized compact renderer stay out of this pass.

The question was whether the span builder could consume row structure more directly and skip work that dense and compact rows share. The answer is yes, but the shared work is not row traversal.

What this run asked ​

A released 253×64 snapshot is about 275-290 µs. Live display_iter is only about 28-49 µs of that. The rest was unattributed framework assembly. Subtractive Criterion stages on the released builder split that remainder:

text
dense snapshot
├── row traversal / cell acquisition
├── span boundary detection
├── attribute comparisons
├── text / combining-character handling
├── span construction / pushes
└── snapshot allocation / initialization

Two probes sit beside that chain. One walks dense rows by Line index instead of GridIterator. The other compares fg / bg / style flags and constructs Style only when a run actually changes. Unit tests require both to match the current snapshot text, spans, and hyperlinks.

Environment ​

Rozi cedd4a6 plus the documentation and experiment archive in this worktree, Rust 1.98.1, and released engine/framework versions 0.26.0/0.9.0. AMD Ryzen 7 5700X3D, 16 threads, CPU scaling reported at 89% of max MHz. Criterion used 20 samples, two-second target measurement, and one-second warm-up. Values below are mean estimates.

135 framework terminal unit tests passed, including the three stage checks. Component timings include benchmark overhead. They lined up closely enough here to treat as an attribution, not as a cycle-accurate partition.

Stage times ​

Same-day measurements. 253×64 live viewport, 5,000 history rows, offset 0. Times in µs except wrap_flags.

Stageplainunicodedense
walk_cells28.028.027.9
walk_viewport32.432.532.3
walk_dense_rows9.19.18.9
map_styles95.695.595.8
detect_runs215.4217.5215.2
detect_runs_cell68.367.767.8
accumulate_text75.276.076.1
build_lines281.6285.0286.9
join_text283.5288.6283.7
wrap_flags0.1850.1850.188
snapshot288.8293.7287.7
lines_dense_rows297.9294.1296.8
lines_attr_keyed161.2162.7161.5
lines_dense_attr132.9133.1132.6

wrap_flags is nanoseconds. Join versus build_lines is in the noise on dense. The "dense" corpus fills every column with a unique letter. Those cells still share the default style, so the builder still emits one span per row.

Attribution ​

Plain snapshot, subtracting adjacent stages:

text
dense snapshot                         288.8 µs
├── GridIterator walk                   28.0
├── point_to_viewport                    4.4
├── style_from_term_cell per cell       63.2
├── Style PartialEq per cell           119.8
├── text, hyperlinks, span pushes       66.2
└── join, wrap flags, cache init         7.2

accumulate_text minus viewport walk is 42.8 µs, so most of the 66 µs after detect_runs is pushing characters. Span allocation, hyperlink scans, and flushes are the rest. Wrap flags are free. Snapshot bookkeeping after the line vec exists is a few microseconds.

That is 1.73% of a 16.7 ms 60 Hz frame for one pane. It is not a user-visible hitch. It is also not row storage. 183 µs of the 289 µs is constructing a 14-field Style and comparing it on every cell, including 200 trailing spaces that never change the run.

Comparing cell attributes without building Style is 68 µs, not 215 µs. The current equality path is the expensive one.

What actually moves the total ​

Dense row indexing makes a bare walk 3.5× faster (9 µs versus 28 µs). Rewriting the full builder to use that walk, while still mapping Style on every cell, does not win. lines_dense_rows is 298 µs against 282 µs for the production loop. Traversal was never the remainder.

Constructing Style only at run boundaries is. lines_attr_keyed is 161 µs, 43% under build_lines, and the tests say the snapshot matches. Combining that with dense rows reaches 133 µs. Once style work is gone, the 19 µs iterator tax shows up, which is when a parts() consumer would start to matter.

I would still not feed compact tails into span assembly yet. The dense builder was spending 120 µs deciding that 16,192 default cells are one run. Released Alacritty pays that too. Fix that boundary first. A compact prefix + repeated tail span is a later, smaller play, and it still has to wait on the snapshot storing every blank column as text.

Decision ​

text
Shared compact/live row
    rejected

Dense viewport + compact history
    strongly promising

Deferred/batch compaction
    rejected

Bulk structural history API
    validated

Snapshot/render traversal
    substantially improved

Remaining blocker
    short-line ingestion conversion cost

Span assembly
    Style-per-cell mapping and PartialEq dominate

Next investigation
    attr-keyed run detection on released tui-lipan

Do not ship an engine fork. Conversion is still the storage blocker. The live 43 µs bulk-read gap from the previous day is small against a frame, and most snapshot time was never compact history. The next useful patch is in tui-lipan, on released dense rows, and it would help Rozi whether or not compact history lands.

Production dependencies and the sibling framework remain unchanged. Follow-up implementation and compact remasurement are in Attr-keyed terminal spans.

MPL-2.0