Skip to content

Attr-keyed terminal spans ​

This run lands the span-assembly follow-up from Dense span-assembly profiling in production tui-lipan, against released alacritty_terminal 0.26.0. Compact history stays frozen. After the renderer change, the compact prototype is remeasured against that new baseline, not against the old per-cell Style builder.

Rozi's committed Cargo.lock still records crates.io tui-lipan 0.9.0. The implementation lives on the sibling branch perf/attr-keyed-terminal-spans and is uncommitted. Measurements below used a local --config path patch. That patch is not a production dependency.

What this run asked ​

A released 253×64 snapshot spent about 183 µs of ~289 µs constructing and comparing a 14-field Style on every cell, then emitting run-oriented spans. The builder should compare the cheap cell attributes that determine those spans, and construct Style only when a run actually changes:

text
cell
  ↓
cheap render-attribute key comparison
  ↓ only when the key changes
construct full Style
  ↓
start / close span

The key has to match snapshot-visible style semantics exactly. Hyperlinks, hidden text, and wide-cell spacers still affect span boundaries, or don't, the same way they did before. The worst case is a style change on every cell. That path should stay close to the old builder, not pay a second Style map on every break.

After that lands, compact history is a different question. The earlier +15% live snapshot penalty was measured against a renderer doing a lot of unnecessary per-cell work. Both paths drop that work now. The relative economics can change.

Environment ​

Rozi cedd4a6 plus the documentation and experiment archive in this worktree. Sibling tui-lipan 32baafe plus the uncommitted attr-keyed renderer. Rust 1.98.1. AMD Ryzen 7 5700X3D, 16 threads, CPU scaling reported at 87% of max MHz. Criterion used 20 samples, two-second target measurement, and one-second warm-up. Values below are mean estimates.

142 tui-lipan terminal unit tests passed, including new checks for default space tails, named versus indexed palette red, underline versus double underline, bold, hidden text, hyperlinks, wide-cell spacers, alternating runs, inverse, and dim. cargo test --features terminal --lib is 2632 passed. The disposable compact tree kept the frozen 133 framework tests and they passed after the same key was ported onto SnapshotRuns only.

The key ​

The hot path does not allocate a RunAttrs struct on every cell. It compares TermColor fg / bg and a six-bit flag mask first. Mapping through the palette happens only when those fields differ. Style is built from the mapped colours already in hand, so a run break does not map twice.

Bits that determine Style identity:

  • BOLD, DIM, ITALIC
  • ALL_UNDERLINES as a boolean, because the snapshot's Style.underline is already a boolean. Single and double underline stay one run.
  • INVERSE, STRIKEOUT

Not in the key, matching current snapshot behaviour:

  • HIDDEN turns the glyph into a space and stays in the run
  • hyperlinks are recorded on the side and do not split style runs
  • WRAPLINE, WIDE_CHAR, and spacers. Spacers are skipped before the key check.
  • underline_color. Style still sets it to None.

Named(Red) and Indexed(1) compare unequal as TermColor and equal after palette mapping, so they merge. That is the case a TermColor-only key would get wrong.

The first version remapped colours in the continue check and again in style_from_term_cell. Alternating styles paid that twice and ran 30% slower. Storing the mapped colours on the open run closed that hole.

Three output shapes ​

Same process, sibling tui-lipan terminal_snapshot bench, 253×64, 5,000 history rows. The per-cell baseline is the 0.9.0 builder. Negative changes mean less time.

CorpusPer-cell µsAttr-keyed µsChange
tails (short log plus default spaces)298151-49.2%
shell (unicode, combining, emoji)288155-46.3%
alternating (red/green every cell)9871039+5.3%

The 43% line-assembly probe from the previous report survives as a full snapshot cut of about 46-49% on realistic rows. Alternating is 16,192 spans. That is not a pane anyone draws on purpose. +5% there is the tax for the key check, and it is close enough.

On Rozi's production snapshot cases, the same path patch, released engine:

CaseReleased 0.9.0 µsAttr-keyed µsChange
render_active/plain275174-37%
render_mix/plain288172-40%
render_deep/plain276162-41%
render_active/unicode275171-38%
render_mix/unicode278179-36%
render_deep/unicode275172-37%
render_active/dense288172-40%
render_mix/dense282187-34%
render_deep/dense281178-37%

The 0.9.0 column is the previous day's render-audit means on this machine (CPU scaling 80%). Today's attr-keyed column is 87%. Direction and size still agree with the same-process tui-lipan numbers. 174 µs is 1.0% of a 16.7 ms 60 Hz frame for one pane. It was 1.7%.

Compact history against the new baseline ​

Frozen patches under tools/experiments/render-history/ were not edited. A disposable copy under /tmp/rozi-attr-compact applied those patches, then ported only the attr-keyed continue check onto SnapshotRuns. Default selector is still bulk RowParts.

Same Rozi terminal_history render_* IDs, same Criterion store, immediately after the attr-keyed dense run. Criterion's change line is compact versus that new production baseline, not versus 0.9.0.

CaseAttr-keyed dense µsCompact + attr µsChange
render_active/plain174171noise
render_mix/plain172103-39%
render_deep/plain16234-79%
render_active/unicode171174noise
render_mix/unicode179106-41%
render_deep/unicode17239-77%
render_active/dense172176noise
render_mix/dense187177noise
render_deep/dense178177noise

The +15% live snapshot result is gone. Live compact and live dense are the same number, inside Criterion noise. That 43 µs leftover from bulk reads was almost entirely per-cell Style work on dense viewport rows. Once both builders skip it, RowParts on a live dense grid is not slower than GridIterator.

Scrollback still wins, and it wins on a faster baseline. Deep plain is 34 µs instead of 162 µs. Mixed short-line history is 103 µs instead of 172 µs. Full-width unique cells still cannot use a repeated tail, so dense stays even.

The old compact live 318 µs versus released 275 µs was a measurement of the wrong renderer. It is not a reason to keep the fork out. Ingest conversion still is: short-line sustained ingest remains 12-25% slower, and that number was never about snapshots.

Decision ​

text
Rendering:
per-cell Style construction     identified bottleneck
attr-keyed run detection        ready to ship in tui-lipan

Compact history:
memory benefit                  proven
scrollback read benefit         proven
live-path contamination         solved
batching                        rejected
remaining blocker               short-line conversion cost

Ship the tui-lipan change. It pays for the investigation whether or not the compact-history fork is ever accepted. Do not put a path patch in Rozi's lockfile. Pick the renderer up on the next framework release.

Leave the compact patches frozen. The next compact question is still ingest conversion, not another snapshot trick. RowParts is the right read API at this baseline. A specialized compact tail span can wait until the snapshot stops storing every blank column as text.

Reproducing ​

Sibling tui-lipan, from perf/attr-keyed-terminal-spans:

bash
cargo test --features terminal --lib widgets::terminal
cargo bench --features terminal --bench terminal_snapshot -- --baseline per-cell --sample-size 20 --measurement-time 2 --warm-up-time 1

Rozi, local path patch only. Restore Cargo.lock afterwards:

bash
cargo --config 'patch.crates-io.tui-lipan.path="../tui-lipan"' \
  bench --bench terminal_history -- 'render_' \
  --sample-size 20 --measurement-time 2 --warm-up-time 1

Compact remasurement used a disposable copy of tools/experiments/render-history/ plus the same attr-keyed continue check on SnapshotRuns. Default ROZI_EXPERIMENT_RENDER is bulk.

MPL-2.0