Skip to content

Render-oriented history reads ​

This experiment keeps the dense viewport / compact history split and changes how snapshots read it. The reproducible patches are not enabled in production. Rozi and the sibling framework retain their normal dependencies. The conversion algorithm is unchanged. Batching stays rejected.

What this run asked ​

The previous day's live render_snapshot() regression was about 22%. That could have been compact storage itself, or the cell iterator the snapshot still used for every visible column. This run separates those.

Three snapshot paths share one candidate executable. ROZI_EXPERIMENT_RENDER is an experiment selector, not a supported setting:

ValueSnapshot constructionLive display_iter
legacyGridIterator plus read_cellhistory-capable
denseGridIteratordense viewport index
unset or bulkrow parts, prefix copy, tail filldense viewport index

Offsets are 0 (live dense rows), 32 (32 compact history rows plus 32 dense viewport rows), and 5,000 (the oldest retained history). Fixture construction, scroll positioning, and snapshot destruction stay outside the timer. The cases measure snapshot construction, not backend drawing or an unchanged cached snapshot.

A compact row is still an exact prefix plus one repeated tail cell. Snapshot construction that indexes width cells pays a representation decision on every column. The bulk path copies the prefix and fills the tail, including the blank spaces the snapshot currently stores.

Environment ​

Rozi cedd4a6 plus the benchmark and documentation changes in this worktree, Rust 1.98.1, and released engine/framework versions 0.26.0/0.9.0 as the dense baseline. AMD Ryzen 7 5700X3D, 16 threads, CPU scaling reported at 80% of max MHz. Criterion used 20 samples, two-second target measurement, and one-second warm-up. Values below are mean estimates. Negative changes mean less time.

The engine differential test, 140 unit tests, 45 reference recordings, and 133 framework terminal unit tests passed, including a new check that bulk and iterator snapshots match at live, mixed, and deep offsets.

Snapshot construction ​

Same-day measurements. The candidate executable is shared; only the selector changes.

CaseReleased dense µsLegacy µsDense-iter µsBulk µs
render_active/plain275356360318
render_mix/plain288350353181
render_deep/plain27634635344
render_active/unicode275346356317
render_mix/unicode278349349182
render_deep/unicode27534634850
render_active/dense288379354324
render_mix/dense282371360320
render_deep/dense281362364320

Change versus released dense:

CaseLegacyDense-iterBulk
render_active/plain+29.8%+29.7%+15.1%
render_mix/plain+23.6%+23.1%-36.3%
render_deep/plain+25.1%+27.3%-84.0%
render_active/unicode+26.1%+29.2%+15.6%
render_mix/unicode+25.2%+25.4%-34.5%
render_deep/unicode+25.6%+26.3%-81.8%
render_active/dense+33.3%+23.0%+18.6%
render_mix/dense+31.5%+28.6%+13.2%
render_deep/dense+28.0%+28.9%+12.2%

The live dense-iter path is not faster than legacy. Switching display_iter to dense indexing does not remove the live regression. The visible rows at offset 0 were already dense; the extra cost was not compact-cell indexing of the screen the user is looking at.

Bulk construction cuts that live gap roughly in half, to about 43 µs (318 µs versus 275 µs on plain). That is 0.26% of a 16.7 ms 60 Hz frame for one pane. It does not recover released live throughput. The leftover is still dense viewport snapshot work under the split grid: style runs, text, and span assembly over 16,192 cells. It is not compact history walking.

Once the view includes compact rows, bulk wins by a lot. Deep plain scrollback is 44 µs instead of 276 µs. Mixed scrollback is 181 µs instead of 288 µs. Those rows already know their repeated tail, so filling it is cheaper than asking every column for a style. Full-width unique cells cannot take that shortcut; their bulk times stay about 12-19% behind released dense.

Isolated row access ​

The engine microbenchmarks clone one 253-cell short row. Fixture construction stays outside the timer. These are not snapshot times.

CaseTime
Dense slice clone329 ns
Compact copy_into398 ns
Compact cell-by-cell index581 ns
Live display_iter over 64×253 cells49.2 µs
Deep display_iter57.5 µs
Live display_row().parts()60 ns
Deep display_row().parts()79 ns

Prefix copy plus tail fill is close to cloning a dense slice. Indexing 253 compact cells is not. display_iter itself is tens of microseconds. Most of a 275 µs snapshot is still framework span construction. parts() without consuming cells is noise.

Decision ​

The 22% live snapshot result was a read-API problem, not a reason to reopen conversion. Pointing display_iter at dense storage is not enough. A bulk history-read API is: scrollback snapshots get cheaper than today's dense engine, and the live gap drops from about 30% to about 15%.

That is not yet a production fork. Short-line ingest is still 12-25% slower, and live snapshots are still not recovered to noise. The remaining live 43 µs is small against a frame budget, but it repeats on every dirty pane. I would only chase it with a profile of dense-row span assembly, not with another storage change.

If a later integration can render compact tails as runs without building a dense intermediate snapshot, do that after the ingest question is settled. The current snapshot still stores every blank column as text, so even an optimal compact read has to emit those spaces until that contract changes.

Production dependencies and the sibling framework remain unchanged. The retained Rozi code change adds mixed and deep snapshot cases to the history benchmark. The storage and read-API work stays in reproducible experimental patches. Follow-up dense span-assembly measurements are in Dense span-assembly profiling. Attr-keyed run detection, and a compact remasurement against that new renderer, are in Attr-keyed terminal spans.

MPL-2.0