Render-oriented history reads
This experiment keeps the dense viewport / compact history split and changes how snapshots read it. The reproducible patches are not enabled in production. Rozi and the sibling framework retain their normal dependencies. The conversion algorithm is unchanged. Batching stays rejected.
What this run asked
The previous day's live render_snapshot() regression was about 22%. That could have been compact storage itself, or the cell iterator the snapshot still used for every visible column. This run separates those.
Three snapshot paths share one candidate executable. ROZI_EXPERIMENT_RENDER is an experiment selector, not a supported setting:
| Value | Snapshot construction | Live display_iter |
|---|---|---|
legacy | GridIterator plus read_cell | history-capable |
dense | GridIterator | dense viewport index |
unset or bulk | row parts, prefix copy, tail fill | dense viewport index |
Offsets are 0 (live dense rows), 32 (32 compact history rows plus 32 dense viewport rows), and 5,000 (the oldest retained history). Fixture construction, scroll positioning, and snapshot destruction stay outside the timer. The cases measure snapshot construction, not backend drawing or an unchanged cached snapshot.
A compact row is still an exact prefix plus one repeated tail cell. Snapshot construction that indexes width cells pays a representation decision on every column. The bulk path copies the prefix and fills the tail, including the blank spaces the snapshot currently stores.
Environment
Rozi cedd4a6 plus the benchmark and documentation changes in this worktree, Rust 1.98.1, and released engine/framework versions 0.26.0/0.9.0 as the dense baseline. AMD Ryzen 7 5700X3D, 16 threads, CPU scaling reported at 80% of max MHz. Criterion used 20 samples, two-second target measurement, and one-second warm-up. Values below are mean estimates. Negative changes mean less time.
The engine differential test, 140 unit tests, 45 reference recordings, and 133 framework terminal unit tests passed, including a new check that bulk and iterator snapshots match at live, mixed, and deep offsets.
Snapshot construction
Same-day measurements. The candidate executable is shared; only the selector changes.
| Case | Released dense µs | Legacy µs | Dense-iter µs | Bulk µs |
|---|---|---|---|---|
render_active/plain | 275 | 356 | 360 | 318 |
render_mix/plain | 288 | 350 | 353 | 181 |
render_deep/plain | 276 | 346 | 353 | 44 |
render_active/unicode | 275 | 346 | 356 | 317 |
render_mix/unicode | 278 | 349 | 349 | 182 |
render_deep/unicode | 275 | 346 | 348 | 50 |
render_active/dense | 288 | 379 | 354 | 324 |
render_mix/dense | 282 | 371 | 360 | 320 |
render_deep/dense | 281 | 362 | 364 | 320 |
Change versus released dense:
| Case | Legacy | Dense-iter | Bulk |
|---|---|---|---|
render_active/plain | +29.8% | +29.7% | +15.1% |
render_mix/plain | +23.6% | +23.1% | -36.3% |
render_deep/plain | +25.1% | +27.3% | -84.0% |
render_active/unicode | +26.1% | +29.2% | +15.6% |
render_mix/unicode | +25.2% | +25.4% | -34.5% |
render_deep/unicode | +25.6% | +26.3% | -81.8% |
render_active/dense | +33.3% | +23.0% | +18.6% |
render_mix/dense | +31.5% | +28.6% | +13.2% |
render_deep/dense | +28.0% | +28.9% | +12.2% |
The live dense-iter path is not faster than legacy. Switching display_iter to dense indexing does not remove the live regression. The visible rows at offset 0 were already dense; the extra cost was not compact-cell indexing of the screen the user is looking at.
Bulk construction cuts that live gap roughly in half, to about 43 µs (318 µs versus 275 µs on plain). That is 0.26% of a 16.7 ms 60 Hz frame for one pane. It does not recover released live throughput. The leftover is still dense viewport snapshot work under the split grid: style runs, text, and span assembly over 16,192 cells. It is not compact history walking.
Once the view includes compact rows, bulk wins by a lot. Deep plain scrollback is 44 µs instead of 276 µs. Mixed scrollback is 181 µs instead of 288 µs. Those rows already know their repeated tail, so filling it is cheaper than asking every column for a style. Full-width unique cells cannot take that shortcut; their bulk times stay about 12-19% behind released dense.
Isolated row access
The engine microbenchmarks clone one 253-cell short row. Fixture construction stays outside the timer. These are not snapshot times.
| Case | Time |
|---|---|
| Dense slice clone | 329 ns |
Compact copy_into | 398 ns |
| Compact cell-by-cell index | 581 ns |
Live display_iter over 64×253 cells | 49.2 µs |
Deep display_iter | 57.5 µs |
Live display_row().parts() | 60 ns |
Deep display_row().parts() | 79 ns |
Prefix copy plus tail fill is close to cloning a dense slice. Indexing 253 compact cells is not. display_iter itself is tens of microseconds. Most of a 275 µs snapshot is still framework span construction. parts() without consuming cells is noise.
Decision
The 22% live snapshot result was a read-API problem, not a reason to reopen conversion. Pointing display_iter at dense storage is not enough. A bulk history-read API is: scrollback snapshots get cheaper than today's dense engine, and the live gap drops from about 30% to about 15%.
That is not yet a production fork. Short-line ingest is still 12-25% slower, and live snapshots are still not recovered to noise. The remaining live 43 µs is small against a frame budget, but it repeats on every dirty pane. I would only chase it with a profile of dense-row span assembly, not with another storage change.
If a later integration can render compact tails as runs without building a dense intermediate snapshot, do that after the ingest question is settled. The current snapshot still stores every blank column as text, so even an optimal compact read has to emit those spaces until that contract changes.
Production dependencies and the sibling framework remain unchanged. The retained Rozi code change adds mixed and deep snapshot cases to the history benchmark. The storage and read-API work stays in reproducible experimental patches. Follow-up dense span-assembly measurements are in Dense span-assembly profiling. Attr-keyed run detection, and a compact remasurement against that new renderer, are in Attr-keyed terminal spans.