Skip to content

Dense viewport and compact history experiment ​

The second experiment separates the live terminal grid from history storage. It is preserved as reproducible engine and framework patches. Rozi's production dependencies remain unchanged. The prototype is not adopted as the default.

Representation boundary ​

The viewport keeps Alacritty's original width-sized Vec<Cell> rows, renamed DenseRow. Ordinary indexing and mutable slices only access dense viewport rows. Negative-line indexing is rejected; history readers must explicitly use read_row or read_cell. Unusual history mutation calls history_row_mut and materializes only the affected row.

History lives in a separate deque. A row becomes either HistoryRow::Dense or HistoryRow::Compact when it scrolls out of the viewport. Compact rows retain an exact cell prefix and one repeated tail cell. Comparing whole cells preserves background colors, flags, hyperlinks, and Unicode data. Rows whose prefix is too long stay dense.

A per-grid pool holds at most four dense buffers and four prefix buffers. Eviction returns buffers to the pool; scrolling reuses them. Width changes release buffers of the old width. Reflow moves row ownership and expands working rows as needed, rather than first expanding the entire history. Rows ending in default cells can resize directly when no text crosses a wrapping boundary.

The framework patch changes history consumers to use the explicit read API. Its text and replay readers test the repeated tail once and iterate the actual prefix. This matters: the first split version still scanned blank tails column by column and was slower at search, copy, and replay.

Measurement boundaries ​

The new terminal_history benchmark fills 5,000 history rows before timing. It then repeatedly processes 10,000 or 100,000 more lines in the same terminal. It covers short ASCII, short Unicode with combining and wide characters, and full-width nonrepeating text at 80×24 and 253×64.

Separate cases measure text scanning with no match, selection export, streaming replay, and height/narrow/wide resize. Fixture setup and destruction stay outside resize timers. The scan case uses the framework text iterator, not Rozi's complete cooperative search implementation.

The engine benchmark separately measures dense-to-compact conversion with cold and warm pools, pooled expansion, eviction, and resizing with and without reflow.

Requested-heap probes count allocations retained by one terminal or production pane. They are not RSS/PSS measurements and do not include a complete server plus attached client. An 86% saving in this short-line workload does not imply an 86% reduction in Rozi's total process memory.

Paired sustained and history-operation results ​

Measured on September 13 with Rust 1.98.1, Rozi 824d925, and published engine/framework versions 0.26.0/0.9.0. The candidate ran first, followed by the unchanged dense benchmark executable on the same machine. Criterion used 10 samples, one second warm-up, and one second target measurement; slow cases automatically ran longer. Values below are mean estimates in milliseconds. Negative changes mean less time. These short local runs establish large effects, not sub-percent rankings.

CaseDense msSplit msChange
boundary/copy/dense4.71584.5998-2.5%
boundary/copy/plain1.80110.5070-71.8%
boundary/copy/unicode2.03560.8053-60.4%
boundary/replay/dense7.57378.0068+5.7%
boundary/replay/plain1.75650.5889-66.5%
boundary/replay/unicode1.81010.6354-64.9%
boundary/resize_height/dense0.28920.0240-91.7%
boundary/resize_height/plain0.26840.0306-88.6%
boundary/resize_height/unicode0.27520.0312-88.7%
boundary/resize_narrow/dense4.48743.8805-13.5%
boundary/resize_narrow/plain3.01521.4137-53.1%
boundary/resize_narrow/unicode3.03011.4179-53.2%
boundary/resize_wide/dense1.53651.2734-17.1%
boundary/resize_wide/plain10.23920.2718-97.3%
boundary/resize_wide/unicode1.53670.2755-82.1%
boundary/search_no_match/dense2.41402.3654-2.0%
boundary/search_no_match/plain1.48760.2553-82.8%
boundary/search_no_match/unicode1.67270.4308-74.2%
sustained/dense/253x64_1000022.518122.7389+1.0%
sustained/dense/253x64_100000227.1488226.7837-0.2%
sustained/dense/80x24_100007.53277.6087+1.0%
sustained/dense/80x24_10000075.386176.0902+0.9%
sustained/plain/253x64_100001.92522.4021+24.8%
sustained/plain/253x64_10000019.416124.1836+24.6%
sustained/plain/80x24_100001.90552.3372+22.7%
sustained/plain/80x24_10000019.104123.3856+22.4%
sustained/unicode/253x64_100003.51184.3045+22.6%
sustained/unicode/253x64_10000034.608642.6055+23.1%
sustained/unicode/80x24_100003.35674.1728+24.3%
sustained/unicode/80x24_10000036.045740.3539+12.0%

The original September 10–11 split trial also found sustained short-line regressions. The revised reader API removes its search/copy regressions without removing the conversion cost. Dense narrow resize is noisy across runs: the earlier dense baseline was 3.29 ms, while this rerun was 4.49 ms. Do not treat the apparent dense narrow-resize improvement as a stable result.

Fresh-terminal comparison ​

The existing terminal_ingest benchmark creates a fresh terminal outside each timer, then times processing and dropping it. It therefore measures a different lifecycle from sustained ingestion. The same-day candidate and dense runs used the same 10-sample settings:

Corpus / viewportDense msSplit msChange
cat_large/200x6014.67218.880+28.7%
cat_large/320x9016.59319.618+18.2%
cat_large/80x2411.07013.034+17.7%
plain_lines/200x6010.7934.835-55.2%
plain_lines/320x9015.9435.215-67.3%
plain_lines/80x247.6272.766-63.7%
scroll_regions/200x604.3004.388+2.0%
scroll_regions/320x905.0775.138+1.2%
scroll_regions/80x243.5783.664+2.4%
sgr_heavy/200x609.3083.608-61.2%
sgr_heavy/320x9013.2013.567-73.0%
sgr_heavy/80x245.4663.600-34.1%
wide_unicode/200x606.7423.469-48.5%
wide_unicode/320x907.0533.529-50.0%
wide_unicode/80x243.1403.274+4.3%

Fresh plain/styled workloads improve substantially, while the large repeated-text corpus regresses. The wide Unicode cases no longer show the earlier shared-row prototype's broad regression. These results do not cancel the sustained short-line regression; they show that fixture lifetime and history population materially change the result.

Conversion costs ​

For one 253-cell row containing short log record, the isolated benchmark measured:

OperationTime estimate interval
Dense to compact, cold pool76.6–77.8 ns
Dense to compact, warm pool65.3–66.5 ns
Compact to dense, pooled363–367 ns
Compact eviction, cold prefix pool44.5–44.6 ns
Split grid resize 253 to 80, no reflow499–501 µs
Dense grid resize 253 to 80, no reflow2.23–2.27 ms
Split grid resize 253 to 80, reflow enabled496–497 µs
Dense grid resize 253 to 80, reflow enabled2.25–2.30 ms

The grid-only resize fixture contains short, unwrapped rows. Enabling reflow here does not force text to wrap. Full-width text in terminal_history exercises actual splitting and joining. The eviction case includes the first prefix-pool metadata allocation; sustained ingestion measures the warmed path. The conversion/eviction cases exclude fixture construction and destruction, but still include moving ownership through the benchmark closure; they are not cycle-level instruction measurements.

Memory and recycling ​

The populated production-pane probe at 253×64 and 5,000 short history rows retains 4,729,837 bytes, down from 33,527,589 bytes in the dense baseline. That is 4.51 MiB versus 31.97 MiB, an 85.9% reduction. At 80×24 it retains 4,041,088 bytes versus 12,061,576 bytes, a 66.5% reduction. Adjacent capacities 4,999/5,000/5,001 increase smoothly in the split design instead of allocating another 1,000 dense rows.

A separate lifecycle probe uses the longer short log record corpus. After filling and warming history, another 10,000 lines perform zero allocation calls. Retained bytes stay at 5,325,696. Narrowing to 80 columns peaks at 6,050,016 bytes and widening afterward to 320 peaks at 5,957,056. Those peaks include the terminal's existing retained allocation, not just incremental scratch space. The direct dense baseline also makes zero allocation calls during steady ingestion, but retains 39,583,136 bytes. Its narrow/wide resize peaks are 39,745,184 and 64,548,184 bytes, with 5,132 and 5,133 allocation calls respectively. The split uses 322 and 201 calls for those same resize stages.

Dense per-cell-colored history retains 33,771,288 bytes in the split prototype versus 39,583,136 bytes in the direct dense baseline. That smaller saving comes from storage allocation/layout rather than compacting the full-width colored cells. Compact storage cannot promise the short-line savings for output that fills the width with different cells. Collected and spooled styled replay both emit 21,286,502 bytes. Spooling's requested-heap peak is 262,441 bytes versus 50,331,860 for collection, matching the existing streaming-export behavior.

Reproducibility ​

The preparation script applies the archived patches to fresh copies of the released crates. A reconstruction check compared all 756 Rust source files and manifests against the tested source trees byte for byte. The script does not edit registry sources, Rozi, or the sibling framework.

The earlier shared-row design remains in the September 10 report. Its performance results must not be attributed to this split implementation. Initial split measurements also precede the reader and resize improvements.

Decision ​

Keep the benchmark suite and reproducible split prototype. Do not enable automatic per-row compaction in production yet. The live-grid boundary works: sustained dense output is essentially unchanged, and explicit history reads make short-line search/copy/replay faster. Bounded recycling also eliminates steady short-line allocation calls.

The remaining cost is substantial on the workloads that save the most memory. At 253 columns, 100,000 plain lines take 24.18 ms instead of 19.42 ms; Unicode takes 42.61 ms instead of 34.61 ms. These are about 4.8 ms and 8.0 ms of extra processing per 100,000 lines per terminal representation. A server and attached clients each process output, so the trade-off repeats across those copies.

This makes the next question a policy question as well as a representation question. A bounded window of dense recent history with less frequent compaction, or compaction driven by a memory budget, could avoid converting every outgoing row immediately. Neither approach has been measured here. They must preserve history semantics and earn their complexity with sustained-state evidence. The current results do not justify shipping a maintained engine fork as an unconditional default.

Validation and limits ​

The split engine passes 138 unit tests, the 400-step differential test, 45 unchanged reference recordings, and its doctest. Focused additions cover exact Unicode/style/hyperlink round trips, bounded pool reuse, dense fallback, mixed-direction iteration, and direct default-tail resizing. The differential test also varies the history limit, including zero.

With the split engine, the framework passes all 132 terminal unit tests, cargo check --features terminal, and cargo clippy --features terminal. Rozi's 27 terminal-filtered library tests pass. Strict engine Clippy reports three existing question_mark findings in upstream src/tty/unix.rs. A repeat with only that lint forced to warnings passes with no other findings.

Final validation with released dependencies passes all 1,933 Rozi tests, formatting, strict all-target Clippy, and cargo build. The documentation site builds successfully, and git diff --check passes. The full test suite needed an unrestricted rerun because the sandbox blocked Unix sockets and isolated runtime paths. Final Clippy/build checks completed on September 14. No dependency override or lockfile source change remains; the sibling framework is unchanged.

These are Linux results. No cross-platform run, long multi-client soak, or separate active-screen rendering throughput measurement was performed. The explicit grid read iterator still supports both live and historical rows for scrolling/rendering. Dense live indexing alone does not prove that every rendering consumer is free of representation-related overhead.

MPL-2.0