Terminal history storage investigation
The compact-history prototype saves substantial memory, but does not pass the ingest performance gate. It is preserved as an experiment and is not enabled in Rozi. Replacing the terminal engine or moving all client terminal state to the server is not justified by this result.
This report includes source review, allocation measurements, a two-pass storage prototype, and ingest comparisons. It is not a leak investigation. No production runtime or dependency changes accompany the report. The prototype and reproduction instructions include engine and framework patches, tests, and a script that prepares disposable source copies.
Scope
- Rozi
824d925, initially clean, using releasedtui-lipan 0.9.0andalacritty_terminal 0.26.0. - Sibling tui-lipan
8084b12, initially clean, inspected without a dependency override. - Linux
7.2.3-arch1-3, x86-64, Rust1.98.1. - Existing
cargo bench --bench terminal_memoryprobe, optimized build with a counting allocator. Its corpus contains short numbered lines with foreground SGR colors.
The earlier 83,056 KiB application PSS result belongs to the September 5 experiment. It is not a fresh process measurement of this revision.
Current allocation results
The existing probe completed successfully. These are retained requested heap bytes for one production client pane, not RSS or PSS. Corpus construction happens outside the measured region.
| Viewport | History limit | Retained bytes | MiB |
|---|---|---|---|
| 80×24 | 0 | 2,199,976 | 2.10 |
| 80×24 | 5,000 | 12,061,576 | 11.50 |
| 253×64 | 0 | 2,897,029 | 2.76 |
| 253×64 | 4,999 | 33,515,133 | 31.96 |
| 253×64 | 5,000 | 33,527,589 | 31.97 |
| 253×64 | 5,001 | 39,587,133 | 37.75 |
The width penalty remains despite the corpus having only 11 printable characters per line. The 5,001-row case also confirms that Alacritty's allocation blocks remain relevant away from the exact boundaries handled by Rozi. The direct, unprimed 253×64 / 5,000 constructor retains 39,583,136 bytes, confirming that Rozi's existing priming optimization still helps.
The separate full-width, per-cell-color replay case exported 21,286,502 bytes. Collecting replay in a vector peaked at 50,331,860 additional requested heap bytes; streaming it to a temporary file peaked at 262,441. The existing streaming work addresses a different cost from retained grid storage. These counts do not include kernel file-cache memory.
Where the bytes go
Rozi retains an authoritative terminal on the server and a terminal on each attached client. Each uses Alacritty's grid. The dominant text cost therefore multiplies with pane count, retained rows, terminal width, and number of clients. Parked attachments also retain their terminal state.
Alacritty's grid/row.rs stores Row<T> as a Vec<T> plus an occupied-cell bound. Construction fills the whole width. That occupied bound accelerates resetting rows; it does not make their allocation sparse. term/cell.rs stores a character, foreground, background, flags, and an optional shared extra object in every cell. Extra objects hold combining characters, underline color, and hyperlinks, so plain ASCII does not need one allocation per character.
At 24 bytes per cell on this target, 5,000 history rows at 253 columns require 30,360,000 bytes of cell arrays per terminal, about 29 MiB. This excludes visible screens, row metadata, cached rows, graphics, snapshots, and application state. Two such histories require about 58 MiB before those other costs. Short lines pay for all 253 columns.
Rozi already avoids the extra 1,000-row allocation at exact history boundaries. The framework's eviction ledger also reserves viewport-sized headroom so it can count dropped lines exactly. Removing that headroom without replacing the accounting would break absolute history positions.
Preferred design to prototype
This section records the design proposed before the implementation experiment. The measured outcome and remaining issues follow below.
Keep the mutable viewport dense. When a row enters history, store an exact cell prefix plus a repeated trailing cell and its count. Store logical width separately from stored prefix length. Only omit cells that are exactly equal to the tail cell, including colors, flags, and extra data. Whitespace with a background, a hyperlink, or a wrap flag is not disposable blank space.
Use a dense fallback when compaction would not save enough. Begin with tail compaction, not a general compression codec or global style interning. It targets the known waste while preserving direct access to the non-repeated cells. Dense, randomly styled output should remain dense.
As an illustration, 11 stored cells plus one tail cell would consume 288 bytes of cell payload instead of 6,072 bytes for a 253-column row. Across 5,000 rows that is about 1.37 MiB instead of 28.95 MiB. This is a payload calculation, not a predicted process result: row metadata, allocation rounding, conversion buffers, and dense fallbacks still need measurement.
History reads should provide cell iteration and indexed access without expanding whole rows. Reflow needs explicit row operations that preserve soft wraps, wide-cell spacers, and combining characters. Rows returning to the live viewport must materialize correctly. A bounded pool of dense rows can avoid allocating a fresh viewport row on every scroll, but it must not retain one full-width allocation per compact history row.
Why this needs engine work
Alacritty's row API returns contiguous immutable and mutable slices, and its grid resize code splits, appends, and mutates history rows. A compact row cannot implement all of those contracts without materialization. The design needs changes to those APIs and their callers, or a separate history store integrated into the grid's scrolling and reflow operations.
Putting another compressed copy next to TerminalScreen would leave the original grid allocated. Removing rows from Alacritty after copying them would require rebuilding its history semantics, including resize, selection, search, replay, and eviction lineage. That is a terminal-engine change disguised as a framework optimization.
The appropriate boundary is therefore Alacritty grid storage, with tui-lipan adapting its history readers and preserving TerminalScreen behavior. Prefer an upstreamable change. A maintained engine fork adds ongoing compatibility work and needs measured savings before adoption. Merely pointing Rozi at the sibling tui-lipan checkout does not cross this boundary.
Decision criteria
The following are proposed acceptance criteria, not measured results:
- At least 40% lower retained terminal allocation for 253-column, 5,000-line short-log workloads. Confirm the benefit with application PSS for a server and one or more clients.
- No more than 5% ingest-throughput regression on repeated controlled runs, including full-width text, SGR-heavy text, Unicode, and sustained output after history reaches capacity.
- Search, replay export, and resize/reflow must be measured too. Reject a design that expands the entire history for these operations or creates history-sized temporary copies.
- Dense history should stay within 5% of baseline retained allocation. Test arbitrary per-cell styles and hyperlinks so the fallback is exercised.
- Differential tests against the existing engine must preserve visible cells, history, cursor, selection text, replay results, and eviction lineage after mixed input and resize sequences. Include background-colored blanks, alternate screens, wide characters at wrap boundaries, combining characters, scrolling regions, and history-limit changes.
These criteria make a bounded engine-storage experiment worthwhile. They do not yet justify shipping a fork. A smaller cell layout could help dense workloads later, but offers less upside for short logs and may replace direct style fields with additional lookup costs.
Prototype outcome
The engine prototype stores logical row width separately from a cell vector whose last cell repeats through that width. Rows entering history compact only when at least half their cell allocation can be released. Dense rows remain dense. History iteration and indexed reads do not allocate. Mutations materialize a compact row when needed; ordinary blank-tail growth and shrink avoid that expansion. Reflow compacts completed rows before retaining them.
The implementation also replaces Alacritty's unsafe, fixed-size row swap with a safe vector swap. Its existing serialized grid format is preserved. The framework adapter changes only two mutable row-slice call sites. The prototype changes Alacritty's public row API, but does not change Rozi's history limit, server ownership, or client parsing architecture.
The first allocation pass reduced the 253×64 / 5,000 production pane from 33,527,589 to 4,670,485 retained bytes, about 86%. At 80×24 / 5,000 it fell from 12,061,576 to 3,965,744 bytes, about 67%. The final revised prototype reproduced those retained-byte counts exactly. This is requested heap allocation, not process PSS. The wide short-line case clearly has enough recoverable memory to justify investigating storage representations.
The CPU gate failed in both passes. The second pass improved cell access and resize behavior, but a paired subset recheck of the already-built binaries still measured:
| Workload | Dense estimate | Revised prototype estimate | Criterion time increase |
|---|---|---|---|
| Scroll regions, 200×60 | 4.337 ms | 4.760 ms | 9.6% |
| Wide Unicode, 320×90 | 7.156 ms | 10.545 ms | 47.0% |
| Long lines, 80×24 | 10.901 ms | 11.961 ms | 9.7% |
Each recheck used 30 samples, one second of warm-up, and a two-second requested measurement window. Criterion extended the window where needed. These are fresh-terminal ingest-and-drop benchmarks, not isolated steady-state ingest. No Cargo build or other benchmark ran concurrently. This was a live development machine, not an isolated benchmark host.
The first full-matrix dense Unicode estimate was 4.192 ms, whereas the later subset estimate was 7.156 ms. That sensitivity means the earlier roughly 153% regression should not be treated as a stable effect size. The paired subset still fails the gate by a wide margin. Some plain-text cases improved, but that does not offset the regressions for adoption purposes.
Correctness checks passed: 137 engine unit tests, all 45 upstream reference recordings, one 400-step differential test against unchanged registry Alacritty, and one doctest. The differential test compares retained cells, colors, flags, combining characters, hyperlinks, cursor position, and display offset after mixed output and resize operations. Passing these checks is useful evidence, not a claim of exhaustive terminal compatibility.
Decision
Do not adopt this implementation or commit a dependency on an engine fork. The temporary framework edits and Cargo overrides are removed at handoff; the patches remain reproducible.
The shared row representation puts storage checks in live cell access and adds conversion work when rows enter history. The measurements do not isolate how much each contributes. A further experiment should separate the dense viewport's hot operations from compact history access and measure bounded recycling of dense row allocations. It should also isolate sustained ingest, search, replay, and reflow memory before repeating the application PSS matrix. That larger design has not been implemented or validated here.
Changes not warranted here
Reducing configured history saves memory by changing user behavior. Reducing server history also changes what attach and resurrection can restore. Neither is a transparent improvement.
Sending server-rendered rows to clients could eliminate duplicated parsing, but introduces remote history requests, cache eviction, bandwidth and latency policy, and new selection/search behavior. That is a much larger project than improving storage shared by both sides.
Shared memory would address only local duplication and would add platform and lifetime management. An allocator replacement cannot remove the full-width cell arrays. Neither is the first move.
Validation and limits
The allocation probes, documentation build, Rozi formatting check, and whitespace checks passed. The archived patches were applied to fresh disposable copies with zero fuzz, and all 184 engine tests and doctests passed again from that reconstructed copy. The preparation script passed its syntax check and successfully prepared the copies without changing the source checkouts.
Strict engine Clippy failed on three question_mark findings in unchanged upstream src/tty/unix.rs with Rust 1.98. A second run forced that specific lint to warnings and passed with no other findings. The upstream file was verified byte-for-byte unchanged; it was not edited to make this experiment's checks pass.
No production Rust changes are retained, so the full Rozi application test suite and application Clippy were not rerun. No process PSS matrix, long-running leak soak, or complete search/replay/reflow performance comparison was performed, because the prototype already failed the ingest acceptance gate.