Compact history ingest decision
The engine benchmarks left compact history with one open question. Converting short lines into compact rows made sustained engine ingest 12–25% slower. That number measures the engine alone. This run asks what the conversion costs in Rozi itself, with a session server, attached clients, and real panes, next to the memory it saves.
It is the last compact-history experiment. Batching, shared rows, and renderer changes are settled in the earlier reports, and nothing here revisits them.
Candidate and baseline
- Dense baseline: Rozi
52cccdc(tui-lipan 0.9.1 from crates.io,alacritty_terminal0.26.0). - Candidate: the same Rozi source built against the render-history engine and its framework changes, ported to released tui-lipan 0.9.1. The bulk
RowPartssnapshot builder now uses the released attr-keyed run helpers. The port passes all 146 framework terminal unit tests. It is archived with a byte-identical reconstruction check intools/experiments/ingest-decision.
Both are release builds. AMD Ryzen 7 5700X3D (8 cores, 16 threads), Linux 7.2.3, Rust 1.98.1, performance governor. The machine was not idle: a browser tab held about one core throughout, and load average stayed near 5.
What was measured
Each run starts an isolated session server and one or two rozi sessions attach clients under script at 253×64, with scrollback = 5000, animations off, and shell integration off. Extra panes come from split, so 4 and 8 panes tile the same screen. When every pane is ready, each one cats the same corpus once:
plain: 2,000,000 lines ofrozi-0000000 plain(19 bytes each)log: 1,000,000 styled log lines, about 76 visible columns and 94 bytes with SGR
The run records server and client CPU time from /proc/PID/stat from the moment output starts until both processes use no more than one tick in 300 ms. It then waits one second and samples PSS and RSS. The child cat and shells are not counted. Dense and compact runs alternate, and the order flips on every repetition.
CPU comes in 10 ms ticks. The smallest case uses about 2 s of CPU, so one tick is 0.5%.
Results
Medians. d/c is dense then compact. The log rows at 1 and 8 panes combine the five matrix repetitions with ten more run for those two rows alone. Every other row has five.
| Panes | Content | Clients | Runs | Server CPU s d/c | Client CPU s d/c | App CPU change | Server PSS MiB d/c | Client PSS MiB d/c | App PSS saved |
|---|---|---|---|---|---|---|---|---|---|
| 1 | plain | 1 | 5 | 1.38 / 1.42 | 0.66 / 0.71 | +90 ms (+4.4%) | 39.8 / 11.6 | 47.5 / 20.6 | 55.1 MiB |
| 4 | plain | 1 | 5 | 9.52 / 8.15 | 1.92 / 1.77 | −1520 ms (−13.3%) | 72.0 / 38.4 | 71.1 / 36.8 | 67.9 MiB |
| 8 | plain | 1 | 5 | 12.89 / 13.15 | 2.59 / 2.71 | +380 ms (+2.5%) | 80.4 / 49.4 | 102.0 / 64.0 | 69.0 MiB |
| 8 | plain | 2 | 5 | 12.74 / 12.57 | 5.85 / 5.52 | −500 ms (−2.7%) | 80.2 / 47.9 | 199.9 / 110.2 | 122.0 MiB |
| 1 | log | 1 | 15 | 1.94 / 2.09 | 0.11 / 0.10 | +140 ms (+6.8%) | 45.7 / 27.2 | 45.0 / 25.4 | 38.0 MiB |
| 4 | log | 1 | 5 | 8.33 / 8.50 | 2.00 / 2.13 | +300 ms (+2.9%) | 65.2 / 62.3 | 92.0 / 85.1 | 9.8 MiB |
| 8 | log | 1 | 15 | 19.98 / 20.24 | 4.49 / 4.70 | +470 ms (+1.9%) | 82.2 / 82.0 | 116.6 / 111.8 | 5.0 MiB |
Client PSS with two clients is the sum of both.
CPU
For short plain lines the difference stays inside run-to-run spread. At 4 panes, compact is 13% faster on the median, which no mechanism here explains. It shows how wide the spread is, not a speed-up. Across plain cases the change runs from −13% to +4%, at most 0.4 s of CPU for 16 million lines.
The styled log at one full-width pane is the one clear cost. Its 15-run server ranges do not overlap (dense 1.89–2.23 s, compact 2.01–2.33 s). The difference is about 140 ms per million lines, or 0.14 µs a line, a +7% change in application CPU.
The first five 8-pane log runs showed +15%. Ten more cut that to +1.9%, so the first figure was noise.
The engine benchmark's 12–25% does not reach the application. Rozi spends most of its ingest CPU reading ptys, framing output for clients, and rendering, and conversion is a small share of that.
Memory
Short lines are where compact history pays. One full-width pane saves 55 MiB of application PSS (28 MiB server, 27 MiB client). Two clients on 8 panes save 122 MiB. Each client holds its own terminal representation, so the saving scales with clients as well as panes.
Styled logs in tiled panes save almost nothing, which is expected. A row compacts only when its content fits in half the pane's width. A 76-column log line needs a pane 152 columns wide, and splitting a 253-column screen leaves no such pane. Those rows fall back to dense after a scan that stops at the first cell. At 4 and 8 panes the candidate pays its small CPU cost for 5–10 MiB.
Two exit bugs found on the way
The first matrix was invalid. Its cleanup killed each client's script wrapper, and the orphaned clients kept running at 55% of a core apiece until 35 had piled up. Reproducing that uncovered two bugs in released tui-lipan 0.9.1. Both trigger when the terminal goes away while a client is exiting, for example a window closed or an ssh connection dropped during the detach after SIGHUP.
- The client spins forever. crossterm 0.29's Unix event source retries a read that returns end-of-file without checking its timeout, and a hung-up pty returns exactly that. A core dump placed the spin in the exit view's cursor-position query, which ran about 2.9 million
readcalls a second on the UI thread. crosstermmasterstill has the loop. - The client aborts. ratatui's terminal
Dropreports a failed cursor restore witheprintln!. Printing to a dead stderr panics, andpanic = "abort"turns that into a core dump.
A tui-lipan fix is ready, not yet released. Host reads wait on the descriptor with poll and stop at POLLHUP, the exit view reads the cursor row the same way, and the runner's terminal cannot panic when dropped. Against eight kill timings it leaves no spin, no core dump, and no surviving process, and the exit view's bytes on a live terminal are unchanged. Rozi picks it up with the next tui-lipan release.
The repaired harness signals only processes it started, checked by executable path and start time, and reaps clients before their wrappers. It refuses to run while any experiment process survives.
Decision
Technically, go. The conversion cost the engine benchmarks flagged is too small to see in the application for short lines. Its one clear cost is 0.14 µs a styled line at full width. In return, short-line history uses 55 MiB less per full-width pane and client pair, and 122 MiB less at 8 panes with 2 clients. Nothing else needs optimizing, and profiling the conversion would not change the answer.
What adoption costs is maintenance. The candidate is a 4,900-line patch to alacritty_terminal, plus a snapshot path in tui-lipan that reads its row API. Shipping it means publishing and maintaining an engine fork, or getting the history split accepted upstream. deny.toml allows no git sources, and a [patch] would be dropped when tui-lipan publishes. That is a product decision, not a performance one, and this report does not make it.
If the fork is not taken on, the experiment stays archived as a proven design. The measurements above are the evidence to reopen it with.
Not measured: resize and reflow in the application, which the engine benchmarks showed faster, and remote sessions.