Pane retained-memory audit
Performance index · Benchmark guide · Reproduction playbook
Verdict
A one-row allocation reserve removes a 1,000-row over-allocation at exact scrollback boundaries. At Rozi's default 5,000-line history, one 253×64 pane with one attached client fell from 96,686 KiB to 83,056 KiB application PSS: −13,630 KiB (−14.1%).
The defect was in terminal-grid allocation, not pane metadata or a leak. Alacritty grows each grid's row storage in 1,000-row blocks. tui-lipan needs temporary history headroom to count evictions exactly. When the exposed history is itself a multiple of 1,000, the first over-limit line can allocate another complete block which is then retained.
Rozi keeps two terminal grids for every pane shown by one client:
- the session server parses output for terminal replies, runtime metadata, replay, and resurrection;
- each attached client parses the same output for rendering and client-local interaction.
The unnecessary block was therefore paid twice in the measured one-client case, and once more for each additional attached client. The fix applies to both server and client grids, preserves the full configured history, and never reconstructs a parser after output has reached it.
No evidence supports an application-level rewrite of terminal rows, a shorter server history, or removing the client parser. Those changes would either belong in the framework or break attach, replay, rendering, and resurrection behavior. Configuring less scrollback remains the direct way to trade history for additional memory savings.
Measurement context
| Item | Value |
|---|---|
| Rozi revision | 1e8f4a0, dirty worktree containing this audit and unrelated in-progress work |
| Framework | tui-lipan 0.7.3, alacritty_terminal 0.26.0 |
| OS | Arch Linux, kernel 7.1.9-arch1-2, x86-64 |
| CPU | AMD Ryzen 7 5700X3D, 8 cores / 16 threads |
| Toolchain | rustc/cargo 1.97.1 |
| Process metric | Linux /proc/<pid>/smaps_rollup, median of five samples 200 ms apart after 2 s quiescence |
| Allocation metric | bracketed System global allocator counters in an optimized benchmark |
Process results are PSS unless stated otherwise. PSS apportions shared mappings and is a better multi-process total than adding RSS. Child shells and workload processes are measured separately from the Rozi server and clients.
The before/after process comparison is one isolated launch per version, with five samples inside that launch rather than five independent cold launches. The effect is much larger than sample variation and is independently confirmed by exact retained-allocation counts.
Architecture and dominant costs
The steady-state memory shape is:
one server grid per pane
+ one rendering grid per pane per attached client
+ viewport snapshots and pane/application state
+ decoded graphics only when a child sends themTerminal rows dominate populated text panes. Alacritty rows own a full cell array at the current column count, so history cost scales primarily with:
panes × retained rows × columns × (server + attached clients)Styles do not create a second history representation; they are cell attributes. Plain and styled process cases therefore had no stable retained-memory difference after quiescence.
Existing safeguards were confirmed:
- server terminal screens disable image storage;
- client screens cap decoded Kitty graphics at 32 MiB per pane instead of the framework's 96 MiB standalone-terminal default;
- output and control queues are bounded;
- closing panes drops their parser, while detached named sessions intentionally retain the server-side parser so they can be reattached.
Secondary contributors are bounded or lifecycle-dependent rather than part of the steady text-grid result:
- every PTY has reader and exit-wait threads in the session server;
- parked client attachments keep their complete pane screens live and continue parsing output for instant switch-back;
- attach seeding holds a 256 KiB replay spool plus a 4 MiB send window and bounded live catch-up;
- resurrection exports dirty pane replay into a background snapshot job;
- closing animations can retain a retired pane briefly before pruning.
Each named session also has its own server process baseline. That makes session count a separate scaling axis from pane and attached-client count.
Reproducibility fixes
The Linux memory matrix had three fidelity defects which were corrected before drawing conclusions:
- It inherited
ROZI_PANE,ROZI_SOCKET, and extension variables when launched from inside Rozi, allowing isolated control commands to target the developer's live session. - Its three-character readiness marker could not survive a 2×1 pane with zero history.
- A “1,000 history” workload emitted only 1,000 lines, failing to fill both the viewport and the requested history. Text cases now emit
history + rows + 1lines.
Child-process discovery now walks the /proc parent graph and records process start times. A validation case measured 674 KiB PSS across the workload shell and its sleeping child instead of incorrectly reporting zero. This does not change application PSS, which always excludes that group.
Root cause
tui-lipan creates an Alacritty terminal with an internal history ceiling of:
configured scrollback + viewport rowsThe extra rows let one terminal handler operation scroll a whole viewport while the ledger still observes exactly how many lines were evicted. After an operation, history is trimmed to the configured limit and the internal ceiling is restored.
Alacritty's storage grows in batches of 1,000 rows and retains up to 1,000 inactive rows. At a 5,000-line limit, the first temporary line above the exposed limit used to cross the next storage boundary. Trimming that line left exactly 1,000 inactive rows, so Alacritty's > 1,000 truncation condition correctly retained the entire block.
The adjacent process control made the discontinuity visible:
| 253×64, one pane, one client | Application PSS |
|---|---|
| 4,999 history lines | 84,886 KiB |
| 5,000 history lines, before | 96,686 KiB |
| Difference for one logical line | +11,800 KiB |
One logical line cannot account for 11.5 MiB. Two extra 1,000-row blocks at 253 columns can.
Fix
At a scrollback limit divisible by 1,000, Rozi constructs the terminal one viewport row taller and immediately shrinks it. The logical viewport remains unchanged, while Alacritty retains one blank row which the ledger can reuse for the temporary over-limit line.
Pane creation initially uses a 120×32 fallback before layout reports the authoritative geometry. A later width resize rebuilds Alacritty's rows and discards that reserve. The final implementation therefore tracks whether each parser has seen output:
- before output, a server resize reconstructs the empty parser at the authoritative dimensions;
- before output, the client does the same even if
SpawnResultalready marked the pane ready; - after output, both sides use ordinary resize and reflow, preserving all terminal state.
Restored server panes are marked non-empty immediately, so replayed state can never be discarded by the optimization.
Exact retained-allocation results
terminal_memory compares an unprimed TerminalScreen with the production pane lifecycle. The production numbers include pane metadata, so the important comparisons are the discontinuities between adjacent history limits.
| Viewport | Constructor | 4,999 live bytes | 5,000 live bytes | Boundary jump |
|---|---|---|---|---|
| 80×24 | unprimed | 12,053,427 | 13,973,427 | +1,920,000 |
| 80×24 | production | 12,057,384 | 12,061,536 | +4,152 |
| 253×64 | unprimed | 33,511,136 | 39,583,136 | +6,072,000 |
| 253×64 | production | 33,515,093 | 33,527,549 | +12,456 |
The same pattern appears at zero and 1,000 lines:
| Viewport / history | Unprimed live bytes | Production live bytes | Difference |
|---|---|---|---|
| 80×24 / 0 | 4,144,051 | 2,199,936 | −1,944,115 |
| 80×24 / 1,000 | 6,096,819 | 4,151,936 | −1,944,883 |
| 253×64 / 0 | 8,984,800 | 2,896,989 | −6,087,811 |
| 253×64 / 1,000 | 15,090,848 | 9,000,989 | −6,089,859 |
At 5,001 lines, production again needs the next block; the optimization does not pretend to remove memory required by a capacity beyond the boundary.
Process PSS result
The saturated default-history scenario was rerun with the same viewport, workload, client count, sampling, and host:
| 253×64, one pane, 5,000 styled lines, one client | Before | After | Change |
|---|---|---|---|
| Client PSS | 53,801 KiB | 46,686 KiB | −7,115 KiB |
| Server PSS | 42,885 KiB | 36,370 KiB | −6,515 KiB |
| Application PSS | 96,686 KiB | 83,056 KiB | −13,630 KiB (−14.1%) |
| Current application RSS | 111,088 KiB | 97,636 KiB | −13,452 KiB |
| Anonymous memory | 82,164 KiB | 69,008 KiB | −13,156 KiB |
| Private memory | 88,024 KiB | 74,232 KiB | −13,792 KiB |
The client and server reductions are similar, as predicted by the duplicate-grid architecture. The after result is also within 1.8 MiB of the 4,999-line control despite preserving the extra line. Exact allocation counts explain the remaining process-level variation more reliably than an additional PSS decimal would.
Scaling matrix
The corrected quick matrix covered 24 combinations:
- terminal canvas: 80×24 and 253×64;
- pane count: 1, 4, and 8;
- history: 0 and 1,000;
- content: plain and styled;
- one attached client.
Representative application PSS:
| Canvas / history / content | 1 pane | 4 panes | 8 panes |
|---|---|---|---|
| 80×24 / 0 / styled | 24,084 KiB | 29,179 KiB | 39,211 KiB |
| 80×24 / 1,000 / styled | 24,421 KiB | 54,414 KiB | 79,522 KiB |
| 253×64 / 0 / styled | 35,391 KiB | 42,724 KiB | 46,760 KiB |
| 253×64 / 1,000 / styled | 36,108 KiB | 73,870 KiB | 108,278 KiB |
These are realistic tiled-canvas totals, not fixed-size per-pane slopes: splitting one canvas makes each additional pane narrower and/or shorter. The allocation benchmark holds each pane's viewport fixed and is the appropriate evidence for row-storage cost per pane.
Options considered
Kept
- Reserve one reusable row only at Alacritty's known 1,000-row boundaries.
- Apply it to server and client constructors.
- Reconstruct only parsers proven not to have consumed output.
- Keep an allocation probe beside the process harness so allocator behavior can be checked after framework upgrades.
Rejected
- Use 4,999 as the default. It avoids this allocator boundary accidentally, changes the user contract, and leaves zero/1,000/other configured boundaries broken.
- Reduce server history. This would truncate attach seeds and resurrection replay.
- Remove client parsing. Rendering, local scrollback, selection, search, damage, and graphics placement require client terminal state.
- Resize a populated terminal through an extra row. That changes reflow, cursor, history, and image semantics; the implementation instead proves the parser is empty before reconstruction.
- Swap the process allocator. The defect is deterministic row capacity, and a previous audit found allocator replacement regressed output workloads.
- Compress rows in Rozi. Alacritty owns the grid representation. Sparse-row or compact-cell work belongs upstream and needs terminal-ingest and resize benchmarks in addition to memory evidence.
Further opportunities
The next meaningful memory reduction would require framework or architecture work:
- Measure a compact/sparse Alacritty row representation against terminal ingest, reflow, and damage tracking. Wide panes with long history have the most leverage.
- Measure real multi-session switching workloads before considering a lower-history or suspended parser policy for parked attachments; the tradeoff is instant, fully current switch-back.
- Revisit direct backpressured replay only if the terminal grid gains an immutable or copy-on-write snapshot. Reading a live grid incrementally while PTY output mutates it cannot preserve the attach baseline invariant.
- Explore an optional lower-fidelity server replay store only if attach and resurrection contracts are deliberately redesigned; it is not a transparent optimization.
- Keep graphics stress testing separate from text history. The existing 32 MiB per-client-pane decoded-image budget is demand-driven and was not responsible for this text case.
- Consider documenting workload-specific scrollback recommendations. Lower limits save memory linearly once allocator boundaries are handled, but choosing the limit is user policy.
Attach-seeding follow-up
The old attach path exported every live pane and queued every encoded frame before the socket could drain. It raised the attaching client's outbox cap from 8 MiB to 64 MiB, so attach memory grew with the whole session snapshot and a valid snapshot larger than that cap disconnected the client.
Attach now records a pane manifest and advances it from Pending through Replaying, CatchingUp, and Live. The server exports no more than one pane per iteration, encodes at most 1 MiB of seed work per iteration, and retains at most 4 MiB of encoded seed frames in the socket outbox. The limit applies to bytes currently queued, not total replay bytes.
The ordering contract is:
For each pane and attaching client, one snapshot establishes the baseline. Output incorporated before that snapshot is not forwarded separately. Output produced after the snapshot is delivered once, after the complete snapshot replay.
Post-manifest control changes wait behind baseline replay, except a resize for a pane still pending. That resize reaches the client before export so its parser has the server's dimensions. Each pane also gets snapshot-point geometry immediately before replay. A resize during replay waits behind it. A pane closed or replaced before export is skipped; its lifecycle delta remains in the catch-up stream.
Live catch-up has an 8 MiB per-client limit. Crossing it disconnects only that attaching client. Heartbeat expiry pauses while replay is in front of the ping. Concurrent attachers share one rotating 1 MiB work budget and at most one pane export per server iteration.
Runtime metrics report current and peak seed queue bytes, current and peak catch-up bytes, panes remaining, cumulative replay bytes, durations, completions, disconnects, and the last disconnect reason. Unit tests cover both sides of the snapshot point, resize and close races, work and resident queue bounds, and catch-up overflow.
TerminalScreen::write_replay_bytes() now produces the same VT stream through an io::Write sink. Attach retains the first 256 KiB in memory; crossing that threshold moves the stream into an unnamed file under Rozi's private cache and the existing pump reads it back in 256 KiB frames. The file closes and disappears when replay completes or the client disconnects. If cache creation or writing fails, Rozi re-exports into memory so an environmental optimization failure cannot prevent attach.
This spool is necessary for exact ordering. A cursor over the live terminal grid would observe post-snapshot PTY mutations, while freezing the server parser until a slow client's 4 MiB network window drains would delay every live client. An immutable grid snapshot would merely replace the replay allocation with a larger cell allocation. The file preserves one atomic export without coupling terminal mutation to client speed.
On a style-dense 253x64 pane with 5,000 history rows, the 21,286,502-byte replay made the collected Vec peak at 50,331,860 allocated bytes because of growth reallocations. The buffered writer peaked at 262,441 bytes, a 99.5% reduction in process heap peak. For a 1,000-row, 120-column style-dense replay, collected export measured 95.0 MiB/s; buffered file creation, write, and read measured 88.2 MiB/s, a 7.4% throughput cost.
Physical-memory follow-up
The replay spool does not use /tmp on the measured Arch host. /tmp is a 16 GiB tmpfs, but Rozi's cache resolves under /home, a disk-backed Btrfs filesystem. A process probe built the same 253x64, 5,000-history style-dense screen, sampled its idle baseline, then retained either the replay Vec or the unnamed cache file. Values below are medians of three isolated cgroup-v2 runs:
| Retained replay | Process PSS delta | Cgroup anon delta | Cgroup file delta | Cgroup total delta | Shmem delta |
|---|---|---|---|---|---|
21,286,502-byte Vec | +21.6 MiB | +21.6 MiB | ~0 | +22.3 MiB | ~0 |
| Unnamed cache file | ~0 | ~0 | +20.3 MiB | +21.0 MiB | ~0 |
Absolute process peak RSS fell from 60.3 MiB to 38.9 MiB, and post-export PSS fell from 58.4 MiB to 36.8 MiB. The bytes did not move into tmpfs/shmem. They did remain charged as disk file cache while the replay file was open, so total cgroup memory fell only about 1.3 MiB in this held-state measurement. Global MemAvailable moved too much between runs to be useful; the cgroup split is the stable evidence.
Large pane replays are spooled to disk-backed cache storage rather than collected in anonymous heap memory on this host. This bounds Rozi's process working set and lets the kernel reclaim replay pages under memory pressure. It does not proportionally reduce instantaneous system-accounted memory while the replay remains cached. Immediately after export, 20.3 MiB of the file charge was still dirty. Requesting 32 MiB of cgroup reclaim wrote it back and reduced the open file's cache charge to 4 KiB. The live replay Vec cannot be reclaimed without swapping the process, so the spool remains worth keeping despite the similar no-pressure total.
Reproduction
cargo bench --bench terminal_memory
cargo bench --bench replay_export
cargo build --release --bench replay_memory
tools/memory-matrix.sh --case 64 253 1 4999 styled 1 \
--output target/memory-investigation/control-h4999
tools/memory-matrix.sh --case 64 253 1 5000 styled 1 \
--output target/memory-investigation/optimized-h5000
tools/memory-matrix.sh --quick \
--output target/memory-investigation/optimized-quick
tools/memory-matrix.sh --smoke \
--output target/memory-investigation/harness-smokeRaw outputs remain under ignored target/memory-investigation/.
Verification
- focused constructor, capacity, pre-output resize, and post-output preservation tests pass;
cargo bench --bench terminal_memorycompletes and reports the expected adjacent-limit shape;- the memory harness smoke lifecycle passes, including session shutdown and PTY-descendant cleanup;
cargo fmt --all -- --check,cargo test,cargo clippy --all-targets -- -D warnings, andgit diff --checkpass.