Skip to content

Element-tree allocation in view and layout ​

Performance index · Benchmark guide · Reproduction playbook · Previous report

Verdict and scope ​

Verdict: five changes land a 32% cut; no architecture change needed.

The 2026-09-03 audit found the view and layout pass allocator-bound and read that as a property of tui-lipan's expansion architecture. Allocation accounting says otherwise. Expansion is the smallest of the three consumers, and the dominant costs are three specific, removable defects rather than the shape of the render model.

Three changes to tui-lipan and two to Rozi's view, none larger than a few dozen lines, cut view_layout/16 from 421 µs to about 285 µs and allocations per frame from 3,249 to 2,473, with no measurable change to ingest. No memoization, retained expansion, or memo_key work was needed, and none is recommended on this evidence.

This audit measures allocation counts and CPU for the view + expand + layout pass only. It does not measure draw, ingest beyond one regression guard, latency, or memory over time.

Measurement context ​

ItemValue
Rozi revision4b8f076, clean worktree plus benches/alloc_probe.rs
Framework revisionf4673d7 plus the three commits below, on perf/element-alloc
OSArch Linux, kernel 7.1.9-arch1-2
CPUAMD Ryzen 7 5700X3D, 8 cores / 16 threads
Profilerdhat 0.3.3, one bracketed frame; plus a counting GlobalAlloc
Criterionpaired baselines, --save-baseline / --baseline

The framework was consumed through the ignored .cargo/config.toml path override during this work. Cargo.lock was restored afterwards, and no committed manifest points at a path or git source.

view_layout/16 measures 415 µs on this tree against the 205 µs the 2026-09-03 report recorded at Rozi c9e9881. That is not a regression. Checking out c9e9881 with its own lockfile and running the same benchmark today gives 390.46 µs, so the historical figure does not reproduce on this machine. Ruled out as causes:

  • Rozi code. src/view/, src/app.rs, src/layout/ and src/state/ are byte-identical between c9e9881 and 4b8f076; the ten intervening commits are tests, CI, docs, a Windows lint fix, a macOS socket-path fix, and client-mailbox coalescing.
  • The framework. Both revisions lock tui-lipan 0.6.1 at the same checksum.
  • The benchmark. view_layout and backend_with_panes are unchanged; the only app_render diff is inbound_drain comments and the added inbound_fairness case.
  • The toolchain. c9e9881 measures 390.46 µs on rustc 1.97.1 and 394.30 µs on 1.90.0. The pass is not toolchain-sensitive, and the audit's 1.98.1 is no longer installed here.
  • CPU scaling. Governor and EPP are performance with boost enabled, as in the audit.

The ratio is also not a constant: 2.36× at 1 pane against 1.90× at 16. What remains is something about the original recording session that cannot be recovered from the report - which notes it ran inside an unrelated rozi pane on the same host, with perf_event_paranoid lowered for sampling.

Treat the 2026-09-03 absolute microseconds as unreliable and its ratios and attribution as sound. Every number below is a paired comparison taken on one machine within one session.

The 390 µs at c9e9881 against 415 µs at 4b8f076 is a further 6% between two binaries whose view and layout code is identical. That is code layout, not behavior, and is not worth chasing.

What a frame allocates ​

benches/alloc_probe.rs brackets exactly one TestBackend::render() with a counting allocator. Counts are exact and repeat run to run. peak_live is peak growth over the state the frame started in, not process RSS.

PanesAllocationsBytesPeak liveTree nodesAllocations per node
1316197,56171,3362413.2
2503307,34394,9603514.4
4944615,443163,4886514.5
81,8111,227,051290,56012514.5
163,2492,126,951506,64820715.7

Allocations and bytes are both linear in pane count, at about 196 allocations and 129 KB per pane. Every allocation is freed within the frame, so this is pure churn rather than growth.

The ratio is the finding. The expanded tree at 16 panes is 207 elements, and the frame makes 3,249 allocations to produce it. The question the previous audit left open - what is allocated once per element per frame - has the answer that nothing is allocated only once: the pass pays roughly sixteen allocations for every element it ends up with.

Where they come from ​

Instrumenting the phases of the pass with a thread-local bucket the counting allocator reads:

PhaseAllocations at 16 panesShare
AppRoot::view (Rozi's own view)1,59749.2%
Reconcile and layout1,27339.2%
expand_element2236.9%
expand_children, output vector1474.5%
Theme carve-out, sweep, other90.3%

This corrects the previous report. High allocator sample counts attributed there primarily to expansion in fact span application view construction and reconciliation; expansion accounts for only about 11% of frame allocations. Half the frame is allocated inside the host's view() before the framework sees an element at all, and most of the rest is reconciliation.

The correction is what kept this work cheap. Sampled CPU profiles name the frames that are hot; they do not say which caller asked for the memory, and reading ComponentRegistry::expand_children at the bottom of an allocator stack as "expansion is the problem" pointed at memoizing the smallest of the three consumers. Allocation accounting is what redirected the effort onto three defects that a sampling profile cannot distinguish from architecture.

Attributing the same frame by call site, with dhat over a debug-info build:

ShapeAllocationsShare of allocationsBytes
Element-sized blocks78724.2%1,048,392
Callback (Arc<dyn Fn>)2979.1%14,856
Key, and the format! string behind it2678.2%small
Everything else1,89858.4%1,063,703

The Key row's bytes are counted under "everything else"; its blocks are tens of bytes each and the interest is in how many there are, not how big.

Half the bytes in a frame are Element blocks, because Element was 1,248 bytes wide. It is as wide as the widest ElementKind payload, and ElementKind::Terminal(Terminal) inlined 1,152 bytes into every node - so a Spacer, a Divider, and a one-word Text each cost 1,248 bytes to move, box, or push.

Three specific defects account for most of the traffic:

DefectAllocationsBytesShare of bytes
reconcile_animated deep-clones its child subtree unconditionally396594,75228.0%
Single-child wrappers route through a Vec round trip~370329,47215.5%
expand_children allocates a fresh output Vec per container147259,58412.2%

The changes ​

All three are in tui-lipan. None changes a public API.

1. reconcile_animated borrows its settled child. It built an owned Element on every path, but only the height-animating branch rewrites a constraint and needs one. The settled branch was deep-copying the whole animated subtree every frame. A host that wraps each of its panes in an Animated - which Rozi does - paid that once per pane per frame.

2. ElementKind's five oversized variants are boxed. Terminal, Slider, Checkbox, Tabs, and ProgressBar now hold a Box. ElementKind drops from 1,152 to 520 bytes and Element from 1,248 to 616 bytes, halving every Box<Element> allocation, every Vec<Element> buffer, and every element move in the pass. It costs one extra allocation per use of those five widgets, against a saving on every node of every other kind. Seventeen match arms changed; deref coercion covered all but the five constructors.

3. A single-child expansion path. Thirteen wrapper kinds - Frame, MouseRegion, Animated, Center, Portal, Group, ThemeProvider, ContextProvider, Memo, EffectScope, DragSource, DropTarget, StatusBarLayout - held exactly one child and reached expansion through expand_children, which cost one Vec to pass the child in and another to carry it back. expand_single skips both. It also removes thirteen copies of a .pop().unwrap_or_else(|| Text::new("").into()) fallback that could never fire, so the change is net negative on line count.

Results ​

Criterion medians, each paired against the same saved baseline on one machine.

CaseBaselineAfter all threeChange
view_layout/148.413 µs37.297 µs−21.5%
view_layout/8238.20 µs166.30 µs−30.0%
view_layout/16421.32 µs289.03 µs−31.2%
inbound_drain/8_panes/642.7827 ms2.8156 msno change (p = 0.15)

Cumulative, in the order applied:

Stepview_layout/16vs baselineMarginal
Baseline421.32 µs——
+ reconcile_animated borrow345.57 µs−16.7%−16.7%
+ boxed ElementKind variants304.59 µs−26.6%−11.9%
+ expand_single289.03 µs−31.2%−5.1%

Allocation counts over the same steps, at 16 panes:

StepAllocationsBytesPeak live
Baseline3,2492,126,951506,648
+ reconcile_animated borrow2,8531,532,199506,648
+ boxed variants and expand_single2,665928,487387,896
Change−18.0%−56.4%−23.4%

The expansion phase specifically fell from 371 allocations to 190, a 48.8% cut, and now accounts for 7.1% of the frame.

Each step cleared the 10% keep threshold on view_layout/16 except expand_single at 5.1%, which is kept because it removes a clearly pathological allocation pair per wrapper node and reduces code.

Two Rozi-side follow-ups ​

With expansion no longer the story, the largest remaining allocation class was Rozi's own: 557 blocks of a frame, 20.9%, spent rebuilding identity strings the application already knows are fixed. Two of those sites are pane-derived and constant for a pane's life, and both were measured against their own paired baselines under a lower keep threshold, since each is a local cleanup rather than a subsystem.

ChangeAllocations removedview_layout/16
Cache body and terminal keys on Pane64−2.25%
Cache the four per-pane chrome animation keys128−4.82%

The second is worth more per key because it removes the format! as well as the Arc<str>. Both needed no framework change: Element::key and Context::animated_color already take impl Into<Key>, and cloning a cached Key is a refcount bump.

Where the pass ended ​

StartEndChange
view_layout/16421.32 µs~285 µs−32%
Allocations per frame at 16 panes3,2492,473−23.9%
inbound_drain/8_panes/642.7827 ms2.8156 msno change (p = 0.15)

No retained tree, no memoization, no caching subsystem, and no allocator swap. Note also that ordinary pane output does not pay this cost: a redraw repaints without rebuilding the view tree, so this is the price of a full structural frame rather than of every frame.

cargo test passes in both repositories - 2,708 framework unit tests plus its integration suites, and Rozi's 1,778, the latter against the released tui-lipan 0.6.1 that Rozi's manifest names - along with cargo clippy --all-targets -- -D warnings and cargo fmt --check.

What allocation counts are, and are not, evidence for ​

Two paired experiments in this pass moved CPU almost exactly in proportion to the allocations they removed: −2.40% of allocations bought −2.25% of view_layout/16, and −4.92% bought −4.82%. That is what makes allocation counts a useful prioritization signal for this workload - the allocator cost here is per call rather than per byte, so a small allocation is worth removing too.

It is not a CPU model, and this report should not be cited as though it were. The sites differ in what they actually do: an Arc<str> from a &str is an allocation and a copy, a formatted key is that plus a String and the formatting itself, and the ratio was not uniform even across pane counts within one experiment. Call-site inspection and a paired CPU measurement remain the acceptance test; the count only decides what to look at first.

The cost of skipping that discipline showed up in this pass. A shape tag in the profile suggested 48 blocks a frame were Key::from(&'static str) in the hot path, which implied an easy LazyLock<Key> win. Reading the actual call sites at the profiled revision showed they were Text::new, a ThemeProvider theme clone, and two empty Text widgets - not keys at all. The optimization did not exist. Only 23 .key( sites exist in the whole view, and nearly all sit in overlays that are closed during the benchmark.

What was not done, and why ​

  • memo_key, subtree reuse, retained expansion. Step 2 of the plan settled this: expansion is 11.4% of frame allocations, and cheap structural fixes took the whole pass down 31%. Retained expansion would target the smallest of the three consumers at by far the highest risk.
  • A process-wide allocator swap. The 2026-09-03 audit measured it as a net regression whenever output flows. That conclusion stands and this work does not revisit it.
  • Pre-sizing or recycling child vectors beyond expand_single. The remaining expand_children output vectors are 57 allocations per frame at 16 panes, 2.1% of the total. A Vec pool on the registry would remove most of them for a persistent field and a checkout/return protocol, which is not worth 2%.
  • Bisecting the 415 µs vs 205 µs gap against the previous report. Resolved under "Measurement context" without bisecting: the historical revision measures 390 µs today, so there is no regression to find.

Backlog, deliberately not worked ​

The pass stopped while sites still existed, because what remained was individually small or wanted plumbing out of proportion to it. Each of these needs a profile of its own before it is touched again, not a search for allocations because allocations exist.

  • Text::new("") allocates an Arc<str> for the empty string, 24 times a frame, for the invisible drag targets in the resize strips. Worth folding into framework work that already touches Text - a shared empty Arc<str>, or Spacer at the call sites - not worth a change of its own.
  • pane_window_key, 32 allocations a frame. It depends on pty_generation, a public field written from 37 places, and a cache behind a field anything may assign is a correctness trap. Revisit only if that field gets encapsulated for an independent design reason; a generation-validating lazy cache is complexity this has not earned.
  • About 166 blocks the profile could not attribute past its backtrace depth. Ignore unless a future CPU profile points there on its own.
  • WorkspaceLayer and titled formatting, 80 blocks between them. Same rule: profile first.

Reproducing ​

bash
cargo bench --bench alloc_probe                        # allocation counts, no framework hook
cargo bench --bench app_render -- --save-baseline base '^app_render/view_layout/(1|8|16)$'
cargo bench --bench app_render -- --baseline base      '^app_render/view_layout/(1|8|16)$'
cargo bench --bench app_render -- --baseline base      '^inbound_drain/8_panes/64$'

Phase attribution and the Element size table need a hook inside the framework, which is not part of the committed set; see the header of benches/alloc_probe.rs for what to add. Call-site attribution used dhat against a CARGO_PROFILE_BENCH_DEBUG=true CARGO_PROFILE_BENCH_STRIP=false build, with the profiler started after setup so only the frame under test was recorded.

MPL-2.0