Element-tree allocation in view and layout
Performance index · Benchmark guide · Reproduction playbook · Previous report
Verdict and scope
Verdict: five changes land a 32% cut; no architecture change needed.
The 2026-09-03 audit found the view and layout pass allocator-bound and read that as a property of tui-lipan's expansion architecture. Allocation accounting says otherwise. Expansion is the smallest of the three consumers, and the dominant costs are three specific, removable defects rather than the shape of the render model.
Three changes to tui-lipan and two to Rozi's view, none larger than a few dozen lines, cut view_layout/16 from 421 µs to about 285 µs and allocations per frame from 3,249 to 2,473, with no measurable change to ingest. No memoization, retained expansion, or memo_key work was needed, and none is recommended on this evidence.
This audit measures allocation counts and CPU for the view + expand + layout pass only. It does not measure draw, ingest beyond one regression guard, latency, or memory over time.
Measurement context
| Item | Value |
|---|---|
| Rozi revision | 4b8f076, clean worktree plus benches/alloc_probe.rs |
| Framework revision | f4673d7 plus the three commits below, on perf/element-alloc |
| OS | Arch Linux, kernel 7.1.9-arch1-2 |
| CPU | AMD Ryzen 7 5700X3D, 8 cores / 16 threads |
| Profiler | dhat 0.3.3, one bracketed frame; plus a counting GlobalAlloc |
| Criterion | paired baselines, --save-baseline / --baseline |
The framework was consumed through the ignored .cargo/config.toml path override during this work. Cargo.lock was restored afterwards, and no committed manifest points at a path or git source.
view_layout/16 measures 415 µs on this tree against the 205 µs the 2026-09-03 report recorded at Rozi c9e9881. That is not a regression. Checking out c9e9881 with its own lockfile and running the same benchmark today gives 390.46 µs, so the historical figure does not reproduce on this machine. Ruled out as causes:
- Rozi code.
src/view/,src/app.rs,src/layout/andsrc/state/are byte-identical betweenc9e9881and4b8f076; the ten intervening commits are tests, CI, docs, a Windows lint fix, a macOS socket-path fix, and client-mailbox coalescing. - The framework. Both revisions lock
tui-lipan0.6.1 at the same checksum. - The benchmark.
view_layoutandbackend_with_panesare unchanged; the onlyapp_renderdiff isinbound_draincomments and the addedinbound_fairnesscase. - The toolchain.
c9e9881measures 390.46 µs on rustc 1.97.1 and 394.30 µs on 1.90.0. The pass is not toolchain-sensitive, and the audit's 1.98.1 is no longer installed here. - CPU scaling. Governor and EPP are
performancewith boost enabled, as in the audit.
The ratio is also not a constant: 2.36× at 1 pane against 1.90× at 16. What remains is something about the original recording session that cannot be recovered from the report - which notes it ran inside an unrelated rozi pane on the same host, with perf_event_paranoid lowered for sampling.
Treat the 2026-09-03 absolute microseconds as unreliable and its ratios and attribution as sound. Every number below is a paired comparison taken on one machine within one session.
The 390 µs at c9e9881 against 415 µs at 4b8f076 is a further 6% between two binaries whose view and layout code is identical. That is code layout, not behavior, and is not worth chasing.
What a frame allocates
benches/alloc_probe.rs brackets exactly one TestBackend::render() with a counting allocator. Counts are exact and repeat run to run. peak_live is peak growth over the state the frame started in, not process RSS.
| Panes | Allocations | Bytes | Peak live | Tree nodes | Allocations per node |
|---|---|---|---|---|---|
| 1 | 316 | 197,561 | 71,336 | 24 | 13.2 |
| 2 | 503 | 307,343 | 94,960 | 35 | 14.4 |
| 4 | 944 | 615,443 | 163,488 | 65 | 14.5 |
| 8 | 1,811 | 1,227,051 | 290,560 | 125 | 14.5 |
| 16 | 3,249 | 2,126,951 | 506,648 | 207 | 15.7 |
Allocations and bytes are both linear in pane count, at about 196 allocations and 129 KB per pane. Every allocation is freed within the frame, so this is pure churn rather than growth.
The ratio is the finding. The expanded tree at 16 panes is 207 elements, and the frame makes 3,249 allocations to produce it. The question the previous audit left open - what is allocated once per element per frame - has the answer that nothing is allocated only once: the pass pays roughly sixteen allocations for every element it ends up with.
Where they come from
Instrumenting the phases of the pass with a thread-local bucket the counting allocator reads:
| Phase | Allocations at 16 panes | Share |
|---|---|---|
AppRoot::view (Rozi's own view) | 1,597 | 49.2% |
| Reconcile and layout | 1,273 | 39.2% |
expand_element | 223 | 6.9% |
expand_children, output vector | 147 | 4.5% |
| Theme carve-out, sweep, other | 9 | 0.3% |
This corrects the previous report. High allocator sample counts attributed there primarily to expansion in fact span application view construction and reconciliation; expansion accounts for only about 11% of frame allocations. Half the frame is allocated inside the host's view() before the framework sees an element at all, and most of the rest is reconciliation.
The correction is what kept this work cheap. Sampled CPU profiles name the frames that are hot; they do not say which caller asked for the memory, and reading ComponentRegistry::expand_children at the bottom of an allocator stack as "expansion is the problem" pointed at memoizing the smallest of the three consumers. Allocation accounting is what redirected the effort onto three defects that a sampling profile cannot distinguish from architecture.
Attributing the same frame by call site, with dhat over a debug-info build:
| Shape | Allocations | Share of allocations | Bytes |
|---|---|---|---|
Element-sized blocks | 787 | 24.2% | 1,048,392 |
Callback (Arc<dyn Fn>) | 297 | 9.1% | 14,856 |
Key, and the format! string behind it | 267 | 8.2% | small |
| Everything else | 1,898 | 58.4% | 1,063,703 |
The Key row's bytes are counted under "everything else"; its blocks are tens of bytes each and the interest is in how many there are, not how big.
Half the bytes in a frame are Element blocks, because Element was 1,248 bytes wide. It is as wide as the widest ElementKind payload, and ElementKind::Terminal(Terminal) inlined 1,152 bytes into every node - so a Spacer, a Divider, and a one-word Text each cost 1,248 bytes to move, box, or push.
Three specific defects account for most of the traffic:
| Defect | Allocations | Bytes | Share of bytes |
|---|---|---|---|
reconcile_animated deep-clones its child subtree unconditionally | 396 | 594,752 | 28.0% |
Single-child wrappers route through a Vec round trip | ~370 | 329,472 | 15.5% |
expand_children allocates a fresh output Vec per container | 147 | 259,584 | 12.2% |
The changes
All three are in tui-lipan. None changes a public API.
1. reconcile_animated borrows its settled child. It built an owned Element on every path, but only the height-animating branch rewrites a constraint and needs one. The settled branch was deep-copying the whole animated subtree every frame. A host that wraps each of its panes in an Animated - which Rozi does - paid that once per pane per frame.
2. ElementKind's five oversized variants are boxed. Terminal, Slider, Checkbox, Tabs, and ProgressBar now hold a Box. ElementKind drops from 1,152 to 520 bytes and Element from 1,248 to 616 bytes, halving every Box<Element> allocation, every Vec<Element> buffer, and every element move in the pass. It costs one extra allocation per use of those five widgets, against a saving on every node of every other kind. Seventeen match arms changed; deref coercion covered all but the five constructors.
3. A single-child expansion path. Thirteen wrapper kinds - Frame, MouseRegion, Animated, Center, Portal, Group, ThemeProvider, ContextProvider, Memo, EffectScope, DragSource, DropTarget, StatusBarLayout - held exactly one child and reached expansion through expand_children, which cost one Vec to pass the child in and another to carry it back. expand_single skips both. It also removes thirteen copies of a .pop().unwrap_or_else(|| Text::new("").into()) fallback that could never fire, so the change is net negative on line count.
Results
Criterion medians, each paired against the same saved baseline on one machine.
| Case | Baseline | After all three | Change |
|---|---|---|---|
view_layout/1 | 48.413 µs | 37.297 µs | −21.5% |
view_layout/8 | 238.20 µs | 166.30 µs | −30.0% |
view_layout/16 | 421.32 µs | 289.03 µs | −31.2% |
inbound_drain/8_panes/64 | 2.7827 ms | 2.8156 ms | no change (p = 0.15) |
Cumulative, in the order applied:
| Step | view_layout/16 | vs baseline | Marginal |
|---|---|---|---|
| Baseline | 421.32 µs | — | — |
+ reconcile_animated borrow | 345.57 µs | −16.7% | −16.7% |
+ boxed ElementKind variants | 304.59 µs | −26.6% | −11.9% |
+ expand_single | 289.03 µs | −31.2% | −5.1% |
Allocation counts over the same steps, at 16 panes:
| Step | Allocations | Bytes | Peak live |
|---|---|---|---|
| Baseline | 3,249 | 2,126,951 | 506,648 |
+ reconcile_animated borrow | 2,853 | 1,532,199 | 506,648 |
+ boxed variants and expand_single | 2,665 | 928,487 | 387,896 |
| Change | −18.0% | −56.4% | −23.4% |
The expansion phase specifically fell from 371 allocations to 190, a 48.8% cut, and now accounts for 7.1% of the frame.
Each step cleared the 10% keep threshold on view_layout/16 except expand_single at 5.1%, which is kept because it removes a clearly pathological allocation pair per wrapper node and reduces code.
Two Rozi-side follow-ups
With expansion no longer the story, the largest remaining allocation class was Rozi's own: 557 blocks of a frame, 20.9%, spent rebuilding identity strings the application already knows are fixed. Two of those sites are pane-derived and constant for a pane's life, and both were measured against their own paired baselines under a lower keep threshold, since each is a local cleanup rather than a subsystem.
| Change | Allocations removed | view_layout/16 |
|---|---|---|
Cache body and terminal keys on Pane | 64 | −2.25% |
| Cache the four per-pane chrome animation keys | 128 | −4.82% |
The second is worth more per key because it removes the format! as well as the Arc<str>. Both needed no framework change: Element::key and Context::animated_color already take impl Into<Key>, and cloning a cached Key is a refcount bump.
Where the pass ended
| Start | End | Change | |
|---|---|---|---|
view_layout/16 | 421.32 µs | ~285 µs | −32% |
| Allocations per frame at 16 panes | 3,249 | 2,473 | −23.9% |
inbound_drain/8_panes/64 | 2.7827 ms | 2.8156 ms | no change (p = 0.15) |
No retained tree, no memoization, no caching subsystem, and no allocator swap. Note also that ordinary pane output does not pay this cost: a redraw repaints without rebuilding the view tree, so this is the price of a full structural frame rather than of every frame.
cargo test passes in both repositories - 2,708 framework unit tests plus its integration suites, and Rozi's 1,778, the latter against the released tui-lipan 0.6.1 that Rozi's manifest names - along with cargo clippy --all-targets -- -D warnings and cargo fmt --check.
What allocation counts are, and are not, evidence for
Two paired experiments in this pass moved CPU almost exactly in proportion to the allocations they removed: −2.40% of allocations bought −2.25% of view_layout/16, and −4.92% bought −4.82%. That is what makes allocation counts a useful prioritization signal for this workload - the allocator cost here is per call rather than per byte, so a small allocation is worth removing too.
It is not a CPU model, and this report should not be cited as though it were. The sites differ in what they actually do: an Arc<str> from a &str is an allocation and a copy, a formatted key is that plus a String and the formatting itself, and the ratio was not uniform even across pane counts within one experiment. Call-site inspection and a paired CPU measurement remain the acceptance test; the count only decides what to look at first.
The cost of skipping that discipline showed up in this pass. A shape tag in the profile suggested 48 blocks a frame were Key::from(&'static str) in the hot path, which implied an easy LazyLock<Key> win. Reading the actual call sites at the profiled revision showed they were Text::new, a ThemeProvider theme clone, and two empty Text widgets - not keys at all. The optimization did not exist. Only 23 .key( sites exist in the whole view, and nearly all sit in overlays that are closed during the benchmark.
What was not done, and why
memo_key, subtree reuse, retained expansion. Step 2 of the plan settled this: expansion is 11.4% of frame allocations, and cheap structural fixes took the whole pass down 31%. Retained expansion would target the smallest of the three consumers at by far the highest risk.- A process-wide allocator swap. The 2026-09-03 audit measured it as a net regression whenever output flows. That conclusion stands and this work does not revisit it.
- Pre-sizing or recycling child vectors beyond
expand_single. The remainingexpand_childrenoutput vectors are 57 allocations per frame at 16 panes, 2.1% of the total. AVecpool on the registry would remove most of them for a persistent field and a checkout/return protocol, which is not worth 2%. - Bisecting the 415 µs vs 205 µs gap against the previous report. Resolved under "Measurement context" without bisecting: the historical revision measures 390 µs today, so there is no regression to find.
Backlog, deliberately not worked
The pass stopped while sites still existed, because what remained was individually small or wanted plumbing out of proportion to it. Each of these needs a profile of its own before it is touched again, not a search for allocations because allocations exist.
Text::new("")allocates anArc<str>for the empty string, 24 times a frame, for the invisible drag targets in the resize strips. Worth folding into framework work that already touchesText- a shared emptyArc<str>, orSpacerat the call sites - not worth a change of its own.pane_window_key, 32 allocations a frame. It depends onpty_generation, a public field written from 37 places, and a cache behind a field anything may assign is a correctness trap. Revisit only if that field gets encapsulated for an independent design reason; a generation-validating lazy cache is complexity this has not earned.- About 166 blocks the profile could not attribute past its backtrace depth. Ignore unless a future CPU profile points there on its own.
WorkspaceLayerandtitledformatting, 80 blocks between them. Same rule: profile first.
Reproducing
cargo bench --bench alloc_probe # allocation counts, no framework hook
cargo bench --bench app_render -- --save-baseline base '^app_render/view_layout/(1|8|16)$'
cargo bench --bench app_render -- --baseline base '^app_render/view_layout/(1|8|16)$'
cargo bench --bench app_render -- --baseline base '^inbound_drain/8_panes/64$'Phase attribution and the Element size table need a hook inside the framework, which is not part of the committed set; see the header of benches/alloc_probe.rs for what to add. Call-site attribution used dhat against a CARGO_PROFILE_BENCH_DEBUG=true CARGO_PROFILE_BENCH_STRIP=false build, with the profiler started after setup so only the frame under test was recorded.