Render and ingest CPU attribution
Performance index · Benchmark guide · Reproduction playbook
Verdict and scope
Verdict: ready with minor improvements.
Idle CPU is already cheap and needs no further work. Sampled profiles attribute the two remaining CPU consumers precisely: the view and layout pass is allocator-bound, with 50.6% of its samples landing in the system allocator, and the inbound ingest path is dominated by terminal cell handling, with 22.5% of samples in one drop glue.
A global allocator swap is a measured trade-off rather than a free win. It reduces view and layout time by 12-16% and increases ingest time by 2.8-7.2%. Ingest is the larger absolute cost whenever output flows, so this audit does not recommend a blanket swap.
This is a focused CPU-attribution audit of one revision. It does not measure memory, latency, lifecycle cleanup, or remote transport, and it does not replace a whole-application audit.
Correction, 2026-09-04. The absolute microseconds in this report do not reproduce. Rerunning
app_render/view_layoutat this same revisionc9e9881, with its own lockfile, gives 390 µs at 16 panes rather than the 205 µs below. The cause was not identified; code, dependency versions, benchmark, toolchain, and CPU scaling were all ruled out. Its ratios, scaling shape, and attribution still hold. See the 2026-09-04 report, which also corrects this report's reading of where the allocator samples come from.
Measurement context
| Item | Value |
|---|---|
| Revision | c9e9881, clean worktree |
| OS | Arch Linux, kernel 7.1.9-arch1-2 |
| CPU | AMD Ryzen 7 5700X3D, 8 cores / 16 threads; frequency boost enabled |
| Power profile | performance |
| Rust | rustc 1.98.1; cargo 1.98.1 |
| Build | release-debug for live sampling; bench profile with debug info kept for benchmark sampling |
| Sampler | samply 0.13.1 at the default 1000 Hz |
| Host policy | kernel.perf_event_paranoid raised from 2 to -1 by the owner; sampling fails at 2 |
| Criterion plotting | gnuplot unavailable; Plotters used |
| Host load | The recording session itself ran inside an unrelated rozi pane on the same host |
Live sessions used isolated HOME, XDG config, state, cache, and data directories. XDG_RUNTIME_DIR was placed in a separate short directory under the per-user runtime base, because a session endpoint is a Unix socket and the full scratch path overruns SUN_LEN.
Idle CPU by pane count
One release-debug client attached to one server at 60x200. Each pane filled 2000 lines of styled scrollback and then slept, so this measures a settled application, not ingest. CPU time came from /proc/<pid>/stat over a single 6-second window per scale, as fractions of one core.
| Panes | Client | Server |
|---|---|---|
| 1 | 0.17% | 0.17% |
| 4 | 0.17% | 0.50% |
| 8 | 0.17% | 0.67% |
| 16 | 0.33% | 0.83% |
No profiler was attached for these samples; see the profiler note below for why that matters. Client idle CPU is flat in pane count. Server idle CPU grows with pane count but stays below 1% of one core at 16 panes. CLK_TCK is 100, so a 6-second window resolves to 0.167%; single samples at this scale are indicative, not precise. These values are consistent with the 2026-08-29 idle server wakeup evidence and leave no meaningful idle headroom.
View and layout
app_render/view_layout measures whole-application view expansion and layout. It does not measure backend drawing or terminal buffer diffing.
| Panes | Populated | Empty |
|---|---|---|
| 1 | 20.118 µs | 20.141 µs |
| 2 | 31.495 µs | 31.333 µs |
| 4 | 58.961 µs | 59.008 µs |
| 8 | 115.36 µs | 114.88 µs |
| 16 | 205.40 µs | 204.77 µs |
Two properties hold across the range. Cost is close to linear in pane count, at roughly 12 µs per pane above a fixed base, with no superlinear term through 16 panes. Populated and empty panes cost the same to within 0.6%, so this pass is priced by tree structure rather than by terminal content.
sidebar_render confirms the same shape from the other side: 116.11 µs hidden against 117.18 µs for the panes tab, 117.23 µs for activity, 117.27 µs for files, and 117.53 µs for git. Sidebar content contributes about 1 µs.
Where the time goes
Sampled with samply record --save-only over a 12-second Criterion --profile-time run of view_layout/16, then symbolized per library with addr2line.
| Share of thread samples | Location |
|---|---|
| 50.6% | leaf frame inside libc.so.6 |
| 49.4% | leaf frame inside the benchmark binary |
Every libc leaf was attributed to its nearest non-libc caller. The callers are <std::alloc::System as GlobalAlloc>::alloc, ::dealloc, and ::realloc, reached from ComponentRegistry::expand_children, ComponentRegistry::expand_element, Element::clone, and Box<Element> writes. The pass is allocator-bound: it rebuilds the element tree each time and pays for the resulting small-object churn.
The largest named Rust self frames agree. Arc<str> drop is 3.24% and its clone is 1.94%, followed by apply_document_theme_carve_out_in_place at 1.38%, FxHasher hashing of SegmentId and Option<u16> at about 2% combined, MouseRegion::clone at 0.81%, and Theme::clone at 0.63%.
Note that expansion, reconciliation, and element cloning are tui-lipan responsibilities. Rozi supplies the view; the allocation traffic is a property of the framework's render architecture.
Inbound ingest
inbound_drain drains a mailbox until empty, so a chunk-size comparison at equal total bytes is valid. The 64 B and 1 KiB cases each carry 256 KiB; the 16 KiB case carries 1 MiB.
| Panes | 64 B chunks | 1 KiB chunks | 16 KiB chunks |
|---|---|---|---|
| 1 | 1.4241 ms | 1.3774 ms | 5.5785 ms |
| 2 | 2.4632 ms | 1.7624 ms | 6.7845 ms |
| 4 | 2.4499 ms | 1.8396 ms | 7.1746 ms |
| 8 | 2.5990 ms | 1.8684 ms | 7.2813 ms |
Normalized per byte at 8 panes, 64 B chunks cost 9.70 ms/MiB against 7.31 ms/MiB for 1 KiB chunks and 7.10 ms/MiB for 16 KiB chunks. Delivering the same bytes in 64 B pieces costs about 33% more CPU than delivering them in 1 KiB pieces, and per-byte cost flattens at 1 KiB and above.
The profiles explain the gap. At 8 panes with 64 B chunks, 23.1% of samples have a libc leaf, of which 8.01% is syscall, and futex mutex lock and unlock together account for 6.17%. At 16 KiB chunks the libc share falls to 2.82%, with no syscall or mutex frame in the top twenty. The excess is per-message fixed cost, not per-byte work.
Cell handling dominates the real work at both sizes. With 16 KiB chunks the top self frames are drop_glue::<Option<Arc<CellExtra>>> at 22.45%, vte::ansi::Handler::input at 15.95%, Row::index_mut at 12.85%, write_at_cursor at 5.62%, and Vec<Cell>::as_mut_slice at 5.46%. Option<Arc<CellExtra>>::clone adds 1.65%. Roughly a quarter of ingest CPU is spent cloning and dropping per-cell extra data that is absent for plain output. This is alacritty_terminal behavior reached through the terminal widget, not rozi code.
Pane-aware coalescing follow-up
A follow-up at revision ca7ed25 tested grouping interleaved output in the client mailbox by (pane_id, local, generation). The queue scans only its trailing output-only segment, retains byte order within each pane, and stops at every control frame. Non-adjacent aggregation is capped at 64 KiB, while the existing adjacent redraw path remains unchanged. The mailbox schedules the first frame immediately, so it adds no buffering delay. This deliberately relaxes relative ordering between independent pane outputs within one segment.
The current-tree baseline differs from the original audit revision, so only the paired change below is meaningful. Values are Criterion median estimates:
| Case | Baseline | Pane-aware | Change |
|---|---|---|---|
inbound_drain/8_panes/64 | 5.2900 ms | 2.9561 ms | -44.1% |
inbound_drain/8_panes/1024 | 3.2265 ms | 3.1491 ms | within noise threshold |
inbound_drain/8_panes/16384 | 22.880 ms | 12.601 ms | -44.9% |
server_fairness/key_round_trip/continuous_pty_ingress | 1.9809 ms | 1.9979 ms | no significant change |
The 64 B result clears the 15-20% keep threshold and neither larger chunk case regressed. The fairness comparison reported p = 0.59, but that benchmark uses the server's channel target and does not traverse the client mailbox. It rules out a server-path regression, not a client-mailbox latency regression. inbound_fairness/hot_plus_quiet later measured 625 µs to drain one 64 KiB hot aggregation plus the quiet pane's frame, which is the extra wait the cap allows.
Allocator experiment
The allocator was substituted with LD_PRELOAD against the same benchmark binary, so no source, manifest, or lockfile changed. Values are Criterion mean estimates.
| Case | glibc | jemalloc | tcmalloc |
|---|---|---|---|
view_layout/8 | 115.67 µs | 101.54 µs (-12.2%) | 96.91 µs (-16.2%) |
view_layout/16 | 206.84 µs | 181.91 µs (-12.1%) | 175.91 µs (-15.0%) |
inbound_drain/8_panes/64 | 2.4825 ms | 2.4921 ms (+0.4%) | 2.6570 ms (+7.0%) |
inbound_drain/8_panes/1024 | 1.8709 ms | 1.9090 ms (+2.0%) | 1.9588 ms (+4.7%) |
inbound_drain/8_panes/16384 | 7.1029 ms | 7.3012 ms (+2.8%) | 7.8271 ms (+10.2%) |
The direction is opposite on the two paths, which matches the attribution: the render pass is allocation-bound and benefits, while ingest is compute-bound and pays the substituted allocator's overhead. Ingest is milliseconds against microseconds for view and layout, so a process-wide swap is a likely net regression whenever output flows.
LD_PRELOAD on a benchmark binary is not equivalent to a #[global_allocator] in a release build. Treat these as directional.
Interpretation
- Confirmed attribution: view and layout spends 50.6% of its samples in the system allocator, reached from element-tree expansion and cloning.
- Confirmed attribution: ingest spends about a quarter of its samples cloning and dropping
Option<Arc<CellExtra>>, and small chunks add about 33% per-byte cost from syscall and lock traffic rather than parsing. - Confirmed optimization: pane-aware client-mailbox coalescing cuts the 8-pane 64 B drain case by 44.1%. The server fairness guard detects no regression but does not exercise the client mailbox.
- Confirmed scaling: view and layout is linear in pane count through 16 panes and independent of pane content.
- Harmless at tested scale: idle CPU, at or below 1% of one core per server at 16 panes.
- Deliberate trade-off, not recommended: a process-wide allocator swap buys 12-16% on view and layout and costs 2.8-7.2% on ingest.
- Confirmed measurement-harness failure, not a product defect: attaching samply to a client hosted by util-linux
scriptstrands the PTY wrapper in a stopped state. See below.
Profiler attachment strands the PTY wrapper
Attaching samply to a harness-hosted client with samply record -p <client-pid> stops that client from serving control requests and from rendering. Every later request fails at the client's 10-second timeout, the process stays alive with all threads parked, CPU falls to 0.00% of one core, the terminal output stream stops growing, and it does not recover after 20 seconds.
This was first mistaken for a wakeup defect in rozi. A controlled comparison rules that out. The harness, workload, pacing, and action sequence were held fixed while only the profiler target changed:
| Profiler target | Result |
|---|---|
| none | 12 of 12 actions served, no stall |
| server process only | 6 of 6 actions served, no stall |
| client process only | stalled on the third action |
| client and server | stalled on the third action |
The stall also occurred in earlier runs where samply exited without profiling because perf_event_paranoid was still 2, so the attach attempt is enough. Descriptor count held steady at 16 descriptors and 7 sockets throughout, no single action is responsible, and the effect is independent of pane content, animation settings, and whether the client's terminal input side is held open.
The cause is the interaction between samply's Linux attach sequence and util-linux script, which the harness uses to give the client a PTY:
- samply 0.13.1 sends
SIGSTOPto the attached PID, opens and enables the perf events, then sendsSIGCONTto that PID. - When
scriptsees its child stop, itscallback_child_sigstophandler stopsscriptitself. It resumes the child only after something else resumesscript. - samply resumes the client but does not know about the
scriptwrapper. The wrapper remains stopped and no longer drains the PTY master. - The client resumes, fills the PTY buffer during a redraw, and blocks the UI thread in its synchronous terminal write. Control connection threads remain alive, but their messages cannot reach the blocked UI thread and time out after 10 seconds.
A standalone reproduction on the same host confirmed the sequence without Rozi. After stopping a writer hosted by util-linux script 2.42.2, both wrapper and child entered a stopped state. Resuming only the child left the wrapper stopped and output fixed at 466,944 bytes for 500 ms. Resuming the wrapper immediately restarted output, which reached 1,257,472 bytes after 200 ms. Repeating the test with samply record -p produced the same stopped wrapper and frozen output, even though samply exited early under perf_event_paranoid=2.
This also explains why paced actions reached the stall in fewer requests than back-to-back actions. Pacing gives each state change time to render a complete frame, so the undrained PTY fills after fewer requests.
Two practical consequences:
- Do not use PID attachment against a client hosted by util-linux
script. Launch the client or a benchmark binary under samply instead. Samply's launch path enables perf events onexecand does not send the stop/resume sequence that strands the wrapper. Sampling the session server is unaffected. - Every client-side profile in this report launched the
app_renderbenchmark binary under samply, so none of the attribution above is affected.
Rozi's adaptive server wait is unrelated. The client runner also polls command messages at least every 50 ms, or at the frame interval while a live terminal exists. The lack of an explicit wake in CommandLink::send can add bounded latency but cannot explain a permanent stall.
Follow-up measurements
- If the render path is worth optimizing, measure a scoped allocation change, such as reusing element storage across expansions, rather than a process-wide allocator swap. Compare against
view_layoutandinbound_draintogether, since they move in opposite directions. - Re-measure ingest after any
alacritty_terminalupdate, since one drop glue carries 22.5% of that path.
Commands
cargo build --locked --profile release-debug
cargo bench --locked --bench app_render
CARGO_PROFILE_BENCH_DEBUG=true CARGO_PROFILE_BENCH_STRIP=false \
cargo bench --locked --bench app_render --no-run
samply record --save-only -n -o target/perf-audit/profiles/view_layout_16.json.gz -- \
target/release/deps/app_render-<hash> --bench 'view_layout/16' --profile-time 12
LD_PRELOAD=/usr/lib/libjemalloc.so \
target/release/deps/app_render-<hash> --bench 'view_layout/(8|16)$'Idle CPU followed the /proc/<pid>/stat method in the reproduction playbook, over one 6-second window at each of 1, 4, 8, and 16 panes.