Skip to content

Bounded batch history compaction ​

This experiment keeps the dense viewport / compact history split and tests when rows should be compacted. The reproducible patches are not enabled in production. Rozi and the sibling framework retain their normal dependencies.

Policies and measurement boundaries ​

PolicyRecent dense rowsBatch sizeMaximum pending rowsDense/prefix buffers per pool
Immediate0104
Window 3232164716
Window 1281283215932

Batching runs synchronously inside ingestion. A bounded recent window delays conversion; the next full batch converts the oldest pending rows. Pool capacity matches batch size so recycled dense allocations survive until subsequent scrolls reuse them. Uncompressible rows remain dense. Resize/reflow uses the preceding prototype's eager conversion path and resets pending bookkeeping.

The ordinary sustained benchmark retains a full 5,000-row history and the same parser across iterations. Every automatic batch runs inside its timer. A separate settled benchmark also drains all pending rows, including the recent window, inside each iteration's timer. This prevents a speedup from merely leaving conversion work outstanding.

Active rendering measures the real framework render_snapshot() for a populated viewport with full history. Fixture construction and snapshot destruction are excluded. It covers snapshot construction, not terminal-backend drawing or an unchanged cached snapshot.

The allocator/latency probe warms history before processing another 100,000 individual lines. It distinguishes regular calls from calls that trigger a compaction batch, reports observed p50/p99/max call latency, and times final debt settlement. Timing overhead and OS scheduling affect these per-call figures; the maximum is an observation, not a worst-case guarantee. Latency and throughput measurements run separately.

Conversion attribution ​

The controlled row benchmarks use 253 columns and vary meaningful length. Values are mean estimates in nanoseconds, from 10 samples with one-second warm-up and target measurement times. The tail scan reproduces the row's occupied-bound scan. It does not scan all 253 cells for ordinary short rows. Component timings include benchmark overhead and must not be summed as cycle attribution.

ContentMeaningful cellsTail scan nsWhole warmed conversion ns
plain00.941.4
plain81.951.6
plain321.990.7
plain801.9172.2
plain2533.433.0
unicode00.941.3
unicode83.052.7
unicode323.092.9
unicode803.0175.7
unicode2533.434.9
styled_tail0489.4485.6
styled_tail8432.1494.3
styled_tail32397.1491.4
styled_tail80313.2480.3
styled_tail2533.432.7
The common short default tail is already cheap to find: about 2–3 ns. A new blank-tail detector
would target little of the measured 52–93 ns whole-conversion cost for 8–32 meaningful cells.
Styled repeated tails can scan hundreds of equal cells and cost much more. Full-width nonrepeating
rows reject compaction quickly and preserve their dense allocation.

The separate plain-prefix swap probe measures about 26 ns at eight cells, 44 ns at 32, and 79 ns at 80. Dense reset measures about 38, 55, and 90 ns for those lengths. Warm eviction measures about 37, 46, and 61 ns. These controlled operations identify length-dependent work; they do not account for every instruction in the full scrolling path or reproduce its exact cache state.

Sustained ingestion and active snapshots ​

Same-day measurements use Rozi cedd4a6 plus the benchmark changes, Rust 1.98.1, and released engine/framework versions 0.26.0/0.9.0 as the dense baseline. All policies use the same candidate executable; the selector changes only construction-time policy. Values are mean milliseconds, using 10 samples and one-second warm-up/target measurement, extended automatically for slow cases.

CaseReleased denseImmediateWindow 32Window 128
boundary/render_active/dense0.2790.3390.3440.339
boundary/render_active/plain0.2750.3360.3340.336
boundary/render_active/unicode0.2740.3380.3410.338
sustained/dense/253x64_1000022.42022.61323.07323.062
sustained/dense/253x64_100000228.972230.880229.529227.496
sustained/dense/80x24_100007.5757.5097.5837.626
sustained/dense/80x24_10000075.21476.83375.96175.842
sustained/plain/253x64_100001.9232.3882.5842.661
sustained/plain/253x64_10000019.20123.73525.93526.674
sustained/plain/80x24_100001.8652.3502.6042.653
sustained/plain/80x24_10000018.70824.26925.98527.187
sustained/unicode/253x64_100003.5424.1264.4114.537
sustained/unicode/253x64_10000033.00042.54643.06643.193
sustained/unicode/80x24_100003.3464.0904.3754.518
sustained/unicode/80x24_10000035.20440.96842.39145.284

A 20-sample repeat with two-second target measurements confirms the live-snapshot regression:

CaseReleased dense msImmediate split msChange
sustained/plain/253x64_10000019.23923.935+24.4%
boundary/render_active/plain0.2780.335+20.4%
boundary/render_active/unicode0.2740.340+23.7%
boundary/render_active/dense0.2800.338+20.7%

Batching does not remove this snapshot overhead. The split's explicit read iterator still handles both live and historical rows. This test does not attribute the regression to one instruction or prove that inlining alone would fix it. It does disprove treating unchanged full-width ingestion as proof that every live-screen consumer is unaffected.

Ingestion with all pending work settled ​

This separate framework executable processes 100,000 lines and calls finish_history_compaction inside each timed iteration. Fixture setup and initial settlement are outside timing. Every iteration starts and ends with zero pending rows. It has its own dependency resolution; compare policies within this executable rather than its absolute times with Rozi's executable.

Corpus / columnsImmediate msWindow 32 msWindow 128 ms
dense/253225.700223.554226.552
dense/8074.68175.61474.832
plain/25324.05525.02425.526
plain/8023.71025.20125.898
unicode/25339.77142.38840.845
unicode/8038.76142.83842.570

Neither batching policy provides a repeatable throughput improvement over immediate compaction. The plain-text cases remain slower even when all delayed conversion is charged to the timer.

Memory and latency ​

The 253×64 engine probe keeps both terminal and parser alive throughout counting. Sample-vector allocations are outside the counting scope. Live bytes before and after 100,000 lines match for every case. The peak below covers steady ingestion; final drain is separately timed. These are requested heap bytes, not process RSS/PSS or combined server/client memory.

WindowContentRetained bytesSteady peak bytesAllocation calls / 100k linesPending rowsDrain µs
0plain5,324,7845,324,784000.03
0unicode6,169,3206,169,39220000000.03
0dense33,764,68833,764,688000.35
0styled_tail5,324,7845,324,784000.03
32plain5,598,5525,598,5520457.10
32unicode6,440,0166,440,088200000459.67
32dense33,764,68833,764,6880450.67
32styled_tail5,598,5525,598,55204529.19
128plain6,242,6486,242,648015733.31
128unicode7,074,8967,074,96820000015733.65
128dense33,764,68833,764,68801572.20
128styled_tail6,242,6486,242,6480157108.66

Plain, full-width dense, and styled-tail output make zero steady allocation calls for all policies. The Unicode corpus makes 200,000 calls under each policy; batching does not eliminate those Unicode-cell allocations. For plain output, the 32-row policy adds about 267 KiB over immediate compaction; the 128-row policy adds about 896 KiB. No complete-history materialization is hidden outside the measurement.

Observed per-call latency in microseconds:

WindowContentRegular p50Regular p99Regular maxBatch p50Batch p99Batch max
0plain0.220.2810.98———
0unicode0.320.50176.72———
0dense1.741.96162.48———
0styled_tail1.131.20175.03———
32plain0.200.2458.690.910.995.55
32unicode0.280.3512.551.141.246.10
32dense1.731.91183.181.992.177.35
32styled_tail0.640.76127.329.1312.7821.21
128plain0.200.25120.201.761.906.82
128unicode0.290.3513.182.172.357.10
128dense1.731.95157.022.222.5113.44
128styled_tail0.640.80164.0817.6622.54175.98

The styled-tail corpus explicitly colors/erases the full row before writing short text. Its long comparison scan makes a 32-row compaction batch much more expensive than a plain batch. Bounding row count does not provide a fixed microsecond budget. Large maxima also occur without batching, so these observations cannot distinguish all scheduler pauses from execution-time variation.

Decision ​

Do not adopt either batching policy. Both keep steady allocation counts bounded, but immediate compaction already does that. They retain more dense memory, create burstier call latency, and fail to reduce total sustained work. Charging final settlement to the timer does not change that conclusion. Trying more window sizes is not justified by these results.

The split remains useful as an experimental base, but its live rendering path is not yet validated as performance-neutral. The newly measured snapshot regression needs attention before considering a production fork. The next useful implementation experiment would isolate dense viewport traversal for rendering and verify its snapshot cost, then examine redundant prefix/reset work if conversion still needs improvement. The ordinary blank-tail scan is already too cheap to explain the regression.

Production dependencies and the sibling framework remain unchanged. The retained Rozi code change adds active snapshot measurement to the history benchmark; the storage and policy work stays in reproducible experimental patches. Follow-up snapshot traversal measurements are in Render-oriented history reads.

MPL-2.0