Bounded batch history compaction
This experiment keeps the dense viewport / compact history split and tests when rows should be compacted. The reproducible patches are not enabled in production. Rozi and the sibling framework retain their normal dependencies.
Policies and measurement boundaries
| Policy | Recent dense rows | Batch size | Maximum pending rows | Dense/prefix buffers per pool |
|---|---|---|---|---|
| Immediate | 0 | 1 | 0 | 4 |
| Window 32 | 32 | 16 | 47 | 16 |
| Window 128 | 128 | 32 | 159 | 32 |
Batching runs synchronously inside ingestion. A bounded recent window delays conversion; the next full batch converts the oldest pending rows. Pool capacity matches batch size so recycled dense allocations survive until subsequent scrolls reuse them. Uncompressible rows remain dense. Resize/reflow uses the preceding prototype's eager conversion path and resets pending bookkeeping.
The ordinary sustained benchmark retains a full 5,000-row history and the same parser across iterations. Every automatic batch runs inside its timer. A separate settled benchmark also drains all pending rows, including the recent window, inside each iteration's timer. This prevents a speedup from merely leaving conversion work outstanding.
Active rendering measures the real framework render_snapshot() for a populated viewport with full history. Fixture construction and snapshot destruction are excluded. It covers snapshot construction, not terminal-backend drawing or an unchanged cached snapshot.
The allocator/latency probe warms history before processing another 100,000 individual lines. It distinguishes regular calls from calls that trigger a compaction batch, reports observed p50/p99/max call latency, and times final debt settlement. Timing overhead and OS scheduling affect these per-call figures; the maximum is an observation, not a worst-case guarantee. Latency and throughput measurements run separately.
Conversion attribution
The controlled row benchmarks use 253 columns and vary meaningful length. Values are mean estimates in nanoseconds, from 10 samples with one-second warm-up and target measurement times. The tail scan reproduces the row's occupied-bound scan. It does not scan all 253 cells for ordinary short rows. Component timings include benchmark overhead and must not be summed as cycle attribution.
| Content | Meaningful cells | Tail scan ns | Whole warmed conversion ns |
|---|---|---|---|
| plain | 0 | 0.9 | 41.4 |
| plain | 8 | 1.9 | 51.6 |
| plain | 32 | 1.9 | 90.7 |
| plain | 80 | 1.9 | 172.2 |
| plain | 253 | 3.4 | 33.0 |
| unicode | 0 | 0.9 | 41.3 |
| unicode | 8 | 3.0 | 52.7 |
| unicode | 32 | 3.0 | 92.9 |
| unicode | 80 | 3.0 | 175.7 |
| unicode | 253 | 3.4 | 34.9 |
| styled_tail | 0 | 489.4 | 485.6 |
| styled_tail | 8 | 432.1 | 494.3 |
| styled_tail | 32 | 397.1 | 491.4 |
| styled_tail | 80 | 313.2 | 480.3 |
| styled_tail | 253 | 3.4 | 32.7 |
| The common short default tail is already cheap to find: about 2–3 ns. A new blank-tail detector | |||
| would target little of the measured 52–93 ns whole-conversion cost for 8–32 meaningful cells. | |||
| Styled repeated tails can scan hundreds of equal cells and cost much more. Full-width nonrepeating | |||
| rows reject compaction quickly and preserve their dense allocation. |
The separate plain-prefix swap probe measures about 26 ns at eight cells, 44 ns at 32, and 79 ns at 80. Dense reset measures about 38, 55, and 90 ns for those lengths. Warm eviction measures about 37, 46, and 61 ns. These controlled operations identify length-dependent work; they do not account for every instruction in the full scrolling path or reproduce its exact cache state.
Sustained ingestion and active snapshots
Same-day measurements use Rozi cedd4a6 plus the benchmark changes, Rust 1.98.1, and released engine/framework versions 0.26.0/0.9.0 as the dense baseline. All policies use the same candidate executable; the selector changes only construction-time policy. Values are mean milliseconds, using 10 samples and one-second warm-up/target measurement, extended automatically for slow cases.
| Case | Released dense | Immediate | Window 32 | Window 128 |
|---|---|---|---|---|
boundary/render_active/dense | 0.279 | 0.339 | 0.344 | 0.339 |
boundary/render_active/plain | 0.275 | 0.336 | 0.334 | 0.336 |
boundary/render_active/unicode | 0.274 | 0.338 | 0.341 | 0.338 |
sustained/dense/253x64_10000 | 22.420 | 22.613 | 23.073 | 23.062 |
sustained/dense/253x64_100000 | 228.972 | 230.880 | 229.529 | 227.496 |
sustained/dense/80x24_10000 | 7.575 | 7.509 | 7.583 | 7.626 |
sustained/dense/80x24_100000 | 75.214 | 76.833 | 75.961 | 75.842 |
sustained/plain/253x64_10000 | 1.923 | 2.388 | 2.584 | 2.661 |
sustained/plain/253x64_100000 | 19.201 | 23.735 | 25.935 | 26.674 |
sustained/plain/80x24_10000 | 1.865 | 2.350 | 2.604 | 2.653 |
sustained/plain/80x24_100000 | 18.708 | 24.269 | 25.985 | 27.187 |
sustained/unicode/253x64_10000 | 3.542 | 4.126 | 4.411 | 4.537 |
sustained/unicode/253x64_100000 | 33.000 | 42.546 | 43.066 | 43.193 |
sustained/unicode/80x24_10000 | 3.346 | 4.090 | 4.375 | 4.518 |
sustained/unicode/80x24_100000 | 35.204 | 40.968 | 42.391 | 45.284 |
A 20-sample repeat with two-second target measurements confirms the live-snapshot regression:
| Case | Released dense ms | Immediate split ms | Change |
|---|---|---|---|
sustained/plain/253x64_100000 | 19.239 | 23.935 | +24.4% |
boundary/render_active/plain | 0.278 | 0.335 | +20.4% |
boundary/render_active/unicode | 0.274 | 0.340 | +23.7% |
boundary/render_active/dense | 0.280 | 0.338 | +20.7% |
Batching does not remove this snapshot overhead. The split's explicit read iterator still handles both live and historical rows. This test does not attribute the regression to one instruction or prove that inlining alone would fix it. It does disprove treating unchanged full-width ingestion as proof that every live-screen consumer is unaffected.
Ingestion with all pending work settled
This separate framework executable processes 100,000 lines and calls finish_history_compaction inside each timed iteration. Fixture setup and initial settlement are outside timing. Every iteration starts and ends with zero pending rows. It has its own dependency resolution; compare policies within this executable rather than its absolute times with Rozi's executable.
| Corpus / columns | Immediate ms | Window 32 ms | Window 128 ms |
|---|---|---|---|
dense/253 | 225.700 | 223.554 | 226.552 |
dense/80 | 74.681 | 75.614 | 74.832 |
plain/253 | 24.055 | 25.024 | 25.526 |
plain/80 | 23.710 | 25.201 | 25.898 |
unicode/253 | 39.771 | 42.388 | 40.845 |
unicode/80 | 38.761 | 42.838 | 42.570 |
Neither batching policy provides a repeatable throughput improvement over immediate compaction. The plain-text cases remain slower even when all delayed conversion is charged to the timer.
Memory and latency
The 253×64 engine probe keeps both terminal and parser alive throughout counting. Sample-vector allocations are outside the counting scope. Live bytes before and after 100,000 lines match for every case. The peak below covers steady ingestion; final drain is separately timed. These are requested heap bytes, not process RSS/PSS or combined server/client memory.
| Window | Content | Retained bytes | Steady peak bytes | Allocation calls / 100k lines | Pending rows | Drain µs |
|---|---|---|---|---|---|---|
| 0 | plain | 5,324,784 | 5,324,784 | 0 | 0 | 0.03 |
| 0 | unicode | 6,169,320 | 6,169,392 | 200000 | 0 | 0.03 |
| 0 | dense | 33,764,688 | 33,764,688 | 0 | 0 | 0.35 |
| 0 | styled_tail | 5,324,784 | 5,324,784 | 0 | 0 | 0.03 |
| 32 | plain | 5,598,552 | 5,598,552 | 0 | 45 | 7.10 |
| 32 | unicode | 6,440,016 | 6,440,088 | 200000 | 45 | 9.67 |
| 32 | dense | 33,764,688 | 33,764,688 | 0 | 45 | 0.67 |
| 32 | styled_tail | 5,598,552 | 5,598,552 | 0 | 45 | 29.19 |
| 128 | plain | 6,242,648 | 6,242,648 | 0 | 157 | 33.31 |
| 128 | unicode | 7,074,896 | 7,074,968 | 200000 | 157 | 33.65 |
| 128 | dense | 33,764,688 | 33,764,688 | 0 | 157 | 2.20 |
| 128 | styled_tail | 6,242,648 | 6,242,648 | 0 | 157 | 108.66 |
Plain, full-width dense, and styled-tail output make zero steady allocation calls for all policies. The Unicode corpus makes 200,000 calls under each policy; batching does not eliminate those Unicode-cell allocations. For plain output, the 32-row policy adds about 267 KiB over immediate compaction; the 128-row policy adds about 896 KiB. No complete-history materialization is hidden outside the measurement.
Observed per-call latency in microseconds:
| Window | Content | Regular p50 | Regular p99 | Regular max | Batch p50 | Batch p99 | Batch max |
|---|---|---|---|---|---|---|---|
| 0 | plain | 0.22 | 0.28 | 10.98 | — | — | — |
| 0 | unicode | 0.32 | 0.50 | 176.72 | — | — | — |
| 0 | dense | 1.74 | 1.96 | 162.48 | — | — | — |
| 0 | styled_tail | 1.13 | 1.20 | 175.03 | — | — | — |
| 32 | plain | 0.20 | 0.24 | 58.69 | 0.91 | 0.99 | 5.55 |
| 32 | unicode | 0.28 | 0.35 | 12.55 | 1.14 | 1.24 | 6.10 |
| 32 | dense | 1.73 | 1.91 | 183.18 | 1.99 | 2.17 | 7.35 |
| 32 | styled_tail | 0.64 | 0.76 | 127.32 | 9.13 | 12.78 | 21.21 |
| 128 | plain | 0.20 | 0.25 | 120.20 | 1.76 | 1.90 | 6.82 |
| 128 | unicode | 0.29 | 0.35 | 13.18 | 2.17 | 2.35 | 7.10 |
| 128 | dense | 1.73 | 1.95 | 157.02 | 2.22 | 2.51 | 13.44 |
| 128 | styled_tail | 0.64 | 0.80 | 164.08 | 17.66 | 22.54 | 175.98 |
The styled-tail corpus explicitly colors/erases the full row before writing short text. Its long comparison scan makes a 32-row compaction batch much more expensive than a plain batch. Bounding row count does not provide a fixed microsecond budget. Large maxima also occur without batching, so these observations cannot distinguish all scheduler pauses from execution-time variation.
Decision
Do not adopt either batching policy. Both keep steady allocation counts bounded, but immediate compaction already does that. They retain more dense memory, create burstier call latency, and fail to reduce total sustained work. Charging final settlement to the timer does not change that conclusion. Trying more window sizes is not justified by these results.
The split remains useful as an experimental base, but its live rendering path is not yet validated as performance-neutral. The newly measured snapshot regression needs attention before considering a production fork. The next useful implementation experiment would isolate dense viewport traversal for rendering and verify its snapshot cost, then examine redundant prefix/reset work if conversion still needs improvement. The ordinary blank-tail scan is already too cheap to explain the regression.
Production dependencies and the sibling framework remain unchanged. The retained Rozi code change adds active snapshot measurement to the history benchmark; the storage and policy work stays in reproducible experimental patches. Follow-up snapshot traversal measurements are in Render-oriented history reads.