Skip to content

Idle server wakeup evidence ​

Performance index · Benchmark guide · Reproduction playbook

Verdict and scope ​

Verdict: ready.

Adaptive waiting with a 4 ms ceiling reduced idle CPU by 50% per attached session at 1, 5, and 10 sessions. The trade-off was an increase in idle-settled key acknowledgement: p50 remained below 2.3 ms and p99 stayed below 4.5 ms at every tested scale. Queue-driven wakeups also improved key acknowledgement under continuous PTY ingress by 13.9%.

This focused audit compares the benchmark-only parent with the complete idle-server wakeup branch. It does not replace a whole-application performance audit.

Measurement context ​

ItemValue
Before4dedfc8a350382315b421a4847fb87cd379bd715 plus the recorded phase-jitter correction
After0638f4175352f5e3ed6721a5ec335c8ea0a3c925 plus the same correction and a 4 ms ceiling
OSArch Linux, kernel 7.1.9-arch1-2
CPUAMD Ryzen 7 5700X3D, 8 cores / 16 threads; frequency boost enabled
Power profileperformance
Rustrustc 1.97.1; cargo 1.97.1
BuildCargo release and bench profiles with --locked
Criterion plottinggnuplot unavailable; Plotters used

Both revisions used isolated target directories. The only before-revision source change was the same benchmark correction used after. All live sessions used isolated HOME, XDG config, state, cache, and runtime directories.

Idle CPU ​

For each scale, release clients remained attached through pseudo-terminals to separate named sessions. After a four-second settle, three simultaneous 10-second samples read server CPU time from /proc/<server-pid>/stat. Percentages are fractions of one core. The table reports the median sample; the reduction compares aggregate medians.

Attached sessionsBefore aggregateAfter aggregateBefore per serverAfter per serverReduction
11.00%0.50%1.00%0.50%50.0%
54.80%2.40%0.96%0.48%50.0%
1010.00%5.00%1.00%0.50%50.0%

Across all repetitions, the before revision used 0.94-1.08% per server and the after revision used 0.48-0.60%. CPU remained approximately linear in the number of attached idle servers, with a 50% lower per-server slope.

Idle-settled latency ​

The permanent idle-latency probe measured 200 key-to-helper acknowledgement round trips. Each sample had 50-66 ms of quiescence using the deterministic schedule 50 + (sequence * 5) % 17. This prevents synchronization with the server's polling cycle while keeping runs reproducible. Additional attached idle sessions supplied the 5- and 10-session host loads. These values are request-latency percentiles, not Criterion estimate intervals.

SessionsBefore p50After p50Before p95After p95Before p99After p99After max
11.722 ms2.259 ms2.176 ms4.179 ms2.951 ms4.420 ms9.562 ms
51.628 ms2.180 ms2.123 ms4.124 ms2.194 ms4.170 ms8.472 ms
101.635 ms2.163 ms2.157 ms4.120 ms2.297 ms4.148 ms8.722 ms

The p50 increase was 31-34%, and p95 increased by about 2 ms. Absolute p99 remained below 4.5 ms and did not degrade as parked-session count rose. Isolated maxima reached 9.6 ms on the after revision and 5.7 ms before, so the wait ceiling is not a whole-request latency bound; transport and host scheduling still contribute.

Rejected 8 ms ceiling ​

The first probe used a fixed 50 ms settle and repeatedly landed near the end of the 8 ms polling interval. It understated the tail. After the phase-jitter correction, the 8 ms experiment measured:

Sessionsp50p95p99Maximum
14.429 ms8.030 ms8.137 ms10.744 ms
54.140 ms8.058 ms8.919 ms8.971 ms
104.082 ms8.007 ms8.811 ms11.505 ms

The 8 ms ceiling reduced median idle CPU by about 60%, but doubled the final 4 ms ceiling's p50 and p95. The 4 ms ceiling gives up ten percentage points of CPU reduction to halve the cold-input tail.

Active ingress ​

continuous_pty_ingress uses a real server-owned PTY emitting paced deterministic output while a client waits for key acknowledgement.

RevisionMean estimateCriterion estimate interval
Before2.3368 ms2.3254-2.3498 ms
After2.0113 ms1.9700-2.0509 ms

The after mean was 13.9% lower. PTY events notify the queue-backed wait directly, so active output does not pay the idle client-IPC fallback interval.

Interpretation ​

  • Confirmed improvement: attached idle server CPU fell by 50% at every tested scale.
  • Deliberate trade-off: idle-settled p95 key latency increased by about 2 ms, while p99 remained below 4.5 ms.
  • Confirmed active behavior: continuous PTY ingress latency improved rather than regressed.
  • Remaining limitation: local IPC still uses bounded polling. A cross-platform readiness mechanism could reduce idle CPU further without the measured cold-input latency increase, but this result does not justify that added complexity.

Commands ​

Permanent benchmark probes:

bash
cargo bench --locked --bench server_fairness -- --idle-latency-probe
cargo bench --locked --bench server_fairness -- continuous_pty_ingress

Idle CPU followed the /proc/<pid>/stat method in the reproduction playbook. For each revision it ran three 10-second samples at 1, 5, and 10 attached named sessions. The latency probe was then run once at each total host load with 200 samples per run.

MPL-2.0