Idle server wakeup evidence
Performance index · Benchmark guide · Reproduction playbook
Verdict and scope
Verdict: ready.
Adaptive waiting with a 4 ms ceiling reduced idle CPU by 50% per attached session at 1, 5, and 10 sessions. The trade-off was an increase in idle-settled key acknowledgement: p50 remained below 2.3 ms and p99 stayed below 4.5 ms at every tested scale. Queue-driven wakeups also improved key acknowledgement under continuous PTY ingress by 13.9%.
This focused audit compares the benchmark-only parent with the complete idle-server wakeup branch. It does not replace a whole-application performance audit.
Measurement context
| Item | Value |
|---|---|
| Before | 4dedfc8a350382315b421a4847fb87cd379bd715 plus the recorded phase-jitter correction |
| After | 0638f4175352f5e3ed6721a5ec335c8ea0a3c925 plus the same correction and a 4 ms ceiling |
| OS | Arch Linux, kernel 7.1.9-arch1-2 |
| CPU | AMD Ryzen 7 5700X3D, 8 cores / 16 threads; frequency boost enabled |
| Power profile | performance |
| Rust | rustc 1.97.1; cargo 1.97.1 |
| Build | Cargo release and bench profiles with --locked |
| Criterion plotting | gnuplot unavailable; Plotters used |
Both revisions used isolated target directories. The only before-revision source change was the same benchmark correction used after. All live sessions used isolated HOME, XDG config, state, cache, and runtime directories.
Idle CPU
For each scale, release clients remained attached through pseudo-terminals to separate named sessions. After a four-second settle, three simultaneous 10-second samples read server CPU time from /proc/<server-pid>/stat. Percentages are fractions of one core. The table reports the median sample; the reduction compares aggregate medians.
| Attached sessions | Before aggregate | After aggregate | Before per server | After per server | Reduction |
|---|---|---|---|---|---|
| 1 | 1.00% | 0.50% | 1.00% | 0.50% | 50.0% |
| 5 | 4.80% | 2.40% | 0.96% | 0.48% | 50.0% |
| 10 | 10.00% | 5.00% | 1.00% | 0.50% | 50.0% |
Across all repetitions, the before revision used 0.94-1.08% per server and the after revision used 0.48-0.60%. CPU remained approximately linear in the number of attached idle servers, with a 50% lower per-server slope.
Idle-settled latency
The permanent idle-latency probe measured 200 key-to-helper acknowledgement round trips. Each sample had 50-66 ms of quiescence using the deterministic schedule 50 + (sequence * 5) % 17. This prevents synchronization with the server's polling cycle while keeping runs reproducible. Additional attached idle sessions supplied the 5- and 10-session host loads. These values are request-latency percentiles, not Criterion estimate intervals.
| Sessions | Before p50 | After p50 | Before p95 | After p95 | Before p99 | After p99 | After max |
|---|---|---|---|---|---|---|---|
| 1 | 1.722 ms | 2.259 ms | 2.176 ms | 4.179 ms | 2.951 ms | 4.420 ms | 9.562 ms |
| 5 | 1.628 ms | 2.180 ms | 2.123 ms | 4.124 ms | 2.194 ms | 4.170 ms | 8.472 ms |
| 10 | 1.635 ms | 2.163 ms | 2.157 ms | 4.120 ms | 2.297 ms | 4.148 ms | 8.722 ms |
The p50 increase was 31-34%, and p95 increased by about 2 ms. Absolute p99 remained below 4.5 ms and did not degrade as parked-session count rose. Isolated maxima reached 9.6 ms on the after revision and 5.7 ms before, so the wait ceiling is not a whole-request latency bound; transport and host scheduling still contribute.
Rejected 8 ms ceiling
The first probe used a fixed 50 ms settle and repeatedly landed near the end of the 8 ms polling interval. It understated the tail. After the phase-jitter correction, the 8 ms experiment measured:
| Sessions | p50 | p95 | p99 | Maximum |
|---|---|---|---|---|
| 1 | 4.429 ms | 8.030 ms | 8.137 ms | 10.744 ms |
| 5 | 4.140 ms | 8.058 ms | 8.919 ms | 8.971 ms |
| 10 | 4.082 ms | 8.007 ms | 8.811 ms | 11.505 ms |
The 8 ms ceiling reduced median idle CPU by about 60%, but doubled the final 4 ms ceiling's p50 and p95. The 4 ms ceiling gives up ten percentage points of CPU reduction to halve the cold-input tail.
Active ingress
continuous_pty_ingress uses a real server-owned PTY emitting paced deterministic output while a client waits for key acknowledgement.
| Revision | Mean estimate | Criterion estimate interval |
|---|---|---|
| Before | 2.3368 ms | 2.3254-2.3498 ms |
| After | 2.0113 ms | 1.9700-2.0509 ms |
The after mean was 13.9% lower. PTY events notify the queue-backed wait directly, so active output does not pay the idle client-IPC fallback interval.
Interpretation
- Confirmed improvement: attached idle server CPU fell by 50% at every tested scale.
- Deliberate trade-off: idle-settled p95 key latency increased by about 2 ms, while p99 remained below 4.5 ms.
- Confirmed active behavior: continuous PTY ingress latency improved rather than regressed.
- Remaining limitation: local IPC still uses bounded polling. A cross-platform readiness mechanism could reduce idle CPU further without the measured cold-input latency increase, but this result does not justify that added complexity.
Commands
Permanent benchmark probes:
cargo bench --locked --bench server_fairness -- --idle-latency-probe
cargo bench --locked --bench server_fairness -- continuous_pty_ingressIdle CPU followed the /proc/<pid>/stat method in the reproduction playbook. For each revision it ran three 10-second samples at 1, 5, and 10 attached named sessions. The latency probe was then run once at each total host load with 200 samples per run.