host scan · current shared-epoch experiment

Target throughput · OPEN v4

50,000 requests/s across the whole cell: 50,000/session/s for single, or exactly 1,562.5/session/s at 32 sessions. All clients share a prepared phase epoch; cell-wide slots are 20 µs apart. Delayed sends retain their intended slots.

workload timed since 2026-09-12: 30 s measured + 5 s warmup × 3 runs, each session offered its exact share of the cell rate — 1,500,000 measured messages per run at any session count. Cells measured earlier are 20,000 warmup + 20,000 measured/session × 3 runs; every cell shows the window it measured. units µs · worst reported percentile/max across sessions
NIC paths and earlier campaign notes (historical context)

Temporary SCAN NIC change (2026-09-08): the Mellanox-only v4 campaign uses Mellanox for new kernel, VMA and SocketXtreme results; Onload/TCPDirect are paused while the Solarflare DAC is removed. Earlier kernel results used Solarflare, including all historical v3 rows. From 2026-09-10 the host runs net.core.busy_read=50 / busy_poll=50 (declared in the machine file; a sitting refuses otherwise), so kernel rows measured from then busy-poll on every scenario. See the campaign note and completed-cell list to identify remeasured v4 rows. The archived C++ kernel and VMA pool results completed at 19:27 and 19:29 UTC carry an additional live CPU-contention qualification: their accept coordinator shared worker 0's core during all three measured phases. They are not clean-pinning comparisons; the latency impact is not quantified. The replacement kernel/VMA pool results at 20:24:51/20:27:20 UTC verify the separate core-19 coordinator, but retain their measured instability qualifications. The rerun's live watcher recorded no placement violation.

Applies to this OPEN v4 target-throughput matrix. Every measured cell shows p50, p99, p99.9 and max for the selected latency. Service is actual client start to completed reply; response starts at the scheduled slot and includes scheduler delay. Missing values show —. All four acceptor models stay visible; highlighting compares the selected metric among visible, qualified results.

Maximum throughput

not measured Highest sustained completed request/reply rate across the whole cell, established by an open-load rate sweep. Report the passing/failing rate bracket and response latency at the passing rate. This measures the configured client/server system, not an isolated engine limit.

No validated maximum-throughput sweep is available yet. The historical closed-loop test allows only one outstanding request per session and is not maximum throughput. Existing target-rate observations are not maximum-throughput results either. Candidate-run instructions and remaining validation work.

Earlier experiments

Historical RTT, reply-gated pacing and aligned-deadline OPEN v3 remain on the separate earlier matrix.

Core placement

scenarioserver coresclient cores
1 · single1015
2 · mono1011-13,15-19
3 · pool10-1315-18
4 · shard10-1315-19

isolated: 10-19 · VMA maintenance (reserved, no app threads): 14 · housekeeping (never measured on): 0-9 · Pool coordinator (separate from workers): 19