host beelink2 · AMD Ryzen 7 8745HS · 8c/16t · no isolated cores · WSL2

Latency matrix: flavour × transport × acceptor model

Every cell is a FIX round trip measured on the wire between two network namespaces. Every percentile and max is the worst reported value across all sessions.

sessions 32 (multi) · 1 (single) build -march=native → znver4, gcc 15.2 units µs recipe 20 000 warmup / 20 000 measured / 3 runs

These numbers are not publishable, and they are not comparable with scan. The certified harness (bench/load/bench.sh) cannot run on this box at all: every cell it launches is sudo -n ip netns exec … against root-owned artifacts under /usr/local/libexec/libhft-bench, and it refuses to place work on non-isolated cores. What is tabulated here came from the same binary (bench_roundtrip) with the same scenario flags and the same recipe, driven directly over loopback. Four things make it indicative:

Applies to Target throughput and the historical diagnostic tables. Every measured cell shows p50, p99, p99.9 and max for the selected latency. Service is actual client start to completed reply; response starts at the scheduled slot and includes scheduler delay. Missing values show —. All four acceptor models stay visible; highlighting compares the selected metric among visible, qualified results.

Temporary values — do not compare flavours on them. Cells marked smoke are behaviour probes, not steady-state latency measurements. The default smoke recipe is 200 warmup / 200 measured / one run, with one session for single and four for mono, pool and shard; the table headings describe the full 32-session recipe.

Use full-recipe measurements for latency comparisons: 20 000 warmup before each run, 20 000 measured, three runs, with the stability and parity gates satisfied.

Maximum throughput

not measured Highest sustained completed request/reply rate across the whole cell, established by an open-load rate sweep. Report the passing/failing rate bracket and response latency at the passing rate. This measures the configured client/server system, not an isolated engine limit.

No validated maximum-throughput sweep is available yet. The historical closed-loop test allows only one outstanding request per session and is not maximum throughput. Existing target-rate observations are not maximum-throughput results either. Candidate-run instructions and remaining validation work.

No target-throughput (open-loop) cell is measured on this host yet. The closed-loop matrix -- every measured cell so far -- is the expanded block below; the open sweep fills the table above as its cells complete.

Historical diagnostics · reply-gated RTT and pacing

These are separate experiments, retained for investigation rather than the two main throughput views. Their rates are not a backend-capacity measurement.

Historical RTT baseline · unpaced, one outstanding request per sessionService latency (µs)
Flavour Transport 1 · Single1 acceptor · 1 session 2 · Mono1 acceptor thread · 32 sessions 3 · PoolThread pool · 32 sessions 4 · ShardSharded acceptors · 32 sessions
Acceptor models measured against a 4-thread client pool — java-pure / kernel, 32 sessions Server threads on cores 10–13, client pool on 15–18, disjoint. Offered load 2000/session (64k/s aggregate), the highest rate the client sustains without backlog. Not comparable to the primary fixed-open table above — a different experiment, not a different day.
Acceptor modelp50p99p99.9maxsessions
thread-per-session (32 thr, 1 core) 2 020 9033 879 632 3 971 2635 287 85632/32
1 multiplexer 219.401706.36 2179.624752.5832/32
4-thread pool (least-loaded) 100.971534.98 7064.2610 050.5232/32

What one number means

A cell is the round-trip time of a FIX NewOrderSingle → ExecutionReport, measured on the wire between two network namespaces on one host. The clock starts at the client's scheduled send, not its actual send, so a client that falls behind is charged for the lateness instead of hiding it.

The parameters, and where they live

All of it is in one versioned file, load-bench-standards.conf, read by the wrapper, by every harness and by the checker. The common recipe is deliberately not per-scenario: the columns are comparable only if the work per message is identical in each.

commonvalue
warmup20 000, discarded
iterations20 000 measured
offered throughput50 000/s aggregate · fixed-rate open
jvm heap192 MB
scenarionos_er
per scenariosessionsacceptorsthreadsruns
1 single1113
2 mono32113
3 pool32143
4 shard3241 each3

The publication harness passes the complete recipe explicitly; bare client or launcher invocations are diagnostics, not substitutes for the harness. Every result carries bench_params (the argument vector), bench_profile (what the server configured) and bench_standards (the defaults in force, with a checksum of the file that supplied them). check-bench-params.sh refuses a table whose cells disagree, or whose profile contradicts the declared expectation.

Pinning policy is machine-independent — cores isolated, server and client sets disjoint, server from the front of the isolated set and clients from the back. The core numbers live in a per-host file, because a standard carrying one box's core list cannot be run on another, and a matrix that cannot be reproduced elsewhere is a single data point.

The four scenarios

1 single
one session, one acceptor. The engine floor, with no concurrency in it. 3 runs.
2 mono
32 sessions, one acceptor on one core. Readiness from poll(2) (or zf_muxer on TCPDirect), drained from a single poll snapshot before re-polling.
3 pool
32 sessions, one acceptor, M threads. Each session is bound to one thread for life, assigned least-loaded at accept.
4 shard
32 sessions, M independent acceptors sharing one port through SO_REUSEPORT, each on its own core, share-nothing. The kernel decides where a connection lands — there is no assignment policy. The standard uses four acceptors; M=1 is a separate diagnostic control.

3 and 4 are not the same thing. The pool and shard columns are separate across all flavours; unsupported native-stack combinations remain explicitly marked no pool / no shard.

When a cell counts

  • All 32 sessions must complete. Where fewer did, the count is printed beside the cell rather than quietly taking the median of whoever connected.
  • Every percentile and max is the worst reported value across all sessions, so one bad session cannot be averaged away.
  • A value shown as ≥ is a saturated clamp, not a reading — latency_u32() tops out at 232−1 ns.
  • Absolute numbers drift between sittings. Retained baselines and newer sweep cells may use separate artifact revisions; consult the handbook's sitting tables and source/artifact manifests before comparing them.

Core pinning

Both sides busy-poll, so a core carrying a server thread and a client starves both. Placement is therefore stated per scenario, not derived at run time.

No machine file for beelink2; the harnesses fall back to deriving cores from /sys/devices/system/cpu/isolated.

Server cores come from the front of the isolated set and client cores from the back, so widening the server — a pool, or more shards — eats into the middle and the two sets stay disjoint by construction rather than by arithmetic that has to be got right each time.

Unpinned is not neutral. isolcpus removes the isolated cores from every process's default affinity mask, so a thread that is never pinned does not land on a random isolated core — it lands on the housekeeping cores beside every interrupt, while the cores reserved for measurement sit idle. The JNI and .NET pool workers shipped that way for an hour; it would have read as "the pool is slow".

Not every shared core is contention. java-pure pins its accept loop and its multiplexer to the same core, and the accept loop blocks in accept() — it burns nothing. pinning-audit.sh --live samples each thread's CPU twice and reports a collision only when more than one of them actually ran.

Not every flavour can run every scenario

flavour1 single2 mono3 pool4 shard
cppyesyesyes*yes
java-pureyesyesyes*yes*
java-jniyesyesyes*yes
dotnet-nativeyesyesyes*yes
dotnet-pureyesyes*†yes*yes
rust-pureyesyesyesyes

*newly capable, not yet measured. †dotnet-pure's mono-thread cells served a thread per session — it had no multiplexed mode until now — so they are not comparable with the rest of that column. Thread-per-session with cores to spare pays no multiplexing cost, so it reads as fast rather than as different: its kernel cell sat at 26.14µs against a single-session 25.61 for exactly that reason. java-pure's mono cells are multiplexed (p99 207, not the 81 002 of its thread-per-session default); it simply had to be asked, and now the standards ask.

Three flavours now have a true pool. C++ gained FixAcceptorThreadPool — one listener, M workers, least-loaded at accept, verified serving 8/8 sessions at 2 per worker — and dotnet-pure gained the equivalent in its bench server.

java-jni reuses the very same serve pass the single-acceptor path runs, with accept disabled, so there is one serve implementation rather than two that drift; verified serving 8/8 sessions at p50 25.79µs.

rust-pure is the FIX engine written in Rust: its own client and server, no hftnet. rust-native, the Rust client driving the C++ engine through the hftnet C ABI, was parked on 2026-09-11: its server was cpp's, so its multi-session cells re-measured the C++ server. All six now pool on kernel TCP, Onload and VMA interposition. dotnet-native uses the native FixAcceptorThreadPool, with per-worker connection ids, status, errors and event rings projected through the existing managed session. Stack listener pools (TCPDirect and SocketXtreme) remain unavailable.

TCPDirect has no reuse-port listener API. SocketXtreme accepts independent SO_REUSEPORT binds, but four process pollers disconnect every client during Logon because they cannot safely own the shared completion stream. Both native-stack shard groups are marked no shard, not measured as one acceptor under a shard label.

Where the cells are not equivalent

Latency comparisons only mean something if the cells do the same work. They do not, and it is invisible in the numbers. All four end up persisting nothing, by four different mechanisms at four different costs:

flavouracceptorstoreaudit
cppFixAcceptorabsent (no template parameter)absent
cpp (--acceptor-impl runner)FixSessionRunnerAcceptornull, compile-timepresent, off
java-jniFixSessionRunnerAcceptorselected at runtime from configpresent, off
dotnet-nativeFixSessionRunnerAcceptorbranches on a null pointer per callpresent, off
dotnet-pureC# FixConnectionseparate implementation—

All flavours now declare this at runtime. Each server prints a bench_profile line stating the acceptor shape, store, audit, schedule, TLS and thread count it actually configured, so the table above is verified rather than read off the source, and check-bench-params.sh asserts it against the declared expectation.

It earns its keep. Wiring java-pure's multiplexer, its thread count was defaulted from the pool scenario's, quietly turning the mono-thread cells into a 4-thread pool. Nothing in the latency showed it; the declaration read threads=4 under a heading that says mono-thread, and that is how it was caught.

best qualified p50 for this model and selected latency > 100 µs > 1 ms — tail failure measuring run in flight not wired capability exists, no bench arm no pool / no shard transport cannot implement that shape n/a not applicable