Updated · 3 October 2026
Latest tests. Per-language CDFs.
The converted-values-v2 codec campaign admits 25 adapters, 320 isolated field-work probes and 554 operation rows. The full-session roster also includes the latest correctness, wire and scheduling-delay tests; a functional pass is not promoted into a latency distribution.
The transport and capacity tables below are retained September campaigns, not newly collected October timings.
Retained reference · 20 September 2026
A comparison with the variables held still.
The client sends a varied FIX NewOrderSingle and completes the sample only after receiving, parsing and extracting the required fields from the derived ExecutionReport. The schedule is open-loop: response time starts at the intended send slot, so scheduler delay is charged rather than hidden.
Swipe the diagram horizontally, or open it at full size.
Understand all four session scenarios → · Codec microbenchmark operation diagrams →
| Host | HP reference host · Intel Xeon Gold 6254 at 3.10 GHz · 18 physical cores · SMT off · one NUMA node |
|---|---|
| Workload | FIX NewOrderSingle (35=D) → ExecutionReport (35=8), one session, one request in flight |
| Offer | 50,000 requests/s; every published cell achieved the target rate with complete delivery |
| Collection | 5-second warm-up followed by three 30-second measured windows; 4.5 million response samples per cell |
| Metric | Response latency from intended schedule slot through reply parsing and business-field extraction |
| Build record | Campaign source revision 162ca9d2; GCC 14 native collection profile; results dated 20 September 2026 |
Pure Java and .NET Pure can use socket interposition through Onload or VMA. TCPDirect and SocketXtreme require vendor APIs and therefore appear on native or native-backed flavors. “Not applicable” is a deployment-model boundary, not a failed benchmark.
At a glance
Lowest measured p50 for each flavor.
| Flavor | Path | p50 | p99 | p99.9 |
|---|---|---|---|---|
| C++ | TCPDirect | 4.702 µs | 4.976 µs | 5.363 µs |
| Java Pure | Onload | 7.195 µs | 7.613 µs | 7.737 µs |
| Java/JNI | TCPDirect | 5.639 µs | 6.549 µs | 6.668 µs |
| .NET Pure | Onload | 11.473 µs | 11.873 µs | 12.004 µs |
| .NET Native | TCPDirect | 6.347 µs | 6.516 µs | 6.610 µs |
| Rust | TCPDirect | 5.869 µs | 6.029 µs | 6.088 µs |
The selected row is the lowest p50 within each flavor. Compare complete per-flavor tables below when tail latency or deployment constraints matter.
Capacity sweep
Passing rates stop where saturation starts.
A separate SCAN campaign stepped the shared kernel-TCP, single-session harness through 50k, 75k and 100k requests/s. The table publishes each flavor’s highest passing target and first failed target; it does not turn an overload point into a sustained result.
| Flavor | Highest sustained | p50 | p99 | p99.9 | First saturated |
|---|---|---|---|---|---|
| C++ | 75,000/s | 22.593 µs | 27.311 µs | 35.050 µs | 100,000/s |
| Java Pure | 50,000/s | 32.633 µs | 36.973 µs | 45.719 µs | 75,000/s |
| Java/JNI | 75,000/s | 26.819 µs | 37.318 µs | 40.839 µs | 100,000/s |
| .NET Pure | 50,000/s | 34.050 µs | 41.463 µs | 56.605 µs | 75,000/s |
| .NET Native | 75,000/s | 27.221 µs | 39.831 µs | 53.862 µs | 100,000/s |
| Rust | 50,000/s in the common HP campaign | 14.504 µs | 15.876 µs | 18.233 µs | Matched sweep pending |
The shared harness found first-fail boundaries at 75k or 100k, so higher probes would not support a defensible sustained-rate chart. Rust is shown from its existing common 50k campaign and remains outside this particular sweep.
Diagnostic SCAN run dated 29 September 2026, source revision 1bec04aa3. BrokerForge housekeeping processes occupied other cores. Read the complete contract and retained saturation evidence →
Workload and measurement points
What does each test measure?
Expand a scenario to follow the FIX messages, locate t0/t1, and inspect a complete captured example for every leg. The shaded interval is timed; grey dashed work is excluded. These diagrams explain the tests, not their performance.
Swipe diagrams horizontally on a small screen, or open them at full size.
What this test measures: Order → acknowledgementnos_er · D → 8 · Client clock
Client clock: t1 − t0. Both timestamps belong to one endpoint; synchronized host clocks are not required.
- t0 · start
- Published transport reference: the intended send slot, before any scheduling delay. Canonical actual-start tests: immediately before order construction and the engine-specific send call.
- t1 · stop
- After the client parses the derived acknowledgement and completes the required business-field extraction and checks.
- Included
- Client order construction/send, both endpoints, the network, server order processing and acknowledgement construction, then client reply processing. The displayed open-loop reference also includes scheduler delay.
- Excluded
- Session establishment and warm-up are outside the headline samples. An acknowledgement is not a trade fill.
Compare the canonical actual-start boundary
Canonical nos_er starts immediately before construction/send, not at the intended schedule slot. Scheduling delay is a separate probe. Do not mix that service-time distribution with the displayed open-loop reference.
Captured sample FIX messages
One complete, correlated chain from the synthetic QuickFIX/C++ functional audit. These are actual captured bytes illustrating the current canonical wire shape, not messages from the older transport campaign or measured latency samples.
The acknowledgement echoes order 11=1, symbol LCOM1, buy side, quantity 1000 and price 10000.00000. ExecType 150=0 and OrdStatus 39=0 mean New: CumQty is 0, LeavesQty is 1000 and AvgPx is 0. No LastQty/LastPx fill fields are present.
| represents SOH (ASCII 0x01). Remove display line breaks and restore SOH to recover the exact message, including valid BodyLength (9) and CheckSum (10). FIX SendingTime (52) is a wall-clock field, not the monotonic measurement clock: do not subtract these timestamps to estimate latency. Tag 60 retains the fixed July fixture time.
Download the original synthetic wire capture → · Download all checked examples and source hashes →
What this test measures: Market-data request → full bookmd · V → W · Client clock
Client clock: t1 − t0. Both timestamps belong to one endpoint; synchronized host clocks are not required.
- t0 · start
- Immediately before constructing and sending the market-data request, at actual task start.
- t1 · stop
- After the client parses the full snapshot and completes the required book-field extraction and checks.
- Included
- Request construction/send, server request parsing and book construction, both network legs, and client parsing/checking of the ten-level snapshot.
- Excluded
- Session establishment and warm-up. No order submission or ExecutionReport leg belongs to this test.
This diagram explains the canonical actual-start test. It does not assert that a matching passing distribution is available for the current chart selection.
Captured sample FIX messages
One complete, correlated chain from the synthetic QuickFIX/C++ functional audit. These are actual captured bytes illustrating the current canonical wire shape, not messages from the older transport campaign or measured latency samples.
Request 262=1 asks for 264=10 levels and both bid/offer types. Its snapshot echoes the request ID and symbol; 268=20 counts twenty entries, not twenty levels. The complete wire example includes all ten bid/offer pairs in order.
| represents SOH (ASCII 0x01). Remove display line breaks and restore SOH to recover the exact message, including valid BodyLength (9) and CheckSum (10). FIX SendingTime (52) is a wall-clock field, not the monotonic measurement clock: do not subtract these timestamps to estimate latency. Tag 60 retains the fixed July fixture time.
Download the original synthetic wire capture → · Download all checked examples and source hashes →
What this test measures: Book → derived order → acknowledgementmd_nos_er · V → W → D → 8 · Client clock
Client clock: t1 − t0. Both timestamps belong to one endpoint; synchronized host clocks are not required.
- t0 · start
- Immediately before constructing and sending the market-data request, at actual task start.
- t1 · stop
- After parsing the final acknowledgement and completing the required report checks.
- Included
- Both request/reply legs, full-book parsing/checks, client-side order derivation from the received book, order construction/send and final acknowledgement processing. One interval covers the entire chain.
- Excluded
- Session establishment and warm-up. A prebuilt constant reaction order is not a valid substitute for deriving it from the parsed snapshot.
This diagram explains the canonical actual-start test. It does not assert that a matching passing distribution is available for the current chart selection.
Captured sample FIX messages
One complete, correlated chain from the synthetic QuickFIX/C++ functional audit. These are actual captured bytes illustrating the current canonical wire shape, not messages from the older transport campaign or measured latency samples.
The received marker 262=1 is odd, so the client derives a buy (54=1). It uses the parsed best ask 10001.00000 for 44 and the ask size 1000 for 38 (within the lot/cap bounds). Order 11=1 carries the same marker; the final acknowledgement echoes these derived values.
| represents SOH (ASCII 0x01). Remove display line breaks and restore SOH to recover the exact message, including valid BodyLength (9) and CheckSum (10). FIX SendingTime (52) is a wall-clock field, not the monotonic measurement clock: do not subtract these timestamps to estimate latency. Tag 60 retains the fixed July fixture time.
Download the original synthetic wire capture → · Download all checked examples and source hashes →
What this test measures: Completed tick → reaction-order arrivaltick_to_trade · W → D · Feed/server clock
Feed/server clock: t1 − t0. Both timestamps belong to one endpoint; synchronized host clocks are not required.
- t0 · start
- Immediately before sending the completed top-of-book tick. Constructing and finalising W happens before this timestamp.
- t1 · stop
- The first instruction of the feed-side application callback after the engine frames/parses the returning order, before benchmark-side marker/extraction checks.
- Included
- Tick send, both network legs, client tick parsing, decision and order derivation/construction/send, and feed-side engine delivery of the returning order.
- Excluded
- Tick construction and finalisation, then post-t1 marker matching and benchmark extraction checks. Those checks can still invalidate the result. There is no acknowledgement/fill leg.
This diagram explains the canonical actual-start test. It does not assert that a matching passing distribution is available for the current chart selection.
Captured sample FIX messages
One complete, correlated chain from the synthetic QuickFIX/C++ functional audit. These are actual captured bytes illustrating the current canonical wire shape, not messages from the older transport campaign or measured latency samples.
The tick is top-of-book only: 268=2, one bid and one offer. Marker 262=1 produces buy order 11=1 at the parsed ask 10001.00000 for quantity 1000. No 35=8 is sent in the headline tick-to-trade workflow.
| represents SOH (ASCII 0x01). Remove display line breaks and restore SOH to recover the exact message, including valid BodyLength (9) and CheckSum (10). FIX SendingTime (52) is a wall-clock field, not the monotonic measurement clock: do not subtract these timestamps to estimate latency. Tag 60 retains the fixed July fixture time.
Download the original synthetic wire capture → · Download all checked examples and source hashes →
Read the complete canonical scenario contract → · Codec operation diagrams and all nine microbenchmark tests →
Detailed reports
Open the evidence for one deployment model.
Each flavor report combines its transport percentile profile, same-language codec microbenchmarks, the full pinned open-source comparison roster, and available rate/capacity evidence.
The dedicated FIX microbenchmark page separates parse, extract and build costs from session latency, with all six LibHFT flavors and their language peers in a qualified 25-adapter comparison with top-five CDFs.
The engine comparison hub shows CDFs and full competition lists per language, with host, transport, session model and workload selectors. Unsupported or non-admitted targets remain unranked.
Browse every target and its current four-flow status → The complete per-language inventory includes not-run, non-parity and excluded rows. Browse every committed benchmark file →
C++
Five retained transport paths and the latest native codec comparison, including llfix encoder results.
Open C++ benchmarks →Java Pure
QuickFIX/J, Philadelphia, Falcon and Artio evidence, including the SCAN rate sweep.
Open Java Pure benchmarks →Java/JNI
Native transport paths and end-to-end comparisons with explicit JNI microbenchmark boundaries.
Open Java/JNI benchmarks →.NET Pure
Managed codec and transport results beside the measured QuickFIX/n evidence.
Open .NET Pure benchmarks →.NET Native
Five native transport paths and the exact P/Invoke comparison boundary.
Open .NET Native benchmarks →Rust
Six codec engines, five transport paths and promoted versus pending session evidence.
Open Rust benchmarks →Flavor 1 of 6
C++
The native reference engine exposes every measured path, from ordinary kernel TCP through both vendor user-space TCP APIs.
| Transport | p50 | p99 | p99.9 | Maximum |
|---|---|---|---|---|
| Kernel TCP | 13.577 µs | 17.935 µs | 23.523 µs | 109.617 µs |
| Solarflare Onload | 5.560 µs | 6.044 µs | 6.593 µs | 19.537 µs |
| Mellanox VMA | 6.142 µs | 6.701 µs | 7.183 µs | 20.516 µs |
| Solarflare TCPDirect | 4.702 µs | 4.976 µs | 5.363 µs | 13.581 µs |
| Mellanox SocketXtreme | 5.788 µs | 6.620 µs | 7.702 µs | 22.113 µs |
Flavor 2 of 6
Java Pure
The independent JVM engine stays entirely in Java. Onload and VMA accelerate its ordinary sockets by interposition without introducing JNI into the engine path.
| Transport | p50 | p99 | p99.9 | Maximum |
|---|---|---|---|---|
| Kernel TCP | 17.063 µs | 18.559 µs | 20.856 µs | 58.363 µs |
| Solarflare Onload | 7.195 µs | 7.613 µs | 7.737 µs | 61.928 µs |
| Mellanox VMA | 7.469 µs | 7.773 µs | 7.912 µs | 36.972 µs |
Flavor 3 of 6
Java/JNI
The Java API over the native C++ core combines JVM integration with access to native vendor transports.
| Transport | p50 | p99 | p99.9 | Maximum |
|---|---|---|---|---|
| Kernel TCP | 15.665 µs | 18.656 µs | 24.764 µs | 117.330 µs |
| Solarflare Onload | 6.047 µs | 6.461 µs | 6.572 µs | 34.939 µs |
| Mellanox VMA | 6.555 µs | 7.001 µs | 7.118 µs | 36.278 µs |
| Solarflare TCPDirect | 5.639 µs | 6.549 µs | 6.668 µs | 40.571 µs |
| Mellanox SocketXtreme | 6.574 µs | 7.514 µs | 8.695 µs | 47.039 µs |
Flavor 4 of 6
.NET Pure
The independent C# engine remains managed end to end. Onload and VMA operate underneath its socket API while the FIX engine stays native-library free.
| Transport | p50 | p99 | p99.9 | Maximum |
|---|---|---|---|---|
| Kernel TCP | 20.794 µs | 24.811 µs | 28.828 µs | 60.730 µs |
| Solarflare Onload | 11.473 µs | 11.873 µs | 12.004 µs | 30.937 µs |
| Mellanox VMA | 11.999 µs | 12.231 µs | 12.368 µs | 32.406 µs |
Flavor 5 of 6
.NET Native
The C# API over the native C++ core provides the complete measured native transport set through the managed boundary.
| Transport | p50 | p99 | p99.9 | Maximum |
|---|---|---|---|---|
| Kernel TCP | 15.481 µs | 16.537 µs | 19.042 µs | 73.840 µs |
| Solarflare Onload | 6.788 µs | 7.203 µs | 7.310 µs | 48.914 µs |
| Mellanox VMA | 7.483 µs | 7.753 µs | 8.316 µs | 49.963 µs |
| Solarflare TCPDirect | 6.347 µs | 6.516 µs | 6.610 µs | 46.143 µs |
| Mellanox SocketXtreme | 7.393 µs | 7.971 µs | 9.310 µs | 51.054 µs |
Flavor 6 of 6
Rust
The independent Rust engine implements the same benchmark workflow directly and exposes all five measured network paths.
| Transport | p50 | p99 | p99.9 | Maximum |
|---|---|---|---|---|
| Kernel TCP | 14.504 µs | 15.876 µs | 18.233 µs | 64.199 µs |
| Solarflare Onload | 6.491 µs | 6.870 µs | 6.961 µs | 19.418 µs |
| Mellanox VMA | 6.979 µs | 7.166 µs | 7.265 µs | 17.736 µs |
| Solarflare TCPDirect | 5.869 µs | 6.029 µs | 6.088 µs | 12.072 µs |
| Mellanox SocketXtreme | 6.905 µs | 7.575 µs | 8.731 µs | 19.533 µs |
What this proves
Comparable latency, not every production question.
This campaign compares the single-session response path at a fixed offer. It does not claim maximum throughput, production topology latency or equivalent resource cost. Multi-session mono, pool and shard results are collected separately because concurrency changes the scheduling and ownership model.
50,000/s is the declared target rate for this latency experiment. The harness reached that rate in every row; it did not search for the highest sustainable rate.