AtrekesLibHFTDocumentationBenchmarks by flavor

Measured comparison

Six flavors. Results with context.

Explore the latest qualified codec comparison, current full-session admission and retained transport references. Every result keeps its campaign date, workload, rate and host attached; unrelated tests are not combined into a synthetic ranking.

6 engine flavors50,000 requests/sp50 · p99 · p99.94.5M samples per cell

Updated · 3 October 2026

Latest tests. Per-language CDFs.

The converted-values-v2 codec campaign admits 25 adapters, 320 isolated field-work probes and 554 operation rows. The full-session roster also includes the latest correctness, wire and scheduling-delay tests; a functional pass is not promoted into a latency distribution.

The transport and capacity tables below are retained September campaigns, not newly collected October timings.

Retained reference · 20 September 2026

A comparison with the variables held still.

The client sends a varied FIX NewOrderSingle and completes the sample only after receiving, parsing and extracting the required fields from the derived ExecutionReport. The schedule is open-loop: response time starts at the intended send slot, so scheduler delay is charged rather than hidden.

Swipe the diagram horizontally, or open it at full size.

Scheduled order-response latency
HP open-loop campaign. The interval begins at the intended scheduled send slot, including delay before the actual send. It ends after reply parsing and required field work. Schematic only; arrows are not to scale. Open diagram ↗

Understand all four session scenarios → · Codec microbenchmark operation diagrams →

HostHP reference host · Intel Xeon Gold 6254 at 3.10 GHz · 18 physical cores · SMT off · one NUMA node
WorkloadFIX NewOrderSingle (35=D) → ExecutionReport (35=8), one session, one request in flight
Offer50,000 requests/s; every published cell achieved the target rate with complete delivery
Collection5-second warm-up followed by three 30-second measured windows; 4.5 million response samples per cell
MetricResponse latency from intended schedule slot through reply parsing and business-field extraction
Build recordCampaign source revision 162ca9d2; GCC 14 native collection profile; results dated 20 September 2026
Read rows horizontally, not as universal engine rankings.

Pure Java and .NET Pure can use socket interposition through Onload or VMA. TCPDirect and SocketXtreme require vendor APIs and therefore appear on native or native-backed flavors. “Not applicable” is a deployment-model boundary, not a failed benchmark.

At a glance

Lowest measured p50 for each flavor.

FlavorPathp50p99p99.9
C++TCPDirect4.702 µs4.976 µs5.363 µs
Java PureOnload7.195 µs7.613 µs7.737 µs
Java/JNITCPDirect5.639 µs6.549 µs6.668 µs
.NET PureOnload11.473 µs11.873 µs12.004 µs
.NET NativeTCPDirect6.347 µs6.516 µs6.610 µs
RustTCPDirect5.869 µs6.029 µs6.088 µs

The selected row is the lowest p50 within each flavor. Compare complete per-flavor tables below when tail latency or deployment constraints matter.

Capacity sweep

Passing rates stop where saturation starts.

A separate SCAN campaign stepped the shared kernel-TCP, single-session harness through 50k, 75k and 100k requests/s. The table publishes each flavor’s highest passing target and first failed target; it does not turn an overload point into a sustained result.

FlavorHighest sustainedp50p99p99.9First saturated
C++75,000/s22.593 µs27.311 µs35.050 µs100,000/s
Java Pure50,000/s32.633 µs36.973 µs45.719 µs75,000/s
Java/JNI75,000/s26.819 µs37.318 µs40.839 µs100,000/s
.NET Pure50,000/s34.050 µs41.463 µs56.605 µs75,000/s
.NET Native75,000/s27.221 µs39.831 µs53.862 µs100,000/s
Rust50,000/s in the common HP campaign14.504 µs15.876 µs18.233 µsMatched sweep pending
No flavor earned a 250k–1M sustained claim.

The shared harness found first-fail boundaries at 75k or 100k, so higher probes would not support a defensible sustained-rate chart. Rust is shown from its existing common 50k campaign and remains outside this particular sweep.

Diagnostic SCAN run dated 29 September 2026, source revision 1bec04aa3. BrokerForge housekeeping processes occupied other cores. Read the complete contract and retained saturation evidence →

Workload and measurement points

What does each test measure?

Expand a scenario to follow the FIX messages, locate t0/t1, and inspect a complete captured example for every leg. The shaded interval is timed; grey dashed work is excluded. These diagrams explain the tests, not their performance.

Swipe diagrams horizontally on a small screen, or open them at full size.

What this test measures: Order → acknowledgementnos_er · D → 8 · Client clock

Client clock: t1 − t0. Both timestamps belong to one endpoint; synchronized host clocks are not required.

Scheduled order-response latency
Published open-loop reference. t0 is the intended send slot; scheduler delay is included. t1 follows reply parsing and required checks. Open diagram ↗
t0 · start
Published transport reference: the intended send slot, before any scheduling delay. Canonical actual-start tests: immediately before order construction and the engine-specific send call.
t1 · stop
After the client parses the derived acknowledgement and completes the required business-field extraction and checks.
Included
Client order construction/send, both endpoints, the network, server order processing and acknowledgement construction, then client reply processing. The displayed open-loop reference also includes scheduler delay.
Excluded
Session establishment and warm-up are outside the headline samples. An acknowledgement is not a trade fill.
Compare the canonical actual-start boundary

Canonical nos_er starts immediately before construction/send, not at the intended schedule slot. Scheduling delay is a separate probe. Do not mix that service-time distribution with the displayed open-loop reference.

Order round trip · nos_er
Client clock: t0 before order construction/send, D to the server, derived ACK 8 back, t1 after client parsing and required extraction/checks. Open diagram ↗

Captured sample FIX messages

One complete, correlated chain from the synthetic QuickFIX/C++ functional audit. These are actual captured bytes illustrating the current canonical wire shape, not messages from the older transport campaign or measured latency samples.

The acknowledgement echoes order 11=1, symbol LCOM1, buy side, quantity 1000 and price 10000.00000. ExecType 150=0 and OrdStatus 39=0 mean New: CumQty is 0, LeavesQty is 1000 and AvgPx is 0. No LastQty/LastPx fill fields are present.

1. CLIENT → SERVER · NewOrderSingle 35=D

8=FIX.4.4|9=132|35=D|34=2|49=CLIENT|52=20261003-00:42:20.433|56=SERVER|
11=1|21=1|38=1000|40=2|44=10000.00000|54=1|55=LCOM1|60=20260703-12:00:00.000|
10=201|

2. SERVER → CLIENT · ExecutionReport · New acknowledgement 35=8

8=FIX.4.4|9=182|35=8|34=2|49=SERVER|52=20261003-00:42:20.433|56=CLIENT|
6=0.00000|11=1|14=0|17=TCP-EXEC|37=TCP-ORDER|38=1000|39=0|44=10000.00000|54=1|55=LCOM1|60=20260703-12:00:00.000|150=0|151=1000|
10=060|

| represents SOH (ASCII 0x01). Remove display line breaks and restore SOH to recover the exact message, including valid BodyLength (9) and CheckSum (10). FIX SendingTime (52) is a wall-clock field, not the monotonic measurement clock: do not subtract these timestamps to estimate latency. Tag 60 retains the fixed July fixture time.

Download the original synthetic wire capture → · Download all checked examples and source hashes →

What this test measures: Market-data request → full bookmd · V → W · Client clock

Client clock: t1 − t0. Both timestamps belong to one endpoint; synchronized host clocks are not required.

Market-data round trip · md
Client clock: t0 before market-data request construction/send, V requests 10 levels, W returns 20 bid/offer entries, t1 after snapshot parsing and required checks. No order is sent. Open diagram ↗
t0 · start
Immediately before constructing and sending the market-data request, at actual task start.
t1 · stop
After the client parses the full snapshot and completes the required book-field extraction and checks.
Included
Request construction/send, server request parsing and book construction, both network legs, and client parsing/checking of the ten-level snapshot.
Excluded
Session establishment and warm-up. No order submission or ExecutionReport leg belongs to this test.

This diagram explains the canonical actual-start test. It does not assert that a matching passing distribution is available for the current chart selection.

Captured sample FIX messages

One complete, correlated chain from the synthetic QuickFIX/C++ functional audit. These are actual captured bytes illustrating the current canonical wire shape, not messages from the older transport campaign or measured latency samples.

Request 262=1 asks for 264=10 levels and both bid/offer types. Its snapshot echoes the request ID and symbol; 268=20 counts twenty entries, not twenty levels. The complete wire example includes all ten bid/offer pairs in order.

1. CLIENT → SERVER · MarketDataRequest 35=V

8=FIX.4.4|9=113|35=V|34=2|49=CLIENT|52=20261003-00:42:26.939|56=SERVER|
146=1|55=LCOM1|262=1|263=0|264=10|265=0|267=2|269=0|269=1|
10=083|

2. SERVER → CLIENT · MarketDataSnapshotFullRefresh 35=W

8=FIX.4.4|9=688|35=W|34=2|49=SERVER|52=20261003-00:42:26.939|56=CLIENT|
55=LCOM1|262=1|268=20|
269=0|270=10000.00000|271=1000|269=1|270=10001.00000|271=1000|
269=0|270=9999.99000|271=1010|269=1|270=10001.01000|271=1010|
269=0|270=9999.98000|271=1020|269=1|270=10001.02000|271=1020|
269=0|270=9999.97000|271=1030|269=1|270=10001.03000|271=1030|
269=0|270=9999.96000|271=1040|269=1|270=10001.04000|271=1040|
269=0|270=9999.95000|271=1050|269=1|270=10001.05000|271=1050|
269=0|270=9999.94000|271=1060|269=1|270=10001.06000|271=1060|
269=0|270=9999.93000|271=1070|269=1|270=10001.07000|271=1070|
269=0|270=9999.92000|271=1080|269=1|270=10001.08000|271=1080|
269=0|270=9999.91000|271=1090|269=1|270=10001.09000|271=1090|
10=218|

| represents SOH (ASCII 0x01). Remove display line breaks and restore SOH to recover the exact message, including valid BodyLength (9) and CheckSum (10). FIX SendingTime (52) is a wall-clock field, not the monotonic measurement clock: do not subtract these timestamps to estimate latency. Tag 60 retains the fixed July fixture time.

Download the original synthetic wire capture → · Download all checked examples and source hashes →

What this test measures: Book → derived order → acknowledgementmd_nos_er · V → W → D → 8 · Client clock

Client clock: t1 − t0. Both timestamps belong to one endpoint; synchronized host clocks are not required.

Market data to order · md_nos_er
Client clock: V request, ten-level W snapshot, client derives D from parsed book values, server returns ACK 8, t1 after report parsing and required checks. Two request/reply legs plus client derivation, not a single round trip. Open diagram ↗
t0 · start
Immediately before constructing and sending the market-data request, at actual task start.
t1 · stop
After parsing the final acknowledgement and completing the required report checks.
Included
Both request/reply legs, full-book parsing/checks, client-side order derivation from the received book, order construction/send and final acknowledgement processing. One interval covers the entire chain.
Excluded
Session establishment and warm-up. A prebuilt constant reaction order is not a valid substitute for deriving it from the parsed snapshot.

This diagram explains the canonical actual-start test. It does not assert that a matching passing distribution is available for the current chart selection.

Captured sample FIX messages

One complete, correlated chain from the synthetic QuickFIX/C++ functional audit. These are actual captured bytes illustrating the current canonical wire shape, not messages from the older transport campaign or measured latency samples.

The received marker 262=1 is odd, so the client derives a buy (54=1). It uses the parsed best ask 10001.00000 for 44 and the ask size 1000 for 38 (within the lot/cap bounds). Order 11=1 carries the same marker; the final acknowledgement echoes these derived values.

1. CLIENT → SERVER · MarketDataRequest 35=V

8=FIX.4.4|9=113|35=V|34=2|49=CLIENT|52=20261003-00:42:35.638|56=SERVER|
146=1|55=LCOM1|262=1|263=0|264=10|265=0|267=2|269=0|269=1|
10=079|

2. SERVER → CLIENT · MarketDataSnapshotFullRefresh 35=W

8=FIX.4.4|9=688|35=W|34=2|49=SERVER|52=20261003-00:42:35.638|56=CLIENT|
55=LCOM1|262=1|268=20|
269=0|270=10000.00000|271=1000|269=1|270=10001.00000|271=1000|
269=0|270=9999.99000|271=1010|269=1|270=10001.01000|271=1010|
269=0|270=9999.98000|271=1020|269=1|270=10001.02000|271=1020|
269=0|270=9999.97000|271=1030|269=1|270=10001.03000|271=1030|
269=0|270=9999.96000|271=1040|269=1|270=10001.04000|271=1040|
269=0|270=9999.95000|271=1050|269=1|270=10001.05000|271=1050|
269=0|270=9999.94000|271=1060|269=1|270=10001.06000|271=1060|
269=0|270=9999.93000|271=1070|269=1|270=10001.07000|271=1070|
269=0|270=9999.92000|271=1080|269=1|270=10001.08000|271=1080|
269=0|270=9999.91000|271=1090|269=1|270=10001.09000|271=1090|
10=214|

3. CLIENT → SERVER · NewOrderSingle 35=D

8=FIX.4.4|9=132|35=D|34=3|49=CLIENT|52=20261003-00:42:35.638|56=SERVER|
11=1|21=1|38=1000|40=2|44=10001.00000|54=1|55=LCOM1|60=20260703-12:00:00.000|
10=216|

4. SERVER → CLIENT · ExecutionReport · New acknowledgement 35=8

8=FIX.4.4|9=182|35=8|34=3|49=SERVER|52=20261003-00:42:35.638|56=CLIENT|
6=0.00000|11=1|14=0|17=TCP-EXEC|37=TCP-ORDER|38=1000|39=0|44=10001.00000|54=1|55=LCOM1|60=20260703-12:00:00.000|150=0|151=1000|
10=075|

| represents SOH (ASCII 0x01). Remove display line breaks and restore SOH to recover the exact message, including valid BodyLength (9) and CheckSum (10). FIX SendingTime (52) is a wall-clock field, not the monotonic measurement clock: do not subtract these timestamps to estimate latency. Tag 60 retains the fixed July fixture time.

Download the original synthetic wire capture → · Download all checked examples and source hashes →

What this test measures: Completed tick → reaction-order arrivaltick_to_trade · W → D · Feed/server clock

Feed/server clock: t1 − t0. Both timestamps belong to one endpoint; synchronized host clocks are not required.

Tick to trade · tick_to_trade
Same feed clock: build the top-of-book tick outside timing; t0 immediately before sending completed W with two entries; client parses and derives D; t1 at feed-side engine delivery of D before benchmark extraction/marker checks. No ExecutionReport leg. Open diagram ↗
t0 · start
Immediately before sending the completed top-of-book tick. Constructing and finalising W happens before this timestamp.
t1 · stop
The first instruction of the feed-side application callback after the engine frames/parses the returning order, before benchmark-side marker/extraction checks.
Included
Tick send, both network legs, client tick parsing, decision and order derivation/construction/send, and feed-side engine delivery of the returning order.
Excluded
Tick construction and finalisation, then post-t1 marker matching and benchmark extraction checks. Those checks can still invalidate the result. There is no acknowledgement/fill leg.

This diagram explains the canonical actual-start test. It does not assert that a matching passing distribution is available for the current chart selection.

Captured sample FIX messages

One complete, correlated chain from the synthetic QuickFIX/C++ functional audit. These are actual captured bytes illustrating the current canonical wire shape, not messages from the older transport campaign or measured latency samples.

The tick is top-of-book only: 268=2, one bid and one offer. Marker 262=1 produces buy order 11=1 at the parsed ask 10001.00000 for quantity 1000. No 35=8 is sent in the headline tick-to-trade workflow.

1. SERVER → CLIENT · MarketDataSnapshotFullRefresh 35=W

8=FIX.4.4|9=138|35=W|34=2|49=SERVER|52=20261003-00:42:39.983|56=CLIENT|
55=LCOM1|262=1|268=2|
269=0|270=10000.00000|271=1000|269=1|270=10001.00000|271=1000|
10=005|

2. CLIENT → SERVER · NewOrderSingle 35=D

8=FIX.4.4|9=132|35=D|34=2|49=CLIENT|52=20261003-00:42:39.984|56=SERVER|
11=1|21=1|38=1000|40=2|44=10001.00000|54=1|55=LCOM1|60=20260703-12:00:00.000|
10=223|

| represents SOH (ASCII 0x01). Remove display line breaks and restore SOH to recover the exact message, including valid BodyLength (9) and CheckSum (10). FIX SendingTime (52) is a wall-clock field, not the monotonic measurement clock: do not subtract these timestamps to estimate latency. Tag 60 retains the fixed July fixture time.

Download the original synthetic wire capture → · Download all checked examples and source hashes →

Read the complete canonical scenario contract → · Codec operation diagrams and all nine microbenchmark tests →

Detailed reports

Open the evidence for one deployment model.

Each flavor report combines its transport percentile profile, same-language codec microbenchmarks, the full pinned open-source comparison roster, and available rate/capacity evidence.

Looking inside the codec?

The dedicated FIX microbenchmark page separates parse, extract and build costs from session latency, with all six LibHFT flavors and their language peers in a qualified 25-adapter comparison with top-five CDFs.

Looking for LibHFT versus other FIX engines?

The engine comparison hub shows CDFs and full competition lists per language, with host, transport, session model and workload selectors. Unsupported or non-admitted targets remain unranked.

Browse every target and its current four-flow status → The complete per-language inventory includes not-run, non-parity and excluded rows. Browse every committed benchmark file →

NATIVE

C++

Five retained transport paths and the latest native codec comparison, including llfix encoder results.

Open C++ benchmarks →
JVM + NATIVE CORE

Java/JNI

Native transport paths and end-to-end comparisons with explicit JNI microbenchmark boundaries.

Open Java/JNI benchmarks →
INDEPENDENT RUST

Rust

Six codec engines, five transport paths and promoted versus pending session evidence.

Open Rust benchmarks →

Flavor 1 of 6

C++

The native reference engine exposes every measured path, from ordinary kernel TCP through both vendor user-space TCP APIs.

Transportp50p99p99.9Maximum
Kernel TCP13.577 µs17.935 µs23.523 µs109.617 µs
Solarflare Onload5.560 µs6.044 µs6.593 µs19.537 µs
Mellanox VMA6.142 µs6.701 µs7.183 µs20.516 µs
Solarflare TCPDirect4.702 µs4.976 µs5.363 µs13.581 µs
Mellanox SocketXtreme5.788 µs6.620 µs7.702 µs22.113 µs

Flavor 2 of 6

Java Pure

The independent JVM engine stays entirely in Java. Onload and VMA accelerate its ordinary sockets by interposition without introducing JNI into the engine path.

Transportp50p99p99.9Maximum
Kernel TCP17.063 µs18.559 µs20.856 µs58.363 µs
Solarflare Onload7.195 µs7.613 µs7.737 µs61.928 µs
Mellanox VMA7.469 µs7.773 µs7.912 µs36.972 µs

Flavor 3 of 6

Java/JNI

The Java API over the native C++ core combines JVM integration with access to native vendor transports.

Transportp50p99p99.9Maximum
Kernel TCP15.665 µs18.656 µs24.764 µs117.330 µs
Solarflare Onload6.047 µs6.461 µs6.572 µs34.939 µs
Mellanox VMA6.555 µs7.001 µs7.118 µs36.278 µs
Solarflare TCPDirect5.639 µs6.549 µs6.668 µs40.571 µs
Mellanox SocketXtreme6.574 µs7.514 µs8.695 µs47.039 µs

Flavor 4 of 6

.NET Pure

The independent C# engine remains managed end to end. Onload and VMA operate underneath its socket API while the FIX engine stays native-library free.

Transportp50p99p99.9Maximum
Kernel TCP20.794 µs24.811 µs28.828 µs60.730 µs
Solarflare Onload11.473 µs11.873 µs12.004 µs30.937 µs
Mellanox VMA11.999 µs12.231 µs12.368 µs32.406 µs

Flavor 5 of 6

.NET Native

The C# API over the native C++ core provides the complete measured native transport set through the managed boundary.

Transportp50p99p99.9Maximum
Kernel TCP15.481 µs16.537 µs19.042 µs73.840 µs
Solarflare Onload6.788 µs7.203 µs7.310 µs48.914 µs
Mellanox VMA7.483 µs7.753 µs8.316 µs49.963 µs
Solarflare TCPDirect6.347 µs6.516 µs6.610 µs46.143 µs
Mellanox SocketXtreme7.393 µs7.971 µs9.310 µs51.054 µs

Flavor 6 of 6

Rust

The independent Rust engine implements the same benchmark workflow directly and exposes all five measured network paths.

Transportp50p99p99.9Maximum
Kernel TCP14.504 µs15.876 µs18.233 µs64.199 µs
Solarflare Onload6.491 µs6.870 µs6.961 µs19.418 µs
Mellanox VMA6.979 µs7.166 µs7.265 µs17.736 µs
Solarflare TCPDirect5.869 µs6.029 µs6.088 µs12.072 µs
Mellanox SocketXtreme6.905 µs7.575 µs8.731 µs19.533 µs

What this proves

Comparable latency, not every production question.

This campaign compares the single-session response path at a fixed offer. It does not claim maximum throughput, production topology latency or equivalent resource cost. Multi-session mono, pool and shard results are collected separately because concurrency changes the scheduling and ownership model.

No maximum-throughput claim.

50,000/s is the declared target rate for this latency experiment. The harness reached that rate in every row; it did not search for the highest sustainable rate.