State model
Sequence continuity is durable business state.
LibHFT tracks inbound and outbound sequence numbers through session lifecycle, reconnect and replay. Durable modes persist the outbound resend window and sequence state so a restarted or promoted process can continue without inventing a reset.
Detect gaps and duplicates
Expected sequence, PossDup handling and recovery state drive whether a message can be admitted.
Persist before success
A durable path records replayable output before treating the application send as safely committed.
Serve the available range
Stored application frames are resent while administrative or unavailable spans become explicit gap fills.
Resume deliberately
Backup endpoints and bounded retry are coupled with persisted sequence restoration and diagnostics.
Gap-recovery flow
From missing sequence to a coherent stream.
- The engine detects an inbound sequence above the expected value.
- It enters a recovery state and issues a bounded ResendRequest according to policy.
- PossDup replay and SequenceReset/GapFill messages are validated against the current gap.
- Expected sequence advances only when continuity is established.
- Buffered or subsequent application messages are released through the declared application boundary.
The precise policy—including chunk size, validation, reset behavior and old-message handling—belongs in the counterparty profile and conformance suite.
Stores
Memory, mapped durability and QuickFIX-compatible modes.
| Mode | Use | Recovery property |
|---|---|---|
| None / memory | Ephemeral sessions, simulation or workloads where durable replay is not required. | Process lifetime defines retained state. |
| Mapped durable store | Persistent sequence and outbound replay window with bounded storage. | Restart and HA promotion can restore state from the same store. |
| QuickFIX-style file store | Migration-compatible file conventions on tracked surfaces. | Behavior must be assessed with the selected runtime and settings. |
Failure semantics
A failed durable write cannot become a successful send.
If the configured durable store refuses an outbound write, LibHFT records a store fault and prevents the engine from pretending it can later satisfy resend. The operator-visible state remains distinct from ordinary disconnection or active resend recovery.
A bounded store needs occupancy, remaining capacity and write-failure visibility. Monitoring those values is part of maintaining resend safety—not an optional dashboard enhancement.
Conformance
Test recovery with real transitions.
- Disconnect after established traffic and reconnect with the expected sequences.
- Request replay across stored, administrative and unavailable ranges.
- Inject high sequence, low sequence, duplicate and malformed reset cases.
- Refuse a durable write and prove no unsafe application frame reaches the peer.
- Restart from persisted state and inspect both protocol behavior and telemetry.