# Maximum throughput and target throughput

These are different experiments. **Target throughput** measures response latency at
a specified open-loop offer. **Maximum throughput** needs a search for the highest
sustained offer under a declared workload and acceptance rule; that search and its
publication path are not implemented yet. There are no verified maximum-throughput
numbers to populate that view today.

The [HTML handbook](load-bench-handbook.html#throughput-benches) contains the same
definitions and quick-start commands. The [scan matrix](latency-matrix-scan.html)
links directly to the instructions for each experiment.

The historical `closed` and reply-gated `paced` tables are archived diagnostics.
Both permit at most one outstanding request per session. Their achieved rates are
useful reply-gated measurements, not unrestricted maximum-throughput results.

## Before running either experiment

Run the commands below from the repository root on `scan`, with exclusive use of
the benchmark cores, namespaces and ports. The publishing driver owns a live
pinning watcher for each cell; direct exploratory `bench.sh` runs still need a
separate watcher and evidence review. Do not run alongside another benchmark session.

Build and stage the intended programs **and their dependencies** first, using the
approved build/install workflow. Cells execute artifacts under
`/usr/local/libexec/libhft-bench/`, not freshly edited source. In particular, staging
the .NET Native flavour's C# program does not stage its native `.so`; JNI also needs its
matching library and jars. The target driver records source/installed hashes but
does not build, install, or establish that the two represent the same build.

Use fresh output directories outside the repository and installed runtime tree.
`run-open-matrix.py` requires a nonexistent directory for a new sitting. Direct
`bench.sh` can erase a previously marked benchmark output directory, so never
reuse a candidate path to retry it.

Each new `bench.sh` evidence directory contains `clock-contract.json`. Latency,
pacing, warm-up and timeout intervals use the flavour's relative monotonic clock
(C++ and Rust use Linux `CLOCK_MONOTONIC` for the shared OPEN v4 schedule and
their raw/`Instant` clocks on legacy paths; Java uses `System.nanoTime`; .NET uses
`Stopwatch`). The JSON's separate Unix-epoch-nanosecond envelope records when the
capture occurred; no UTC timestamp is read or subtracted in a measured interval.
The publication audit requires a complete contract containing the measured
flavour. If the harness is interrupted, the retained contract remains
`in-progress` and cannot pass that gate.

These canonical runs are local round trips across namespaces and do not claim
cross-host one-way latency. The default contract consequently records
`one_way_latency_qualified: false`. A future one-way campaign must declare a
verified PTP/PHC or external synchronisation method and its maximum offset;
ordinary NTP is not promoted to a precision guarantee.

## Target throughput: canonical 50,000/s offer

Read-only plan, then the full supported matrix:

```bash
taskset -c 0-9 python3 bench/load/run-open-matrix.py --host scan --schedule-version 4 --list
taskset -c 0-9 python3 bench/load/run-open-matrix.py --host scan --schedule-version 4 \
  --out "/home/yann/libhft-bench-runs/target-open-$(date -u +%Y%m%dT%H%M%SZ)-$$"
```

Alternatively, select one cell, using `flavour/transport/scenario` order:

```bash
taskset -c 0-9 python3 bench/load/run-open-matrix.py --host scan --schedule-version 4 \
  --only '^cpp/kernel/single$' \
  --out "/home/yann/libhft-bench-runs/target-cpp-kernel-$(date -u +%Y%m%dT%H%M%SZ)-$$"
```

The driver explicitly selects `open`, `nos_er`, raw capture and a **50,000/s cell
aggregate**, with the shared-epoch v4 schedule and the standard workload. It removes inherited recipe overrides such as
`TPUT`, `ITERS` and `LOOP_MODE`; setting them cannot turn this into a rate sweep.
Single has a 50,000/session/s target. Each 32-session cell has four client
processes × eight sessions, with exactly **1,562.5/session/s**. All processes use
one prepared phase epoch; cell slots are uniformly spaced at 20 µs before runtime
scheduler delay. See [the v4 contract](OPEN-SCHEDULE-V4.md).

The version flag is deliberate: omitting it retains the historical v3 schedule
(integer 1,562/session/s, 49,984/s effective aggregate and locally aligned batches).
V4 writes only `results-scan-open-v4.tsv` for data. Its HTML is published at both
`latency-matrix-scan.html` (the current main page) and `latency-matrix-scan-open-v4.html`.
Missing cells never borrow v3 results. `gen-matrix.py scan` selects the newest available
schedule; `gen-matrix.py scan --schedule-version 3` regenerates the separate
`latency-matrix-scan-legacy-v3.html` archive without replacing the current page.

After each completed cell that passes the evidence audits, the driver imports its
result and updates the target matrix and handbook HTML. An audit failure stops the
sitting before import and preserves its evidence. This command **does publish**.
Source and installed-artifact changes are checked before every launch and before
publication. Use a fresh sitting for a retry; `--resume` is for an interrupted,
unchanged sitting.

The driver must itself be pinned entirely within the machine file's explicit
housekeeping set; its watcher inherits that affinity. Each attempt retains
`live-pinning-watch.log`, `live-pinning-console.log` and, on a passing gate,
`live-pinning-summary.json`. Before import it requires at least one completed
positive sample, no observed placement violation, valid complete sample records,
and a watcher that remained alive until the driver requested its graceful stop.
Startup/teardown INCONCLUSIVE samples remain visible, not relabelled as PASS.
A graceful stop finishes and retains the current sample. A crashed, stuck, empty,
entirely inconclusive or malformed watch rejects publication and keeps the old
matrix row. This is discrete global FIX-process sampling, not continuous residency
or proof that every client and measured phase was sampled. Per-role startup
readbacks remain a separate mandatory audit. Direct importer invocations do not
automatically create this live evidence; use the driver for the publishing sweep.

**Response time** runs from intended slot to completion (`resp_*`); **service time**
runs from actual task start to completion. The v4 matrix defaults to Response and lets
you select Service; the historical v3 archive still opens on Service. Both show
p50/p99/p99.9/max. Read deficits, shedding, scheduler
slip, stability and pinning qualifications alongside either latency view. A numeric
saturated observation is qualified evidence, not a passing load-capacity claim.
The current matrix separates incomplete delivery, target missed, complete with latency
variation, and other audit failures/warnings. CV% is sample standard deviation / mean
across the three per-run percentile values, computed per session; each percentile shows
the worst session's CV. The dropdown selects service or response CVs along with latency.
For the current v4 matrix, **complete** requires response p50 and p99 CV **both strictly
below 10%**; either CV at or above 10% gives **complete · variation**. Response p99.9 has
no CV threshold; service CVs are descriptive only, even in Service view. This rule
applies only after delivery, target and audit checks; it cannot clear their warnings or
failures. An unavailable CV is shown as “—” and **complete · CV unavailable**, never
assumed zero or included in best-result highlighting. Original TSV statuses, run-time
stability verdicts and historical matrices remain unchanged; the label is a separate
publication policy, not a benchmark rerun or a claim that the measured variation vanished.

## Maximum throughput: exploratory candidate trials only

There is no automated bracketing/refinement search, duration-based measurement
mode, maximum-rate acceptance policy, or maximum-results importer yet. The existing
`bench.sh` can collect candidate open rates without updating public results.
For example, these three **exploratory** C++/kernel/mono trials use fresh evidence
paths and approximately 30 seconds of intended measured schedule per run:

```bash
candidate_root=$(mktemp -d /home/yann/libhft-bench-runs/max-candidates-XXXXXX)
candidate_index=0
for candidate_rate in 50000 75000 100000; do
  candidate_iterations=$(( (candidate_rate * 30 + 31) / 32 ))
  candidate_port=$(( 48000 + 2 * candidate_index ))
  env -i HOME="$HOME" USER="$USER" LC_ALL=C \
    PATH=/usr/local/sbin:/usr/local/bin:/usr/sbin:/usr/bin:/sbin:/bin \
    BENCH_MACHINE=scan BENCH_STANDARDS="$PWD/bench/load/load-bench-standards.conf" \
    BENCH_RAW_SAMPLES=1 OPEN_SCHEDULE=cell-v4 LOOP_MODE=open WIRE=nos_er TPUT="$candidate_rate" \
    N=32 CLIENT_MODE=thread-per-core CLIENT_PROCS=4 \
    WARMUP=20000 ITERS="$candidate_iterations" RUNS=3 PORT_BASE="$candidate_port" \
    taskset -c 0-9 bash bench/load/bench.sh --scenario mono --only '^cpp/kernel$' \
      --out "$candidate_root/mono-cpp-kernel-$candidate_rate"
  candidate_index=$((candidate_index + 1))
done
```

`TPUT` is aggregate; `ITERS` is **per session, per measured run**. For `single`,
use division by 1, `N=1` and `CLIENT_PROCS=1`; for `pool` or `shard`, retain 32/4
and select that scenario. Use the same shape and arrival policy at every candidate.
`SECS` is a server lifetime budget, **not** a measured-duration control. Scaling
iterations equalizes intended schedule duration, not actual completion time or
steady-state behaviour. Use `ceil(rate * seconds / sessions)` and retain the exact
intended duration `sessions * iterations / rate`. Larger rates also increase raw
storage/memory costs: the current raw replay has a 16-million-measured-sample cell
limit, and the 192 MiB Java heap is not suitable for arbitrary long single-session
candidates. Preflight and disclose memory before a search; resource exhaustion is
not an engine-capacity result.

Inspect each candidate's recipe, client logs, raw samples, `rows.csv` and pinning
evidence. A harness exit code alone is not acceptance. Bracket a passing rate with
a higher failing rate, refine that interval, and repeat the proposed boundary.
Require complete validated replies, no shedding, compliant offered rate/scheduler
slip, repeatability and no accumulating backlog; final draining alone does not
prove sustained operation. Apply a response-latency limit **only if an SLO was
declared in advance**, and then call the result maximum throughput *within that SLO*.

Keep the in-flight cap and client resources fixed and disclosed. A client-generator
or cap limit is a limit of this configured experiment, not proof of engine capacity.
Report the last-pass/first-fail bracket rather than an exact universal maximum.
Result-line aggregate `achieved_per_sec` is the **minimum per-session rate**, not
cell-total throughput; do not relabel it as a total.

**Do not import candidate CSVs into `results-scan-open.tsv`.** Do not run
`ingest-rows.py ... scan --mode open` on them. The generic importer is not a
50,000-only firewall: consistent experimental rates can replace a target row.
The v4 target importer does enforce its canonical 50k/20k/20k/three-run profile;
do not route candidates through the historical importer to bypass that gate.
Keep these trials external until a separate maximum-results schema, admission
policy and publication workflow exist.

Completed v4 candidates can now be inspected with the read-only
[window and exact-metric verifiers](MAX-CANDIDATE-EVIDENCE.md). They check raw
chronology, per-run quantiles and logged aggregation without relaxing the target
importer. A separate state-only rate planner accepts externally audited verdicts;
none of these tools launches a sweep or turns numeric eligibility into admission.

## Temporary SCAN Mellanox-only selection

While the Solarflare DAC is removed (2026-09-08), SCAN's machine file explicitly
sets `KERNEL_NIC=mlx`. Kernel uses ordinary sockets on `mlxA/mlxB`, without VMA;
VMA and SocketXtreme retain their Mellanox routes. Earlier SCAN kernel results used
Solarflare, so retain the sitting's machine file and identify remeasured rows rather
than treating the NIC-path change as an engine-only comparison.

Select the original five flavours and omit Solarflare transports:

```bash
taskset -c 0-9 python3 bench/load/run-open-matrix.py --host scan --schedule-version 4 \
  --only '^(cpp|java-pure|java-jni|dotnet-pure|dotnet-native)/(kernel|vma|socketxtreme)/' \
  --list
```

This lists 46 cells. Replace `--list` with `--out /absolute/path/to/a/new/sitting`
to run them with per-cell HTML publication. When restoring the Solarflare kernel
comparison, set `KERNEL_NIC=sfc` in `load-bench-machine-scan.conf`, stage it with
`sudo /usr/local/sbin/libhft-bench-install bench-machine`, and start a fresh sitting.
Never change machine configuration or installed artifacts during a sitting.

## Refresh the HTML without running a benchmark

After changing labels or run notes, regenerate the scan matrix and handbook from
their templates and existing evidence:

```bash
taskset -c 0-9 python3 bench/load/gen-matrix.py scan
taskset -c 0-9 python3 bench/load/gen-matrix.py scan --schedule-version 4
taskset -c 0-9 python3 bench/load/gen-handbook.py
```

These commands render existing results; they do not run, validate or import a new
benchmark. Edit `matrix-template.html` and `load-bench-handbook-template.html`, not
their generated HTML files. Historical `closed` and `paced` data stay under
**Historical diagnostics**; renaming a view must not change a result's experiment.
