Read this first. fw100 targets 100 Gbps, but a target is not a result. This document separates (a) what was actually measured, (b) on what hardware, and (c) what is still a projection. Anything not measured on a real NIC is labelled NOT MEASURED.
fwbench drives the full C pipeline (parse → flowcache → ACL → conntrack →
NAT/PAT → rate-limit → LPM → checksum) using an in-memory PMD. There is
no NIC, no kernel, no DPDK in this loop — it measures pipeline logic only.
Measurement host (the machine that produced the numbers below):
| Item | Value |
|---|---|
| CPU | Intel Core i5-8350U @ 1.70 GHz (4 cores / 8 threads, max turbo 3.6 GHz) |
| TSC | invariant, calibrated at ~1.896 GHz |
| RAM / NUMA | 7.3 GB, single socket / single NUMA node, no hugepages |
| Governor | set to performance for the run (cores at ~3.6 GHz); normally powersave |
| NIC | none (in-memory PMD) |
| Packet size | 64 B (synthetic UDP/TCP) |
| Build | -O3 -march=native, sanitizers off |
| Pinning | taskset -c 2 (one core) |
Results (single core, ./build-rel/fwbench, performance governor; stable
across 3 runs):
| Scenario | cycles/pkt | Mpps (synthetic) | Notes |
|---|---|---|---|
| cached allow | ~116 | ~16.3 | established flow, flowcache hit |
| cached SNAT | ~146 | ~13.0 | established flow, NAT reply resolve |
| new-flow SNAT | ~210 | ~8.9 | conntrack create + PAT alloc (miss path) |
| drop flood | ~144 | ~13.2 | SYN flood to a closed port, ACL drop |
These are NOT MEASURED on a NIC. They are the cost of the pipeline logic in isolation. Real throughput adds NIC RX/TX, DMA, cache misses on real packet bytes, memory bandwidth and (for DPDK) poll-mode overhead.
Note: under the default
powersavegovernor the same build measured ~20–40% slower and noisier (e.g. cached-allow ~132 cyc/pkt). Always benchmark with theperformancegovernor and pinned cores, and never on battery.
The TSC is invariant, so tsc_hz ticks == one wall-clock second. That makes
Mpps = tsc_hz / cycles_per_pkt a real wall-clock, single-core rate for the
pipeline logic (still with the in-memory PMD — not a NIC number). Example:
116 cycles → 1.896e9 / 116 ≈ 16.3 Mpps.
Line-rate packet rates (worst case = smallest packets):
| Frame size | Wire rate | Packets/sec |
|---|---|---|
| 64 B | 100 Gbps | 148.8 Mpps |
| 512 B | 100 Gbps | ~23.5 Mpps |
| 1500 B | 100 Gbps | ~8.16 Mpps |
(100 Gbps = 100e9 / ((preamble+IFG 20 B + frame) × 8).)
Using the synthetic per-core rates above as an upper bound on what a real core could sustain (optimistic — real cores will do worse):
| Scenario | 64 B (148.8 Mpps) | 1500 B (8.16 Mpps) |
|---|---|---|
| cached allow (~16.3 Mpps/core) | ~10 cores | < 1 core |
| cached SNAT (~13.0 Mpps/core) | ~12 cores | < 1 core |
| new-flow SNAT (~8.9 Mpps/core) | ~17 cores | ~1 core |
These are projections, NOT MEASURED. They assume the pipeline cost is the only cost, which is false on real hardware. Treat them as a floor on the number of cores, not a promise.
A hand-tuned DPDK forwarder can hit ~20–50 cycles/pkt for trivial forwarding. fwbench reports ~166–430 because it runs the full stateful pipeline (conntrack, NAT, ACL, LPM, checksums), not plain forwarding. Closing that gap is future work (see §5).
A 100 Gbps claim is only valid when all of the following are reported together with the number:
- Hardware: CPU model + frequency, RAM, NUMA topology, NIC model + driver.
- DPDK PMD in use (not tap/afpacket), RSS across ≥ the number of worker cores.
- Packet size and whether it is 64 B worst case.
- Measured packets/sec and Gbps (both), with the formula used.
- CPU utilization per core and whether cores are pinned/isolated.
- Packet loss (rx vs tx) over the measurement window.
- Connection count / NAT configuration (empty tables flatter the numbers).
- Tool + duration (e.g.
dpdk-testpmd, TRex,pktgen, ≥ 30 s).
Until a run produces all eight on real hardware, the claim is NOT MEASURED.
cmake -B build-rel -DCMAKE_BUILD_TYPE=Release # -O3, sanitizers OFF
cmake --build build-rel -j"$(nproc)"
# Pin to full frequency + one isolated core for stable, comparable numbers:
for c in /sys/devices/system/cpu/cpu*/cpufreq/scaling_governor; do
echo performance | sudo tee "$c" >/dev/null
done
taskset -c 2 ./build-rel/fwbench
# (restore your normal governor afterwards)Sanitizer builds (ASan/UBSan/TSan) are for correctness, not performance —
never benchmark them. The -march=native flag is used when supported; a
portable fallback build is available if the toolchain rejects it.
Ordered by expected impact, all currently NOT MEASURED:
- DPDK PMD + real NIC — the single biggest gap; tap/afpacket copy through the kernel and are not a 100 Gbps path.
- Batch checksums / vectorize parse — the parser and checksum fixups are scalar; SIMD would cut the cached-path cost.
- Hugepage mempools + NUMA-local pools — reduce TLB misses under load.
- Flow-cache sizing / 2-level cache — the direct-mapped cache can thrash on adversarial 5-tuple spreads.
- Busy-poll tuning + core isolation (
isolcpus,SCHED_FIFO).
- Measured today: pipeline-logic cost only, on an i5-8350U laptop, in-memory
PMD,
performancegovernor, pinned core — ~116 cycles/pkt cached-allow, ~16 Mpps/core synthetic; ~146 cycles/pkt cached-SNAT. NOT MEASURED on a NIC. - 100 Gbps at 64 B needs 148.8 Mpps; at 1500 B it needs ~8.16 Mpps.
- The design (per-core lock-free path, flow cache, RCU/shard locks only on miss) is built to scale out across cores, but the 100 Gbps number itself is a target, not yet a measurement. It becomes a measurement only when §3 is satisfied on DPDK + real hardware.