Skip to content

Latest commit

 

History

History
144 lines (110 loc) · 6.17 KB

File metadata and controls

144 lines (110 loc) · 6.17 KB

Performance

Read this first. fw100 targets 100 Gbps, but a target is not a result. This document separates (a) what was actually measured, (b) on what hardware, and (c) what is still a projection. Anything not measured on a real NIC is labelled NOT MEASURED.


1. What was measured (synthetic pipeline cost)

fwbench drives the full C pipeline (parse → flowcache → ACL → conntrack → NAT/PAT → rate-limit → LPM → checksum) using an in-memory PMD. There is no NIC, no kernel, no DPDK in this loop — it measures pipeline logic only.

Measurement host (the machine that produced the numbers below):

Item Value
CPU Intel Core i5-8350U @ 1.70 GHz (4 cores / 8 threads, max turbo 3.6 GHz)
TSC invariant, calibrated at ~1.896 GHz
RAM / NUMA 7.3 GB, single socket / single NUMA node, no hugepages
Governor set to performance for the run (cores at ~3.6 GHz); normally powersave
NIC none (in-memory PMD)
Packet size 64 B (synthetic UDP/TCP)
Build -O3 -march=native, sanitizers off
Pinning taskset -c 2 (one core)

Results (single core, ./build-rel/fwbench, performance governor; stable across 3 runs):

Scenario cycles/pkt Mpps (synthetic) Notes
cached allow ~116 ~16.3 established flow, flowcache hit
cached SNAT ~146 ~13.0 established flow, NAT reply resolve
new-flow SNAT ~210 ~8.9 conntrack create + PAT alloc (miss path)
drop flood ~144 ~13.2 SYN flood to a closed port, ACL drop

These are NOT MEASURED on a NIC. They are the cost of the pipeline logic in isolation. Real throughput adds NIC RX/TX, DMA, cache misses on real packet bytes, memory bandwidth and (for DPDK) poll-mode overhead.

Note: under the default powersave governor the same build measured ~20–40% slower and noisier (e.g. cached-allow ~132 cyc/pkt). Always benchmark with the performance governor and pinned cores, and never on battery.

Interpreting the cycles/packet

The TSC is invariant, so tsc_hz ticks == one wall-clock second. That makes Mpps = tsc_hz / cycles_per_pkt a real wall-clock, single-core rate for the pipeline logic (still with the in-memory PMD — not a NIC number). Example: 116 cycles → 1.896e9 / 116 ≈ 16.3 Mpps.

2. What 100 Gbps actually requires

Line-rate packet rates (worst case = smallest packets):

Frame size Wire rate Packets/sec
64 B 100 Gbps 148.8 Mpps
512 B 100 Gbps ~23.5 Mpps
1500 B 100 Gbps ~8.16 Mpps

(100 Gbps = 100e9 / ((preamble+IFG 20 B + frame) × 8).)

Core-count projection from the synthetic numbers — NOT MEASURED

Using the synthetic per-core rates above as an upper bound on what a real core could sustain (optimistic — real cores will do worse):

Scenario 64 B (148.8 Mpps) 1500 B (8.16 Mpps)
cached allow (~16.3 Mpps/core) ~10 cores < 1 core
cached SNAT (~13.0 Mpps/core) ~12 cores < 1 core
new-flow SNAT (~8.9 Mpps/core) ~17 cores ~1 core

These are projections, NOT MEASURED. They assume the pipeline cost is the only cost, which is false on real hardware. Treat them as a floor on the number of cores, not a promise.

Why the synthetic cycles/pkt is higher than an ideal DPDK datapath

A hand-tuned DPDK forwarder can hit ~20–50 cycles/pkt for trivial forwarding. fwbench reports ~166–430 because it runs the full stateful pipeline (conntrack, NAT, ACL, LPM, checksums), not plain forwarding. Closing that gap is future work (see §5).

3. What must be true to claim 100 Gbps (the checklist)

A 100 Gbps claim is only valid when all of the following are reported together with the number:

  1. Hardware: CPU model + frequency, RAM, NUMA topology, NIC model + driver.
  2. DPDK PMD in use (not tap/afpacket), RSS across ≥ the number of worker cores.
  3. Packet size and whether it is 64 B worst case.
  4. Measured packets/sec and Gbps (both), with the formula used.
  5. CPU utilization per core and whether cores are pinned/isolated.
  6. Packet loss (rx vs tx) over the measurement window.
  7. Connection count / NAT configuration (empty tables flatter the numbers).
  8. Tool + duration (e.g. dpdk-testpmd, TRex, pktgen, ≥ 30 s).

Until a run produces all eight on real hardware, the claim is NOT MEASURED.

4. How to reproduce the synthetic numbers

cmake -B build-rel -DCMAKE_BUILD_TYPE=Release     # -O3, sanitizers OFF
cmake --build build-rel -j"$(nproc)"

# Pin to full frequency + one isolated core for stable, comparable numbers:
for c in /sys/devices/system/cpu/cpu*/cpufreq/scaling_governor; do
    echo performance | sudo tee "$c" >/dev/null
done
taskset -c 2 ./build-rel/fwbench
# (restore your normal governor afterwards)

Sanitizer builds (ASan/UBSan/TSan) are for correctness, not performance — never benchmark them. The -march=native flag is used when supported; a portable fallback build is available if the toolchain rejects it.

5. Known optimization headroom (future work)

Ordered by expected impact, all currently NOT MEASURED:

  1. DPDK PMD + real NIC — the single biggest gap; tap/afpacket copy through the kernel and are not a 100 Gbps path.
  2. Batch checksums / vectorize parse — the parser and checksum fixups are scalar; SIMD would cut the cached-path cost.
  3. Hugepage mempools + NUMA-local pools — reduce TLB misses under load.
  4. Flow-cache sizing / 2-level cache — the direct-mapped cache can thrash on adversarial 5-tuple spreads.
  5. Busy-poll tuning + core isolation (isolcpus, SCHED_FIFO).

6. Bottom line

  • Measured today: pipeline-logic cost only, on an i5-8350U laptop, in-memory PMD, performance governor, pinned core — ~116 cycles/pkt cached-allow, ~16 Mpps/core synthetic; ~146 cycles/pkt cached-SNAT. NOT MEASURED on a NIC.
  • 100 Gbps at 64 B needs 148.8 Mpps; at 1500 B it needs ~8.16 Mpps.
  • The design (per-core lock-free path, flow cache, RCU/shard locks only on miss) is built to scale out across cores, but the 100 Gbps number itself is a target, not yet a measurement. It becomes a measurement only when §3 is satisfied on DPDK + real hardware.