Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 1 addition & 1 deletion ARCHITECTURE.md
Original file line number Diff line number Diff line change
Expand Up @@ -110,7 +110,7 @@ The source Region stays pinned until accepted replacement writes finish. Conditi

### I/O

Reads and writes use independent bounded engine pools. Reclaim has separate read lanes. POSIX uses positioned worker I/O; optional io_uring uses fixed-depth rings. Buffered I/O is the default. Direct mode aligns runtime record I/O and keeps control, recovery, and unavoidable remainder operations buffered. Locks cover bounded in-memory work and release before device I/O.
Reads and writes use independent bounded engine pools. Reclaim has separate read lanes. POSIX uses positioned worker I/O; optional io_uring uses fixed-depth rings. Ring count and in-flight depth are independent: extra rings split the same admission bound and do not add slots. Buffered I/O is the default. Direct mode aligns runtime record I/O and keeps control, recovery, and unavoidable remainder operations buffered. SQPOLL and IOPOLL are advanced per-pool opt-ins with different requirements; they stay off until selected explicitly. Locks cover bounded in-memory work and release before device I/O.

Each lane uses one concrete `IoEngine` for admission, submission, cancellation, statistics, and shutdown. Driver-specific constructors start POSIX workers or an io_uring driver behind the same bounded command and completion protocol. Callers share the engine through `Arc`; its final owner joins the workers. Submitted requests retain their buffers and capacity until actual completion, independently of the caller's wait deadline.

Expand Down
10 changes: 5 additions & 5 deletions BENCHMARK.md
Original file line number Diff line number Diff line change
Expand Up @@ -108,19 +108,19 @@ Repeat sizes in `CACHE_SOAK_VALUE_BYTES` to weight a production distribution. Us
| Backend | Variable suffix | Default (read / write / reclaim) |
| --------- | -------------------------------- | --------------------------------- |
| POSIX | `_POSIX_<ROLE>_WORKERS` | 4 / 4 / 1 |
| io_uring | `_IO_URING_<ROLE>_RINGS` | 4 / 4 / 1 |
| io_uring | `_IO_URING_<ROLE>_MAX_IN_FLIGHT` | 256 / 256 / 1 |
| io_uring | `_IO_URING_<ROLE>_RINGS` | 1 / 1 / 1 |
| io_uring | `_IO_URING_<ROLE>_MAX_IN_FLIGHT` | 64 / 64 / 1 |
| io_uring | `_IO_URING_<ROLE>_IOPOLL` | false |
| io_uring | `_IO_URING_<ROLE>_SQPOLL_MS` | absent (disabled) |
| io_uring | `_IO_URING_<ROLE>_SQPOLL_CPU` | absent (unpinned) |

Optional fill admission uses `_FILL_CONTROL=disabled|observe|adaptive` under the same three prefixes. Enabled modes require explicit `_FILL_BYTES_PER_SECOND` and `_FILL_OPERATIONS_PER_SECOND` ceilings. The harness prints the effective options and adds a `type=fill_control` report beside cache records. Compare the modes in alternating order with identical ceilings and traffic; record rejected and hypothetical fills alongside completed throughput. High ceilings help measure instrumentation overhead, while lower ceilings and injected storage stalls exercise admission behavior. Buffered macOS results do not qualify Linux NVMe or cgroup throttling.

Select the backend with `_IO_ENGINE=posix|io-uring`. POSIX workers bound concurrent operations. io_uring ring count and aggregate in-flight limit are independent; changing one does not rewrite the other. Ring count must not exceed the in-flight limit. IOPOLL requires `_IO_MODE=direct`; SQPOLL CPU requires an idle timeout. Only the selected backend's settings are read. Each harness prints the resulting `IoEngineOptions`, and buffer estimates use its actual concurrency. The request benchmark's default read-wait capacity follows the selected read pool's in-flight limit. POSIX Immediate L2 reads still cap `CACHE_BENCH_CLIENTS` at the POSIX read-worker count, so a worker sweep must set clients at least as high as the worker count.
Select the backend with `_IO_ENGINE=posix|io-uring`. POSIX workers bound concurrent operations. io_uring ring count and aggregate in-flight limit are independent; changing one does not rewrite the other. Start with one ring per pool and set `_IO_URING_READ_MAX_IN_FLIGHT` to the concurrent L2 get count (`CACHE_BENCH_CLIENTS` in the request benchmark). Leave write and reclaim unset so they follow library defaults. Extra rings split the same depth and do not add slots. Ring count must not exceed the in-flight limit. SQPOLL and IOPOLL environment variables are advanced follow-ups with different requirements: IOPOLL requires `_IO_MODE=direct`; SQPOLL CPU requires an idle timeout. Do not enable them in the same sweep as rings or in-flight depth, and do not enable both as a pair. Only the selected backend's settings are read. Each harness prints the resulting `IoEngineOptions`, and buffer estimates use its actual concurrency. The request benchmark's default read-wait capacity follows the selected read pool's in-flight limit. POSIX Immediate L2 reads still cap `CACHE_BENCH_CLIENTS` at the POSIX read-worker count, so a worker sweep must set clients at least as high as the worker count.

`CACHE_BENCH_STATS` is now `CACHE_BENCH_ACTIVITY_COUNTERS`. Machine-readable reports use `version=2`, renaming the cache record field `statistics_enabled` to `activity_counters_enabled`; the counter population is unchanged.

The previous `_READ_IO_WORKERS`, `_WRITE_IO_WORKERS`, and `_RECLAIM_WORKERS` variables are rejected with a migration error. For POSIX, use `_POSIX_<ROLE>_WORKERS`. For io_uring, specify ring count and total in-flight depth separately; to reproduce a previous non-default read/write worker value of N, use N rings and 64 × N in-flight requests. The default effective topology is unchanged.
The previous `_READ_IO_WORKERS`, `_WRITE_IO_WORKERS`, and `_RECLAIM_WORKERS` variables are rejected with a migration error. For POSIX, use `_POSIX_<ROLE>_WORKERS`. For io_uring, specify ring count and total in-flight depth separately. Unset io_uring knobs follow `IoUringOptions::default()`: one ring and 64 in-flight requests for read and write, and one ring and one in-flight request for reclaim. To reproduce a previous POSIX-shaped default of 4 workers, set 4 rings and 256 in-flight requests.

Capacity variables follow the library's terms: `_CAPACITY_MIB` is L2, `_L1_CAPACITY_MIB` is L1, `_MANAGED_MEMORY_LIMIT_MIB` is the overall managed-memory budget, and `_REGION_SIZE_MIB` selects Region size where supported. The former `_MEMORY_MIB`, `_L1_MIB`, `_L2_MIB`, and `_REGION_MIB` names are rejected with their replacements. `recovery_scale` uses the same L1 naming under `CACHE_RECOVERY`.

Expand All @@ -132,7 +132,7 @@ Capacity variables follow the library's terms: `_CAPACITY_MIB` is L2, `_L1_CAPAC
/var/tmp/cache2-qualification
```

The runner records the revision, machine, filesystem, block device, complete configuration, raw logs, medians, and checksums. It exercises buffered and Direct POSIX I/O, a POSIX Direct worker sweep with matching client depth, an experimental io_uring Direct in-flight sweep (one ring at 16/32/64 concurrent gets), mixed turnover, and final warm recovery. Qualification builds with `--features io-uring` and sets `CACHE_BENCH_ACTIVITY_COUNTERS=true` so `result phase=read_io` records in-flight peak, slot wait, payload bytes, and read amplification. IOPOLL is an optional follow-up (`CACHE_BENCH_IO_URING_READ_IOPOLL=true` with Direct I/O), not part of the default matrix.
The runner records the revision, machine, filesystem, block device, complete configuration, raw logs, medians, and checksums. It exercises buffered and Direct POSIX I/O, a POSIX Direct worker sweep with matching client depth, an experimental io_uring Direct in-flight sweep (one ring at 16/32/64 concurrent gets), mixed turnover, and final warm recovery. Qualification builds with `--features io-uring` and sets `CACHE_BENCH_ACTIVITY_COUNTERS=true` so `result phase=read_io` records in-flight peak, slot wait, payload bytes, and read amplification. SQPOLL and IOPOLL are advanced host follow-ups, not part of the default matrix; IOPOLL still needs `CACHE_BENCH_IO_URING_READ_IOPOLL=true` with Direct I/O.

A release pass requires:

Expand Down
1 change: 1 addition & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -14,6 +14,7 @@
### Improvements

- Background write and reclaim timeouts now enter a reversible `CacheHealth::Recovering` state: new cache fills return overload while existing reads and deletes remain available. Original requests keep their bounded buffers and Regions and are never resubmitted; all affected work must complete validation and publication before fills resume. `RuntimeOptions::io_recovery_timeout` defaults to `None` for recovery until completion or close; use `Some(duration)` to bound recovery or `Some(Duration::ZERO)` for immediate cancellation. Recovery checks at fixed one-second intervals. Close interrupts recovery and preserves the existing unfenced-write safeguards; drain may wait indefinitely. Actual I/O errors and invalid completions remain terminal.
- The I/O source for an unavailable `IoEngineOptions::IoUring` selection now names the missing crate feature, operating system, or architecture.

## v0.5.0 (2026-09-16)

Expand Down
40 changes: 27 additions & 13 deletions CONFIGURATION.md
Original file line number Diff line number Diff line change
Expand Up @@ -269,26 +269,34 @@ Buffered POSIX I/O is the portable production baseline. It benefits from the ker

Direct mode is Linux-only and requires `O_DIRECT` for aligned record I/O. It reduces page-cache duplication but can amplify small reads and exposes aligned I/O failures instead of silently falling back. Control and recovery operations remain buffered. Necessarily unaligned runtime remainders use the buffered compatibility path.

io_uring is feature-gated and experimental. Its three pools configure physical rings and aggregate execution bounds independently:
io_uring is feature-gated and experimental. Enable `cache2` with `features = ["io-uring"]` on a supported Linux target. `RuntimeOptions::io_mode` stays Buffered unless Direct is chosen separately. Start from `IoUringOptions::default()` and set only the read depth:

```rust
use cache2::{
IoEngineOptions, IoMode, IoUringOptions, IoUringSqPollOptions, RuntimeOptions,
};
use cache2::{IoEngineOptions, IoUringOptions, RuntimeOptions};

let mut sq_poll = IoUringSqPollOptions::new(2_000);
sq_poll.cpu = Some(4);
let mut io = IoUringOptions::default();
io.read.max_in_flight = 128;
io.read.sq_poll = Some(sq_poll);
io.read.max_in_flight = 32;
let mut runtime = RuntimeOptions::default();
runtime.io_engine = IoEngineOptions::IoUring(io);
runtime.io_mode = IoMode::Direct;
```

`rings` controls driver-thread and kernel-ring count. `max_in_flight` is the aggregate admission bound and is divided as evenly as possible across those rings. SQPOLL's idle value is milliseconds; optional CPU affinity applies to each ring in that pool. SQPOLL defaults off; requested flags fail explicitly when the kernel cannot provide them.
`rings` is driver-thread and kernel-ring count. `max_in_flight` is the aggregate admission bound and is divided as evenly as possible across those rings. Extra rings do not add execution slots. POSIX `read_workers` is both thread count and depth; map that depth to `read.max_in_flight` with `rings = 1`, not to `rings`. Configuration construction fails with `ErrorKind::Unsupported` if the crate feature, OS, or architecture is missing; the I/O source names the requirement. Kernel opcode and NODROP checks happen at open.

| Pool | `rings` | `max_in_flight` | Pattern |
| ------- | ------- | --------------------------------------- | ---------------------------------------------------- |
| Read | 1 | Concurrent L2 `get` futures | Start here; this is admission depth |
| Write | 1 | Leave the default 64 | Sequential fill; do not copy the read depth |
| Reclaim | 1 | Leave the default 1 | Raise only for observed Free-Region lag |

Keep `rings = 1` until a single driver thread is CPU-bound, then raise `rings` while holding `max_in_flight` fixed. `rings` must not exceed that pool's in-flight limit; each ring receives a floor or ceiling share of the same depth. Under `ReadAdmission::Immediate`, extra gets beyond read depth are busy misses, so do not set `max_in_flight` far above caller concurrency.

### Advanced io_uring polling

IOPOLL is an additional explicit per-pool opt-in through the `IoUringPoolOptions::io_poll` field. It requires `IoMode::Direct` and a filesystem and block device that support polling. While requests are outstanding the driver busy-polls the device and consumes CPU, and a cancellation stays advisory until the polled operation completes, so profile IOPOLL on the target host before adopting it.
Leave `IoUringPoolOptions::sq_poll` and `io_poll` off. They are independent advanced flags with different requirements and are easy to combine incorrectly; do not enable either while sweeping rings or in-flight depth, and do not turn both on because one of them helped.

SQPOLL starts a kernel submission thread. The idle value is milliseconds. Optional CPU affinity applies to every ring in that pool, so a multi-ring pool pins every polling thread to the same CPU. SQPOLL does not require Direct I/O. Open fails if the kernel lacks non-fixed SQPOLL files or cannot provide the requested flags.

IOPOLL busy-polls completions and consumes CPU while requests are outstanding. It requires `IoMode::Direct` and a filesystem and block device that support polling; configuration construction returns `InvalidInput` otherwise. Cancellation stays advisory until the polled operation completes. Profile IOPOLL alone on the target host.

### Statistics

Expand Down Expand Up @@ -367,11 +375,17 @@ Use this profile when the data set exceeds host RAM and the goal is device behav

- Keep Buffered POSIX as the portable production default; switch to Direct only after measuring amplification and warm-close cost on the target filesystem.
- Size POSIX `read_workers` to the application's concurrent L2 get depth. Worker count is both thread count and admission depth.
- For experimental io_uring, start with one read ring and set `max_in_flight` to that same concurrent get depth. Extra rings do not add execution slots.
- For experimental io_uring, do not scale `rings` with vCPU count. Keep one read ring and set `read.max_in_flight` to concurrent L2 gets. Leave write and reclaim at their defaults.
- Keep write execution modest; more write workers cannot exceed sequential device fill.
- Profile IOPOLL separately with `IoUringPoolOptions::io_poll` and Direct I/O on a polling-capable filesystem; do not enable it from ring or in-flight sweeps.
- Leave SQPOLL and IOPOLL off during ring and in-flight sweeps. They are independent advanced flags; profile at most one of them, and only after depth is set. IOPOLL still requires Direct I/O on a polling-capable filesystem.
- Watch `requests_in_flight_peak`, `slot_wait_ns`, `l2_read_busy_misses`, and Direct versus served-byte amplification instead of throughput alone.

Keep the library default of one ring per pool. Set `read.max_in_flight` to concurrent L2 gets; use 32 or 64 when that is the caller depth, and lower it when the caller issues fewer gets. Do not raise `rings` with vCPU count. Leave write and reclaim at their defaults. Leave SQPOLL and IOPOLL off.

| Read rings | Read `max_in_flight` | Write / reclaim | SQPOLL / IOPOLL |
| ---------- | ---------------------------- | ---------------------------------- | --------------- |
| 1 | Concurrent L2 gets, 32 or 64 | Library default (1 / 64 and 1 / 1) | Off |

## Diagnostic map

Use `Cache::snapshot()` for regular telemetry and `Cache::detailed_snapshot()` for periodic diagnosis.
Expand Down
17 changes: 16 additions & 1 deletion README.md
Original file line number Diff line number Diff line change
Expand Up @@ -89,7 +89,22 @@ Changing the append-shard count rebinds recovered Active Regions during a warm o

The default `ReadAdmission::Immediate` returns a miss under read-engine or buffer pressure. `ReadAdmission::Wait` enables a queue bounded by `max_waiters` and a positive `timeout`. Queued requests retain their read descriptor and allocate a buffer after admission. Queue saturation, memory pressure, and timeout return explicit overload.

Buffered POSIX I/O is the production path. Direct I/O is an explicit Linux mode. io_uring requires the `io-uring` feature and remains experimental; its ring count and aggregate in-flight limit are independent, so size `max_in_flight` (or POSIX `read_workers`) to concurrent L2 gets rather than adding rings. SQPOLL and IOPOLL are explicit per-pool opt-ins: SQPOLL adds kernel submission polling with configurable idle time and optional CPU affinity, while IOPOLL adds completion polling and requires direct I/O on polling-capable storage.
Buffered POSIX I/O is the production path. Direct I/O is an explicit Linux mode. io_uring requires the `io-uring` crate feature on a supported Linux target and remains experimental. Keep one ring per pool; set `read.max_in_flight` to concurrent L2 gets, and leave write and reclaim at their defaults. Extra rings split the same depth across driver threads and do not add slots.

```toml
cache2 = { version = "0.5", features = ["io-uring"] }
```

```rust
use cache2::{IoEngineOptions, IoUringOptions, RuntimeOptions};

let mut io = IoUringOptions::default();
io.read.max_in_flight = 32; // concurrent L2 get() futures
let mut runtime = RuntimeOptions::default();
runtime.io_engine = IoEngineOptions::IoUring(io);
```

Configuration construction fails with `ErrorKind::Unsupported` when the feature, OS, or architecture is missing, and the I/O source names which one. SQPOLL and IOPOLL are advanced per-pool flags with different requirements; leave both off unless a host profile needs one of them. See [I/O engine and mode](CONFIGURATION.md#io-engine-and-mode).

### Platform support

Expand Down
15 changes: 7 additions & 8 deletions benchmarks/src/config.rs
Original file line number Diff line number Diff line change
Expand Up @@ -50,9 +50,9 @@ pub fn io_engine_from_env(prefix: &str) -> io::Result<IoEngineOptions> {
}
"io-uring" => {
let mut options = IoUringOptions::default();
options.read = io_uring_pool(prefix, "READ", 4, 256)?;
options.write = io_uring_pool(prefix, "WRITE", 4, 256)?;
options.reclaim = io_uring_pool(prefix, "RECLAIM", 1, 1)?;
options.read = io_uring_pool(prefix, "READ", options.read)?;
options.write = io_uring_pool(prefix, "WRITE", options.write)?;
options.reclaim = io_uring_pool(prefix, "RECLAIM", options.reclaim)?;
Ok(IoEngineOptions::IoUring(options))
}
value => Err(invalid(format!("unsupported {name}: {value}"))),
Expand Down Expand Up @@ -196,13 +196,12 @@ pub fn parse_l1_eviction_policy(name: &str) -> io::Result<L1EvictionPolicy> {
fn io_uring_pool(
prefix: &str,
role: &str,
rings: usize,
max_in_flight: usize,
mut options: IoUringPoolOptions,
) -> io::Result<IoUringPoolOptions> {
let prefix = format!("{prefix}_IO_URING_{role}");
let mut options = IoUringPoolOptions::default();
options.rings = setting(&format!("{prefix}_RINGS"))?.unwrap_or(rings);
options.max_in_flight = setting(&format!("{prefix}_MAX_IN_FLIGHT"))?.unwrap_or(max_in_flight);
options.rings = setting(&format!("{prefix}_RINGS"))?.unwrap_or(options.rings);
options.max_in_flight =
setting(&format!("{prefix}_MAX_IN_FLIGHT"))?.unwrap_or(options.max_in_flight);
options.io_poll = env_bool(&format!("{prefix}_IOPOLL"), false)?;
let idle = setting(&format!("{prefix}_SQPOLL_MS"))?;
let cpu = setting(&format!("{prefix}_SQPOLL_CPU"))?;
Expand Down
Loading