Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
10 changes: 9 additions & 1 deletion ARCHITECTURE.md
Original file line number Diff line number Diff line change
Expand Up @@ -80,7 +80,7 @@ This layout converts small foreground mutations into ordered, batched Region wri
5. Return after bounded in-memory admission.
6. The shard worker writes sealed batches and publishes their L2 mappings only after write completion.

Full staging or short-path contention returns structured `ErrorKind::Overloaded`. Admission is shard-local. Success means accepted staging; `put_l2` becomes visible when publication completes. `drain` fences all mutations accepted before its operation barrier and waits for their Region writes and L2 publication, but does not issue the recovery durability syncs.
Full staging or short-path contention returns structured `ErrorKind::Overloaded`. Staging admission is shard-local; optional adaptive fill budgets are shared across shards. Success means accepted staging; `put_l2` becomes visible when publication completes. `drain` fences all mutations accepted before its operation barrier and waits for their Region writes and L2 publication, but does not issue the recovery durability syncs.

### `get`

Expand Down Expand Up @@ -114,6 +114,14 @@ Reads and writes use independent bounded engine pools. Reclaim has separate read

Each lane uses one concrete `IoEngine` for admission, submission, cancellation, statistics, and shutdown. Driver-specific constructors start POSIX workers or an io_uring driver behind the same bounded command and completion protocol. Callers share the engine through `Arc`; its final owner joins the workers. Submitted requests retain their buffers and capacity until actual completion, independently of the caller's wait deadline.

### Pre-timeout fill pressure

`io::fill_control` owns optional background observations and adaptive fill admission. `Disabled` retains the existing path. Enabled modes preallocate a fixed observation table sized by append and reclaim worker counts. Background workers register before engine admission and retain their observations through completion validation and publication. A full table skips that request and counts `dropped_observations` instead of failing I/O. Workers checkpoint at one quarter of the I/O deadline, capped at 500 ms, and pause new fills without entering timeout recovery. Aggregate validated throughput provides an approximate drain estimate for snapshots and does not drive pause decisions. Pressure is `Healthy` or `Paused` and does not change terminal health or prove a device fault.

Adaptive foreground `put`/`put_l2` load pause-holder state and otherwise only compete for staging. Byte and record ceilings are instance-wide, shared by every shard worker; they pace non-essential background flush. Urgent, drain, and rotation flushes always proceed. A flush that cannot take a span refunds the consumed budget and retries. `Observe` does not delay flush or reject fills and counts pause refusals as `would_reject`; it still checkpoints waits so pause is observable. Adaptive rejects new fills immediately while paused and skips optional reinsertion. Pause is released when the slow I/O completes, not after later publication. Reads, deletes, and essential reclaim bypass fill budgets. Close stops admission independently of outstanding I/O.

See [adaptive fill admission](CONFIGURATION.md#adaptive-fill-admission) for rate ceilings, pause conditions, bounded bursts, and tuning limits.

### Memory

The managed-memory limit covers the index mapping, heat bits, L1, append buffers, reclaim buffers, metadata, cache-owned thread stacks, recovery scratch, and transient reads. Total deployment memory additionally includes allocator metadata, Tokio, process overhead, and the kernel page cache. `CacheConfig::new` rejects invalid or insufficient memory budgets before file access; actual allocation can still fail during open.
Expand Down
2 changes: 2 additions & 0 deletions BENCHMARK.md
Original file line number Diff line number Diff line change
Expand Up @@ -114,6 +114,8 @@ Repeat sizes in `CACHE_SOAK_VALUE_BYTES` to weight a production distribution. Us
| io_uring | `_IO_URING_<ROLE>_SQPOLL_MS` | absent (disabled) |
| io_uring | `_IO_URING_<ROLE>_SQPOLL_CPU` | absent (unpinned) |

Optional fill admission uses `_FILL_CONTROL=disabled|observe|adaptive` under the same three prefixes. Enabled modes require explicit `_FILL_BYTES_PER_SECOND` and `_FILL_OPERATIONS_PER_SECOND` ceilings. The harness prints the effective options and adds a `type=fill_control` report beside cache records. Compare the modes in alternating order with identical ceilings and traffic; record rejected and hypothetical fills alongside completed throughput. High ceilings help measure instrumentation overhead, while lower ceilings and injected storage stalls exercise admission behavior. Buffered macOS results do not qualify Linux NVMe or cgroup throttling.

Select the backend with `_IO_ENGINE=posix|io-uring`. POSIX workers bound concurrent operations. io_uring ring count and aggregate in-flight limit are independent; changing one does not rewrite the other. Ring count must not exceed the in-flight limit. IOPOLL requires `_IO_MODE=direct`; SQPOLL CPU requires an idle timeout. Only the selected backend's settings are read. Each harness prints the resulting `IoEngineOptions`, and buffer estimates use its actual concurrency. The request benchmark's default read-wait capacity follows the selected read pool's in-flight limit.

`CACHE_BENCH_STATS` is now `CACHE_BENCH_ACTIVITY_COUNTERS`. Machine-readable reports use `version=2`, renaming the cache record field `statistics_enabled` to `activity_counters_enabled`; the counter population is unchanged.
Expand Down
4 changes: 4 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -7,6 +7,10 @@
- All optional statistics now use `RuntimeOptions::stats`: replace `runtime.statistics` with `runtime.stats.activity_counters` and `CacheSnapshot::statistics_enabled` with `activity_counters_enabled`. Activity counters, terminal request counters, and latency collection remain independent and disabled by default; counter semantics and the on-disk format are unchanged. See [the configuration migration guide](CONFIGURATION.md#migrating-from-05).
- The benchmarking `RegionIndexTurnoverReport` field `config` is renamed to `options`, matching the `RegionIndexTurnoverOptions` input it carries.

### Features

- Optional `RuntimeOptions::fill_control` observes background request age and estimated drain time before timeout. `Observe` reports pause refusals as `would_reject`. `Adaptive` pauses new `put`/`put_l2` fills when a background worker checkpoints an old I/O (one quarter of the deadline, capped at 500 ms) or timeout recovery is active, releases checkpoint pause when that I/O completes, paces non-essential background flush by encoded bytes and record count, and suppresses optional reinsertion while paused. Foreground admission loads pause-holder state and otherwise only competes for staging. A full observation table skips that request instead of failing I/O. Reads, deletes, accepted writes, and essential reclaim retain their paths. `CacheSnapshot::fill_control` exposes pressure and accounting independently of statistics. Controller storage is included in managed memory. Disabled by default; see [configuration](CONFIGURATION.md#adaptive-fill-admission) for ceilings and measurement limits.

### Improvements

- Background write and reclaim timeouts now enter a reversible `CacheHealth::Recovering` state: new cache fills return overload while existing reads and deletes remain available. Original requests keep their bounded buffers and Regions and are never resubmitted; all affected work must complete validation and publication before fills resume. `RuntimeOptions::io_recovery_timeout` defaults to `None` for recovery until completion or close; use `Some(duration)` to bound recovery or `Some(Duration::ZERO)` for immediate cancellation. Recovery checks at fixed one-second intervals. Close interrupts recovery and preserves the existing unfenced-write safeguards; drain may wait indefinitely. Actual I/O errors and invalid completions remain terminal.
Expand Down
Loading