Offline evaluation of decision policies that are never executed.
Your model recommends an action. Something else — a human, a legacy rule — takes a different one. The outcome of your recommendation is never observed, so RMSE and AUC cannot tell you whether the recommendation was better. This library gives you a metric that can.
from regret_eval import DecisionFrame, compare, persistence, shrinkage
frame = DecisionFrame(
observations, # one row per (context, action actually exercised)
group_col="group", # the context in which exactly one action is chosen
candidate_col="carrier",
utility_col="utility", # realised outcome, higher is better
weight_col="volume", # exposure of the group, for volume-weighted aggregation
)
compare(
frame,
{
"shrinkage": shrinkage(history, exposure_col="calls"),
"persistence": persistence(history, window=pd.Timedelta(days=1)),
},
reference="persistence",
)examples/quickstart.py runs exactly that on 600 synthetic decision groups. The
generator is seeded, so this output is reproducible with python examples/quickstart.py:
groups: 600 · decidable perimeter: 87.8%
equal_weighted volume_weighted n_groups coverage skill_equal_weighted skill_volume_weighted
shrinkage 0.080 0.075 477.0 0.905 0.189 0.124
persistence 0.099 0.085 469.0 0.890 0.000 0.000
random 0.150 0.149 527.0 1.000 -0.513 -0.745
Three things to read in that table. shrinkage removes 19% of persistence's regret
equal-weighted but only 12% volume-weighted — positive on both, which is the minimum
bar for deployability, and the gap between the two columns is the honest story:
shrinkage helps most where exposure is thin, and thin groups are precisely the ones
volume-weighting discounts. random sits 51–75% worse than the reference, which is
the only evidence that the metric discriminates at all. And coverage is below 1 for
the two fitted policies but exactly 1 for random: the fitted ones sometimes pick a
candidate that carried no traffic that day, and those groups are reported unobserved
rather than silently scored.
The posterior oracle. For each decision context you observe, after the fact, the realised utility of every candidate that was genuinely exercised during the window. The best of them is the oracle choice. The regret of a policy is the gap between the oracle's utility and the utility of the candidate the policy would have picked. Lower is better; zero means the policy matched the best available outcome.
The decidable perimeter. Where only one candidate was ever exercised, every
policy is forced into the same choice and scores zero. Including those contexts
inflates every policy — including a random one — and destroys the metric's
ability to discriminate. decidable_perimeter() restricts scoring to contexts
with at least two exercised candidates. This is the unflattering choice and the
only honest one, so report the share of contexts and the share of volume it
covers next to your results.
Two aggregations, always both. Equal-weighted answers does it decide well on
average?. Volume-weighted answers does it decide well where it costs?. A
policy can win one and lose the other; that policy is not deployable. Reporting
only the favourable column is the most common way these evaluations mislead, so
aggregate_regret() always returns both and compare() always prints both.
I built a routing decision-support system for a wholesale telecom carrier. It recommended, every fifteen minutes, which partner to send traffic to; a human approved. Because recommendations were never auto-applied, nothing about the recommended path was ever observed.
The first version predicted a quality rate by regression and ranked by predicted value. Prediction error looked fine. Then I compared it, on regret, to a deliberately naive reference — carry forward yesterday's best partner, no learning at all.
The heuristic won — the learned model's regret came out about 10% higher.
That result reoriented the whole project. Predicting well is not deciding well: a model can estimate every value accurately and still get the ordering wrong, and only the ordering determines the quality of a decision. Re-optimising the same model against regret rather than prediction error made it 7% worse at predicting and 10.2% better at deciding. Reformulating it as a learning-to-rank problem finally beat the reference.
None of that is visible if you measure prediction error. This library is the measurement apparatus, extracted and rewritten from scratch so it carries no domain data.
The four locks that keep an offline evaluation from being optimistic — an optimistic evaluation is more dangerous than no evaluation, because it licenses a deployment:
| Lock | Rule | Enforced by |
|---|---|---|
| Identical universe | Train on exactly the candidate set inference will see. Same filters, same eligibility rules. | your pipeline |
| Temporal validation | Train on the past, test on the immediate future. Never shuffle. | splits.py |
| Held-out policy selection | Any "use model A here, model B there" rule is learned on a slice of training data, never on the test set. | your pipeline |
| Explicit perimeter | Score only where a decision is genuinely possible, and say what fraction that is. | decidable_perimeter() |
And one diagnostic that costs nothing: the reversed split. Training on the future to test the past has no operational meaning. Its job is to tell you where your advantage comes from. If it comes from a drift the model learned and extrapolated, reversing time destroys it. If the advantage survives reversal, it comes from structure in the problem rather than from the period you happened to sample. In my case it not only survived, it was strongest there — which is what convinced me the effect was real.
assert_no_overlap() is written to be wired into CI. A split is logic like any
other and it will break silently the day someone reindexes a frame upstream.
| Baseline | What it is | Why it's there |
|---|---|---|
random_choice |
Uniform pick among exercised candidates | A control, not a competitor. If random scores near your model, the metric is broken and every comparison built on it is void. |
persistence |
Carry forward the trailing window's best | Three lines, no features. Startlingly hard to beat. Losing to it tells you something exact. |
shrinkage |
James–Stein style pull toward a pooled prior, in proportion to how thin the exposure is | Often the right production answer on low-exposure segments. Knowing when not to use a model is the skill. |
pip install -e ".[dev]"
pytest
python examples/quickstart.pyPython ≥ 3.10 · numpy · pandas. No other dependencies.
- Regret is computed against realised utilities, which carry their own noise. The absolute number therefore includes an irreducible component; only relative comparisons between policies judged by the same oracle are meaningful.
- The oracle sees only candidates that were actually exercised. A candidate nobody ever tried cannot be evaluated, and no amount of statistics fixes that — it needs deliberate exploration.
- Utilities must be comparable across contexts for equal-weighted aggregation to mean anything. Normalise before you aggregate.
MIT licensed.