Skip to content

Latest commit

 

History

1 Commit

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

regret-eval

Offline evaluation of decision policies that are never executed.

Your model recommends an action. Something else — a human, a legacy rule — takes a different one. The outcome of your recommendation is never observed, so RMSE and AUC cannot tell you whether the recommendation was better. This library gives you a metric that can.

from regret_eval import DecisionFrame, compare, persistence, shrinkage

frame = DecisionFrame(
    observations,              # one row per (context, action actually exercised)
    group_col="group",         # the context in which exactly one action is chosen
    candidate_col="carrier",
    utility_col="utility",     # realised outcome, higher is better
    weight_col="volume",       # exposure of the group, for volume-weighted aggregation
)

compare(
    frame,
    {
        "shrinkage": shrinkage(history, exposure_col="calls"),
        "persistence": persistence(history, window=pd.Timedelta(days=1)),
    },
    reference="persistence",
)

examples/quickstart.py runs exactly that on 600 synthetic decision groups. The generator is seeded, so this output is reproducible with python examples/quickstart.py:

groups: 600 · decidable perimeter: 87.8%

             equal_weighted  volume_weighted  n_groups  coverage  skill_equal_weighted  skill_volume_weighted
shrinkage             0.080            0.075     477.0     0.905                 0.189                  0.124
persistence           0.099            0.085     469.0     0.890                 0.000                  0.000
random                0.150            0.149     527.0     1.000                -0.513                 -0.745

Three things to read in that table. shrinkage removes 19% of persistence's regret equal-weighted but only 12% volume-weighted — positive on both, which is the minimum bar for deployability, and the gap between the two columns is the honest story: shrinkage helps most where exposure is thin, and thin groups are precisely the ones volume-weighting discounts. random sits 51–75% worse than the reference, which is the only evidence that the metric discriminates at all. And coverage is below 1 for the two fitted policies but exactly 1 for random: the fitted ones sometimes pick a candidate that carried no traffic that day, and those groups are reported unobserved rather than silently scored.


The idea in three paragraphs

The posterior oracle. For each decision context you observe, after the fact, the realised utility of every candidate that was genuinely exercised during the window. The best of them is the oracle choice. The regret of a policy is the gap between the oracle's utility and the utility of the candidate the policy would have picked. Lower is better; zero means the policy matched the best available outcome.

The decidable perimeter. Where only one candidate was ever exercised, every policy is forced into the same choice and scores zero. Including those contexts inflates every policy — including a random one — and destroys the metric's ability to discriminate. decidable_perimeter() restricts scoring to contexts with at least two exercised candidates. This is the unflattering choice and the only honest one, so report the share of contexts and the share of volume it covers next to your results.

Two aggregations, always both. Equal-weighted answers does it decide well on average?. Volume-weighted answers does it decide well where it costs?. A policy can win one and lose the other; that policy is not deployable. Reporting only the favourable column is the most common way these evaluations mislead, so aggregate_regret() always returns both and compare() always prints both.


Why this exists

I built a routing decision-support system for a wholesale telecom carrier. It recommended, every fifteen minutes, which partner to send traffic to; a human approved. Because recommendations were never auto-applied, nothing about the recommended path was ever observed.

The first version predicted a quality rate by regression and ranked by predicted value. Prediction error looked fine. Then I compared it, on regret, to a deliberately naive reference — carry forward yesterday's best partner, no learning at all.

The heuristic won — the learned model's regret came out about 10% higher.

That result reoriented the whole project. Predicting well is not deciding well: a model can estimate every value accurately and still get the ordering wrong, and only the ordering determines the quality of a decision. Re-optimising the same model against regret rather than prediction error made it 7% worse at predicting and 10.2% better at deciding. Reformulating it as a learning-to-rank problem finally beat the reference.

None of that is visible if you measure prediction error. This library is the measurement apparatus, extracted and rewritten from scratch so it carries no domain data.


Guard rails worth stealing

The four locks that keep an offline evaluation from being optimistic — an optimistic evaluation is more dangerous than no evaluation, because it licenses a deployment:

Lock Rule Enforced by
Identical universe Train on exactly the candidate set inference will see. Same filters, same eligibility rules. your pipeline
Temporal validation Train on the past, test on the immediate future. Never shuffle. splits.py
Held-out policy selection Any "use model A here, model B there" rule is learned on a slice of training data, never on the test set. your pipeline
Explicit perimeter Score only where a decision is genuinely possible, and say what fraction that is. decidable_perimeter()

And one diagnostic that costs nothing: the reversed split. Training on the future to test the past has no operational meaning. Its job is to tell you where your advantage comes from. If it comes from a drift the model learned and extrapolated, reversing time destroys it. If the advantage survives reversal, it comes from structure in the problem rather than from the period you happened to sample. In my case it not only survived, it was strongest there — which is what convinced me the effect was real.

assert_no_overlap() is written to be wired into CI. A split is logic like any other and it will break silently the day someone reindexes a frame upstream.


Baselines you have to beat

Baseline What it is Why it's there
random_choice Uniform pick among exercised candidates A control, not a competitor. If random scores near your model, the metric is broken and every comparison built on it is void.
persistence Carry forward the trailing window's best Three lines, no features. Startlingly hard to beat. Losing to it tells you something exact.
shrinkage James–Stein style pull toward a pooled prior, in proportion to how thin the exposure is Often the right production answer on low-exposure segments. Knowing when not to use a model is the skill.

Install

pip install -e ".[dev]"
pytest
python examples/quickstart.py

Python ≥ 3.10 · numpy · pandas. No other dependencies.

Limits, stated up front

  • Regret is computed against realised utilities, which carry their own noise. The absolute number therefore includes an irreducible component; only relative comparisons between policies judged by the same oracle are meaningful.
  • The oracle sees only candidates that were actually exercised. A candidate nobody ever tried cannot be evaluated, and no amount of statistics fixes that — it needs deliberate exploration.
  • Utilities must be comparable across contexts for equal-weighted aggregation to mean anything. Normalise before you aggregate.

MIT licensed.

About

Offline evaluation of decision policies that are never executed: posterior oracle, regret, temporal splits.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages