Skip to content

Latest commit

 

History

10 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

propensity-scoring

CI Python 3.10+ License: MIT

Lead scoring that ends in a decision rather than a leaderboard.

pip install -e . && propensity train

Most propensity write-ups stop at an AUC. That number cannot tell you who to call. This one carries through to the part that decides whether the model was worth building: which leads do we contact, and what does that earn?

  policy                     threshold  contacted  precision   recall        profit
  ---------------------------------------------------------------------------------
  contact everyone              0.0000      2,969      0.107    1.000       -62,550
  threshold 0.5                 0.5000          3      1.000    0.009         3,150
  F1-optimal                    0.1290        649      0.280    0.571       121,050
  breakeven (deployable)        0.1250        674      0.274    0.580       120,900
  best possible (oracle)        0.1221        695      0.272    0.592       122,550

Contacting every lead loses 62,550. The default 0.5 threshold contacts three people out of nearly three thousand and is effectively the model switched off. The rule derived from the economics returns 120,900 — 98.6% of what a perfectly chosen threshold would have made.


Where the threshold comes from

Contacting a lead costs c. A conversion is worth v. Contacting a lead with conversion probability p has expected value p·v − c, which is positive exactly when

p > c / v

That is the whole derivation. At a 150 contact cost and 1,200 conversion value the threshold is 0.125 — nowhere near 0.5, and it moves the moment the economics move.

This is the argument for caring about calibration. The rule is a comparison against a probability, so it is only correct if the model's output is a probability. Applied to an uncalibrated score it is a threshold on a number that means nothing, and the decision is wrong in a direction AUC cannot show you.

Why not just optimise F1

Because F1 never sees the economics, so it cannot respond to them.

$ propensity sensitivity
    cost  breakeven  F1 thresh    F1 profit   breakeven profit      oracle
  ------------------------------------------------------------------------
      60      0.050      0.109      123,780            172,560     174,660
     100      0.083      0.109       92,700             95,700      99,400
     150      0.125      0.109       53,850             48,750      53,850
     200      0.167      0.109       15,000             19,400      22,600
     250      0.208      0.109      -23,850             11,850      12,000
     300      0.250      0.109      -62,700              4,800       5,700

The F1 threshold is 0.109 in every row. Quintuple the cost of a contact and it does not move, because balancing precision and recall equally is a claim about the economics made without reference to them. By the bottom two rows it is losing money while the breakeven rule is still profitable.

There is a test that pins this, so it cannot quietly stop being true:

cheap     = f1_optimal_threshold(y, p, Economics(1200, 60)).threshold
expensive = f1_optimal_threshold(y, p, Economics(1200, 300)).threshold
assert cheap == pytest.approx(expensive)

Ranking quality

  PR-AUC              0.296   (2.76x the 0.107 base rate)
  ROC-AUC             0.771   flattering at this imbalance
  Brier               0.0869
  Lift @ top 10%      3.13x
  Recall @ top 10%    0.313

PR-AUC leads and ROC-AUC follows, deliberately. At a ~10% base rate the false positive rate has the whole negative class in its denominator, so a few hundred bad leads barely move ROC-AUC while the shortlist handed to a sales team fills with people who will never convert. Precision-recall uses no true negatives, so it degrades exactly when the shortlist does.

PR-AUC's baseline is the base rate, not 0.5, which is why the raw number is meaningless on its own — 0.296 against a 0.107 base rate is 2.76x random.

Calibration

  expected calibration error 0.0238

           bin      n  predicted  observed     gap
  ------------------------------------------------
   0.02-0.03     297      0.028     0.007   0.021
   0.05-0.06     297      0.051     0.047   0.004
   0.06-0.07     296      0.061     0.061   0.000
   0.13-0.19     297      0.158     0.226   0.068
   0.19-0.54     297      0.274     0.337   0.062

Bins are quantile-based, because a skewed score distribution leaves uniform-width bins nearly empty at the top and an ECE computed over three samples is noise.

The model is under-confident at the top end — it says 0.27 for leads that convert 34% of the time. That is visible here and completely invisible to AUC, since any monotone transform of the scores leaves the ranking untouched. There is a test for that too.

Is the booster earning its complexity?

$ propensity compare-models
  model                  PR-AUC   ROC-AUC    Brier  lift@10%      ECE
  ------------------------------------------------------------------
  logistic                0.287     0.766   0.0872      2.88   0.0224
  gradient_boosting       0.296     0.771   0.0869      3.13   0.0238

+0.009 PR-AUC. Real, but modest — and the verdict printed underneath that table is computed from the numbers, not written into the source, so it cannot drift into a claim the data contradicts.

What a random split would have told you

$ propensity compare-splits
  both scored on the same 1,484 held-out leads from the final 3 months
  leaky training set contains 1,274 leads from that same period

  training set                   PR-AUC   ROC-AUC  lift@10%
  --------------------------------------------------------
  earlier months only             0.296     0.761      3.46
  random sample of all months     0.318     0.773      3.46

Both models are scored on the same rows. The only difference is that one was allowed to train on leads from the period it is being judged on. That is worth +0.022 PR-AUC, and it is exactly what a random train/test split hands the model before reporting it as skill.

Holding the evaluation set fixed matters. Comparing a temporal holdout against a random holdout compares two different test distributions and measures nothing.


The data

Synthetic, so every number above is reproducible with one command — and so none of them are evidence about real leads. It is built to contain the problems that make this more than a fit call:

  • Imbalance — about 1 in 10 converts, so accuracy is useless
  • Temporal drift — paid search degrades over the year, referral improves
  • Non-linearity — engagement matters more for large accounts, and contact attempts help then start to hurt
  • Missing not at random — company size is absent more often for self-serve signups, and that absence predicts
  • Six noise features — so feature importance has a chance to embarrass itself

The pipeline

Everything that learns lives inside one scikit-learn Pipeline. Not tidiness — an imputer fitted before the split has already seen the holdout's median, and the resulting optimism is unmeasurable afterwards.

  • Median imputation with a missingness indicator, because the absence is signal
  • Rare categories grouped below 1% frequency; unseen values route there at inference instead of raising
  • HistGradientBoostingClassifier — the same histogram algorithm as LightGBM and XGBoost's hist mode, keeping dependencies to numpy and scikit-learn

Commands

propensity train                          # metrics, calibration, decision
propensity sensitivity                    # how thresholds respond to economics
propensity compare-models                 # booster against a linear baseline
propensity compare-splits                 # the cost of training on the future
propensity card -o MODEL_CARD.md          # generated, never hand-written

--value and --cost set the economics; --rows, --seed, --holdout-months and --model control the rest.

The model card

propensity card generates it from the fitted model, so it cannot drift away from what it describes. A hand-written model card is accurate exactly once. It records the training data, the temporal split, performance, the calibration curve, the decision rule, and — the section people skip — what the model is not licensed for.

Honest limitations

  • Synthetic data. The generating process is known, which makes the numbers reproducible and none of them external evidence.
  • The oracle threshold uses holdout outcomes. It is an upper bound for comparison, not a deployable policy.
  • Economics are constant per lead. Real contract values vary; a per-lead value estimate would move the threshold per lead, which is a better system and a harder one.
  • The booster's margin is small. On this data a linear model is nearly as good and far easier to explain to a stakeholder. That is a real finding, not a failure.
  • No monitoring. Drift is present by construction, so in production the temporal holdout and the calibration curve would both need refreshing on a schedule.

Development

pip install -e ".[dev]"
pytest

60 tests. The load-bearing ones assert the economics: that breakeven is c/v, that the F1 threshold does not move when the economics do, that breakeven beats F1 when contacts are expensive, that miscalibration is invisible to AUC, and that preprocessing cannot leak.

License

MIT

About

Lead propensity modelling that ends in a decision: calibrated probabilities, an expected-value threshold derived from the economics, and a generated model card.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages