Lead scoring that ends in a decision rather than a leaderboard.
pip install -e . && propensity trainMost propensity write-ups stop at an AUC. That number cannot tell you who to call. This one carries through to the part that decides whether the model was worth building: which leads do we contact, and what does that earn?
policy threshold contacted precision recall profit
---------------------------------------------------------------------------------
contact everyone 0.0000 2,969 0.107 1.000 -62,550
threshold 0.5 0.5000 3 1.000 0.009 3,150
F1-optimal 0.1290 649 0.280 0.571 121,050
breakeven (deployable) 0.1250 674 0.274 0.580 120,900
best possible (oracle) 0.1221 695 0.272 0.592 122,550
Contacting every lead loses 62,550. The default 0.5 threshold contacts three people out of nearly three thousand and is effectively the model switched off. The rule derived from the economics returns 120,900 — 98.6% of what a perfectly chosen threshold would have made.
Contacting a lead costs c. A conversion is worth v. Contacting a lead with conversion probability p has expected value p·v − c, which is positive exactly when
p > c / v
That is the whole derivation. At a 150 contact cost and 1,200 conversion value the threshold is 0.125 — nowhere near 0.5, and it moves the moment the economics move.
This is the argument for caring about calibration. The rule is a comparison against a probability, so it is only correct if the model's output is a probability. Applied to an uncalibrated score it is a threshold on a number that means nothing, and the decision is wrong in a direction AUC cannot show you.
Because F1 never sees the economics, so it cannot respond to them.
$ propensity sensitivity
cost breakeven F1 thresh F1 profit breakeven profit oracle
------------------------------------------------------------------------
60 0.050 0.109 123,780 172,560 174,660
100 0.083 0.109 92,700 95,700 99,400
150 0.125 0.109 53,850 48,750 53,850
200 0.167 0.109 15,000 19,400 22,600
250 0.208 0.109 -23,850 11,850 12,000
300 0.250 0.109 -62,700 4,800 5,700The F1 threshold is 0.109 in every row. Quintuple the cost of a contact and it does not move, because balancing precision and recall equally is a claim about the economics made without reference to them. By the bottom two rows it is losing money while the breakeven rule is still profitable.
There is a test that pins this, so it cannot quietly stop being true:
cheap = f1_optimal_threshold(y, p, Economics(1200, 60)).threshold
expensive = f1_optimal_threshold(y, p, Economics(1200, 300)).threshold
assert cheap == pytest.approx(expensive) PR-AUC 0.296 (2.76x the 0.107 base rate)
ROC-AUC 0.771 flattering at this imbalance
Brier 0.0869
Lift @ top 10% 3.13x
Recall @ top 10% 0.313
PR-AUC leads and ROC-AUC follows, deliberately. At a ~10% base rate the false positive rate has the whole negative class in its denominator, so a few hundred bad leads barely move ROC-AUC while the shortlist handed to a sales team fills with people who will never convert. Precision-recall uses no true negatives, so it degrades exactly when the shortlist does.
PR-AUC's baseline is the base rate, not 0.5, which is why the raw number is meaningless on its own — 0.296 against a 0.107 base rate is 2.76x random.
expected calibration error 0.0238
bin n predicted observed gap
------------------------------------------------
0.02-0.03 297 0.028 0.007 0.021
0.05-0.06 297 0.051 0.047 0.004
0.06-0.07 296 0.061 0.061 0.000
0.13-0.19 297 0.158 0.226 0.068
0.19-0.54 297 0.274 0.337 0.062
Bins are quantile-based, because a skewed score distribution leaves uniform-width bins nearly empty at the top and an ECE computed over three samples is noise.
The model is under-confident at the top end — it says 0.27 for leads that convert 34% of the time. That is visible here and completely invisible to AUC, since any monotone transform of the scores leaves the ranking untouched. There is a test for that too.
$ propensity compare-models
model PR-AUC ROC-AUC Brier lift@10% ECE
------------------------------------------------------------------
logistic 0.287 0.766 0.0872 2.88 0.0224
gradient_boosting 0.296 0.771 0.0869 3.13 0.0238+0.009 PR-AUC. Real, but modest — and the verdict printed underneath that table is computed from the numbers, not written into the source, so it cannot drift into a claim the data contradicts.
$ propensity compare-splits
both scored on the same 1,484 held-out leads from the final 3 months
leaky training set contains 1,274 leads from that same period
training set PR-AUC ROC-AUC lift@10%
--------------------------------------------------------
earlier months only 0.296 0.761 3.46
random sample of all months 0.318 0.773 3.46Both models are scored on the same rows. The only difference is that one was allowed to train on leads from the period it is being judged on. That is worth +0.022 PR-AUC, and it is exactly what a random train/test split hands the model before reporting it as skill.
Holding the evaluation set fixed matters. Comparing a temporal holdout against a random holdout compares two different test distributions and measures nothing.
Synthetic, so every number above is reproducible with one command — and so none of them are evidence about real leads. It is built to contain the problems that make this more than a fit call:
- Imbalance — about 1 in 10 converts, so accuracy is useless
- Temporal drift — paid search degrades over the year, referral improves
- Non-linearity — engagement matters more for large accounts, and contact attempts help then start to hurt
- Missing not at random — company size is absent more often for self-serve signups, and that absence predicts
- Six noise features — so feature importance has a chance to embarrass itself
Everything that learns lives inside one scikit-learn Pipeline. Not tidiness — an imputer fitted before the split has already seen the holdout's median, and the resulting optimism is unmeasurable afterwards.
- Median imputation with a missingness indicator, because the absence is signal
- Rare categories grouped below 1% frequency; unseen values route there at inference instead of raising
HistGradientBoostingClassifier— the same histogram algorithm as LightGBM and XGBoost'shistmode, keeping dependencies to numpy and scikit-learn
propensity train # metrics, calibration, decision
propensity sensitivity # how thresholds respond to economics
propensity compare-models # booster against a linear baseline
propensity compare-splits # the cost of training on the future
propensity card -o MODEL_CARD.md # generated, never hand-written--value and --cost set the economics; --rows, --seed, --holdout-months and --model control the rest.
propensity card generates it from the fitted model, so it cannot drift away from what it describes. A hand-written model card is accurate exactly once. It records the training data, the temporal split, performance, the calibration curve, the decision rule, and — the section people skip — what the model is not licensed for.
- Synthetic data. The generating process is known, which makes the numbers reproducible and none of them external evidence.
- The oracle threshold uses holdout outcomes. It is an upper bound for comparison, not a deployable policy.
- Economics are constant per lead. Real contract values vary; a per-lead value estimate would move the threshold per lead, which is a better system and a harder one.
- The booster's margin is small. On this data a linear model is nearly as good and far easier to explain to a stakeholder. That is a real finding, not a failure.
- No monitoring. Drift is present by construction, so in production the temporal holdout and the calibration curve would both need refreshing on a schedule.
pip install -e ".[dev]"
pytest60 tests. The load-bearing ones assert the economics: that breakeven is c/v, that the F1 threshold does not move when the economics do, that breakeven beats F1 when contacts are expensive, that miscalibration is invisible to AUC, and that preprocessing cannot leak.
MIT