Skip to content

Repository files navigation

NextLM B2B Propensity Benchmark

CI License: MIT DOI

Cite this work

Anzalone, C., & Soni, H. (2026). NextLM Savant 3.5 Wins Top-Decile Capture and Lift Against Nine Frontier Flagship Models at Orders of Magnitude Lower Cost. Zenodo. https://doi.org/10.5281/zenodo.22076381

This repository publishes the grading methodology behind NextLM's blinded B2B propensity benchmark. It lets readers inspect and run the exact prompt contract, metric definitions, equal-business aggregation, and paired bootstrap used in the report:

NextLM Savant 3.5 Wins Top-Decile Capture and Lift Against Nine Frontier Flagship Models at Orders of Magnitude Lower Cost: A Blinded 18-Business B2B Propensity Benchmark

It does not contain the proprietary system, customer data, model weights, training recipe, or production implementation. This is the scorecard, not the scorer.

What is public

  • The exact zero-shot system instruction, record template, batch envelope, and response contract.
  • The exact eight-shot injection format and fold-safe example-selection rule, with fabricated examples in the same wire shape.
  • Tie-aware ROC-AUC and non-interpolated average precision (PR-AUC).
  • Lift@10 and Capture@10 with deterministic top-decile selection.
  • Equal-business macro-averaging across the 11 evaluable businesses.
  • The paired 5,000-draw business bootstrap with seed 20260812.
  • A deterministic 528-row dummy CSV using the real evaluator-facing schema and covering 11 wholly fabricated businesses.
  • A hashed Python 3.12 dependency lock plus an explicit PCG64/linear-quantile statistical runtime contract.
  • Tests for metric, tie, aggregation, eligibility, and bootstrap behavior.

What is intentionally private

  • The 36,704 real prospect records and sealed conversion outcomes.
  • Real seller/ICP context and the contents of the eight training-only examples.
  • Tenant identities, per-business benchmark results, and scoring outputs.
  • Savant model weights, Nemotron fine-tuning or adaptation recipes, structured enrichment, calibration, residual logic, Mamba-2/Mixture-of-Experts integration, deployment configuration, and product source code.

The real records and outcomes are proprietary to NextLM and its customers. The dummy data is fabricated and makes no empirical claim about any model.

Study basis

The current cohort contains 36,704 prospects across 18 businesses. A business enters the common four-metric comparison only when it has at least 20 positive and 20 negative outcomes; 11 businesses meet that frozen rule. Every system is evaluated on those same businesses. Metrics are calculated independently within each business and then averaged with equal weight, so a larger customer cannot dominate the result.

The nine frontier systems were evaluated zero-shot and eight-shot on blinded, pre-outcome inputs. Outcomes were joined only after predictions were frozen. The report identifies the evaluated model versions and presents both prompting conditions for the current cohort.

Metric definitions

For each eligible business, let k = ceil(0.10 * n). Rows sort by score descending. Exact score ties use a frozen hashed tie key ascending.

  • Capture@10 = positive outcomes in the top k / all positive outcomes.
  • Lift@10 = precision in the top k / business base rate.
  • ROC-AUC uses average ranks for tied scores.
  • PR-AUC is non-interpolated average precision with tied thresholds handled as score blocks.

The published value for each metric is the arithmetic mean of its 11 business-level values. The code uses the actual ceil-defined top decile; therefore Capture@10 is approximately, but not algebraically forced to equal, 0.10 * Lift@10 when a business row count is not divisible by ten.

Run the demonstration

Python 3.10 or newer is required.

python3 --version  # must report 3.10 or newer
python3 -m venv .venv
source .venv/bin/activate
python -m pip install -e '.[dev]'

python scripts/evaluate.py \
  --input data/dummy_predictions.csv \
  --score savant_score \
  --score frontier_score

python scripts/paired_bootstrap.py \
  --input data/dummy_predictions.csv \
  --reference savant_score \
  --candidate frontier_score

pytest
python scripts/verify_reproducibility.py
python scripts/verify_release.py

The default python3 on some macOS installations is Python 3.9. Install a newer Python first, or use uv python install 3.12 and uv venv --seed --python 3.12 .venv. The --seed flag installs pip, which the commands above use to install the project.

For the frozen release environment used to compare byte-for-byte output, use Python 3.12 and the hashed lock file:

python3.12 -m venv .venv
source .venv/bin/activate
python -m pip install --require-hashes -r requirements-reproducible.txt
python -m pip install --no-deps --no-build-isolation -e .

The bootstrap resamples 11 paired business rows with replacement in each of 5,000 draws and calculates the mean reference - candidate margin. Its default seed is 20260812, matching the paper. The 95% interval is the 2.5th to 97.5th percentile of those draws. Pairing is essential: both systems were scored on the same independent business units. The code explicitly freezes NumPy's PCG64 bit generator and the linear quantile method; requirements-reproducible.txt freezes the Python 3.12 dependency versions and hashes.

Repository map

prompts/                    Exact prompt contracts and fabricated examples
src/b2b_propensity_benchmark/
  metrics.py                Per-business metrics and macro averaging
  bootstrap.py              Paired equal-business percentile bootstrap
  prompts.py                Executable exact prompt construction
scripts/
  evaluate.py               Runnable metric CLI
  paired_bootstrap.py       Runnable 5,000-draw bootstrap CLI
  generate_dummy_data.py    Deterministic fabricated-data generator
  verify_reproducibility.py Frozen dummy hash, metric, and CI comparison
  verify_release.py         Public-release boundary audit
data/dummy_predictions.csv  Fabricated runnable example only
data/expected_dummy_results.json  Frozen demonstration outputs
tests/                      Mathematical and reproducibility checks

Interpretation

This code verifies how scores are graded; it does not reproduce the paper's proprietary observations or train any scoring model. Eleven businesses are the correct unit count for the current paired analysis, so confidence intervals can remain wide even when point-estimate margins are positive. A point-estimate lead must not be described as statistically significant when its interval crosses zero.

The two dummy score columns deliberately use equal outcome-signal weights and independent frozen noise. Neither is perfect, and the fabricated example is not constructed to prove that Savant wins.

Reproducibility scope

An outsider can reproduce the prompt construction, dummy-data bytes, all four metric implementations, equal-business aggregation, and seeded bootstrap behavior. The public release cannot independently reproduce the paper's numerical point estimates or confidence intervals because the real prospect rows, outcomes, and per-business result table are proprietary. This repository therefore verifies the grading method, not the private empirical observations.

The model IDs in prompts/models.md are the requested OpenRouter IDs used in the report. Provider aliases and backend revisions can change after a run; this repository does not claim that a future call to an alias recreates the original provider snapshot.

Governance

The default branch is protected: direct pushes are disabled, CI is required, and changes require code-owner review. External contributions should use a fork and pull request. See CONTRIBUTING.md and SECURITY.md.

License

The evaluation code and prompt documentation are released under the MIT License. No license or right is granted to NextLM customer data, Savant model artifacts, or other proprietary technology, none of which is included here.

About

Reproducible prompts, metrics, and paired business bootstrap for the NextLM B2B propensity benchmark

Topics

Resources

Contributing

Security policy

Stars

3 stars

Watchers

1 watching

Forks

Releases

Packages

Used by

Contributors

Languages