You are given a randomly-initialized ReLU MLP and a FLOP budget. Predict the per-neuron mean activation under N(0, 1) input, without spending the whole budget on forward passes. Your score is the error on the final layer only (final_layer_mse, against N=1e9 Monte-Carlo ground truth), multiplied by max(0.1, C_m / B_m) — the fraction of the FLOP budget you spent, floored at 10%. Once you are under 10% of budget the multiplier stays at 0.1, so spending less does not change your score and only a lower final_layer_mse does; every bundled example is already at that floor. Lower is better.
Every MLP in the suite has width 1024 and depth 16, so predict() returns a (16, 1024) array, and each MLP carries its own budget of 2**41 FLOPs (roughly 6.5e4 forward passes).
Per-neuron mean activations of a small MLP (width 4, depth 5) after Monte-Carlo ground truth, exactly what your estimator predicts. Generate your own at the hosted WhestBench Explorer.
The kit is a five-stage ladder: each stage adds one more part of the harness. Stage 1 needs only local math and no CLI knowledge; Stage 5 produces a packaged submission.
git clone https://github.com/AIcrowd/whest-starterkit.git
cd whest-starterkituv sync && uv run python estimator.py
That run printed a Monte-Carlo convergence table: n_samples, the FLOPs the sampler spent, the FLOPs your predict() spent, then all_layers_mse and final_layer_mse. The leaderboard ranks the latter. Both compare your prediction against a fresh Monte-Carlo estimate at that sample count, not against ground truth, so read down the column: n=10 is mostly sampling noise, and only the bottom row is a reliable measure of your estimator. (The grader compares against baked N=1e9 ground truth instead.) To experiment, edit predict() in estimator.py and re-run.
Compare against a bundled baseline:
uv run python estimator.py --baseline mean_propagation
uv run python examples/02_mean_propagation.py
uv run python examples/03_covariance_propagation.py
uv run python examples/04_shipped_weights.py
For the curriculum table, see examples/README.md.
The Tutorial has a walkthrough for each stage.
Stage 1 — Iterate locally · the math: estimator vs Monte Carlo.
uv run python estimator.pyStage 2 — Validate the contract · contract correctness (shapes, types).
uv run whest validate --estimator estimator.pyStage 3 — Run on the public set · real scoring against the public Mini split (100 MLPs), in-process, debuggable with pdb.
uv run whest run --estimator estimator.py \
--dataset hf://aicrowd/arc-whestbench-public-2026@v2-phase2 \
--split mini \
--runner localStage 4 — Subprocess runner · isolation; closer to the grader environment.
uv run whest run --estimator estimator.py \
--dataset hf://aicrowd/arc-whestbench-public-2026@v2-phase2 \
--split mini \
--runner subprocessStage 5 — Package and Submit · build the submission artifact, then submit it (run whest login once first; see Submit to AIcrowd below).
uv run whest package --estimator estimator.py # build & inspect the tarball
uv run whest submit --estimator estimator.py # ship it (also packages, in one step)These package only
estimator.py, the common case. To embed weights or split across modules, point--estimatorat a folder instead: see Stage 5 → Embedding weights or multiple modules.
Once you reach Stage 5, submit from the CLI. To authenticate, log in once with your
AIcrowd API key, set AICROWD_API_KEY,
or pass whest login --api-key <key>:
uv run whest loginThen package + submit in one step (add --watch to follow the submission until it is scored):
uv run whest submit --estimator estimator.pyYour score and per-MLP detail appear on the challenge leaderboard. Full walkthrough: Stage 5 → Submit to AIcrowd.
uv run whest doctor
whest doctor prints a 6-row health check. For what each row means and how to fix
warnings, see docs/reference/whest-doctor.md.
For other symptoms, see Troubleshooting.
Phase 2 fair accounting: the official Challenge Rules govern eligibility. Packing independent values into one machine element to reduce billed work is not permitted, even when all operations are metered. See Fair accounting and packing.
Send rules questions, a FLOP price that looks wrong, or a residual-cap exception request to arc-whestbench@aicrowd.com.
Past Stage 1, the documentation has six sections. Pick whichever matches your task. For the full map and guided reading paths, see Documentation.
🪜 Tutorial — Climb the 5-stage ladder above
- Stage 1: Iterate locally — The math;
flopscope+local_engine.py, nowhestCLI. - Stage 2: Validate the contract — Class resolves,
setup()runs, shape, finite values. - Stage 3: Run locally — Real scoring against the grader's MLP suite, in-process.
- Stage 4: Subprocess runner — Process isolation over the grader's transport. Catches dirty imports, stdout writes, and (on Linux) runaway memory. One worker serves the whole suite, so state does not reset between MLPs.
- Stage 5: Package and Submit — Build the AIcrowd submission tarball and ship it.
📖 Concepts — Why this challenge exists, what's measured, how ground truth works
- Problem Setup — MLP architecture, He init, the research question, further reading.
- Scoring Model — Pipeline diagram,
adjusted_final_layer_score/all_layers_mseformulas, calibration table. - Ground Truth — How the evaluator computes reference values via Monte Carlo.
- Allowed Code — What a submission may use, the prohibition list, the data-file carve-out, and how the rule is enforced.
🔧 How-to — Recipes: write, debug, optimize, submit
Writing and iterating
- Write an Estimator — Minimal structure, contract checklist, common first failure.
- Inspect MLP Structure — Traversing the
MLPobject. - Validate, Run, Package — The standard local loop, plus a useful-flags table.
- Use Evaluation Datasets — Pre-create datasets for fast, reproducible iteration.
Optimizing
- Algorithm Ideas — Monte Carlo, mean propagation, covariance, hybrid, plus open directions.
- Manage FLOP Budget — Where your FLOPs go; line-by-line walkthrough of
examples/02. - Performance Tips — Matmul placement, free ops, env-var knobs.
Debugging and shipping
- Debugging Checklist — Tiered procedure for when a result looks wrong.
- Pre-Submission Checklist — One-screen gate before you click submit.
- Ship Weights and Multi-File Submissions — Precompute offline, load in
setup()viasubmission_dir, folder-mode packaging, the 50 MiB / 50-file caps.
📚 Reference — Exact contracts, schemas, lookup material
The round
- Competition Rounds — Every round side by side: shape, budget, wall caps, and how residual time was charged. If a number you saw elsewhere doesn't match your local run, or if you need to reproduce a score from an earlier round, read this first.
Estimator API
- Estimator Contract —
predict/setup/teardownsignatures,SetupContext, failure-semantics table, lifecycle diagram. - Code Patterns —
flopscopepatterns, ReLU expectation derivation, when the Gaussian assumption breaks. local_engineAPI — Stage 1's MLP factory and Monte-Carlo helpers.
FLOP and scoring details
- Flopscope Primer —
BudgetContextownership, attribute reference, op cost table. - Score Report Fields — Every field you'll see in
whest runoutput.
CLI
- CLI Reference — Pointer at the upstream
whestCLI. whest doctor— The 6 install/env checks and how to fix WARN/FAIL rows.
🚑 Troubleshooting — When something breaks
- Common Participant Errors — Symptom → cause → fix-now → verify.
- FAQ — Quick answers; includes "local score great, submission 10x worse".
🔬 Advanced — Deeper tooling
- Profile Simulation — Benchmark the flopscope backend's correctness and wall-clock scaling across network sizes.
- WhestBench Explorer — Hosted interactive visualizer at aicrowd.github.io/whestbench-explorer for inspecting MLPs and ground truth.
├── estimator.py ← The participant's entry point; every stage operates on this file.
├── local_engine.py ← Single-file re-implementation of the harness; safe to read end-to-end.
├── examples/ ← Numbered reference estimators (01–04) with a curriculum table.
├── docs/ ← Full documentation; start at docs/README.md.
├── tests/ ← Drift gates: README commands, local_engine parity, flopscope billing facts.
├── .whestignore ← Controls what `whest package --estimator .` ships. Read it before your first submission.
└── CHANGELOG.md ← Rule, parameter and FLOP-figure changes are announced here.
Released under the MIT License. For the release process, see docs/RELEASING.md.

