Skip to content

Repository files navigation

ARC WhiteBox Estimation Challenge 2026 logo

ARC WhiteBox Estimation Challenge 2026 - Starter Kit

Challenge whestbench Starter Kit MLP Explorer flopscope Hugging Face License: MIT
CI

whest-starterkit walkthrough: clone, sync, estimate, validate, score

🎬 60-second overview

You are given a randomly-initialized ReLU MLP and a FLOP budget. Predict the per-neuron mean activation under N(0, 1) input, without spending the whole budget on forward passes. Your score is the error on the final layer only (final_layer_mse, against N=1e9 Monte-Carlo ground truth), multiplied by max(0.1, C_m / B_m) — the fraction of the FLOP budget you spent, floored at 10%. Once you are under 10% of budget the multiplier stays at 0.1, so spending less does not change your score and only a lower final_layer_mse does; every bundled example is already at that floor. Lower is better.

Every MLP in the suite has width 1024 and depth 16, so predict() returns a (16, 1024) array, and each MLP carries its own budget of 2**41 FLOPs (roughly 6.5e4 forward passes).

A small ReLU MLP (width 4, depth 5) shown as a layer-by-layer heatmap of per-neuron mean activations after Monte-Carlo ground-truth estimation; rows are layers, columns are neurons, color intensity is mean activation
Per-neuron mean activations of a small MLP (width 4, depth 5) after Monte-Carlo ground truth, exactly what your estimator predicts. Generate your own at the hosted WhestBench Explorer.

The kit is a five-stage ladder: each stage adds one more part of the harness. Stage 1 needs only local math and no CLI knowledge; Stage 5 produces a packaged submission.

🚀 Your first 5 minutes (Stage 1: just python)

git clone https://github.com/AIcrowd/whest-starterkit.git
cd whest-starterkit
uv sync && uv run python estimator.py

That run printed a Monte-Carlo convergence table: n_samples, the FLOPs the sampler spent, the FLOPs your predict() spent, then all_layers_mse and final_layer_mse. The leaderboard ranks the latter. Both compare your prediction against a fresh Monte-Carlo estimate at that sample count, not against ground truth, so read down the column: n=10 is mostly sampling noise, and only the bottom row is a reliable measure of your estimator. (The grader compares against baked N=1e9 ground truth instead.) To experiment, edit predict() in estimator.py and re-run.

Compare against a bundled baseline:

uv run python estimator.py --baseline mean_propagation

🧪 Try the examples (still Stage 1)

uv run python examples/02_mean_propagation.py
uv run python examples/03_covariance_propagation.py
uv run python examples/04_shipped_weights.py

For the curriculum table, see examples/README.md.

🪜 Climb the ladder (Stages 2-5)

The Tutorial has a walkthrough for each stage.

Stage 1 — Iterate locally · the math: estimator vs Monte Carlo.

uv run python estimator.py

Stage 2 — Validate the contract · contract correctness (shapes, types).

uv run whest validate --estimator estimator.py

Stage 3 — Run on the public set · real scoring against the public Mini split (100 MLPs), in-process, debuggable with pdb.

uv run whest run --estimator estimator.py \
    --dataset hf://aicrowd/arc-whestbench-public-2026@v2-phase2 \
    --split mini \
    --runner local

Stage 4 — Subprocess runner · isolation; closer to the grader environment.

uv run whest run --estimator estimator.py \
    --dataset hf://aicrowd/arc-whestbench-public-2026@v2-phase2 \
    --split mini \
    --runner subprocess

Stage 5 — Package and Submit · build the submission artifact, then submit it (run whest login once first; see Submit to AIcrowd below).

uv run whest package --estimator estimator.py   # build & inspect the tarball
uv run whest submit  --estimator estimator.py   # ship it (also packages, in one step)

These package only estimator.py, the common case. To embed weights or split across modules, point --estimator at a folder instead: see Stage 5 → Embedding weights or multiple modules.

🏁 Submit to AIcrowd

Once you reach Stage 5, submit from the CLI. To authenticate, log in once with your AIcrowd API key, set AICROWD_API_KEY, or pass whest login --api-key <key>:

uv run whest login

Then package + submit in one step (add --watch to follow the submission until it is scored):

uv run whest submit --estimator estimator.py

Your score and per-MLP detail appear on the challenge leaderboard. Full walkthrough: Stage 5 → Submit to AIcrowd.

🚑 When something breaks

uv run whest doctor

whest doctor prints a 6-row health check. For what each row means and how to fix warnings, see docs/reference/whest-doctor.md.

For other symptoms, see Troubleshooting.

Phase 2 fair accounting: the official Challenge Rules govern eligibility. Packing independent values into one machine element to reduce billed work is not permitted, even when all operations are metered. See Fair accounting and packing.

Send rules questions, a FLOP price that looks wrong, or a residual-cap exception request to arc-whestbench@aicrowd.com.

📚 Documentation

Past Stage 1, the documentation has six sections. Pick whichever matches your task. For the full map and guided reading paths, see Documentation.

🪜 Tutorial — Climb the 5-stage ladder above
📖 Concepts — Why this challenge exists, what's measured, how ground truth works
  • Problem Setup — MLP architecture, He init, the research question, further reading.
  • Scoring Model — Pipeline diagram, adjusted_final_layer_score / all_layers_mse formulas, calibration table.
  • Ground Truth — How the evaluator computes reference values via Monte Carlo.
  • Allowed Code — What a submission may use, the prohibition list, the data-file carve-out, and how the rule is enforced.
🔧 How-to — Recipes: write, debug, optimize, submit

Writing and iterating

Optimizing

Debugging and shipping

📚 Reference — Exact contracts, schemas, lookup material

The round

  • Competition Rounds — Every round side by side: shape, budget, wall caps, and how residual time was charged. If a number you saw elsewhere doesn't match your local run, or if you need to reproduce a score from an earlier round, read this first.

Estimator API

  • Estimator Contract — predict/setup/teardown signatures, SetupContext, failure-semantics table, lifecycle diagram.
  • Code Patterns — flopscope patterns, ReLU expectation derivation, when the Gaussian assumption breaks.
  • local_engine API — Stage 1's MLP factory and Monte-Carlo helpers.

FLOP and scoring details

CLI

  • CLI Reference — Pointer at the upstream whest CLI.
  • whest doctor — The 6 install/env checks and how to fix WARN/FAIL rows.
🚑 Troubleshooting — When something breaks
  • Common Participant Errors — Symptom → cause → fix-now → verify.
  • FAQ — Quick answers; includes "local score great, submission 10x worse".
🔬 Advanced — Deeper tooling

📁 Repo layout

├── estimator.py     ← The participant's entry point; every stage operates on this file.
├── local_engine.py  ← Single-file re-implementation of the harness; safe to read end-to-end.
├── examples/        ← Numbered reference estimators (01–04) with a curriculum table.
├── docs/            ← Full documentation; start at docs/README.md.
├── tests/           ← Drift gates: README commands, local_engine parity, flopscope billing facts.
├── .whestignore     ← Controls what `whest package --estimator .` ships. Read it before your first submission.
└── CHANGELOG.md     ← Rule, parameter and FLOP-figure changes are announced here.

⚖️ License & contributing

Released under the MIT License. For the release process, see docs/RELEASING.md.

About

Starter kit for the ARC Whitebox Estimation Challenge — clone, run python estimator.py, climb the ladder.

Resources

Stars

9 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages