An open specification for scoring night sky quality for amateur astrophotography — plus a working implementation of it.
NSQI turns the mess of atmospheric, celestial, and site conditions into a single number from 0 (unusable) to 100 (world-class). The algorithm is specified as an open standard, so any language can implement it, and a conformance suite of 118 test vectors verifies that an implementation got it right.
v0.3.0-draft · 39 factors · 27 independent signals · 6 domains · Lean 4 proofs · CC BY-SA 4.0 (spec) / MIT (code)
| Range | Label | Conditions |
|---|---|---|
| 0–10 | Unusable | Active precipitation, dense fog, heavy overcast, or daytime |
| 10–25 | Poor | Bright moon with clouds, heavy urban light pollution, thick haze |
| 25–50 | Fair | Suburban skies, partial clouds, moderate humidity |
| 50–70 | Good | Rural skies, mostly clear, stable atmosphere, little or no moon |
| 70–85 | Excellent | Dark site, exceptional transparency and seeing, no moon |
| 85–100 | World-class | High-altitude, arid, pristine dark site under perfect conditions |
DarkCast, the reference web application, is the easiest way to see NSQI in action.
Dashboard — every site scored and ranked. Sites reading 0 are not broken: it is daytime there, and the daylight gate drives the score to zero regardless of how good the rest of the conditions are.
Site detail — the score, the seeing and coherence figures derived from it, and the observing windows over the next four nights. The orthographic globe is rendered server-side from the site's coordinates.
Domains — the score decomposed into its six domains, then into individual factors with their normalized value, raw measurement, and confidence. Every number here is traceable back to a specific data source.
Forecast — 96 hours ahead. The diurnal cycle is the dominant feature: the score collapses at dawn and recovers at dusk, and the shaded bands mark the stable observing windows worth planning around.
Pipeline health — freshness, cadence, and row counts for every data source, with a live log. CAMS shows as missing here because it is the one source that needs credentials.
Most "is it clear tonight?" tools answer a meteorological question: will there be clouds. That is necessary but nowhere near sufficient. An astrophotographer under a cloudless sky can still be defeated by a full moon, a jet-stream-driven seeing collapse, high-altitude cirrus that a cloud-cover percentage never reports, airglow, zodiacal light, or smoke aloft from a fire two states away.
NSQI's premise is that these are all measurable, all forecastable from public data, and all composable into one score — provided you are explicit about how they combine. The specification is the contribution; the code is a demonstration that the specification is implementable.
- Deterministic — identical inputs always produce identical output, within a stated floating-point tolerance.
- Source-agnostic — the spec defines what is computed, not where the data comes from. Acquisition, APIs, and caching are implementation concerns.
- Composable — factors aggregate into domains, domains into the final score. New factors can be added without restructuring the algorithm.
- Confidence-aware — missing or low-quality inputs degrade the score's confidence gracefully instead of silently poisoning it.
- Decorrelated — correlated measurements (cirrus fraction and total cloud fraction, r ≈ 0.55) are weight-penalized so they cannot double-count.
- Gated — some conditions are disqualifying, not merely bad. Rain does not get averaged away by an otherwise excellent night.
The score is a weighted geometric mean across six domains. A geometric mean is deliberate: it punishes a single catastrophic factor far harder than an arithmetic mean would, which is the correct behavior for a night that is perfect except for the clouds.
| Domain | Weight | Factors |
|---|---|---|
| Atmospheric transparency | 0.24 | 10 |
| Atmospheric seeing | 0.20 | 6 |
| Light pollution | 0.19 | 5 |
| Celestial geometry | 0.15 | 5 |
| Meteorological dynamics | 0.14 | 8 |
| Local environment | 0.08 | 5 |
| Total | 1.00 | 39 |
formal/NSQI.lean states 23 results (16 theorems, 7 lemmas) about the algorithm in Lean 4, using dependent types so that invalid states are unrepresentable. 21 are proven; 2 are stated but not yet discharged (perfect_score and gate_only_transparent_when_healthy still carry sorry).
Proven properties include:
- The score is in [0, 100] for any valid input (
score_bounded) - Every normalization curve produces a value in [0, 1] (
all_normalizations_bounded) - The weighted geometric mean is bounded by the min and max of its inputs (
wgm_le_max,wgm_ge_min) - Improving any single factor never decreases the score (
wgm_monotone_in_component) - Decorrelation conserves total weight (
effectiveWeights_conserve_sum) - Gate factors dominate when triggered (
gate_dominance) - Domain weights are positive and sum to 1 (
domainWeights_pos,domainWeights_sum)
Built with Lean v4.27.0 and Mathlib:
cd formal && lake buildThe normative documents, all readable directly on GitHub. See
spec/ for an index, the v0.2.0 archive, and what changed in v0.3.0.
| Document | Status | What it covers |
|---|---|---|
| Technical Specification | Normative | The algorithm end to end: domains, factor taxonomy, the five normalization curves, decorrelation, aggregation, gates, and conformance requirements. |
| Derived Factor Specifications | Normative | Formulas for the 14 factors computed from other measurements, plus two fallback estimators, with inputs, units, and physical reasoning. |
| Data Source Mapping | Informational | Concrete endpoints and extraction paths per factor, in priority order. |
| Glossary | Informational | Atmospheric, optical, and astronomical terms, plus acronyms. |
Machine-readable companions: factor registry (39 factors, units, weights, curves) · conformance suite (118 vectors) · Lean 4 proofs
spec/ Normative specification documents (see the table above)
registry/ Machine-readable factor catalog
conformance/ Conformance suite
formal/ Lean 4 formalization and proofs
ref-impl/ Reference implementation (Python) — the readable version of the spec
pipeline/ Production scoring pipeline (FastAPI, fetchers, loaders, scoring engine)
darkcast/ DarkCast — Phoenix/LiveView web application
Both v0.2.0 and v0.3.0 of the spec, registry, and conformance suite are kept in-tree so existing implementations have a stable target while v0.3.0 is still a draft.
These serve different purposes and it is worth not confusing them:
ref-impl/is the spec made executable — pure Python overnumpyandskyfield, no I/O, no database, optimized for being read alongside the specification. If you are implementing NSQI in another language, read this.pipeline/is the production system — fetches GFS, CAMS, GOES, IMS, HMS, SWPC, AirNow, OpenWeatherMap, and CelesTrak data into PostGIS and scores sites continuously. It is what actually runs DarkCast.
Everything here is a working draft and may change substantially.
| Component | Version | Status |
|---|---|---|
| Technical specification | 0.3.0-draft | Working draft |
| Factor registry | 0.3.0 | 39 factors |
| Derived factor specs | 0.3.0 | Formulas for all derived factors |
| Conformance suite | 0.3.0 | 118 vectors |
| Lean 4 formalization | 0.3.0 | 21 of 23 results proven |
| Data source mapping | 0.3.0 | All 39 factors mapped |
| Reference implementation | 0.3.0 | Python; validated against the 0.3.0 registry and conformance suite |
- Read
spec/nsqi-spec-v0.3.0.mdfor the algorithm, thenspec/nsqi-derived-factors-v0.3.0.mdfor the derived-factor formulas. - Load
registry/nsqi-registry-v0.3.0.jsonfor canonical factor names, units, weights, and normalization curves. Do not hardcode these — the registry is the source of truth. - Validate against
conformance/nsqi-conformance-v0.3.0.json. Vectors cover normalization curves (39), full pipeline runs (53), decorrelation (7), weighted geometric mean (7), gates (6), and algebraic properties (6). Tolerance is 0.01 absolute on the 0–100 scale. - Consult
ref-impl/when the prose is ambiguous, and please open an issue — that ambiguity is a spec bug.
You need mise (for the toolchain pinned in
.tool-versions — Erlang 28.3.2, Elixir 1.19.4, Node 25.6.0, Python 3.12.8),
uv, and a container runtime for the database.
make setup # system libs, toolchain, dependencies, database, schema
make seed # register the built-in sites
make api # scoring API -> http://localhost:8000
make web # DarkCast -> http://localhost:5001make seed downloads a Copernicus DEM tile per site to compute horizon and
terrain factors, so the first run takes a while and needs a few GB of disk.
Scores appear once a GFS cycle has been ingested.
make test # pipeline, ref-impl, and darkcast suites
make doctor # check that the required tooling is presentmake setup installs the system libraries (GDAL and the PostGIS client tools,
via scripts/install-system-deps.sh, which handles brew/apt/dnf/pacman/apk),
the pinned toolchain from .tool-versions, the Python and Elixir dependencies,
and then starts PostGIS and applies pipeline/schema/*.sql in order.
Data source credentials are optional — copy .env.example to .env and fill
in what you want. Every source that needs a key degrades to "missing" on the
/admin/pipeline page rather than breaking anything. GFS, HRRR, GOES, IMS,
HMS, SWPC, and CelesTrak need no credentials at all; only CAMS strictly
requires one (a free Copernicus ADS key).
make doctor reports which tools are actually present if something looks wrong.
Deployment configuration is intentionally not published in this repository.
- Specification, registry, conformance suite, derived factors — CC BY-SA 4.0. Share and adapt, with attribution, under the same license. (full text)
- All code (
ref-impl/,pipeline/,darkcast/,formal/) — MIT. (full text)
The split matters if you are reusing this: the prose and data files are share-alike, so a derived specification stays open, while the code carries no such obligation.
A personal project in active development, and the spec is a draft that may change substantially. Feedback is genuinely welcome — especially conformance failures, ambiguities in the specification text, and disagreements about factor weights or correlation coefficients. Open an issue.
Jeff Smith




