From ec57d34cd5fe9b114ebb789fd42c4444c290096b Mon Sep 17 00:00:00 2001 From: Cuuper22 <97895153+Cuuper22@users.noreply.github.com> Date: Thu, 13 Aug 2026 23:27:32 -0700 Subject: [PATCH] Rewrite the public site prose in a plain teaching voice Finishes the documentation rewrite by covering the two files it missed. Earlier work rewrote the project markdown and the JavaScript comments, but the prose a visitor actually reads on docs/index.html and docs/observatory.html was still in the old style. Visible text only. Technical terms now carry a defining clause where they first appear: held-out split, regret, resolver, scope, root input, dependency cone, scenario fixture, HBM, MFU, PUE, collective, preregistered, prior, residual, durable frontier, PDU, and WAN. Jargon gave way to plain equivalents where the plain phrasing says the same thing, so "epistemic ranking regions" became "ranking under uncertainty" and "inadmissible" became "cannot support a claim". Nothing structural changed, and no number changed. The docs-stats gate parses the stat-grid digits against values computed live from the registry, so an altered digit would both fail CI and publish a false fact. Strings that JavaScript matches on were identified first and frozen, which is why the "Collective" legend item keeps its exact wording, with the term defined in the surrounding prose instead. Verified: docs-stats prints OK; an ordered comparison of the tag skeleton shows 757 of 757 tags on index.html and 2058 of 2058 on observatory.html with zero differences, proving no tag, attribute, class, id, data attribute, or href changed; every numeric literal matches (182 and 576); and the full suite reports 1029 passed with only the 10 pre-existing Windows temp-dir errors that main already produces. Co-Authored-By: Claude Opus 5 (1M context) --- docs/index.html | 83 ++++++++++++++------------- docs/observatory.html | 130 +++++++++++++++++++++--------------------- 2 files changed, 107 insertions(+), 106 deletions(-) diff --git a/docs/index.html b/docs/index.html index cacfeea..b8b7b91 100644 --- a/docs/index.html +++ b/docs/index.html @@ -85,7 +85,7 @@

gpu_stack/readme/frontdoor.txt

gpu_stack

- A virtual AI datacenter you can interrogate. It predicts what a training run does to time, power, and money, says how sure it is, and can show you what every one of its numbers is made of. + A virtual AI datacenter you can question. It estimates what a training run costs in time, power, and money, states how sure it is, and shows you what every number is built from.

Open observatory @@ -96,9 +96,9 @@

gpu_stack

@@ -111,7 +111,7 @@

Read it like a receipt, not a magic answer.

@@ -130,17 +130,17 @@

Read it like a receipt, not a magic answer.

equations run cost, tokens, power - The graph walks the ancestry. + The graph walks upstream through what defines it.
root debt 619 named inputs - Unknowns stay inspectable. + Every unknown keeps its name.
scenario explicit assignments - Fixtures are anchors, not market claims. + A fixture is a test anchor, not a price.
@@ -164,7 +164,7 @@

Read it like a receipt, not a magic answer.

token_journey.exe

Follow one token from math to money.

-

Every station is a layer of the model. The lamp lights as the token passes through, and the trip always ends at an invoice.

+

Each station is one layer of the model. The lamp lights as the token passes through, and the trip always ends at an invoice.

@@ -208,14 +208,14 @@

Follow one token from math to money.

Symbolic causal backbone

-

The old engine becomes useful when measurements can prove it wrong.

+

A model earns trust when a measurement can prove it wrong.

- The registry keeps equations, units, references, constraints, scenario assumptions, and unresolved boundaries attached. The research layer now adds observations, held-out splits, temporal events, interventions, uncertainty, residuals, and decision regret so graph completeness is no longer mistaken for scientific progress. + The registry keeps every equation next to its units, sources, constraints, scenario assumptions, and the boundaries it has not resolved. The research layer adds the things that can contradict it: real observations, timed events, interventions, uncertainty, residuals, held-out splits, and regret. A held-out split is data the model never saw while it was being fitted, so a prediction on that data is a real test. Regret is what the choice cost you compared with the choice you would have made knowing the outcome. A complete graph is not the same thing as a tested one.

  • Only universal physics constants belong in Constant.
  • Everything else stays a Variable: clocks, voltages, tariffs, GPU counts, batch sizes, and facility assumptions.
  • -
  • A root input is visible modeling debt. That is much better than hidden modeling debt with a haircut.
  • +
  • A root input is modeling debt you can see. That beats hidden debt papered over with a fudge factor.
@@ -245,7 +245,7 @@

The old engine becomes useful when measurements can prove universal constants, named unknowns

-

One output, many upstream obligations. Click a layer to see what it owes.

+

One answer at the top rests on every layer beneath it. Click a layer to see what it owes.

@@ -282,8 +282,8 @@

Choose a question and follow what it depends on.

Cost per token is not a lone price. It depends on run cost, token count, facility power, throughput, utilization, hardware choices, and root assumptions that still need better evidence.

    -
  • Use this as a mental model for the resolver, not as a live numerical solver.
  • -
  • The moving gold segment marks the direction of dependency pressure.
  • +
  • This is a picture of how the resolver works, the code that walks the graph and computes a value. It is not a live solver.
  • +
  • The moving gold segment shows which way the dependencies run.
@@ -297,7 +297,7 @@

Choose a question and follow what it depends on.

-

The graph is useful because each hop keeps its label. If a hop cannot be resolved from equations or scenario assignments, it comes back as a named missing boundary.

+

Every hop keeps its label, and that is what makes the graph useful. If a hop cannot be worked out from equations or scenario assignments, it comes back as a named missing boundary instead of a guess.

synthetic fixture resolves 4 of 4 targets @@ -314,15 +314,16 @@

Choose a question and follow what it depends on.

cone_browser.exe

Inspect any variable's upstream cone.

- Choose a target and click any node to expand its direct dependencies. - Each hop shows the unit, scope, and whether it is an equation, a root - input, or a physics constant. Root inputs carry a gold badge because - they are the visible modeling debt. + Pick a target, then click any node to open its direct dependencies. + Each hop shows its unit, its scope, meaning the part of the system it + belongs to, and whether it is an equation, a root input, or a physics + constant. Root inputs carry a gold badge because they are the debt you + can still see.

    -
  • The tree loads from a pre-generated JSON snapshot of the registry.
  • -
  • Click a node button to toggle its direct dependencies open or closed.
  • -
  • Constants are marked to distinguish universal physics values.
  • +
  • The tree loads from a JSON snapshot of the registry, generated ahead of time.
  • +
  • Click a node to open or close its direct dependencies.
  • +
  • Constants carry their own mark, so universal physics values are easy to tell apart.
@@ -346,7 +347,7 @@

Inspect any variable's upstream cone.

root_debt.dat

Root inputs are the visible unpaid invoices.

- root-debt ranks unresolved root inputs by downstream blast radius. The point is not to pretend the largest family is bad. The point is to know which unknowns are currently expensive. + root-debt ranks unresolved root inputs by how much depends on them. A large family is not automatically a bad one. The ranking exists so you know which unknowns are expensive right now.

  • Total roots in the observed summary: 619.
  • @@ -362,7 +363,7 @@

    Root inputs are the visible unpaid invoices.

    physical.mosfet
    weight 1866, roots 18
    physical.process
    weight 1293, roots 8
-

These bars normalize the five README weights against the top family. No new metric is being invented here.

+

The bars scale the five README weights against the heaviest family. Nothing new is being measured here.

@@ -371,7 +372,7 @@

Root inputs are the visible unpaid invoices.

CLI.exe

Use the command line as a microscope.

- The package is still closer to a research instrument than a polished app. That is useful right now. Ask it what exists, what is unresolved, and where a claim bottoms out. + This is a research instrument, not a polished app, and right now that is the useful thing to be. Ask it what exists, what is still unresolved, and where a claim bottoms out.

python -m gpu_stack.cli stats
@@ -387,52 +388,52 @@ 

Use the command line as a microscope.

Root input

-

A variable with no defining value relation yet. It might be a real scenario boundary, or it might be physics that still needs decomposition.

+

A variable the model cannot yet compute from anything else. It may be a genuine edge of the scenario, or physics that nobody has broken down yet.

Dependency cone

-

The upstream set of variables, equations, assumptions, and constants needed to explain one target.

+

Everything upstream of one number: the variables, equations, assumptions, and constants you need in order to explain it.

Scenario fixture

-

A named set of explicit assignments used to resolve targets reproducibly. Synthetic fixtures are test anchors, not market claims.

+

A named set of fixed input values, so a target resolves the same way every time. Synthetic fixtures are test anchors, not claims about real prices.

MFU

-

Model FLOPs Utilization: how much of the theoretical model compute is actually useful during training.

+

Model FLOPs Utilization: the share of the machine's peak math rate that training actually uses. When it is low, the run is usually waiting on memory instead of doing math, which is exactly what a roofline chart is drawn to show.

HBM

-

High Bandwidth Memory: the fast memory sitting close to the accelerator package, often a ceiling for throughput.

+

High Bandwidth Memory: the fast memory stacked right next to the accelerator. Moving data in and out of it is often the real speed limit, not the math.

PUE

-

Power Usage Effectiveness: total facility power divided by IT equipment power. Cooling and overhead show up here.

+

Power Usage Effectiveness: total facility power divided by IT equipment power. Cooling and everything else the building draws shows up here.

Good for now

-

Tracing a result into mechanics and evidence. Replaying explicit compute, communication, checkpoint, outage, and recovery events. Keeping observations, assumptions, priors, modeled values, and unmeasured claims visibly separate.

+

Tracing a result back to mechanics and evidence. Replaying compute, communication, checkpoint, outage, and recovery events one at a time. Keeping observations, assumptions, priors, modeled values, and unmeasured claims clearly apart.

Not finished yet

-

E001 now has explicit recovery mechanics and three measured learning stages. LC3 held canonical work equal: adaptive continuation preserved learning and saved attempted work and opportunity ticks, but failed its frozen device-energy bound. This is not evidence that a frontier-scale multi-site run converges, and local GPU energy is not facility energy.

+

E001 now has explicit recovery mechanics and three measured learning stages. LC3 held useful work equal across policies: adaptive continuation kept the learning and saved both attempted work and opportunity ticks, the scheduled time it no longer needed, but it went past its frozen device-energy bound. None of that shows a frontier-scale multi-site run converges, and energy measured at one GPU is not energy measured at a facility.

next_work.exe

-

The frontier program is six falsifiable questions.

+

The frontier program is six questions a result could prove wrong.

-
1. Beyond One Datacenter · energy gate failedAt equal canonical work, adaptive continuation was learning-noninferior and saved work and scheduled time, but its measured RTX energy interval exceeded the frozen bound. Scale remains blocked.
-
2. Shape the Power Waveform · nextFactor checkpoint cadence from survivor continuation, attribute phase-level power, and test whether dependency-safe scheduling removes the energy penalty without losing the measured learning, work, or time gains.
-
3. Semantic Fault ToleranceAllocate canaries, replay, and redundancy by counterfactual learning harm instead of fault label.
-
4. Fluid Inference TopologyMeasure interaction gains and regime crossings when serving topology changes per request.
-
5. Architecture as a Datacenter VariableCo-design model modules and heterogeneous hardware under one facility power and time envelope.
-
6. Firm Grid-responsive InferenceMeasure meter-verified demand response with quality, tail latency, rebound, and hidden work inside the same boundary.
+
1. Beyond One Datacenter · energy gate failedAt equal useful work, adaptive continuation learned just as well and saved work and scheduled time, but its measured RTX energy interval went past the frozen bound. Scaling stays blocked.
+
2. Shape the Power Waveform · nextSeparate checkpoint cadence from survivor continuation, attribute power phase by phase, and test whether scheduling that respects dependencies removes the energy penalty without giving back the measured learning, work, or time gains.
+
3. Semantic Fault ToleranceSpend canaries, replay, and redundancy on the faults that would really have hurt learning, not on whatever the fault is labeled.
+
4. Fluid Inference TopologyMeasure what is gained, and where the behavior flips, when the serving topology changes for every request.
+
5. Architecture as a Datacenter VariableDesign the model's parts and the mixed hardware together, under a single facility power and time budget.
+
6. Firm Grid-responsive InferenceMeasure meter-verified demand response with output quality, tail latency, rebound, and hidden work all counted inside the same boundary.
diff --git a/docs/observatory.html b/docs/observatory.html index 67339b0..46c9860 100644 --- a/docs/observatory.html +++ b/docs/observatory.html @@ -87,7 +87,7 @@

Can one observable controller safely spend semantic small-model learning + virtual datacenter -

Loading the held-out semantic-consistency result. No conclusion is shown until its evidence boundary is known.

+

Loading the held-out semantic-consistency result, the test run on data the controller never saw. No conclusion appears here until the page knows what evidence stands behind it.

Reading data/e001-semantic-consistency-v1.json…

@@ -108,16 +108,16 @@

One controller, six unseen stress familie

Result loading

-

The observatory will state whether adaptive switching preserved held-out learning and beat the frozen fixed comparator.

+

The page will say whether adaptive switching kept the held-out learning and beat the fixed comparator that was frozen before the run.

- Infrastructure remains modeled unless the artifact says otherwise. + The infrastructure is simulated, not measured, unless the artifact says otherwise.
-

Every policy receives the same training examples and must reach the same useful-work target. The result will show whether changing consistency mode helped, what it saved, and where the controller refused to guess.

-

The fixed comparator is selected on calibration only. Six untouched stress families pair adaptive and fixed policies against a hindsight whole-policy envelope; sampling intervals and infrastructure uncertainty remain separate.

-

Exact optimizer commits, mode transitions, membership, lineage, work, WAN accounting, assumptions, uncertainty, and missing evidence appear below.

+

Every policy gets the same training examples and has to reach the same amount of useful work. The result shows whether switching consistency mode helped, what it saved, and where the controller refused to guess.

+

The fixed comparator is chosen from calibration data only. Six untouched stress families pair the adaptive and fixed policies against the best whole-policy choice in hindsight. Sampling intervals and infrastructure uncertainty stay separate rather than being merged into one number.

+

Exact optimizer commits, mode changes, membership, lineage, work, WAN accounting for traffic on the long link between sites, assumptions, uncertainty, and missing evidence all appear below.

@@ -127,7 +127,7 @@

Result loading

Paired held-out effects

-

Learning, time, communication, work, and energy stay adjacent

+

Learning, time, communication, work, and energy stay side by side

six evaluation families · 90% paired intervals
@@ -137,14 +137,14 @@

Learning, time, communication, work,
-

Epistemic ranking regions

-

Where the result survives uncertainty, and where it abstains

+

Ranking under uncertainty

+

Where the result survives uncertainty, and where it refuses to answer

direct labels · no hover required
-

Each held-out stress family is divided into labeled infrastructure-uncertainty regions where adaptive wins, a fixed policy wins, ranking reverses, or the controller abstains.

+

Each held-out stress family is split into labeled regions of infrastructure uncertainty, showing where adaptive wins, where a fixed policy wins, where the ranking flips, and where the controller declines to rank them at all.

Accessible ranking-region table @@ -169,7 +169,7 @@

What the controller saw, chose, and
-

Aligned adaptive and comparator mode intervals with membership, WAN, abstention, merge, and rejoin events.

+

The adaptive and comparator mode intervals on one shared axis, marked with membership, WAN, abstention, merge, and rejoin events.

Accessible controller-decision timeline @@ -183,7 +183,7 @@

What the controller saw, chose, and

Untouched family ledger

No family disappears inside the median

- adaptive versus calibration-selected fixed policy + adaptive versus the fixed policy chosen on calibration data
@@ -234,9 +234,9 @@

Measured learning, exact accounting
- Load on-demand exact epoch trace + Load the exact epoch trace on demand
-

Raw optimizer-commit trace has not been loaded. Opening this disclosure in Full trace fetches the separately bound artifact.

+

The raw optimizer-commit trace is not loaded yet. Opening this panel in Full trace fetches the separate artifact it is bound to.

@@ -248,11 +248,11 @@

Measured learning, exact accounting

Causal reading

  1. 1

    An assumed site interruption beginsAvailability is an input, not an observed outage rate.

  2. -
  3. 2

    Membership response is missingReactive outage membership is not implemented.

  4. -
  5. 3

    Sync cadence changesThe controller acts after a completed communication cycle.

  6. -
  7. 4

    Traffic and elapsed time changeMechanics are modeled by event ordering and resources.

  8. -
  9. 5

    Learning response is a priorOne-step delay observations seed an unfitted sensitivity.

  10. -
  11. 6

    Every gate stays explicitScalar falsifiers and structured evidence requirements can fail or remain unresolved independently.

  12. +
  13. 2

    Nothing changes who is in the groupDropping a site from the group during an outage is not implemented.

  14. +
  15. 3

    Sync cadence changesThe controller acts only after a finished collective, the round where every site exchanges and combines results.

  16. +
  17. 4

    Traffic and elapsed time changeThe mechanics come from the order of events and the resources they use.

  18. +
  19. 5

    Learning response is a priorA prior is a starting guess, not a fit. One-step delay observations seed it.

  20. +
  21. 6

    Every gate stays explicitNumeric falsifiers and evidence requirements can each fail, or stay unresolved, on their own.

@@ -282,7 +282,7 @@

Three-site scenario

Causal field

- Values follow the artifact; evidence class does not collapse into color. + Values come from the artifact. Color alone never tells you what class of evidence a number is.
@@ -366,8 +366,8 @@

Four policies, one failure clock

The same failure hits all four policies. The shared clock shows what stops, what is restored, what must be redone, and when useful work catches up.

-

All policy tracks share the artifact time domain. Recovery completion is mechanical: each policy reaches the same durable frontier after explicit preemption, restore, replay, and membership transitions.

-

Every mark is projected from a persisted recovery episode. Event IDs, nanosecond bounds, work dispositions, checkpoint lineage, and matched-frontier hashes remain available in the structured trace.

+

Every policy track shares one time axis taken from the artifact. Recovery is mechanical: each policy reaches the same durable frontier, the last point safely written to storage, after explicit preemption, restore, replay, and membership changes.

+

Every mark is drawn from a saved recovery episode. Event IDs, nanosecond bounds, what happened to each unit of work, checkpoint lineage, and matched-frontier hashes stay available in the structured trace.

@@ -396,8 +396,8 @@

Work conservation

-

Inter-site byte composition

- six disjoint traffic classes +

Where the inter-site bytes went

+ six traffic classes that never overlap
@@ -416,7 +416,7 @@

Mechanical completion

Recovery-v2 learning was unresolved

This mechanics artifact contains no held-out learning observation.

LC1 below adds measured local small-model learning. It does not turn the recovery-v2 mechanics run into real multi-site learning evidence.

-

Recovery-v2 has no recovery-quality observation IDs; LC1 is a separate attached calibration artifact.

+

Recovery-v2 records no recovery-quality observation IDs. LC1 is a separate calibration artifact attached to it.

@@ -426,7 +426,7 @@

Recovery-v2 learning was unresolved

Measured small-model learning

-

Did adaptive recovery preserve learning efficiency?

+

Did adaptive recovery keep learning just as efficient?

Reading learning artifact…
@@ -436,13 +436,13 @@

Did adaptive recovery preserve learning efficiency?LC1 · paired interrupted calibration

The learning result is loading

-

The observatory will state the measured answer once the persisted artifact is available.

+

The page will state the measured answer once the saved artifact loads.

Lower held-out loss means the model predicts better. The cards show how much training each policy attempted to get there, and the lines show what it learned over the same failure clock.

-

Policy medians summarize matched strata. The paired effect subtracts fixed-local progress per FLOP from adaptive progress per FLOP within each stratum, so zero means no retained-efficiency difference.

+

Policy medians summarize matched strata. Inside each stratum, the paired effect subtracts fixed-local progress per FLOP from adaptive progress per FLOP, so zero means the two are equally efficient.

Artifact, dataset, runtime, paired evaluation, and per-run records remain attached below.

@@ -482,9 +482,9 @@

Adaptive minus fixed progress per FLOP

Preregistered decision

-

Falsifier gate strip

+

The gates that could fail this claim

- every gate remains individually visible + every gate was written down before the run and stays visible
@@ -537,13 +537,13 @@

What survives when useful work is held equal?

Same 524,288 useful tokens · six held-out pairs

Equal-work result loading

-

The observatory will state the result after loading the persisted LC3 artifact.

+

The page will state the result once the saved LC3 artifact loads.

-

Useful work is the fair ruler here. Both policies finish with the same 524,288 tokens actually kept by the model. Adaptive avoided redo work and finished the schedule sooner without a meaningful learning loss, but its GPU consumed too much energy to pass.

-

LC3 removes LC1’s stopping-early denominator trap. Six paired evaluation strata share an equal canonical-work endpoint. Read learning as paired NLL noninferiority, physical work as attempted-FLOP savings, schedule time as opportunity ticks, and energy as a paired device-energy ratio.

+

Useful work is the fair ruler here. Both policies end with the same 524,288 tokens actually kept by the model. Adaptive avoided work it would have had to redo and finished the schedule sooner with no meaningful loss of learning, but its GPU drew too much energy to pass.

+

LC3 closes the trap in LC1, where stopping early shrank the denominator and flattered the result. Six paired evaluation strata all end at the same amount of canonical work. Read learning as paired NLL noninferiority, physical work as attempted-FLOP savings, schedule time as opportunity ticks, and energy as a paired device-energy ratio.

Exact pairs, run records, protocol predecessors, hashes, and the modeled-mechanics boundary remain attached below.

@@ -586,7 +586,7 @@

Likely causal clue, not proof

Frozen LC3 decision

Seven gates passed; energy failed

- the failed gate remains visually dominant + the failed gate stays the loudest thing here
@@ -654,7 +654,7 @@

Did checkpoint cadence cause LC3’s energy p

Result loading

-

The observatory will state the measurement boundary after loading the persisted artifact.

+

The page will state the measurement boundary once the saved artifact loads.

No policy claim is available yet.
@@ -662,7 +662,7 @@

Result loading

The experiment ran; the measuring stick failed. All 32 planned comparisons started from the exact same trained model. Asking the sensor every 20 milliseconds did not make it answer that fast: its reading effectively changed about every 495 milliseconds. That is too slow to decide whether sparse continuation saves energy.

-

The 2×2 cadence-by-continuation design remains intact, but its energy estimand is inadmissible. Both frozen update-count invalidators fired and the logger calibration hit its boundary. Arm medians and contrasts remain visible as diagnostics, never as mechanism evidence.

+

The 2×2 cadence-by-continuation design is intact, but the energy quantity it was meant to estimate cannot be used. Both frozen update-count invalidators fired and the logger calibration hit its limit. Arm medians and contrasts stay visible as diagnostics, never as evidence about mechanism.

The 32-run ledger, per-phase metrics, bound source identities, and lazy raw telemetry are attached below.

@@ -689,7 +689,7 @@

Two active invalidators stopped the cla

Observed arm medians

The intended 2×2 comparison completed

- energy values are descriptive and inadmissible + energy values describe the runs; they cannot support a claim
@@ -709,11 +709,11 @@

State, work, and final NLL matched exac

Raw diagnostic contrasts

-

Visible, but explicitly inadmissible

+

Shown here, but they cannot support a claim

measurement invalidators take precedence
-
DO NOT INTERPRET AS A POLICY PASS. These numbers describe the rejected telemetry stream; they cannot establish that sparse continuation fixed LC3’s energy failure.
+
DO NOT READ THIS AS A PASS. These numbers describe the telemetry stream that was rejected. They cannot show that sparse continuation fixed LC3’s energy failure.
@@ -757,7 +757,7 @@

Local board telemetry is not facility p
Selected-run phase metrics -

Phase energy is retained for diagnosis, but remains inadmissible for a causal policy comparison because the logger invalidators fired.

+

Phase energy is kept for diagnosis. It cannot support a causal comparison between policies, because the logger invalidators fired.

@@ -772,9 +772,9 @@

Local board telemetry is not facility p
- Load on-demand raw point trace + Load the raw point trace on demand
-

Raw telemetry has not been loaded. Opening this disclosure in Full trace fetches the separate point artifact.

+

Raw telemetry is not loaded yet. Opening this panel in Full trace fetches the separate point artifact.

PhaseDurationIdle-subtracted board energyEnergy / canonical tokenEffective update equivalents
@@ -802,15 +802,15 @@

What caused the local GPU-energy penalty?Result loading

-

The observatory will state the mechanism after loading the persisted artifact.

+

The page will state the mechanism once the saved artifact loads.

Mechanism estimate loading.
-

Checkpointing less often fixed the measured energy problem without giving back the recovery win. Both policies kept exactly 524,288 useful tokens. Sparse continuation attempted less work, finished 40 scheduling ticks sooner, and stayed inside the frozen GPU-energy limit. An estimated idle-baseline subtraction was less certain and could not rule out zero interaction.

-

PW2 repeats PW1’s frozen 2×2 design with a supported cumulative-energy counter. No measurement invalidator fired. Positive total and checkpoint-group interactions pass the frozen primary attribution gates while sparse continuation preserves the learning, work, schedule, and energy gates. The sensitivity-only idle-subtracted interaction crossed zero, so the result is not insensitive to baseline treatment.

+

Checkpointing less often fixed the measured energy problem without giving back the recovery win. Both policies kept exactly 524,288 useful tokens. Sparse continuation attempted less work, finished 40 scheduling ticks sooner, and stayed inside the frozen GPU-energy limit. One version of the calculation subtracts an estimated idle baseline; that version is less certain and cannot rule out that the effect is zero.

+

PW2 repeats PW1’s frozen 2×2 design, this time with a cumulative-energy counter the hardware actually supports. No measurement invalidator fired. Positive total and checkpoint-group interactions pass the frozen primary attribution gates, and sparse continuation holds the learning, work, schedule, and energy gates. The sensitivity-only, idle-subtracted interaction crossed zero, so the result does depend on how the idle baseline is handled.

The 32-run ledger, selected-run counter phases, bound identities, and lazy raw counter points are attached below.

@@ -832,7 +832,7 @@

The cumulative counter cleared the fro

Paired causal contrasts

-

Primary attribution passed; baseline sensitivity is unresolved

+

The main attribution passed; the baseline question is still open

six frozen evaluation blocks · 90% intervals
@@ -843,7 +843,7 @@

Primary attribution passed; baseline se

Counter support by phase

-

Checkpoint groups are supported; rare recovery phases are exploratory

+

Checkpoint groups have enough counter data; rare recovery phases do not

effective counter-update equivalents
@@ -908,9 +908,9 @@

Local RTX + TinyStories result only
- Load on-demand raw cumulative-counter trace + Load the raw cumulative-counter trace on demand
-

Raw counter telemetry has not been loaded. Opening this disclosure in Full trace fetches the separate point artifact.

+

Raw counter telemetry is not loaded yet. Opening this panel in Full trace fetches the separate point artifact.

@@ -927,7 +927,7 @@

Local RTX + TinyStories result only

E002-PW3 · physical rack mechanism

-

Can recovery slack become a rack-power control surface?

+

Can slack in recovery be used to steer rack power?

Reading rack experiment artifact… @@ -938,14 +938,14 @@

Can recovery slack become a rack-power control

Result loading

-

The observatory will explain what moved, what the rack meter saw, and whether learning stayed fixed.

- No rack result is inferred from the earlier laptop experiment. +

The page will explain what moved, what the rack meter saw, and whether learning stayed fixed.

+ Nothing about this rack is inferred from the earlier laptop experiment.
-

The same training and recovery work runs several ways. The shaped run may move only operations that have real timing slack, then the rack meter tells us whether separating those operations actually reduced the electrical shock.

-

Paired blocks hold useful work, state generations, failures, and learning commitments fixed while changing only the legal release policy for checkpoint and rejoin flows.

+

The same training and recovery work runs several ways. The shaped run is allowed to move only the operations that have real timing slack. The rack meter then says whether spreading those operations out really did soften the electrical spike.

+

Paired blocks hold useful work, state generations, failures, and learning commitments fixed. The only thing that changes is when checkpoint and rejoin traffic is allowed to be released.

Exact event intervals, clock alignment, sensor coverage, semantic obligations, and chunk hashes appear below.

@@ -972,7 +972,7 @@

What moved, and what the rack meter saw
-

Paired rack-PDU power traces aligned above job-level checkpoint and recovery event rails.

+

Paired power traces from the rack PDU, the unit that distributes power to the rack, drawn above the checkpoint and recovery event rails for the job.

@@ -981,7 +981,7 @@

What moved, and what the rack meter saw

Paired physical effects

-

Electrical benefit and systems cost stay adjacent

+

Electrical benefit and systems cost stay side by side

six evaluation blocks · 90% paired intervals @@ -1018,7 +1018,7 @@

Every electrical and semantic gate counts
-

Rack mechanism, not facility transfer

+

A rack result, not a whole-facility result

Evidence boundary loading.

@@ -1091,7 +1091,7 @@

Three-policy comparison

E001 v1 screening evidence chain

What the v1 learning prior did and did not know

- Evidence is immutable + Evidence cannot be edited
@@ -1107,7 +1107,7 @@

What the v1 learning prior did and did not know

Seed observations

-

Published source measurements are read from the observatory artifact. Publication-rounding intervals remain distinct from run-to-run variance.

+

Published source measurements are read from the observatory artifact. Uncertainty that comes from rounding in the paper is tracked separately from real run-to-run variation.

@@ -1117,7 +1117,7 @@

Seed observations

Transfer boundary

-

The attached literature records cover a narrow delay setting. They do not identify progress per FLOP, longer local-update intervals, frontier-scale transfer, multi-site interruption behavior, or an active-outage controller.

+

The attached literature covers one narrow delay setting. It says nothing about progress per FLOP, longer local-update intervals, frontier-scale transfer, how a multi-site run behaves when it is interrupted, or a controller that acts during an outage.

@@ -1139,15 +1139,15 @@

Prediction vs observation

No held-out multi-site learning observation - Prediction requires observed data. No residual can be computed. + A residual is the gap between a prediction and an observation, so with nothing observed there is nothing to compute.

What would resolve this?

    -
  1. 1

    Repeated small-model delay calibrationMeasure multiple delay intervals and optimizers rather than extrapolating one step.

  2. -
  3. 2

    Held-out optimizer, model, and site combinationsEvaluate combinations excluded from prior construction.

  4. -
  5. 3

    Controlled 30B to 100B-plus multi-site runVary delay and cadence under a defined policy with identical evaluation accounting.

  6. +
  7. 1

    Repeated small-model delay calibrationMeasure several delay intervals and optimizers instead of extrapolating from a single step.

  8. +
  9. 2

    Held-out optimizer, model, and site combinationsTest combinations that were deliberately left out when the prior was built.

  10. +
  11. 3

    Controlled 30B to 100B-plus multi-site runVary delay and cadence under one stated policy, counting the evaluation the same way every time.

@@ -1156,7 +1156,7 @@

What would resolve this?

Policy decision ledger

- The controller reads a completed communication cycle, then queues the next epoch. + The controller reads one finished communication cycle, then queues the next epoch.
ObservationValueEvidenceAction