Rewrite forecast-accuracy claims around a defensible benchmark#88
Merged
Conversation
The site claimed the SVAR "beats the random walk clearly on CPI at every horizon" and had "useful inflation signal". A driftless random walk on a trending log price level forfeits the whole trend as forecast error, so that comparison is close to uninformative. Against a random walk with drift the CPI ratio goes from 0.63 to 0.83 at h=1 and from 0.67 to 1.03 at h=8 -- no better than naive. The tell was in the original numbers: the model won only on the two trending price series and tied or lost on the four series where a random walk is genuinely hard to beat. That is the signature of a benchmark artefact, not of inflation signal. What survives is Bank Rate -- the one non-trending series, where no-change is the right naive -- which improves under the harder benchmark to 0.79 at h=1 (p=0.018). The page now says that is the defensible forecasting claim. Also corrects the GDP framing in the other direction: 1.06-1.09 is not significant at any horizon, and drops to 0.77 excluding six Covid-target origins. Both figures published, neither preferred. The AR(1) comparison is shown as unusable rather than as a 3:1 win, since 94.5% of its eight-step squared error for UK GDP comes from one origin. Predictive validation for boe-svar downgraded moderate -> weak, which is what the evidence now supports. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Two items the persona reviews flagged independently. Coverage: "all fourteen outturns fall inside the 68% bands" was listed among the headline validation statistics. A calibrated 68% interval should contain about 9.5 of 14; containing all of them means the bands are ~3x wider than a 0.3pp RMSE warrants. The model is under-confident and the figure was being scored as a win. It now says so -- and adds that the fourteen points are seven quarters x two variables from a single origin, so coverage cannot be established from this sample in either direction. The page now makes no coverage claim. Multiplier: the OBR emulator's impact multiplier for current spending is ~1.0 by construction against the OBR's own published 0.6 -- a two-thirds overstatement against the institution being replicated, applying to every spending-side figure on the page. It was disclosed on the model page and absent from the evidence page. Also names the three exogenous channels (imports, Bank Rate, CPI) that bias the estimate upward and the dead dividends channel that biases the other way, and states that they do not cancel. RMSE rounded 0.32pp -> 0.3pp: two significant figures from seven overlapping observations implied precision that is not there. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
|
The latest updates on your projects. Learn more about Vercel for GitHub.
|
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Follows boe-var-model#6. Three claims on the validation page did not survive scrutiny; two of them were flattering and one was unflattering.
1. The CPI "win" was a benchmark artefact
The page said the SVAR "beats the random walk clearly on CPI at every horizon" and has "useful inflation signal". Against a random walk with drift — the textbook naive for a trending log level — CPI goes 0.63 → 0.83 at h=1 and 0.67 → 1.03 at h=8, i.e. no better than naive.
What survives is Bank Rate: the one non-trending series, which improves under the harder benchmark to 0.79 at h=1 (p=0.018). The page now states that is the defensible forecasting claim, not inflation.
2. The GDP "loss" was overstated in the other direction
1.06–1.09 is not statistically distinguishable from a random walk at any horizon (p = 0.38–0.67), and drops to 0.77 excluding six Covid-target origins. Both published; neither preferred. Over-claiming a loss reads as humility but is the same error as over-claiming a win.
3. Coverage was scored backwards
"All fourteen outturns fall inside the 68% bands" was listed as a headline validation statistic. A calibrated 68% band should contain ~9.5 of 14. Containing all of them means the intervals are ~3× wider than a 0.3pp RMSE warrants — over-dispersion, presented as a pass. And the fourteen points are seven quarters × two variables from a single origin, so coverage isn't establishable either way. The page now makes no coverage claim.
4. The multiplier gap was missing from the evidence page
The OBR emulator's impact multiplier for current spending is ~1.0 by construction against the OBR's own published 0.6 — a two-thirds overstatement against the institution being replicated, applying to every spending-side figure. It was on the model page and absent here. Now stated, along with the three exogenous channels (imports, Bank Rate, CPI) that bias upward and the dead dividends channel that biases down — and that they do not cancel.
predictive_validationfor boe-svar downgraded moderate → weak, which is what the evidence now supports. 183 integration tests pass.🤖 Generated with Claude Code