You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
/api/v1/scores shows eval_count: 5–6 per model and the page says "Evaluated weekly". Host inventory (2026-09-24) of evaluations.trace_id (sweep-v1.0-<8hex>/…) shows every one of those evaluations comes from one evening: US ran six sweep invocations between 04:40 and 05:46 UTC on 2026-02-16 (586e3662, 03808eaf, 33367ba8, 01098633, 340be9c7, 8a2b3989); EU ran four more between 06:04 and 08:07 (4f3270f4, 8bcf80c3, 39a48a7f, fe94fd7b). Those are an operator retrying the same sweep after provider errors (bad model IDs, Gemini 429/MAX_TOKENS, Anthropic credit), not five independent weekly runs.
So today: eval_count, avg/min/max/stddev_accuracy, the 5-eval publish threshold, and the 10-eval trend gate all treat retries of one sweep as independent samples. The threshold that gates publication was crossed by retrying, not by evidence.
Asks
Aggregate by sweep window (e.g. ISO week, or a sweep_id column derived from the trace_id prefix), not by raw row count: eval_count = number of distinct sweep windows with a completed run; within a window, take the latest completed run (or the median) as that window's sample.
Publish threshold and trend gate should count windows, not rows.
What the published numbers actually are
/api/v1/scoresshowseval_count: 5–6per model and the page says "Evaluated weekly". Host inventory (2026-09-24) ofevaluations.trace_id(sweep-v1.0-<8hex>/…) shows every one of those evaluations comes from one evening: US ran six sweep invocations between 04:40 and 05:46 UTC on 2026-02-16 (586e3662,03808eaf,33367ba8,01098633,340be9c7,8a2b3989); EU ran four more between 06:04 and 08:07 (4f3270f4,8bcf80c3,39a48a7f,fe94fd7b). Those are an operator retrying the same sweep after provider errors (bad model IDs, Gemini 429/MAX_TOKENS, Anthropic credit), not five independent weekly runs.So today:
eval_count,avg/min/max/stddev_accuracy, the 5-eval publish threshold, and the 10-eval trend gate all treat retries of one sweep as independent samples. The threshold that gates publication was crossed by retrying, not by evidence.Asks
sweep_idcolumn derived from thetrace_idprefix), not by raw row count:eval_count= number of distinct sweep windows with a completed run; within a window, take the latest completed run (or the median) as that window's sample.first_evaluated_at/last_evaluated_at/windowson/api/v1/scoresso the site can say "1 sweep, Feb 2026" instead of implying weekly cadence — or the site should stop saying "Evaluated weekly" until Frontier sweep cannot run: flag off, no beat scheduler exists, 15 evaluations stuck in status='running' since 2026-02-16 #34 makes it true.Related: #34 (sweep can't run), CIRISAI/CIRISCore#2 (US/EU split of the same night's retries).