For status='failed' rows there is no dedicated error field. The failure is stored as scenario_results = {"error": …, "first_error": …} with total/correct/errors NULL — the same column that holds per-scenario results for successful runs. Examples from 2026-02-16 (host inventory 2026-09-24):
gemini-2.5-pro: "Catastrophic error rate: 277/300", first_error "Empty response (finishReason=MAX_TOKENS)"
openrouter-llama-4-maverick: HTTP 400 "meta-llama/llama-4-maverick-instruct is not a valid model ID" (300/300)
together-deepseek-v3.1: HTTP 404 "Unable to access model deepseek-ai/DeepSeek-V3.1-0324" (300/300)
gemini-2.5-pro ×4 total: HTTP 429 quota exceeded
claude-sonnet-4-5 / claude-opus-4-6 (EU): HTTP 400 "Your credit balance is too low to access the Anthropic API"
Every one of these is an operational failure (bad model ID, quota, billing) that nobody saw for seven months because the only place it's recorded is a JSONB blob on a row the scores endpoint filters out.
Asks
- Add
error_code / error_message / error_class (provider_config | quota | billing | model_output | timeout) columns; keep scenario_results for results.
- Surface failures:
/api/v1/admin/frontier-sweeps/{id} should list failed models with error_class, and the admin UI Frontier page should show them red instead of just "N completed".
- Sweep summary notification (email/Discord —
services/notifications.py already exists) when any model fails with provider_config or billing — those are the two classes that will fail identically next week.
Related: #34, #39.
For
status='failed'rows there is no dedicated error field. The failure is stored asscenario_results = {"error": …, "first_error": …}withtotal/correct/errorsNULL — the same column that holds per-scenario results for successful runs. Examples from 2026-02-16 (host inventory 2026-09-24):gemini-2.5-pro: "Catastrophic error rate: 277/300", first_error "Empty response (finishReason=MAX_TOKENS)"openrouter-llama-4-maverick: HTTP 400 "meta-llama/llama-4-maverick-instruct is not a valid model ID" (300/300)together-deepseek-v3.1: HTTP 404 "Unable to access model deepseek-ai/DeepSeek-V3.1-0324" (300/300)gemini-2.5-pro×4 total: HTTP 429 quota exceededclaude-sonnet-4-5/claude-opus-4-6(EU): HTTP 400 "Your credit balance is too low to access the Anthropic API"Every one of these is an operational failure (bad model ID, quota, billing) that nobody saw for seven months because the only place it's recorded is a JSONB blob on a row the scores endpoint filters out.
Asks
error_code/error_message/error_class(provider_config | quota | billing | model_output | timeout) columns; keepscenario_resultsfor results./api/v1/admin/frontier-sweeps/{id}should list failed models witherror_class, and the admin UI Frontier page should show them red instead of just "N completed".services/notifications.pyalready exists) when any model fails withprovider_configorbilling— those are the two classes that will fail identically next week.Related: #34, #39.