You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Two of the 13 seeded frontier_models rows failed 300/300 on 2026-02-16 with provider-side "not a valid model ID" errors:
model_id
provider
error
openrouter-llama-4-maverick
OpenRouter
HTTP 400 meta-llama/llama-4-maverick-instruct is not a valid model ID
together-deepseek-v3.1
Together
HTTP 404 Unable to access model deepseek-ai/DeepSeek-V3.1-0324
Migrations 017_seed_frontier_models.sql, 019_fix_openrouter_model_id.sql, 020_fix_together_model_ids.sql touch these IDs — check whether 019/020 actually fixed the default_model_name these rows carry in production (the failures may predate 019/020, but the seed should be verified against the providers' current catalogues rather than assumed). Note the EU database later recorded 4 completed runs for both models, so a working ID existed at some point that evening — compare the target_model/default_model_name used by the later successful rows.
Asks:
A startup or pre-sweep model-ID validation step: call each provider's /models (or a 1-token completion) for every active frontier_models row and mark invalid ones active=false with the provider error, before burning 300 scenarios each.
Fix the two IDs (or retire the models) in a new migration; do not edit 017–020 in place.
Two of the 13 seeded
frontier_modelsrows failed 300/300 on 2026-02-16 with provider-side "not a valid model ID" errors:openrouter-llama-4-maverickmeta-llama/llama-4-maverick-instruct is not a valid model IDtogether-deepseek-v3.1Unable to access model deepseek-ai/DeepSeek-V3.1-0324Migrations
017_seed_frontier_models.sql,019_fix_openrouter_model_id.sql,020_fix_together_model_ids.sqltouch these IDs — check whether 019/020 actually fixed thedefault_model_namethese rows carry in production (the failures may predate 019/020, but the seed should be verified against the providers' current catalogues rather than assumed). Note the EU database later recorded 4 completed runs for both models, so a working ID existed at some point that evening — compare thetarget_model/default_model_nameused by the later successful rows.Asks:
/models(or a 1-token completion) for every activefrontier_modelsrow and mark invalid onesactive=falsewith the provider error, before burning 300 scenarios each.gemini-2.5-flash4,gemini-2.5-pro2,o34,openrouter-llama-4-maverick4,together-deepseek-v3.14) are all there because of provider errors or the 12 orphanedrunningrows (Frontier sweep cannot run: flag off, no beat scheduler exists, 15 evaluations stuck in status='running' since 2026-02-16 #34) — fixing IDs + the reaper + one clean sweep gets all 13 publishing.Related: #34, #39, #40.