A benchmark for spec-gaming resistance in LLMs: do models honor embedded policy rules when a plausible user request creates pressure to bend them?
-
Updated
Jun 25, 2026 - Python
A benchmark for spec-gaming resistance in LLMs: do models honor embedded policy rules when a plausible user request creates pressure to bend them?
Free 7-point checklist to red-team any 'green' AI/computational result — plus a teardown of our own retracted flagship result.
Rules in prompts are requests. EVAV makes them contracts. 13 SOTA integrity checks + behavior verification + signed AI provenance. One command before you ship.
让测试绿了不等于你做对了 — 双轴收敛纪律 (spec-gaming orthogonal axis) 的 DeepSeek Harness 原生实现 | Two-axis convergence discipline for coding agents
Add a description, image, and links to the spec-gaming topic page so that developers can more easily learn about it.
To associate your repository with the spec-gaming topic, visit your repo's landing page and select "manage topics."