feat(B.7): benchmark de calidad con 10 casos etiquetados ejecutado contra el servicio real - #7
Merged
Merged
Conversation
… service B.7: tests/fixtures/prompt_cases/cases.json holds the ten labeled cases from the product spec (ambiguous coding, Spanish safety constraints, unfilled variables, contradictory, empty-input controlled failure, research-with-sources, JSON contract, sensitive data, already-excellent, unknown target). run_case_benchmark()/evaluate_case() in benchmark_quality.py run each case through compile_prompt and verify six invariants per case: 1 goal preservation (distinctive goal tokens survive compilation) 2 gap detection (ambiguities present iff the case expects them) 3 no invention (the B.4 USER-REQUIREMENTS layer never contains a critical constraint the input did not declare — the compiled TEXT intentionally carries the product-policy set, which is labeled policy, not invention) 4 executability (output contract present when required) 5 critical-constraint retention (dash/underscore normalized) 6 non-degradation (case 9: compiled score >= input's own score - 5) The benchmark IS the test: tests/test_benchmark_cases.py executes all ten (11 test functions) against the real service. Hallazgos reales que este benchmark destapó y corrige: - compile_prompt accepted UNKNOWN model targets silently and echoed them into NSL artifacts (compile_for_gui validated since P1-2 — a parity gap). compile_prompt now raises the same clear ValueError. - case09 initially failed because the 'excellent' prompt did not name a model target — B.2 legitimately flags that as a gap; the fixture now represents a genuinely complete prompt. Suite: 297 passed + 10 subtests, ruff clean.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Eje B — B.7
tests/fixtures/prompt_cases/cases.json: los 10 casos del spec (ambiguo, seguridad explícita, variables sin rellenar, contradictorio, vacío→fallo controlado, investigación con fuentes, JSON con esquema, datos sensibles, ya excelente, target desconocido).run_case_benchmark()/evaluate_case()ejecutan cada caso contracompile_promptreal y verifican 6 invariantes por caso: preservación del goal, detección de huecos, no-invencción (en la capa de REQUISITOS B.4 — la política de producto en el texto es etiquetada, no invención), ejecutabilidad, retención de restricciones críticas, y no-degradación del caso 9 (score compilado ≥ score del input − 5).Hallazgos reales que el benchmark destapó (y este PR corrige)
compile_promptaceptaba targets de modelo desconocidos en silencio y los propagaba al NSL (TARGET=modelo_fantasma_xyz). Paridad con P1-2 restaurada: mismo ValueError claro quecompile_for_gui.El benchmark es el test: 11 funciones de test ejecutan los 10 casos. Suite: 297 passed + 10 subtests, ruff limpio.