Skip to content

feat(B.7): benchmark de calidad con 10 casos etiquetados ejecutado contra el servicio real - #7

Merged
klssxx merged 1 commit into
mainfrom
feat/b7-quality-benchmark
Sep 25, 2026
Merged

klssxx merged 1 commit into
mainfrom
feat/b7-quality-benchmark

Conversation

@klssxx

@klssxx klssxx commented Sep 25, 2026

Copy link
Copy Markdown
Owner

Eje B — B.7

tests/fixtures/prompt_cases/cases.json: los 10 casos del spec (ambiguo, seguridad explícita, variables sin rellenar, contradictorio, vacío→fallo controlado, investigación con fuentes, JSON con esquema, datos sensibles, ya excelente, target desconocido).

run_case_benchmark() / evaluate_case() ejecutan cada caso contra compile_prompt real y verifican 6 invariantes por caso: preservación del goal, detección de huecos, no-invencción (en la capa de REQUISITOS B.4 — la política de producto en el texto es etiquetada, no invención), ejecutabilidad, retención de restricciones críticas, y no-degradación del caso 9 (score compilado ≥ score del input − 5).

Hallazgos reales que el benchmark destapó (y este PR corrige)

  1. compile_prompt aceptaba targets de modelo desconocidos en silencio y los propagaba al NSL (TARGET=modelo_fantasma_xyz). Paridad con P1-2 restaurada: mismo ValueError claro que compile_for_gui.
  2. El caso 9 falló al principio porque el prompt "excelente" no nombraba target — B.2 lo señala legítimamente; el fixture ahora es genuinamente completo.

El benchmark es el test: 11 funciones de test ejecutan los 10 casos. Suite: 297 passed + 10 subtests, ruff limpio.

… service

B.7: tests/fixtures/prompt_cases/cases.json holds the ten labeled cases
from the product spec (ambiguous coding, Spanish safety constraints,
unfilled variables, contradictory, empty-input controlled failure,
research-with-sources, JSON contract, sensitive data, already-excellent,
unknown target). run_case_benchmark()/evaluate_case() in
benchmark_quality.py run each case through compile_prompt and verify six
invariants per case:

1 goal preservation (distinctive goal tokens survive compilation)
2 gap detection (ambiguities present iff the case expects them)
3 no invention (the B.4 USER-REQUIREMENTS layer never contains a
  critical constraint the input did not declare — the compiled TEXT
  intentionally carries the product-policy set, which is labeled
  policy, not invention)
4 executability (output contract present when required)
5 critical-constraint retention (dash/underscore normalized)
6 non-degradation (case 9: compiled score >= input's own score - 5)

The benchmark IS the test: tests/test_benchmark_cases.py executes all
ten (11 test functions) against the real service.

Hallazgos reales que este benchmark destapó y corrige:
- compile_prompt accepted UNKNOWN model targets silently and echoed them
  into NSL artifacts (compile_for_gui validated since P1-2 — a parity
  gap). compile_prompt now raises the same clear ValueError.
- case09 initially failed because the 'excellent' prompt did not name a
  model target — B.2 legitimately flags that as a gap; the fixture now
  represents a genuinely complete prompt.

Suite: 297 passed + 10 subtests, ruff clean.
@klssxx
klssxx merged commit 43f95e5 into main Sep 25, 2026
1 check passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant