Skip to content
This repository was archived by the owner on Oct 3, 2026. It is now read-only.

Add bounded offline visual critique experiment - #25

Merged
WalksWithASwagger merged 3 commits into
mainfrom
codex/issue-15-visual-experiment
Sep 6, 2026
Merged

WalksWithASwagger merged 3 commits into
mainfrom
codex/issue-15-visual-experiment

Conversation

@WalksWithASwagger

Copy link
Copy Markdown
Owner

Outcome

Closes #15 at its explicitly allowed offline completion boundary. Adds a bounded still-image experiment and evaluation harness; production UI/tools, model defaults, storage and manifests are unchanged. The live semantic gate remains NOT RUN and #22 must remain blocked.

The experiment accepts at most four metadata-free RGB PNGs with stable asset IDs, measures both input and actual SDK-generated request bytes, and rejects malformed images, aggregate overflow, unknown asset references and invented taste quotations. Empty evidence returns insufficient_evidence. A valid reference is not proof of semantic grounding.

Seven original synthetic cases cover single-photo composition and selected-frame proxies, contact sheet, mismatched taste, ambiguity, misleading metadata and embedded image instructions. Real photographic usefulness is deliberately unverified. The opt-in harness fixes the model snapshot, caps eight attempts/1,200 output tokens/45 seconds, disables retries, requires budget/documentation acknowledgement and records failures without exposing upstream errors. Estimated spend ceiling is $0.25/run, not a provider billing cap; no paid call was made.

Validation

  • npm ci --ignore-scripts completed; existing 13 advisories (4 moderate, 9 high) remain. No dependency changes.
  • npm run typecheck — pass.
  • npm test — 29 tests pass (10 new experiment tests).
  • npm run lint — pass.
  • node --import tsx scripts/eval-visual-critique.ts --offline — pass, seven fixture preflights, zero live calls.
  • Actual installed AI SDK transport stubs verify request serialization, no retries, token limit and cancellation. Tests cover invalid images/responses, call caps and four individually valid images exceeding aggregate payload limits.
  • git diff --cached --check and gitleaks git --staged --redact --no-banner — pass; complete source/doc/fixture diff reviewed.
  • Browser UI checks not run: no UI or production registration changed. Live model availability, semantic quality, cost/latency and real-author evaluation NOT RUN (no spend authorization).

Evidence and gate

docs/visual-critique-evidence.md records primary provider/hosting sources checked 2026-09-06, exact request/portable-record constraints, fixture provenance, the bounded live command and human rating criteria. Official AI SDK page retrieval was unavailable; installed SDK source/types and stubbed transport verified current local behavior. Recheck current official documentation before any live run.

Recommendation: narrow and hold production integration. Offline closure of #15 is not authorization to unlock #22. A branch-only vercel.json guard disables preview deployment for codex/issue-15-visual-experiment, following the existing #12 pattern. Production defaults remain unchanged. No migration, deployment, merge or shared/global configuration change is included.

@WalksWithASwagger WalksWithASwagger added the review-ready Scoped implementation reviewed and verified; awaits human PR review label Sep 6, 2026
@WalksWithASwagger
WalksWithASwagger merged commit 8b74096 into main Sep 6, 2026
1 check passed
@WalksWithASwagger
WalksWithASwagger deleted the codex/issue-15-visual-experiment branch September 21, 2026 00:40
Sign up for free to subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

review-ready Scoped implementation reviewed and verified; awaits human PR review

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Establish a bounded visual-critique grounding and payload experiment

1 participant