This repository was archived by the owner on Oct 3, 2026. It is now read-only.
Add bounded offline visual critique experiment - #25
Merged
Merged
Conversation
5 tasks
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to subscribe to this conversation on GitHub.
Already have an account?
Sign in.
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Outcome
Closes #15 at its explicitly allowed offline completion boundary. Adds a bounded still-image experiment and evaluation harness; production UI/tools, model defaults, storage and manifests are unchanged. The live semantic gate remains NOT RUN and #22 must remain blocked.
The experiment accepts at most four metadata-free RGB PNGs with stable asset IDs, measures both input and actual SDK-generated request bytes, and rejects malformed images, aggregate overflow, unknown asset references and invented taste quotations. Empty evidence returns insufficient_evidence. A valid reference is not proof of semantic grounding.
Seven original synthetic cases cover single-photo composition and selected-frame proxies, contact sheet, mismatched taste, ambiguity, misleading metadata and embedded image instructions. Real photographic usefulness is deliberately unverified. The opt-in harness fixes the model snapshot, caps eight attempts/1,200 output tokens/45 seconds, disables retries, requires budget/documentation acknowledgement and records failures without exposing upstream errors. Estimated spend ceiling is $0.25/run, not a provider billing cap; no paid call was made.
Validation
npm ci --ignore-scriptscompleted; existing 13 advisories (4 moderate, 9 high) remain. No dependency changes.npm run typecheck— pass.npm test— 29 tests pass (10 new experiment tests).npm run lint— pass.node --import tsx scripts/eval-visual-critique.ts --offline— pass, seven fixture preflights, zero live calls.git diff --cached --checkandgitleaks git --staged --redact --no-banner— pass; complete source/doc/fixture diff reviewed.Evidence and gate
docs/visual-critique-evidence.mdrecords primary provider/hosting sources checked 2026-09-06, exact request/portable-record constraints, fixture provenance, the bounded live command and human rating criteria. Official AI SDK page retrieval was unavailable; installed SDK source/types and stubbed transport verified current local behavior. Recheck current official documentation before any live run.Recommendation: narrow and hold production integration. Offline closure of #15 is not authorization to unlock #22. A branch-only
vercel.jsonguard disables preview deployment forcodex/issue-15-visual-experiment, following the existing #12 pattern. Production defaults remain unchanged. No migration, deployment, merge or shared/global configuration change is included.