Images whose caption failed are described on the next pass - #368
Merged
Merged
Conversation
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Images whose caption failed are described on the next pass (#368)
An uploaded image got no text in the index and never would. The configured vision model refused the caption request (403, in 0.2 s), OCR correctly found no text in a photograph, and the indexer cached the empty result: images cache even when empty, so a picture with nothing to transcribe is not re-sent to the vision model on every pass. That rule could not tell an image with nothing to say from one that was never described, so the empty entry outlived the failure, including after the vision model setting was changed to one that works. The failure itself was invisible:
enrich.pyreports it on stderr, which the server discarded whenever the pass as a whole succeeded.Retry.
caption_imagenow returnsNonewhen the request fails and''when captioning is off or the image is too large. A failed caption is not cached, so the next pass asks again; the OCR text, if any, is still indexed for that pass. While a vision model is configured, an empty cached image result is treated as missing and re-extracted: a successful caption is never empty, so such an entry was written by a failed request or before captioning was on. With captioning off, empty entries stand as before.warm_image_cache.pyfollows the same rules.Visible. Lines from the indexer's stderr that report a failure (
caption failed ...,extract failed ...) go to the server log through the sweep's logger, prefixedindex:, even when the pass succeeds.Tested with the real
enrich.pyagainst a stand-in gateway in the test process: a refused caption leaves no cache entry and the next pass, with the gateway answering, caches the description; a good entry is reused without a request; an empty entry is retried with a model configured and kept with none. All server and web tests pass.