This project is a Streamlit web application that leverages OpenAI's GPT-4o to generate descriptions for uploaded images
-
Updated
Jul 19, 2024 - Python
This project is a Streamlit web application that leverages OpenAI's GPT-4o to generate descriptions for uploaded images
AI Image Description Generator accurately extracts the key elements from images and interprets the creative purposes behind them, which can be applied in fields such as scientific research, artistic creation, and the mutual search between images and texts.
Image identification with Kosmos2 model, drawing and cutting bbox with object detection
为 DeepSeek V4.0 等单模态大模型装上眼睛和耳朵 —— 本地视觉+音频 MCP 服务器,基于 Ollama + MiniCPM-V 4.6 + faster-whisper,图片描述·视频分析·语音转文字
Dsh-visual-plugin.Give your text-only model eyes: forward user images to any OpenAI-compatible vision model and see the results in a Web UI right panel
Experimenting with mastodon.social client alt-text usage dataset.
DeepSeek Harness 识图插件:为不具备原生识图能力的模型提供识图能力(阿里云百炼 qwen3.5-omni-plus,失败自动切换智谱 glm-4.6v-flash)。由 claude-vision-skill 移植适配。 | Vision tool for DeepSeek Harness
给 DeepSeek Harness 纯文本模型装上原生视觉(Windows):粘贴即看图——预注入描述,模型首轮就看见,不用选模型、不用调工具;see_image 精查;自定义视觉后端(任意 OpenAI 兼容模型)+ 四后端容灾;换主模型视觉自动跟随。| Give text-only DeepSeek Harness models native-feeling vision on Windows: paste and the model just sees it — pre-injected descriptions, see_image tool, custom backends, 4-backend failover.
A see_image vision tool plugin for DeepSeek Harness — describe images through any OpenAI-compatible vision model (GitHub Copilot, OpenAI, Ollama, vLLM, LM Studio).
279K image alt-text pairs from 489 Bluesky accounts — curated for quality, validated at 90%+ alt-text rate
A lightweight console utility that uses an LLM to generate descriptions and keywords for images.
Give vision to non-vision models in OpenCode using Google's free Gemini API. Transparently replaces image parts with detailed text descriptions — vision models stay untouched.
A new package that processes user-submitted text descriptions of images or videos containing watermarks and returns structured, watermark-free descriptions. It uses an LLM to reinterpret the content w
An intelligent assistant powered by the ReAct framework, leveraging LangChain for tool-based reasoning and Gradio for a user-friendly interface. Supports tasks like weather queries, PDF summarization, image descriptions, and more.
It is an innovative repository housing a sophisticated Large Language Model (LLM) project, showcasing the intersection of advanced natural language processing and cutting-edge artificial intelligence. This repository serves as a comprehensive platform for the development, experimentation, and application of state-of-the-art language models.
AI-Powered-Solution-for-Assisting-Visually-Impaired-Individuals
A DeepSeek Harness tool plugin that lets text-only agents "see" local images — auto-detects the real format and returns a detailed text description via any OpenAI-compatible vision model.
CLI tool for generating image metadata in bulk — powered by Gemini for smarter, more context-aware descriptions.
DSH multimodal plugin: drag-image auto-describe with a configurable OpenAI-compatible vision model (host patch + agent preset + optional adapter). MIT.
GoldenLeaf is a Python application for creating an image search system using the CLIP model. It generates descriptions for herbarium images using the Llava model, enhancing multi-modal search capabilities. The system allows automated image description generation, multi-modal data augmentation, and customizable configurations for efficient training.
Add a description, image, and links to the image-description topic page so that developers can more easily learn about it.
To associate your repository with the image-description topic, visit your repo's landing page and select "manage topics."