Stanislav Panev1
Minhyek Jeon1
Vaishnavi Khindkar1
Ahish Deshpande1
Celso de Melo2
Shuowen Hu2
Shayok Chakraborty1,3
Fernando De la Torre1
1Carnegie Mellon University 2DEVCOM Army Research Lab 3Florida State University
ECCV 2026
AVODDiag is a suite of tools for generating synthetic aerial top-down view image benchmarks for diagnosing vehicle object detectors using commercial or opensource foundational text-to-image generative models, large language models (LLMs), and visual language models (VLMs).
Important
Unfortunately, Google has deprecated all versions of Imagen in their API. Gemini 3x (Nano Banana) models still can be used for image generation and editing.
-
Google GenAI API
-
Image generation
- Imagen 3 (Deprecated)
- Imagen 4 (Deprecated)
- Gemini 3 Pro Image (Nano Banana Pro)
- Gemini 3.1 Flash Image (Nano Banana 2)
-
Image editing
- Gemini 2.5 Flash Image (Nano Banana)
-
Attribute extraction and image annotations
- Gemini 2.5 Flash
- Gemini 2.5 Flash Lite
-
-
OpenAI API
- Prompt Composition
- GPT-5
- Prompt Composition
Create Anaconda Environment
This project is based on Python 3.11+.
(base) $ conda create -n avoddiag python=3.11
(base) $ conda activate avoddiagClone Project's GitHub Repository
(avoddiag) $ git clone https://github.com/humansensinglab/AVODDiag.git
(avoddiag) $ cd ./AVODDiagInstall Requirements
(avoddiag) $ python -m pip install \
--build-constraint build-constraints.txt \
-r requirements.txtThe config folder contains an example TOML config file example.toml. The config files currently hold the following parameters:
- API keys for the Google and OpenAI providers.
- Folder and file paths related to the bounding box approval process.
This package provides seven notebooks related to steps for generating synthetic diagnostic aerial-view image datasets from a pre-defined attribute taxonomy. This is a minimal working example and the workflow is as follows:
- Use
001_ImageGeneration.ipynbto generate the primary synthetic dataset based on pre-defined taxonomy attributes. - Use
002_ExtractImageAttributes.ipynbto extract the taxonomy attributes of the generated images. - Use
003_AnalyzeAttributes.ipynbto analyze and process the input and the generated attributes of the primary dataset. - Use
004_ImageEditing.ipynbto edit the primary dataset in order to enrich. - Use
003_AnalyzeAttributes.ipynbagain to analyze and process the generated attributes of the secondary (image editing) dataset. - Use
005_MergeDatasets.ipynbto merge the primary and secondary datasets. - Use
006_ExtractImageAnnotations.ipynbto extract automatic annotations for all generated images. - Use
007_ApproveAnnotations.ipynbto start a web-based application to approve or reject the automatic annotations.
Below we provide image count information about the synthetic diagnostic and the three supplementary real datasets we used in our paper.
| Dataset Name | Type | Train Split | Test Split | Total |
|---|---|---|---|---|
| Imagen 3 | Synthetic | – | 5,453 | 5,453 |
| Urban (Miami) | Real | 2,284 | – | 2,284 |
| Industrial (LA) | Real | 2,000 | – | 2,000 |
| Desert (Phoenix) | Real | 2,000 | – | 2,000 |
- Download link
- SHA256:
d4ba7002fe5714ebb129a9d4cc528646b93b50a8a94510dbaaca3608e1b4dce2
Folder structure:
Synthetic-Diagnostic_Imagen3.zip/
├─ annotations_coco/
│ └─ matched_gemini-2.5-flash-lite_vikhyatk+moondream2_car_square_bboxes.json
├─ images/
└─ metadata/
Each real supplementary dataset .zip file contains a QGIS project file and two subfolders—Layers and Python. To recreate each dataset, complete the steps below in the following order:
- Download and install the QGIS application on your device, if unavailable. We used version 3.44 to produce the project files, but the newer ones should also work fine.
- Unzip the archive.
- Open the provided
.qgzproject file with QGIS. - Open the Python console within QGIS.
- Open and run
01_ExportRasterTiles.pylocated inPythonfolder to download and save the raster image tiles inimagessubfolder, which will be automatically created. - Open and run
02_ExportCOCOAnnotations.pylocated inPythonfolder to generate COCO format bounding box annotations for the "small vehicle" class as a JSON file.
- Download link
- SHA256:
3c4ba7ee08dcf89fb57d0d9386b955d547df8c82c5c719948865b258fd741d9f
- Download link
- SHA256:
1b1af7e8745e5f6bd397a66564613cada89003ea05f805f977277e2e24434516
Coming soon
arXiv
@misc{panev2026diagnosingaerialviewobjectdetectors,
title={Diagnosing Aerial-View Object Detectors with Foundational Image Generative Models},
author={Stanislav Panev and Minhyek Jeon and Vaishnavi Khindkar and Ahish Deshpande and Celso M de Melo and Shuowen Hu and Shayok Chakraborty and Fernando De la Torre},
year={2026},
eprint={2607.02718},
archivePrefix={arXiv},
primaryClass={cs.CV},
url={https://arxiv.org/abs/2607.02718},
}