Skip to content

Repository files navigation

Diagnosing Aerial-View Object Detectors with Foundational Image Generative Models

Stanislav Panev1  Minhyek Jeon1  Vaishnavi Khindkar1  Ahish Deshpande1 
Celso de Melo2  Shuowen Hu2  Shayok Chakraborty1,3  Fernando De la Torre1

1Carnegie Mellon University  2DEVCOM Army Research Lab  3Florida State University

ECCV 2026

arXiv Project Page

Table of contents

Overview

AVODDiag is a suite of tools for generating synthetic aerial top-down view image benchmarks for diagnosing vehicle object detectors using commercial or opensource foundational text-to-image generative models, large language models (LLMs), and visual language models (VLMs).

APIs and supported models

Important

Unfortunately, Google has deprecated all versions of Imagen in their API. Gemini 3x (Nano Banana) models still can be used for image generation and editing.

  • Google GenAI API

    • Image generation

      • Imagen 3 (Deprecated)
      • Imagen 4 (Deprecated)
      • Gemini 3 Pro Image (Nano Banana Pro)
      • Gemini 3.1 Flash Image (Nano Banana 2)
    • Image editing

      • Gemini 2.5 Flash Image (Nano Banana)
    • Attribute extraction and image annotations

      • Gemini 2.5 Flash
      • Gemini 2.5 Flash Lite
  • OpenAI API

    • Prompt Composition
      • GPT-5

Setup

Create Anaconda Environment

This project is based on Python 3.11+.

(base) $ conda create -n avoddiag python=3.11
(base) $ conda activate avoddiag

Clone Project's GitHub Repository

(avoddiag) $ git clone https://github.com/humansensinglab/AVODDiag.git
(avoddiag) $ cd ./AVODDiag

Install Requirements

(avoddiag) $ python -m pip install \
    --build-constraint build-constraints.txt \
    -r requirements.txt

Config files

The config folder contains an example TOML config file example.toml. The config files currently hold the following parameters:

  • API keys for the Google and OpenAI providers.
  • Folder and file paths related to the bounding box approval process.

Usage

This package provides seven notebooks related to steps for generating synthetic diagnostic aerial-view image datasets from a pre-defined attribute taxonomy. This is a minimal working example and the workflow is as follows:

  1. Use 001_ImageGeneration.ipynb to generate the primary synthetic dataset based on pre-defined taxonomy attributes.
  2. Use 002_ExtractImageAttributes.ipynb to extract the taxonomy attributes of the generated images.
  3. Use 003_AnalyzeAttributes.ipynb to analyze and process the input and the generated attributes of the primary dataset.
  4. Use 004_ImageEditing.ipynb to edit the primary dataset in order to enrich.
  5. Use 003_AnalyzeAttributes.ipynb again to analyze and process the generated attributes of the secondary (image editing) dataset.
  6. Use 005_MergeDatasets.ipynb to merge the primary and secondary datasets.
  7. Use 006_ExtractImageAnnotations.ipynb to extract automatic annotations for all generated images.
  8. Use 007_ApproveAnnotations.ipynb to start a web-based application to approve or reject the automatic annotations.

Data

Below we provide image count information about the synthetic diagnostic and the three supplementary real datasets we used in our paper.

Dataset Name Type Train Split Test Split Total
Imagen 3 Synthetic – 5,453 5,453
Urban (Miami) Real 2,284 – 2,284
Industrial (LA) Real 2,000 – 2,000
Desert (Phoenix) Real 2,000 – 2,000

Synthetic Diagnostic Data

Imagen 3

  • Download link
  • SHA256: d4ba7002fe5714ebb129a9d4cc528646b93b50a8a94510dbaaca3608e1b4dce2

Folder structure:

Synthetic-Diagnostic_Imagen3.zip/
 ├─ annotations_coco/
 │  └─ matched_gemini-2.5-flash-lite_vikhyatk+moondream2_car_square_bboxes.json
 ├─ images/
 └─ metadata/

Real Supplementary Datasets

Each real supplementary dataset .zip file contains a QGIS project file and two subfolders—Layers and Python. To recreate each dataset, complete the steps below in the following order:

  1. Download and install the QGIS application on your device, if unavailable. We used version 3.44 to produce the project files, but the newer ones should also work fine.
  2. Unzip the archive.
  3. Open the provided .qgz project file with QGIS.
  4. Open the Python console within QGIS.
  5. Open and run 01_ExportRasterTiles.py located in Python folder to download and save the raster image tiles in images subfolder, which will be automatically created.
  6. Open and run 02_ExportCOCOAnnotations.py located in Python folder to generate COCO format bounding box annotations for the "small vehicle" class as a JSON file.

Urban Environment (Miami, FL, USA)

  • Download link
  • SHA256: 3c4ba7ee08dcf89fb57d0d9386b955d547df8c82c5c719948865b258fd741d9f

Industrial Environment (Los Angeles, CA, USA)

  • Download link
  • SHA256: 1b1af7e8745e5f6bd397a66564613cada89003ea05f805f977277e2e24434516

Desert Environment (Phoenix, AZ, USA)

Coming soon

BibTeX Citations

arXiv

@misc{panev2026diagnosingaerialviewobjectdetectors,
      title={Diagnosing Aerial-View Object Detectors with Foundational Image Generative Models}, 
      author={Stanislav Panev and Minhyek Jeon and Vaishnavi Khindkar and Ahish Deshpande and Celso M de Melo and Shuowen Hu and Shayok Chakraborty and Fernando De la Torre},
      year={2026},
      eprint={2607.02718},
      archivePrefix={arXiv},
      primaryClass={cs.CV},
      url={https://arxiv.org/abs/2607.02718}, 
}

About

A suite of tools for generating synthetic aerial top-down view image benchmarks for diagnosing object detectors.

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages