Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
33 changes: 29 additions & 4 deletions ARCHITECTURE.md
Original file line number Diff line number Diff line change
Expand Up @@ -147,6 +147,23 @@ layouts). `vp sync` reconciles the index idempotently (content-addressed, so
re-running skips what's present) and records per-asset provenance; `vp sync
--dry-run` previews found/matched/unmatched/classes per source.

**Multi-provider targets.** With a `target:` and `copy` mode, objects land in a
content-addressed sink. Same-provider transfers are server-side (S3 CopyObject /
GCS rewrite — bytes never transit the client). Cross-provider transfers
(local→S3, S3→GCS, …) *relay* the bytes sync already read to compute the sha256:
one read (needed anyway) + one upload (size-verified against the landed
object), never a second download. Either way an unchanged re-sync stays
metadata-only.

**Request economy & resilience.** Target membership is resolved with one prefix
listing per run (rclone-style fast-list), never a per-object existence check.
Every remote call is wrapped in exponential-backoff retries
(`sources/retry.py`); exhaustion surfaces as a per-object `IngestFailure`, so a
flaky bucket degrades a sync instead of aborting it. Ingest concurrency is
sized for latency-bound object stores (16+ workers for remote sources) and
tunable with `vp sync --jobs`. Labels fetched during class inference are cached
and replayed at ingest, so no remote label is ever read twice.

### Geometry model & task coverage (`core/models.py`, `formats/`)
The tagged geometry above lets one dataset model cover the common CV tasks.
Classification uses the ImageFolder convention; detection uses YOLO/COCO bbox;
Expand Down Expand Up @@ -223,11 +240,17 @@ The roadmap is sequenced so each phase unblocks the next.
- [x] `vp pack --profile training` (WebDataset shards)
- [ ] DuckDB index — deferred by decision (see Storage & index)

### Multi-source ingestion ✅ (remote backends pending)
### Multi-source ingestion ✅
- [x] declarative `sources:` + `vp sync` (+ `--dry-run`), local backends, joins,
provenance, class reconciliation, idempotent re-sync
- [ ] remote backends via fsspec behind extras (s3/gcs/azure/git, pinned by ref)
and COCO-format sources, plugging into the same resolver layer
- [x] remote YOLO/COCO/ImageFolder sources via fsspec behind extras
(s3/gcs/azure), metadata-only re-sync, content-addressed cloud target
- [x] cross-provider `copy` targets (local↔S3, S3↔GCS, …): server-side copy when
providers match, single-pass verified byte relay when they don't
- [x] hardening: retries with backoff on every remote call, fast-list target
membership (no per-object HEAD), `--jobs` concurrency control,
single-read remote labels
- [ ] git sources pinned by ref

### Task coverage ✅ (beyond detection)
- [x] tagged geometry model (bbox | polygon | keypoints | none), backward compatible
Expand Down Expand Up @@ -259,8 +282,10 @@ The roadmap is sequenced so each phase unblocks the next.
- [ ] Hugging Face Datasets export

### Phase C — Reporting & polish
- [x] machine-readable output: every pipeline command takes `--json` and prints
a schema-versioned envelope (`cli/output.py`) — the stable contract for
driving VisionPack from services/UIs/CI; errors are structured too
- [ ] HTML validation / stats / drift reports
- [ ] JSON report output for stats and diff
- [ ] richer terminal output with `rich`
- [ ] move CLI plumbing from `argparse` to `typer` once commands stabilize

Expand Down
16 changes: 12 additions & 4 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -9,7 +9,7 @@ leak-free, ready-to-train dataset.
![Python](https://img.shields.io/badge/python-3.11%2B-blue)
![License](https://img.shields.io/badge/license-Apache--2.0-blue)
![Status](https://img.shields.io/badge/status-active%20development-orange)
![Tests](https://img.shields.io/badge/tests-118%20passing-brightgreen)
![Tests](https://img.shields.io/badge/tests-143%20passing-brightgreen)

[Documentation](https://caiowing.github.io/VisionPack/) ·
[Install](https://caiowing.github.io/VisionPack/installation/) ·
Expand Down Expand Up @@ -176,8 +176,12 @@ See the [Cloud Sync guide](https://caiowing.github.io/VisionPack/cloud-sync/).
extra dependencies, scale-proof via LSH bucketing; surfaced in `vp validate`.
- **Multi-source sync** — declarative `sources:` + `vp sync`, with per-asset
provenance; idempotent re-sync that only pulls what's new.
- **Cloud-native** — sync from and to S3/GCS/Azure without downloading the whole
dataset; server-side `copy` into a content-addressed target, streaming export.
- **Cloud-native, multi-provider** — sync YOLO, COCO, and ImageFolder sources
from S3/GCS/Azure without downloading the whole dataset; server-side `copy`
into a content-addressed target when source and target share a provider,
single-pass verified relay across providers (S3→GCS, local→S3, …); one
fast-list instead of per-object lookups, retries with backoff on every remote
call, tunable concurrency (`--jobs`); streaming export.
- **Content-addressed snapshots & diff** — reproducible versions; compare any two.
- **Strong validation** — unreadable images, missing/orphan labels, unknown
classes, invalid/out-of-bounds boxes, exact + near duplicates, split leakage.
Expand All @@ -195,6 +199,10 @@ See the [Cloud Sync guide](https://caiowing.github.io/VisionPack/cloud-sync/).
(WebDataset shards); exports hardlink from the CAS or stream from the cloud.
- **Interoperable I/O** — YOLO (incl. YOLO-seg), COCO, ImageFolder in and out;
semantic masks out.
- **Machine-readable everything** — every pipeline command takes `--json` and
prints a stable, schema-versioned envelope on stdout, so services, UIs, and CI
can drive VisionPack without scraping text. See the
[JSON Output guide](https://caiowing.github.io/VisionPack/json-output/).

Full command reference and per-command options live in the
[CLI guide](https://caiowing.github.io/VisionPack/usage/).
Expand Down Expand Up @@ -246,7 +254,7 @@ multi-source ingestion (local and cloud) → validation → deterministic splits
snapshots → ready-to-train export/packing → evaluation (`vp eval`) and
model-in-the-loop labeling (`vp autolabel` / `vp queue`) — works end-to-end
across classification, detection, instance/semantic segmentation, and keypoints,
with 118 passing tests. APIs may still shift; feedback and contributions are
with 143 passing tests. APIs may still shift; feedback and contributions are
welcome.

```bash
Expand Down
27 changes: 24 additions & 3 deletions docs/cloud-sync.md
Original file line number Diff line number Diff line change
Expand Up @@ -33,10 +33,22 @@ pip install "visionpack[azure]" # Azure Blob
cause a mismatch.
- **Re-sync is metadata-only.** Re-running lists object metadata, sees the etags
match, and does nothing — no downloads, no copies.
- **One listing beats many lookups.** Source metadata comes from paginated
LISTs (never a per-object HEAD), and the target CAS is checked with **one
prefix listing per run** instead of a per-object existence check — at 100k
objects that's ~100 LISTs, not 100k HEADs.
- **Transient errors are retried.** Every remote call gets exponential-backoff
retries (throttling, dropped connections, 5xx). If an object still fails, it
is recorded as a per-object failure and the rest of the sync proceeds — the
command exits non-zero so CI can gate on it.

{: .note }
v1 is **same-provider** (S3↔S3 or GCS↔GCS). Cross-cloud transfer (S3↔GCS) is on
the roadmap.
Transfers are **server-side within one provider** (S3↔S3, GCS↔GCS) — the bytes
never touch your machine. **Cross-provider** targets (S3→GCS, local→S3, S3→local
…) also work: sync *relays* the bytes it already read to compute the `sha256`,
so a cross-provider copy still costs exactly **one read + one upload**, never a
second download. Relayed uploads are verified against the landed object's
metadata before the index points at them.

## Declare remote sources

Expand Down Expand Up @@ -75,8 +87,17 @@ Then reconcile. Re-running is idempotent — unchanged objects are skipped entir
```bash
vp sync
vp sync --source camera-A # just one source
vp sync --jobs 32 # concurrent transfers per source
```

Remote sources default to **16+ concurrent transfers** (object-store throughput
is latency-bound, so the CPU-derived default would undersize it); tune with
`--jobs`.

All three source formats work remotely: **YOLO** (images + label dir),
**COCO** (`labels:` points at the instances JSON), and **ImageFolder**
(`root:` is the directory of class subfolders).

## A content-addressed target

Set a `target:` and `copy` mode lands objects in a self-sufficient,
Expand Down Expand Up @@ -111,7 +132,7 @@ Pick how each source materializes its bytes with `copy:`.

| Mode | What it does | Use when |
|------|--------------|----------|
| `copy` | Server-side copy into the `target:` content-addressed store. Target is self-sufficient; global dedup. | The common cloud case. |
| `copy` | Copy into the `target:` content-addressed store — server-side when source and target share a provider, single-pass relay when they don't. Target is self-sufficient; global dedup. | The common cloud case. |
| `reference` | No copy — the index points straight at the source object. | You control the source bucket and want zero extra storage. |
| `ingest` | Download into the **local** CAS (`.vp/objects/`). | Offline / edge work on a remote dataset. |

Expand Down
12 changes: 10 additions & 2 deletions docs/index.md
Original file line number Diff line number Diff line change
Expand Up @@ -60,25 +60,33 @@ packs, and exports.
|------|--------|----------|
| **Classification** | ImageFolder (folder-per-class) | whole-image label |
| **Detection** | YOLO, COCO | bounding box |
| **Instance segmentation** | COCO | polygon |
| **Instance segmentation** | YOLO-seg, COCO | polygon |
| **Semantic segmentation (export)** | — | class-index mask PNGs (`--format masks`) |
| **Keypoints / pose** | COCO | keypoints |

## What's in the box

- **Deterministic, lockable splits** — `stratified` / `random` / `hash`, captured in snapshots.
- **Near-duplicate & cross-split leakage detection** — perceptual-hash tier, scale-proof via LSH bucketing.
- **Multi-source sync** — declarative `sources:` + `vp sync`, with per-asset provenance.
- **Cloud-native** — sync from and to S3/GCS/Azure without downloading the whole dataset; see {% link cloud-sync.md %}.
- **Cloud-native, multi-provider** — sync YOLO/COCO/ImageFolder sources from and to
S3/GCS/Azure without downloading the whole dataset, including cross-provider
targets; see [Cloud Sync]({% link cloud-sync.md %}).
- **Content-addressed snapshots & diff** — reproducible versions; compare any two.
- **Strong validation** — unreadable images, missing/orphan labels, unknown classes, bad boxes, exact + near duplicates, split leakage.
- **Benchmarking & model-in-the-loop** — `vp eval` (mAP / accuracy on a locked split),
`vp autolabel` (confident predictions become labels), `vp queue` (what to label next).
- **Packing & export** — `archive` (`.tar.zst`) and `training` (WebDataset shards); byte-free exports via hardlinks / streaming manifests.
- **Machine-readable output** — every pipeline command takes `--json` and prints a
stable, schema-versioned envelope; see [JSON Output]({% link json-output.md %}).

## Next steps

- [Install VisionPack]({% link installation.md %}) — `pip install visionpack`
- [Quickstart]({% link quickstart.md %}) — a dataset in 60 seconds
- [CLI Guide]({% link usage.md %}) — every command and its options
- [Cloud Sync]({% link cloud-sync.md %}) — S3 / GCS / Azure datasets
- [JSON Output]({% link json-output.md %}) — drive VisionPack from other programs

{: .note }
VisionPack is in **active development** (early but usable). The end-to-end
Expand Down
5 changes: 3 additions & 2 deletions docs/installation.md
Original file line number Diff line number Diff line change
Expand Up @@ -43,8 +43,9 @@ pip install "visionpack[gcs]" # Google Cloud Storage (gcsfs)
pip install "visionpack[azure]" # Azure Blob (adlfs)
```

You can combine them: `pip install "visionpack[s3,gcs]"`. See {% link cloud-sync.md %}
for declaring remote sources and a cloud target.
You can combine them: `pip install "visionpack[s3,gcs]"`. See
[Cloud Sync]({% link cloud-sync.md %}) for declaring remote sources and a cloud
target.

## Develop from source

Expand Down
105 changes: 105 additions & 0 deletions docs/json-output.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,105 @@
---
title: JSON Output
nav_order: 6
---

# JSON Output — drive VisionPack from other programs
{: .no_toc }

Every pipeline command accepts `--json` and prints **exactly one
machine-readable JSON document to stdout** — no progress bars, no prose. This
is the stable contract for wrapping VisionPack in another program: a backend
service, a UI, CI, a notebook.

1. TOC
{:toc}

## The envelope

Success:

```json
{
"schema": 1,
"command": "sync",
"data": { "...": "command-specific payload" }
}
```

Failure (the process also exits non-zero):

```json
{
"schema": 1,
"command": "diff",
"error": { "type": "VisionPackError", "message": "No snapshot named 'v99'." }
}
```

Rules a consumer can rely on:

- `schema` bumps **only on a breaking change** to the envelope or an existing
`data` shape. New fields may appear without a bump — parse leniently.
- Success has `data`; failure has `error` (never both). Check `error` +
the exit code, not stderr.
- Exit codes keep their CLI meaning: `0` success, `1` domain failure (e.g.
validation errors, ingest failures), `2` command error.

## Commands and payloads

| Command | `data` highlights |
|---|---|
| `vp sync --json` | `summaries[]` (per source: `assets_added`, `assets_existing`, `annotations`, `objects`, `failures[]`), `total_assets_added`, `total_failures` |
| `vp sync --dry-run --json` | `plans[]` (per source: `images_found`, `labels_found`, `matched`, `class_names[]`) |
| `vp import ... --json` | `assets`, `annotations`, `objects`, `classes_added`, `recorded_source`, `failures[]` |
| `vp validate --json` | `ok`, `errors`, `warnings`, `issues[]` (severity, code, message, asset_id, path) |
| `vp stats --json` | `stats` (counts, `class_distribution`, `resolutions`), `splits` (per-split breakdowns) |
| `vp split create/lock/list/show --json` | `id`, `strategy`, `locked`, `sets` (name → count); `show` adds `asset_ids` |
| `vp snapshot create/list/show --json` | snapshot records (`version`, `message`, `created_at`, `stats`) |
| `vp diff v1 v2 --json` | `assets_added/removed`, `annotations_added/removed/modified`, `classes_added/removed`, `splits_changed` |
| `vp export --json` | `format`, `output`, per-format counts (`images`, `objects`, `sets`, `streamed`) |
| `vp pack --json` | `profile`, `format`, shard/archive counts and paths |
| `vp fsck --json` | `ok`, `mode`, `checked_assets`, `checked_objects`, `issues[]` |
| `vp eval ... --json` | full result: `task`, `scope`, `metrics` (mAP@50, mAP@50-95, accuracy…), `per_class` |
| `vp autolabel ... --json` | `labeled`, `objects`, `skipped_existing`, `skipped_low_confidence`, `unmatched`, `unknown_classes[]` |
| `vp queue --json` | `total`, `items[]` (`asset_id`, `score`, `reasons[]`) |

## Example: a pipeline from a script

```bash
set -e
vp sync --json > sync.json
vp validate --json > validate.json || echo "validation found errors"
vp split create --json > split.json
vp snapshot create -m "auto $(date -I)" --json > snapshot.json

jq '.data.total_assets_added' sync.json
jq '.data.errors' validate.json
jq '.data.version' snapshot.json
```

Or from Python, without parsing text:

```python
import json, subprocess

def vp(*argv: str) -> dict:
proc = subprocess.run(["vp", *argv, "--json"], capture_output=True, text=True)
envelope = json.loads(proc.stdout)
if "error" in envelope:
raise RuntimeError(envelope["error"]["message"])
return envelope["data"]

added = vp("sync")["total_assets_added"]
report = vp("validate")
```

{: .note }
Prefer the JSON contract over importing `visionpack` internals when driving the
tool from another service: the CLI + envelope is the supported integration
surface, and the `schema` field is the compatibility signal.

## See also

- [CLI Guide]({% link usage.md %}) — the same commands, human-readable.
- [Cloud Sync]({% link cloud-sync.md %}) — remote sources and targets.
6 changes: 6 additions & 0 deletions docs/usage.md
Original file line number Diff line number Diff line change
Expand Up @@ -35,6 +35,11 @@ uv run python -m visionpack --help # or run the module directly
The examples below use the bare `vp` command; prefix with `uv run` when working
from a source checkout.

{: .tip }
Every pipeline command also takes `--json` and prints one machine-readable,
schema-versioned document to stdout — the supported way to drive VisionPack
from another program. See [JSON Output]({% link json-output.md %}).

## Initialize A Dataset

Create a VisionPack project in the current directory:
Expand Down Expand Up @@ -119,6 +124,7 @@ sources:
vp sync --dry-run # preview found / matched / unmatched / classes per source
vp sync # ingest; idempotent, records per-asset provenance
vp sync --source camera-A # sync just one source
vp sync --jobs 32 # concurrent transfers per source (remote defaults to 16+)
```

Sources can also live in object stores. Remote URIs go anywhere a local path
Expand Down
Loading
Loading