Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
38 changes: 27 additions & 11 deletions ARCHITECTURE.md
Original file line number Diff line number Diff line change
Expand Up @@ -32,7 +32,8 @@ visionpack/
cli/
main.py # argparse wiring; registers every subcommand
commands/ # one module per command (init, import, sync, validate,
# fsck, stats, split, snapshot, diff, export, pack, annotate)
# fsck, stats, split, snapshot, diff, export, pack,
# annotate, eval, autolabel, queue)
core/
project.py # Project: manifest + index + object store handle
manifest.py # Manifest dataclass + pydantic schema (sources, classes…)
Expand All @@ -45,7 +46,8 @@ visionpack/
index/
json_index.py # local JSON index with cached, O(1) annotation lookups
formats/
yolo.py coco.py classification.py # import + export per format
yolo.py coco.py classification.py # import + export per format (yolo covers YOLO-seg)
masks.py # semantic-mask export (class-index PNGs)
sources/
resolver.py schema.py join.py importer.py # declarative multi-source sync
validation/
Expand All @@ -57,6 +59,10 @@ visionpack/
split.py # deterministic split creation / resolution
stats.py # dataset + per-split statistics
snapshot.py # content-addressed snapshots
predictions.py # model-prediction loading (vp/COCO/YOLO) resolved to assets
eval.py # vp eval: AP/mAP, accuracy, confusion matrix vs a split set
autolabel.py # vp autolabel: persist confident predictions as annotations
curation.py # vp queue: active-learning ranking + label-quality audit
```

---
Expand Down Expand Up @@ -185,7 +191,7 @@ with a self-describing `dataset.json`.

## Testing

`unittest` suite under `tests/` (55 tests). Run with:
`unittest` suite under `tests/`. Run with:

```bash
uv run python -m unittest discover -s tests -q
Expand Down Expand Up @@ -227,22 +233,27 @@ The roadmap is sequenced so each phase unblocks the next.
- [x] tagged geometry model (bbox | polygon | keypoints | none), backward compatible
- [x] classification: ImageFolder import/export
- [x] instance segmentation (polygons) and keypoints via COCO
- [ ] semantic segmentation (per-class mask PNGs)
- [ ] YOLO-seg / YOLO-pose import-export; dedicated keypoint importer
- [ ] `--format auto` task/format detection
- [x] semantic segmentation (per-class mask PNGs via `vp export --format masks`)
- [x] YOLO-seg import-export (polygon label lines; `--seg` / segmentation-task default)
- [ ] YOLO-pose import-export; dedicated keypoint importer
- [ ] `--format auto` task/format detection on import (predictions already auto-detect)

### Phase B — Differentiators
- [x] near-duplicate & cross-split leakage detection (perceptual-hash tier)
- [ ] optional embedding tier (CLIP/DINOv2) for semantic near-duplicates
- [ ] label-health audit (`vp audit`): duplicate/degenerate/edge-pinned boxes,
aspect-ratio outliers, class imbalance
- [ ] model-in-the-loop quality (confident detections with no matching label)
- [x] model-in-the-loop quality (`vp queue --include-labeled`: confident
detections with no matching label, and labels the model never finds)
- [ ] distribution-drift diff between snapshots (per-class deltas / KL)
- [ ] dataset → model lineage (`vp snapshot tag v4 trained:<run-id>`)

### Benchmarking (planned)
- [ ] `vp eval` — score predictions against a locked test split (mAP, accuracy,
confusion matrix), turning a dataset into a reproducible benchmark
### Benchmarking
- [x] `vp eval` — score predictions against a split set (per-class AP@50,
mAP@50, mAP@50-95, precision/recall; accuracy + confusion matrix for
classification), turning a dataset into a reproducible benchmark.
Predictions load from vp JSON, COCO JSON, or YOLO txt (`predictions.py`)
- [ ] mask IoU for segmentation eval (bbox IoU today)
- [ ] benchmark objects (snapshot + split + protocol) and a leaderboard
- [ ] dataset/benchmark cards (`vp card`) for publishable, citable artifacts
- [ ] Hugging Face Datasets export
Expand All @@ -253,8 +264,13 @@ The roadmap is sequenced so each phase unblocks the next.
- [ ] richer terminal output with `rich`
- [ ] move CLI plumbing from `argparse` to `typer` once commands stabilize

### Model-in-the-loop ✅ (first slice)
- [x] `vp autolabel` — persist confident predictions as annotations
(`source.type = "model"`, `--min-confidence`, `--replace` opt-in)
- [x] active-learning queue (`vp queue`) — rank unlabeled images by model
uncertainty; audit labeled ones for GT/prediction disagreement

### Later
- [ ] `vp annotate prepare` / `ingest`; CVAT and Label Studio packages
- [ ] active-learning queue (rank unlabeled images by model uncertainty)
- [ ] remote storage integrations (S3/GCS/Azure) for assets
- [ ] optional PyTorch dataset helpers
28 changes: 27 additions & 1 deletion CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -4,7 +4,33 @@ All notable changes to VisionPack are tracked here.

## [Unreleased]

- Nothing yet.
- Model-in-the-loop foundation: a shared predictions loader
(`visionpack/predictions.py`) reads model output in three formats — vp-native
JSON, COCO results/instances JSON, and YOLO txt directories (what Ultralytics
`predict` writes with `save_txt`/`save_conf`) — and resolves images to assets
by asset id or original filename.
- `vp eval`: score predictions against a split's labels, turning a locked split
+ snapshot into a reproducible benchmark. Detection/segmentation/keypoints get
COCO-style AP (per-class AP@50, mAP@50, mAP@50-95, precision/recall at a
confidence threshold); classification gets accuracy, per-class P/R/F1 and a
confusion matrix. `--json` for machine-readable output.
- `vp autolabel`: persist confident predictions as annotations with
`source.type = "model"` (auditable, distinguishable from human labels). Only
unlabeled assets are touched unless `--replace`; `--min-confidence` filters
objects.
- `vp queue`: active-learning queue that ranks images by annotation value —
unlabeled images first (by model uncertainty when predictions are given), and
with `--include-labeled` audits labeled images for ground-truth/prediction
disagreement (possible missing or stale labels).
- YOLO-seg: YOLO imports now accept polygon label lines
(`class x1 y1 x2 y2 ...`) as instance-segmentation geometry, and `vp export
--format yolo` writes YOLO-seg labels for segmentation projects (or with
`--seg`); plain boxes degrade to four-corner polygons.
- Semantic masks: `vp export --format masks` rasterizes polygon (and box)
annotations into 8-bit class-index PNGs (0 = background), split-aware, with a
`classes.txt` documenting the pixel-value mapping.
- `vp --version` now reports the installed package version (was hardcoded and
out of sync with `pyproject.toml`); removed duplicated README sections.
- Cloud sync foundation (PR 1 of docs/SPEC-cloud-sync.md): `FsspecResolver` for
remote schemes (`s3://`, `gs://`, `az://`) behind optional extras
(`visionpack[s3]`, `[gcs]`, `[azure]`); `Resolver.stat` for metadata-only
Expand Down
79 changes: 23 additions & 56 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -9,7 +9,7 @@ leak-free, ready-to-train dataset.
![Python](https://img.shields.io/badge/python-3.11%2B-blue)
![License](https://img.shields.io/badge/license-Apache--2.0-blue)
![Status](https://img.shields.io/badge/status-active%20development-orange)
![Tests](https://img.shields.io/badge/tests-100%20passing-brightgreen)
![Tests](https://img.shields.io/badge/tests-118%20passing-brightgreen)

[Documentation](https://caiowing.github.io/VisionPack/) ·
[Install](https://caiowing.github.io/VisionPack/installation/) ·
Expand All @@ -25,6 +25,11 @@ vp validate # catch corrupt images, bad boxes, near-duplicate leakag
vp split create # deterministic, reproducible train/val/test
vp snapshot create -m "baseline"
vp export --format yolo --split # ready-to-train layout

# ...train / predict with any framework, then close the loop:
vp eval runs/predict/labels --format yolo # mAP on the locked test set
vp autolabel preds.json --min-confidence 0.6 # confident predictions -> labels
vp queue --predictions preds.json # what should a human label next?
```

---
Expand Down Expand Up @@ -109,7 +114,8 @@ tasks you actually use:
|------|--------|----------|
| **Classification** | ImageFolder (folder-per-class) | whole-image label |
| **Detection** | YOLO, COCO | bounding box |
| **Instance segmentation** | COCO | polygon |
| **Instance segmentation** | YOLO-seg, COCO | polygon |
| **Semantic segmentation (export)** | — | class-index mask PNGs (`--format masks`) |
| **Keypoints / pose** | COCO | keypoints |

```bash
Expand Down Expand Up @@ -177,9 +183,18 @@ See the [Cloud Sync guide](https://caiowing.github.io/VisionPack/cloud-sync/).
classes, invalid/out-of-bounds boxes, exact + near duplicates, split leakage.
- **Comparable metrics** — per-split stats so class balance stays auditable as data
grows.
- **Benchmarking (`vp eval`)** — score model predictions (vp/COCO JSON or YOLO
txt) against a split's labels: per-class AP@50, mAP@50, mAP@50-95,
precision/recall — or accuracy + confusion matrix for classification. A locked
split + a snapshot = a reproducible benchmark.
- **Model-in-the-loop** — `vp autolabel` persists confident predictions as
annotations (recorded as `source.type = "model"`, never silently overwriting
human labels); `vp queue` ranks what a human should label next and audits
existing labels against the model (missing/stale-label signals).
- **Packing & byte-free export** — `archive` (`.tar.zst`) and `training`
(WebDataset shards); exports hardlink from the CAS or stream from the cloud.
- **Interoperable I/O** — YOLO, COCO, ImageFolder in and out.
- **Interoperable I/O** — YOLO (incl. YOLO-seg), COCO, ImageFolder in and out;
semantic masks out.

Full command reference and per-command options live in the
[CLI guide](https://caiowing.github.io/VisionPack/usage/).
Expand Down Expand Up @@ -211,56 +226,6 @@ release commit or tag:

---

## Release process

Releases are prepared locally, reviewed as a GitHub Release draft, and published
to PyPI only when the GitHub Release is published.

```powershell
.\scripts\prepare-release.ps1 0.1.1
git push origin HEAD
git push origin v0.1.1
gh release create v0.1.1 --draft --title "v0.1.1" --notes-file CHANGELOG.md
```

Review the draft release notes in GitHub. Publishing the release triggers
`.github/workflows/publish.yml`, which builds the package, validates the
artifacts with `twine check`, and publishes to PyPI through Trusted Publishing.

Use `-NoCommit -NoTag` to update files and run the checks without creating the
release commit or tag:

```powershell
.\scripts\prepare-release.ps1 0.1.1 -NoCommit -NoTag
```

---

## Release process

Releases are prepared locally, reviewed as a GitHub Release draft, and published
to PyPI only when the GitHub Release is published.

```powershell
.\scripts\prepare-release.ps1 0.1.1
git push origin HEAD
git push origin v0.1.1
gh release create v0.1.1 --draft --title "v0.1.1" --notes-file CHANGELOG.md
```

Review the draft release notes in GitHub. Publishing the release triggers
`.github/workflows/publish.yml`, which builds the package, validates the
artifacts with `twine check`, and publishes to PyPI through Trusted Publishing.

Use `-NoCommit -NoTag` to update files and run the checks without creating the
release commit or tag:

```powershell
.\scripts\prepare-release.ps1 0.1.1 -NoCommit -NoTag
```

---

## How it works

VisionPack is **manifest-driven** and **content-addressed**: `visionpack.yaml`
Expand All @@ -278,9 +243,11 @@ vision, see **[docs/DESIGN.md](docs/DESIGN.md)**.

VisionPack is in **active development** (early but usable). The core workflow —
multi-source ingestion (local and cloud) → validation → deterministic splits →
snapshots → ready-to-train export/packing — works end-to-end across
classification, detection, instance segmentation, and keypoints, with 100 passing
tests. APIs may still shift; feedback and contributions are welcome.
snapshots → ready-to-train export/packing → evaluation (`vp eval`) and
model-in-the-loop labeling (`vp autolabel` / `vp queue`) — works end-to-end
across classification, detection, instance/semantic segmentation, and keypoints,
with 118 passing tests. APIs may still shift; feedback and contributions are
welcome.

```bash
uv run python -m unittest discover -s tests -q
Expand Down
62 changes: 57 additions & 5 deletions docs/usage.md
Original file line number Diff line number Diff line change
Expand Up @@ -241,6 +241,58 @@ content-addressed store, and **cloud-backed** images are written to a
`manifest.jsonl` (image → object URI) for streaming instead of being downloaded.
See [Cloud Sync]({% link cloud-sync.md %}#export-for-training-streaming).

Segmentation projects export YOLO-seg polygon labels automatically (force either
way with `--seg` / `--no-seg`), and `--format masks` writes semantic masks —
8-bit PNGs whose pixel value is the class index (0 = background, mapping recorded
in `classes.txt`):

```bash
vp export --format yolo --output exports/seg-v1 --split # YOLO-seg labels
vp export --format masks --output exports/masks-v1 --split # class-index PNGs
```

## Evaluate A Model (vp eval)

Score model predictions against the labels of a split's set (the test set by
default) — a locked split plus a snapshot makes the number a reproducible
benchmark:

```bash
vp eval runs/predict/labels --format yolo # Ultralytics save_txt output
vp eval predictions.json # vp-native or COCO JSON
vp eval predictions.json --split default --set val --json
```

Detection, segmentation, and keypoints report per-class AP@50, mAP@50,
mAP@50-95, and precision/recall at a confidence threshold (`--conf`);
classification reports accuracy, per-class precision/recall/F1, and a confusion
matrix. Predictions reference images by asset id (what exports name files) or by
original filename; unresolvable references are reported, never dropped silently.

## Autolabel & The Annotation Queue

`vp autolabel` persists confident predictions as annotations. Model labels are
recorded with `source.type = "model"`, so they stay distinguishable from human
labels. Only unlabeled assets are touched unless you pass `--replace`:

```bash
vp autolabel predictions.json --min-confidence 0.6
```

`vp queue` ranks images by how much a human label would help (active learning):
unlabeled images first — ordered by model uncertainty when predictions are given
— and with `--include-labeled` it audits existing labels for ground-truth /
prediction disagreement (possible missing or stale labels):

```bash
vp queue --predictions predictions.json --limit 20
vp queue --predictions predictions.json --include-labeled --json
```

Together these close the model-in-the-loop cycle: export → train/predict →
`vp eval` (measure) → `vp autolabel` (label the easy images) → `vp queue`
(send the hard ones to humans) → re-train.

## Pack Archive

Create a compressed archive package:
Expand Down Expand Up @@ -279,14 +331,14 @@ More stable SDK methods will be added as the internal workflows settle.

## Current Limitations

- semantic segmentation (per-class mask PNGs) is not yet supported; instance
segmentation (polygons) and keypoints are supported via COCO
- YOLO import/export is detection-only (no YOLO-seg / YOLO-pose yet)
- `--format auto` detection is not implemented; pass `--format` explicitly
- segmentation metrics in `vp eval` use each polygon's enclosing box (mask IoU
is planned); YOLO-pose import/export and a dedicated keypoint importer are
not implemented yet
- `--format auto` detection on import is not implemented; pass `--format`
explicitly (predictions for `vp eval`/`autolabel`/`queue` *are* auto-detected)
- `vp annotate` is scaffolded but not implemented yet
- cloud sync (S3/GCS/Azure) is **same-provider** in v1 — cross-cloud transfer
(S3↔GCS) and remote COCO/ImageFolder sync are planned; `pack` is local-only
- `vp eval` (scoring predictions into benchmark metrics) is planned
- the local index is SQLite (`index.db`); `stats` and exports stream records (flat
RAM at scale), while `validate`, `split`, dedup, and the WebDataset pack still load
the full set (next streaming pass — see
Expand Down
Loading
Loading