Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
26 changes: 18 additions & 8 deletions ARCHITECTURE.md
Original file line number Diff line number Diff line change
Expand Up @@ -9,7 +9,9 @@ original product vision see [docs/DESIGN.md](docs/DESIGN.md).
## Design principles

1. **CLI-first.** The primary interface is the `vp` command, scriptable in
notebooks, training servers, and CI. No web app in the core.
notebooks, training servers, and CI. No web app in the core. For Python
callers the same workflow is exposed as a facade (`visionpack.sdk`) with the
CLI's locking and result shapes.
2. **The manifest is the source of truth.** `visionpack.yaml` declares the dataset
(classes, sources, splits, validation policy, pack profiles). Behaviour is
driven by the manifest + the internal index, never by "a folder with the right
Expand All @@ -32,8 +34,10 @@ visionpack/
cli/
main.py # argparse wiring; registers every subcommand
commands/ # one module per command (init, import, sync, validate,
# fsck, stats, split, snapshot, diff, export, pack,
# annotate, eval, autolabel, queue)
# audit, fsck, stats, split, snapshot, diff, export,
# pack, annotate, eval, autolabel, queue)
sdk/ # VisionPackClient: the Python facade over the whole
# workflow (same locking + result shapes as the CLI)
core/
project.py # Project: manifest + index + object store handle
manifest.py # Manifest dataclass + pydantic schema (sources, classes…)
Expand Down Expand Up @@ -63,6 +67,8 @@ visionpack/
eval.py # vp eval: AP/mAP, accuracy, confusion matrix vs a split set
autolabel.py # vp autolabel: persist confident predictions as annotations
curation.py # vp queue: active-learning ranking + label-quality audit
audit.py # vp audit: label-health findings (duplicate/tiny/edge boxes,
# aspect outliers, class imbalance)
```

---
Expand Down Expand Up @@ -259,17 +265,21 @@ The roadmap is sequenced so each phase unblocks the next.
- [x] semantic segmentation (per-class mask PNGs via `vp export --format masks`)
- [x] YOLO-seg import-export (polygon label lines; `--seg` / segmentation-task default)
- [ ] YOLO-pose import-export; dedicated keypoint importer
- [ ] `--format auto` task/format detection on import (predictions already auto-detect)
- [x] `--format auto` format detection on import (now the default; structural
detection of YOLO/COCO/ImageFolder, explicit `--format` when ambiguous)

### Phase B — Differentiators
- [x] near-duplicate & cross-split leakage detection (perceptual-hash tier)
- [ ] optional embedding tier (CLIP/DINOv2) for semantic near-duplicates
- [ ] label-health audit (`vp audit`): duplicate/degenerate/edge-pinned boxes,
aspect-ratio outliers, class imbalance
- [x] label-health audit (`vp audit`): duplicate/degenerate/edge-pinned boxes,
aspect-ratio outliers, class imbalance (advisory by default;
`--fail-on-findings` for CI)
- [x] model-in-the-loop quality (`vp queue --include-labeled`: confident
detections with no matching label, and labels the model never finds)
- [ ] distribution-drift diff between snapshots (per-class deltas / KL)
- [ ] dataset → model lineage (`vp snapshot tag v4 trained:<run-id>`)
- [x] distribution-drift diff between snapshots (`vp diff --drift`: per-class
share deltas, smoothed KL + Jensen–Shannon divergence)
- [x] dataset → model lineage (`vp snapshot tag v4 trained:<run-id>`, free-form
tags; `snapshots_by_tag` lookup in the SDK)

### Benchmarking
- [x] `vp eval` — score predictions against a split set (per-class AP@50,
Expand Down
34 changes: 34 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -4,6 +4,40 @@ All notable changes to VisionPack are tracked here.

## [Unreleased]

- Distribution drift between snapshots (`vp diff v1 v2 --drift`, SDK
`ds.drift("v1", "v2")`): per-class object counts and distribution-share
deltas (biggest movers first) plus smoothed KL and Jensen–Shannon divergence,
computed from the stats frozen in each snapshot — reproducible forever.
- Dataset → model lineage (`vp snapshot tag v4 trained:run-812`): free-form
tags on snapshots (add/`--remove`, shown in `snapshot list`/`show`); SDK
`tag_snapshot`/`untag_snapshot`/`snapshots_by_tag` (a bare `key:` prefix
matches any value).
- `vp import --format auto` (now the default): the import format is detected
from the dataset's structure — instances-style JSON means COCO (a directory
with the JSON next to the images works too), `.txt` labels or
classes/data.yaml furniture means YOLO, folder-per-class means ImageFolder;
ambiguous layouts ask for an explicit `--format` instead of guessing. The
SDK's `import_dir` defaults to `auto` as well.
- Python SDK (`visionpack.sdk`): the whole dataset lifecycle behind one class,
`VisionPackClient` — init/open, import, sync, validate, audit, stats, splits,
snapshots (including read-only `checkout(version)` views), export, and the
model loop (evaluate/autolabel/annotation queue). Mutating methods take the
same project lock the CLI takes, and summaries come back as the same
JSON-friendly shapes the `--json` envelopes carry.
- `vp audit`: label-health audit (roadmap Phase B) — duplicate same-class boxes,
degenerate (tiny) boxes, edge-pinned and whole-image boxes, aspect-ratio
outliers, rare classes, and class imbalance. Findings are advisory by default
(`--fail-on-findings` gates CI); thresholds configurable via flags or
`validation.audit` in `visionpack.yaml`; `--json` supported.
- Security: class names arriving from imported data (folder names, COCO
categories, `classes.txt`) are sanitized before being used as export path
components, so a hostile name like `../../x` can no longer write outside the
export directory.
- Robustness: a decompression-bomb image (header claiming absurd dimensions)
now records a per-file ingest failure instead of aborting the whole
import/sync batch.
- Performance: sync/import now read only asset *ids* when checking which assets
already exist (`SELECT id`), instead of materializing every asset record.
- Model-in-the-loop foundation: a shared predictions loader
(`visionpack/predictions.py`) reads model output in three formats — vp-native
JSON, COCO results/instances JSON, and YOLO txt directories (what Ultralytics
Expand Down
34 changes: 30 additions & 4 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -83,11 +83,12 @@ uv run vp --help
# 1. create a project (the manifest is visionpack.yaml)
vp init --name factory-defects --task detection

# 2. bring in a YOLO dataset
vp import ./raw --format yolo
# 2. bring in a dataset (YOLO/COCO/ImageFolder auto-detected)
vp import ./raw

# 3. check it for real problems
vp validate
vp validate # invalid labels, duplicates, leakage
vp audit # suspicious labels: double boxes, tiny boxes, imbalance

# 4. a deterministic, reproducible split
vp split create --train 0.8 --val 0.1 --test 0.1 --strategy stratified
Expand All @@ -103,6 +104,20 @@ vp stats --by split
vp export --format yolo --split
```

Prefer Python? The same workflow is one class away
([Python SDK docs](https://caiowing.github.io/VisionPack/sdk/)):

```python
from visionpack.sdk import VisionPackClient

ds = VisionPackClient.open(".")
ds.validate()
ds.create_split(train=0.8, val=0.1, test=0.1, strategy="stratified")
ds.snapshot("baseline")
ds.export("./exports/yolo", format="yolo", split="default")
metrics = ds.evaluate("runs/predict/labels", format="yolo")
```

---

## Works across the common CV tasks
Expand Down Expand Up @@ -182,9 +197,20 @@ See the [Cloud Sync guide](https://caiowing.github.io/VisionPack/cloud-sync/).
single-pass verified relay across providers (S3→GCS, local→S3, …); one
fast-list instead of per-object lookups, retries with backoff on every remote
call, tunable concurrency (`--jobs`); streaming export.
- **Content-addressed snapshots & diff** — reproducible versions; compare any two.
- **Content-addressed snapshots & diff** — reproducible versions; compare any
two, including class-distribution drift (`vp diff v1 v2 --drift`: per-class
share deltas, KL/JS divergence).
- **Dataset → model lineage** — `vp snapshot tag v4 trained:<run-id>` links a
training run to the exact dataset version it consumed.
- **Strong validation** — unreadable images, missing/orphan labels, unknown
classes, invalid/out-of-bounds boxes, exact + near duplicates, split leakage.
- **Label-health audit (`vp audit`)** — the labels that are valid but usually
wrong: the same object boxed twice, tiny/degenerate boxes, boxes pinned to
image borders or covering the whole image, extreme aspect ratios, rare
classes and class imbalance. Advisory by default, `--fail-on-findings` for CI.
- **Python SDK (`visionpack.sdk`)** — the whole workflow programmatically:
`VisionPackClient` wraps import/sync/validate/audit/split/snapshot/export and
the model loop with the same locking and result shapes as the CLI.
- **Comparable metrics** — per-split stats so class balance stays auditable as data
grows.
- **Benchmarking (`vp eval`)** — score model predictions (vp/COCO JSON or YOLO
Expand Down
6 changes: 4 additions & 2 deletions docs/json-output.md
Original file line number Diff line number Diff line change
Expand Up @@ -53,10 +53,12 @@ Rules a consumer can rely on:
| `vp sync --dry-run --json` | `plans[]` (per source: `images_found`, `labels_found`, `matched`, `class_names[]`) |
| `vp import ... --json` | `assets`, `annotations`, `objects`, `classes_added`, `recorded_source`, `failures[]` |
| `vp validate --json` | `ok`, `errors`, `warnings`, `issues[]` (severity, code, message, asset_id, path) |
| `vp audit --json` | `ok`, `findings`, `by_code` (code → count), `class_counts`, `items[]` (code, message, asset_id, path, class_id) |
| `vp stats --json` | `stats` (counts, `class_distribution`, `resolutions`), `splits` (per-split breakdowns) |
| `vp split create/lock/list/show --json` | `id`, `strategy`, `locked`, `sets` (name → count); `show` adds `asset_ids` |
| `vp snapshot create/list/show --json` | snapshot records (`version`, `message`, `created_at`, `stats`) |
| `vp diff v1 v2 --json` | `assets_added/removed`, `annotations_added/removed/modified`, `classes_added/removed`, `splits_changed` |
| `vp snapshot create/list/show --json` | snapshot records (`version`, `message`, `created_at`, `stats`, `tags[]`) |
| `vp snapshot tag ... --json` | `version`, `tag`, `removed`, `tags[]` (the snapshot's tags after the change) |
| `vp diff v1 v2 --json` | `assets_added/removed`, `annotations_added/removed/modified`, `classes_added/removed`, `splits_changed`; with `--drift`: `drift` (`classes[]` share deltas, `kl_divergence`, `js_divergence`) |
| `vp export --json` | `format`, `output`, per-format counts (`images`, `objects`, `sets`, `streamed`) |
| `vp pack --json` | `profile`, `format`, shard/archive counts and paths |
| `vp fsck --json` | `ok`, `mode`, `checked_assets`, `checked_objects`, `issues[]` |
Expand Down
90 changes: 90 additions & 0 deletions docs/sdk.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,90 @@
---
title: Python SDK
nav_order: 7
---

# Python SDK

Everything the CLI does is available from Python through the SDK — the same
project on disk, the same locking, the same results the `--json` envelopes
carry, without subprocess plumbing. Use it from notebooks, training scripts,
labeling services, or CI jobs:

```python
from visionpack.sdk import VisionPackClient

ds = VisionPackClient.init("./factory-defects", task="detection") # or .open(".")
ds.import_dir("./raw", format="yolo")

report = ds.validate() # ValidationReport (errors/warnings/ok)
audit = ds.audit() # AuditReport (label-health findings)
print(ds.stats()["class_distribution"])

ds.create_split(train=0.8, val=0.1, test=0.1, strategy="stratified")
ds.lock_split()
ds.snapshot("baseline")
ds.export("./exports/yolo", format="yolo", split="default")
```

Close the model-in-the-loop cycle with the same handle:

```python
metrics = ds.evaluate("runs/predict/labels", format="yolo") # mAP on the test set
ds.autolabel("preds.json", min_confidence=0.6) # confident preds -> labels
for item in ds.annotation_queue("preds.json")[:20]: # what to label next
print(item["score"], item["path"])
```

## The client

| Area | Methods |
|------|---------|
| Lifecycle | `VisionPackClient.init(root, name=..., task=...)`, `VisionPackClient.open(root)` (also `sdk.init` / `sdk.open`) |
| Ingest | `import_dir(source, format="yolo"\|"coco"\|"imagefolder", images=..., copy_mode=...)`, `sync(source=..., jobs=...)`, `plan_sync()` |
| Quality | `validate(strict=...)`, `audit(**thresholds)`, `stats()`, `split_stats()` |
| Splits | `create_split(...)`, `lock_split()`, `split()` |
| Versions | `snapshot(message)`, `snapshots()`, `get_snapshot(version)`, `checkout(version)`, `diff(v1, v2)`, `drift(v1, v2)`, `tag_snapshot(v, tag)`, `untag_snapshot(v, tag)`, `snapshots_by_tag(tag)` |
| Outputs | `export(output, format="yolo"\|"coco"\|"imagefolder"\|"masks", split=..., seg=...)` |
| Model loop | `load_predictions(...)`, `evaluate(...)`, `autolabel(...)`, `annotation_queue(...)` |
| Data access | `assets()`, `annotations()`, `samples()` (streaming iterator), `len(ds)`, `for asset, ann in ds:` |

## Guarantees

- **Safe next to the CLI.** Every mutating method takes the same project lock
`vp` takes, so an SDK caller and a CLI process can never corrupt each
other's writes — the second writer fails fast with a clear error.
- **Stable, JSON-friendly returns.** Summaries come back as plain dicts that
mirror the [`--json` contract]({% link json-output.md %}), so a service can
switch between shelling out to `vp` and importing the SDK without
re-parsing anything.
- **Read-only snapshot views.** `ds.checkout("v2")` returns a client pinned to
that snapshot: exports, stats, and evaluation reflect the frozen state, and
mutating methods raise instead of silently writing into live history.

## Versions, drift, and lineage

```python
ds.snapshot("after week-30 batch") # v5
print(ds.drift("v4", "v5")["js_divergence"]) # did the class mix shift?

# after training, link the run to the exact dataset version it consumed:
ds.tag_snapshot("v5", "trained:run-812")
ds.snapshots_by_tag("trained:") # every version a model trained on
```

## Streaming a dataset

`samples()` iterates straight off the index without materializing it, so a
training-adjacent script can walk 100k+ assets in bounded memory:

```python
for asset, annotation in ds.samples():
if annotation is None:
continue
for obj in annotation.objects:
box = obj.bbox # enclosing BBox for any geometry
...
```

The lower-level `Project` object stays reachable as `ds.project` for anything
the facade doesn't cover yet.
Loading
Loading