Base URL (default): http://localhost:8000
All request and response bodies are JSON (Content-Type: application/json).
Interactive, always-up-to-date docs are served by the running server at /docs
(Swagger UI) and /redoc (ReDoc); the OpenAPI schema is at /openapi.json.
- Conventions
- Authentication
- Errors
- Filter syntax
- Data types
- Index types & metrics
- Health
- Collections
- Documents
- Search
- Document ids are always strings. If you omit
idon insert/upsert, the server generates one (a UUID hex) and returns it in the write results. - A document has named vector fields (
vectors) and named scalar fields (fields). The vector field names must match the collection schema. - The server stores client-supplied vectors only. It does not generate embeddings.
Authentication is off by default. When the server is started with
ZVEC_SERVER_AUTH_ENABLED=true and a ZVEC_SERVER_API_KEY, every endpoint
except the health probes (/healthz, /readyz) requires the key as a bearer
token:
Authorization: Bearer <api_key>curl http://localhost:8000/collections \
-H "Authorization: Bearer $ZVEC_SERVER_API_KEY"A missing or invalid key returns 401 Unauthorized with the standard error
envelope (code: "unauthorized") and a WWW-Authenticate: Bearer response
header. When auth is enabled, the interactive docs (/docs, /openapi.json)
also require the key. See
CONFIGURATION.md.
Errors use a consistent envelope:
{
"error": {
"code": "collection_not_found",
"message": "Collection 'articles' not found.",
"details": { "name": "articles" }
}
}code is a stable, machine-readable string (snake_case); details is optional.
Status codes and their code values:
| Status | code |
When |
|---|---|---|
| 400 | invalid_argument |
Argument rejected by the server or engine (e.g. a malformed filter). |
| 401 | unauthorized |
Auth enabled and the API key was missing or invalid. |
| 404 | collection_not_found |
No such collection. |
| 404 | document_not_found |
No such document (single-document fetch). |
| 409 | collection_already_exists |
A collection with that name already exists. |
| 422 | schema_validation_error |
Invalid schema (bad dtype/index/metric). |
| 422 | validation_error |
Request body failed validation (ids+filter both/neither, etc.). |
| 500 | zvec_operation_error / internal_error |
Unexpected server / Zvec engine error. |
| 503 | collection_unavailable |
Collection registered but not currently open on disk. |
The filter string used by search and delete-by-filter uses Zvec's
SQL-like expression syntax. It is passed through to Zvec verbatim.
- Use a single
=for equality — not Python's==. - Quote string literals with single quotes:
'tech'. - Operators:
=,<,>,<=,>=,AND,OR,NOT,IN,BETWEEN,LIKE.
Examples:
category = 'tech'
category = 'tech' AND year > 2020
year BETWEEN 2018 AND 2022
category IN ('tech', 'science')
NOT (category = 'sports')
title LIKE 'intro%'
A malformed filter results in a 400 (invalid_argument).
Provide dtypes by name (case-sensitive, as Zvec defines them).
Vector dtypes (for vectors[].dtype):
| Name | Meaning |
|---|---|
VECTOR_FP32 |
32-bit float dense vector (default). |
VECTOR_FP16 |
16-bit float dense vector. |
VECTOR_FP64 |
64-bit float dense vector. |
VECTOR_INT8 |
8-bit int dense vector. |
SPARSE_VECTOR_FP16 |
16-bit float sparse vector. |
SPARSE_VECTOR_FP32 |
32-bit float sparse vector. |
Scalar dtypes (for fields[].dtype):
| Name | Meaning |
|---|---|
INT32, INT64 |
Signed integers. |
UINT32, UINT64 |
Unsigned integers. |
FLOAT, DOUBLE |
Floating point. |
STRING |
Text. |
BOOL |
Boolean. |
ARRAY_* |
Array variants of the above. |
Index types (for vectors[].index): hnsw (default), flat, ivf,
hnsw_rabitq, ivf_rabitq. The RaBitQ variants use RaBitQ quantization and
are only available when the server runs on Linux x86_64 (the published
Docker image does); elsewhere creating one returns 422. They have not been
benchmarked here yet — see the quantization note below.
Optional per-index tuning goes in vectors[].params:
| Index | Recognized params |
|---|---|
hnsw |
m, ef_construction |
ivf |
n_list, n_iters, use_soar |
flat |
(none) |
hnsw_rabitq |
m, ef_construction, total_bits, num_clusters, sample_count |
ivf_rabitq |
n_list, total_bits, sample_count |
Quantization (hnsw, flat, ivf): quantize_type — fp16, int8,
or int4 — makes the index search over quantized vectors; enable_rotate: true
applies a random rotation before quantizing. Zvec keeps the original
full-precision vectors alongside (fetch returns them unchanged), so the
quantized index is additional storage, not a replacement.
Measure before adopting. On SIFT1M (1M × 128) with Zvec 0.7.0 — server defaults, mmap on — every quantized variant used more disk and memory than FP32 and was no faster:
fp16kept recall (0.995) at +37% disk / +35% RSS;int8lost ~1 point of recall@10 at +20% disk / +23% RSS;int4lost ~28 points, and rotation madeint4worse (0.54). Results may differ for higher-dimensional embeddings. Run the quantization sweep on your own data:python -m benchmarks quant(seebenchmarks/README.md).
{ "name": "embedding", "dim": 768, "index": "hnsw",
"params": { "m": 16, "quantize_type": "int8", "enable_rotate": true } }Metrics (for vectors[].metric): cosine (default), ip (inner product),
l2 (Euclidean). Aliases such as dot / inner_product and euclidean are
accepted.
Liveness probe.
Response 200
{ "status": "ok" }Readiness probe with collection counts.
Response 200
{
"status": "ready",
"collections_loaded": 2,
"collections_unavailable": 0
}Create a collection.
Request body (CreateCollectionRequest):
| Field | Type | Required | Notes |
|---|---|---|---|
name |
string | yes | Matches ^[A-Za-z0-9_-]{3,64}$. |
vectors |
array of VectorFieldSpec |
yes | At least one. |
fields |
array of ScalarFieldSpec |
no | Defaults to []. |
options |
object | no | { "enable_mmap": bool }. |
embedding_model |
string | null | no | Free-form metadata label only; not used to compute embeddings. |
VectorFieldSpec:
| Field | Type | Default | Notes |
|---|---|---|---|
name |
string | — | Vector field name. |
dim |
int (> 0) | — | Vector dimension. |
dtype |
string | VECTOR_FP32 |
See data types. |
index |
string | hnsw |
hnsw / flat / ivf. |
metric |
string | cosine |
cosine / ip / l2. |
params |
object | null | null |
Index tuning params. |
ScalarFieldSpec:
| Field | Type | Default | Notes |
|---|---|---|---|
name |
string | — | Scalar field name. |
dtype |
string | — | See scalar dtypes. |
nullable |
bool | false |
Whether the field may be null. |
indexed |
bool | false |
Build an inverted index (faster filters). |
Example request
{
"name": "articles",
"vectors": [
{
"name": "embedding",
"dim": 4,
"dtype": "VECTOR_FP32",
"index": "hnsw",
"metric": "cosine",
"params": { "m": 16, "ef_construction": 200 }
}
],
"fields": [
{ "name": "category", "dtype": "STRING", "indexed": true },
{ "name": "year", "dtype": "INT32", "indexed": true }
],
"options": { "enable_mmap": true }
}Response 201 (CollectionInfo)
{
"name": "articles",
"path": "/data/collections/articles",
"schema_version": 1,
"embedding_dimension": 4,
"embedding_model": null,
"vectors": [
{ "name": "embedding", "data_type": "VECTOR_FP32", "dimension": 4 }
],
"fields": [
{ "name": "category", "data_type": "STRING", "nullable": false },
{ "name": "year", "data_type": "INT32", "nullable": false }
],
"options": { "enable_mmap": true },
"stats": { "doc_count": 0, "index_completeness": {} },
"available": true,
"created_at": "2026-06-23T12:00:00+00:00",
"updated_at": "2026-06-23T12:00:00+00:00"
}
409(collection_already_exists) if a collection with that name exists.422(schema_validation_error) for an invalid dtype/index/metric or bad name.
List collections.
Response 200 (CollectionListResponse)
{
"collections": [
{
"name": "articles",
"embedding_dimension": 4,
"embedding_model": null,
"doc_count": 3,
"created_at": "2026-06-23T12:00:00+00:00"
}
]
}Get full collection info, including live stats. Returns the same
CollectionInfo shape as create. 404 if not found.
Drop a collection. This deletes its data on disk.
Response 200 (MessageResponse)
{ "message": "Collection 'articles' deleted" }Flush pending writes to disk.
Response 200 (MessageResponse)
{ "message": "Collection 'articles' flushed" }Optimize the collection's indexes.
Response 200 (MessageResponse)
{ "message": "Collection 'articles' optimized" }All document endpoints are under /collections/{name}.
Insert documents. Use upsert to insert-or-replace and update to modify
existing documents; all three share the WriteRequest body and WriteResponse
shape.
Request body (WriteRequest):
{
"docs": [
{
"id": "a1",
"vectors": { "embedding": [0.10, 0.20, 0.30, 0.40] },
"fields": { "category": "tech", "year": 2021 }
},
{
"vectors": { "embedding": [0.90, 0.10, 0.05, 0.02] },
"fields": { "category": "science", "year": 2019 }
}
]
}DocIn:
| Field | Type | Default | Notes |
|---|---|---|---|
id |
string | null | null |
Auto-generated if omitted. |
vectors |
object (name → list of float) | {} |
Keys must match the schema's vector fields. |
fields |
object (name → value) | {} |
Scalar field values. |
Response 200 (WriteResponse)
{
"results": [
{ "id": "a1", "ok": true, "code": "OK", "message": "" },
{ "id": "f1c2...hex", "ok": true, "code": "OK", "message": "" }
],
"success_count": 2,
"error_count": 0
}Insert or replace documents. Same body/response as insert.
Update existing documents. Same body/response as insert.
Delete by either ids or filter — exactly one must be set
(422 otherwise).
Request body (DeleteRequest) — by ids:
{ "ids": ["a1", "a2"] }By filter (see filter syntax):
{ "filter": "category = 'science' AND year < 2020" }Response 200 (DeleteResponse) — by ids:
{
"ok": true,
"results": [
{ "id": "a1", "ok": true, "code": "OK", "message": "" },
{ "id": "a2", "ok": true, "code": "OK", "message": "" }
],
"filter": null,
"message": null
}By filter:
{
"ok": true,
"results": null,
"filter": "category = 'science' AND year < 2020",
"message": null
}A malformed filter returns
400(invalid_argument).
Fetch documents by id. Missing ids are simply omitted from the result map.
Request body (FetchRequest):
{
"ids": ["a1", "a2"],
"output_fields": ["category", "year"],
"include_vector": false
}| Field | Type | Default | Notes |
|---|---|---|---|
ids |
array of string | — | At least one. |
output_fields |
array of string | null | null |
Restrict returned scalar fields. |
include_vector |
bool | false |
Include vector values in the result. |
Response 200 (FetchResponse)
{
"docs": {
"a1": {
"id": "a1",
"score": null,
"vectors": null,
"fields": { "category": "tech", "year": 2021 }
}
}
}Fetch a single document by id. Returns 404 (document_not_found) if missing.
Query parameters
| Name | Type | Default | Notes |
|---|---|---|---|
include_vector |
bool | false |
Include vector values. |
output_fields |
repeated string | (all) | Restrict returned scalar fields. |
Example
GET /collections/articles/docs/a1?include_vector=true
Response 200 (DocOut)
{
"id": "a1",
"score": null,
"vectors": { "embedding": [0.10, 0.20, 0.30, 0.40] },
"fields": { "category": "tech", "year": 2021 }
}Stream every document as NDJSON (application/x-ndjson, one JSON object
per line) — for backups, migrations, or re-indexing. Each line is shaped like a
DocIn (id, vectors, fields; no score), so an export can be posted back
to /docs/insert in batches unchanged.
| Query param | Type | Default | Notes |
|---|---|---|---|
include_vector |
bool | true |
Include vectors (needed to re-import). |
output_fields |
string (repeatable) | all | Restrict scalar fields, e.g. ?output_fields=year. |
curl -s localhost:8000/collections/articles/export > articles.ndjson{"id":"a1","vectors":{"embedding":[0.1,0.2,0.3,0.4]},"fields":{"category":"tech","year":2021}}
{"id":"a2","vectors":{"embedding":[0.2,0.1,0.0,0.9]},"fields":{"category":"news","year":2019}}
- Consistent snapshot. The export reflects the collection as of the request; writes made while it streams are not included, and are not blocked by it (the collection lock is taken per batch, not for the whole stream).
- Errors. Bad input (unknown
output_fields, missing collection) returns a normal JSON error before streaming starts. If the export is cut short after it has started — e.g. the collection is dropped or the server shuts down — the last line is an error envelope ({"error": {...}}) instead of a document. Check for it before treating an export as complete. - Order is unspecified.
Vector similarity search. Each query searches by an explicit vector or by
an existing document id (exactly one per query).
Request body (SearchRequest):
| Field | Type | Default | Notes |
|---|---|---|---|
queries |
array of QuerySpec |
— | At least one. |
topk |
int (1–1000) | 10 |
Number of results per query. |
filter |
string | null | null |
SQL-like scalar filter (see above). |
include_vector |
bool | false |
Include vectors in hits. |
output_fields |
array of string | null | null |
Restrict returned scalar fields. |
QuerySpec (exactly one of vector / id):
| Field | Type | Notes |
|---|---|---|
field |
string | Vector field to search. |
vector |
list of float | null | Query vector. |
id |
string | null | Search by an existing document's vector. |
params |
object | null | Query tuning for the field's index (below). |
Query params by index type (unknown keys return 400):
| Index | Params |
|---|---|
hnsw, hnsw_rabitq |
ef (int), radius (float), is_linear (bool), is_using_refiner (bool) |
ivf |
nprobe (int) |
ivf_rabitq |
nprobe, radius, is_linear, is_using_refiner, scale_factor (float) |
flat |
none |
ef / nprobe trade speed for recall; is_linear: true forces an exact
brute-force scan; radius limits hits to a distance threshold.
Example request
{
"queries": [
{ "field": "embedding", "vector": [0.12, 0.22, 0.29, 0.41], "params": { "ef": 64 } }
],
"topk": 3,
"filter": "category = 'tech' AND year > 2020",
"include_vector": false,
"output_fields": ["category", "year"]
}Response 200 (SearchResponse)
{
"results": [
{
"id": "a1",
"score": 0.0123,
"vectors": null,
"fields": { "category": "tech", "year": 2021 }
}
]
}Results are a flat list of hits. For a single query they are ordered by score. A malformed filter returns
400(invalid_argument).
Run one query, bucket hits by the value of a scalar field, and return the
best topk_per_group hits from each of the best group_count groups. Typical
RAG use: chunks stored with a doc_id field, grouped so one long document
can't crowd every other document out of the results.
| Field | Type | Default | Notes |
|---|---|---|---|
query |
QuerySpec |
— | Same shape as a search query. |
group_by |
string | — | Scalar field defining the groups. |
group_count |
int (1–1000) | 10 |
Maximum groups returned. |
topk_per_group |
int (1–1000) | 3 |
Maximum hits per group. |
filter |
string | null | null |
SQL-like scalar filter. |
include_vector |
bool | false |
Include vectors in hits. |
output_fields |
array of string | null | null |
Restrict returned scalar fields. |
Example request
{
"query": { "field": "embedding", "vector": [0.12, 0.22, 0.29, 0.41] },
"group_by": "doc_id",
"group_count": 5,
"topk_per_group": 2
}Response 200 (GroupSearchResponse)
{
"groups": [
{ "value": "doc-7", "results": [{ "id": "doc-7#3", "score": 0.011, "fields": { "doc_id": "doc-7" } }] },
{ "value": "doc-2", "results": [{ "id": "doc-2#0", "score": 0.019, "fields": { "doc_id": "doc-2" } }] }
]
}Group
values are always strings (anINT64field's3comes back as"3"), and documents with a nullgroup_byvalue form a group with value"". An unknowngroup_byfield or a malformed filter returns400.