diff --git a/doc/glossary/index.rst b/doc/glossary/index.rst new file mode 100644 index 000000000..89e742eb9 --- /dev/null +++ b/doc/glossary/index.rst @@ -0,0 +1,98 @@ +Glossary +======== + +Data containers +--------------- + +.. glossary:: + + NDArray (:ref:`API `) + A compressed, chunked multidimensional data array. Supports NumPy-like + slicing and broadcasting, and out-of-core computation. Backed by an SChunk. + Use for array workloads, especially when too large to fit in memory + uncompressed. + + CTable (:ref:`API `) + A columnar table for structured data. Columns are stored, compressed, and + queried independently, with SUMMARY indexes available when eligible. Use with + structured data that benefits from compression, speed, and persistence. + + LazyArray (:ref:`API `) + API to store an expression or function and delay computation until the + value is explicitly requested. Executes and stores results chunk-by-chunk. + + LazyExpr (:ref:`API `) + Object that stores an expression consisting of at least one NDArray + object. Follows the LazyArray API for storage and deferred computation. + +Compression +----------- + +.. glossary:: + + codec (:class:`API `) + A compressor and matching decompressor that implements a particular + compression method. Codec choice affects compression/decompression speed and size of the compressed data. + + clevel + An integer from 0 (no compression) to 9 (the highest compression level) that controls compression effort. Higher levels generally take longer to compress and may produce smaller results, depending on the codec and data. + + filters (:class:`API `) + Transformations applied to data before the codec compresses it in order to expose patterns that can improve compression. Some filters are reversible (such as SHUFFLE), while others deliberately discard precision (such as TRUNC_PREC). + +Low-level data structures +------------------------- + +.. glossary:: + + SChunk (:ref:`API `) + The foundational container for managing a sequence of individual, + compressed chunks. NDArrays and CTable columns are built on top of SChunk. + Use when you want to directly manipulate raw compressed data and metadata. + + chunk + The unit of storage and compression, stored within a SChunk. Chunks are + sized to fit disk/network I/O, typically 1-64MB. + + block + The unit of decompression, stored within a chunk. Blocks are sized to + fit CPU caches, typically 32-512KB. + + subblock + An indexing segment within a block. One eighth the length of a block. + + frame + A serialized format for storing chunks along with metadata. Frames may + be contiguous (CFrame) or sparse (SFrame). + + CFrame + Frame format for storing chunks contiguously either in-memory or on-disk. + Short for contiguous frame. See `CFrame format `_. + + SFrame + Frame format for storing chunks non-contiguously on-disk. Short for sparse + frame. See `SFrame format `_. + +Indexes +------- + +.. glossary:: + + index (:ref:`API `) + Auxiliary data attached to an NDArray or CTable to speed up queries. + Allows queries to skip chunks, blocks, or rows of data that do not match. + + SUMMARY + Lightweight index that stores per-segment minimum and maximum values to skip segments that cannot match a query. + + BUCKET + Stores values sorted within each chunk and groups their positions into buckets to locate possible query matches. + + PARTIAL + Stores values sorted separately within each chunk, along with their exact positions, to find matches. + + FULL + Stores values sorted together across all chunks, along with their exact positions, to find matches and speed up sorting. + + OPSI + Uses repeated ordering cycles to improve filtering and provide exact matching positions for checking conditions on other columns, but is not intended to converge to a globally sorted index. diff --git a/doc/python-blosc2.rst b/doc/python-blosc2.rst index e932d3a7c..b817b6d94 100644 --- a/doc/python-blosc2.rst +++ b/doc/python-blosc2.rst @@ -224,5 +224,6 @@ Tutorials Guides API Reference + Glossary Development Release Notes