Compression configuration

A Config selects the compression scheme and its encoding and decoding parameters. execution_backend defaults to "auto", which selects the execution backend automatically.

Choose a configuration

Config Compression Input requirements
DFloat11Config() Lossless BF16 compression BF16 tensors
TileANSConfig() Lossless compression for multiple dtypes Supported tensor dtypes
LatticeRANSConfig(target_bpp=...) Lossy compression at a target bitrate Nonempty, finite, two-dimensional tensors

Both lossless schemes reproduce every input bit. Their compressed size depends on the tensor’s data distribution. TileANSConfig also supports BF16, so either lossless scheme can be used for that dtype. LatticeRANSConfig accepts finite targets with 0.001 <= target_bpp <= 11 bits per element, including non-integer values. Its actual stored rate is available through CompressedTensor.actual_bpp.

Supported tensor dtypes

Tensor dtype DFloat11Config TileANSConfig LatticeRANSConfig
float32 — lossless lossy
float16 — lossless lossy
bfloat16 lossless lossless lossy
float8_e4m3fn — lossless lossy
float8_e4m3fnuz — lossless lossy
float8_e5m2 — lossless lossy
float8_e5m2fnuz — lossless lossy
int64 — lossless lossy
int32 — lossless lossy
int16 — lossless lossy
int8 — lossless lossy
uint64 — lossless lossy
uint32 — lossless lossy
uint16 — lossless lossy
uint8 — lossless lossy
bool — lossless lossy

The lossless schemes accept tensors of different shapes, while lattice quantization requires two-dimensional input. BF16, FP16, and FP8 refer to the corresponding PyTorch dtypes above. Packed four-bit formats, FP64, and complex dtypes are not supported.

Parameter conventions

Changes to encode settings affect subsequent compression, not existing compressed data. Decode settings take effect when restoring a tensor. Most settings can retain their defaults. For lossy compression, target_bpp sets the target rate and a positive row_rdo_iterations enables per-row rate–distortion optimization (RDO).

CompressionConfig

These execution settings apply to all three configuration classes.

Field Type Default Stage Meaning
execution_backend str or None "auto" Both "auto" or None prefers CUDA and falls back to the PyTorch implementation if unavailable. "cuda" requires CUDA. "eager" selects the PyTorch fallback for tensor encoding and decoding.

DFloat11Config

Lossless BF16 compression.

Field Type Default Stage Meaning
execution_backend str or None "auto" Both Backend selection as above
bytes_per_thread Positive int or None 16 Encode Encoded bytes processed per thread. Affects compression ratio and decoding parallelism.
threads_per_block Positive int or None 128 Encode Threads per block during encoding

TileANSConfig

Lossless compression of the supported tensor dtypes.

Field Type Default Stage Meaning
execution_backend str or None "auto" Both Backend selection as above
tile_elements int in [0, 2^31 − 1] 0 Encode Elements per compressed tile. Affects compression ratio and decoding parallelism. 0 selects automatically.
probability_bits 0, 9, 10, 11, 12 0 Encode Probability-table precision. 0 selects automatically.
raw_lane_threshold float in [0, 8] 7.9 Encode Threshold for storing hard-to-compress data directly, measured in estimated encoded bits per input byte.
threads_per_block Positive int or None None Both GPU block width. None selects automatically.

LatticeRANSConfig

Lossy compression of finite, two-dimensional tensors.

At encoding time, the effective target is the smaller of target_bpp and the dtype limit: 1 bpp for bool, 8 bpp for int8, uint8, and all supported FP8 dtypes, and 11 bpp for other supported dtypes. The config retains the requested value. These limits apply to the encoding target; actual_bpp includes metadata and may exceed them. Tensor distribution and metadata overhead can make low targets unattainable. Compression of bool remains lossy.

Field Type Default Stage Meaning
execution_backend str or None "auto" Both Backend selection as above
target_bpp Finite float in [0.001, 11] 4.0 Encode Target bits per input element, capped by dtype at encoding time. Non-integer targets are supported. Inspect actual_bpp for the stored rate.
prob_bits int in [9, 15], 0, or None None Encode Probability-table precision. None or 0 selects automatically.
tile_elements Positive int or None None Encode Elements per compressed tile. Affects compression ratio and decoding parallelism. None selects automatically.
row_rdo_iterations int in [0, 8] 0 Encode Per-row rate–distortion refinement sweeps. 0 disables refinement. More sweeps increase compression time.
row_rdo_candidates Positive int 5 Encode Number of candidate quantizations per row for RDO. More candidates increase compression time.
scale_search_iterations Positive int 12 Encode Number of quantization-scale search iterations
scale_search_max_vectors Positive int 262144 Encode Sample limit for rate search, in vectors of eight elements
threads_per_block Positive int or None None Decode GPU block width. None selects automatically.
l2_prefetch bool True Decode Enable GPU L2 cache prefetching during decoding

Compression principles

The encoding and decoding processes are described in DFloat11, Tile-ANS, and Lattice-rANS.