Compressed Linear usage
CompressedLinear replaces a PyTorch linear layer with one that uses compressed weights.
Call layer(x) as usual: the input’s last dimension changes from in_features to
out_features, and the other dimensions stay the same.
Use a Config to select the compression scheme and its parameters:
DFloat11Config, TileANSConfig, or LatticeRANSConfig. The selected scheme must
support the weight dtype.
A CUDA GPU and the matching CuPy package are required. See Quick start for installation.
Replace an existing layer
from_linear compresses an existing layer’s weights and returns a new layer with a copy
of the original bias. For pretrained models, load the checkpoint before calling this method:
import torch
import entropack as ep
linear = torch.nn.Linear(256, 256, dtype=torch.bfloat16, device="cuda")
config = ep.LatticeRANSConfig(target_bpp=4.0)
layer = ep.CompressedLinear.from_linear(linear, config=config)
x = torch.randn(8, 256, dtype=torch.bfloat16, device="cuda")
with torch.inference_mode():
output = layer(x)
print(output.shape, f"{layer.compressed_bits:.2f} bits per weight")
Assign the returned layer to the corresponding module attribute to use it in the model.
Replace several layers in a model
This example replaces ordinary linear layers recursively and keeps the output layer in its
original format. Names in skip are module paths, as reported by named_modules().
import torch
import entropack as ep
model = torch.nn.Sequential(
torch.nn.Linear(256, 256),
torch.nn.GELU(),
torch.nn.Linear(256, 64),
).to(device="cuda", dtype=torch.bfloat16).eval()
config = ep.LatticeRANSConfig(target_bpp=4.0)
def compress_linears(module, config, skip=(), prefix=""):
for name, child in list(module.named_children()):
path = f"{prefix}.{name}" if prefix else name
if path in skip:
continue
if type(child) is torch.nn.Linear:
replacement = ep.CompressedLinear.from_linear(child, config=config)
setattr(module, name, replacement.train(child.training))
else:
compress_linears(child, config, skip, path)
compress_linears(model, config, skip={"2"})
x = torch.randn(8, 256, device="cuda", dtype=torch.bfloat16)
with torch.inference_mode():
output = model(x)
print(output.shape, type(model[0]).__name__, type(model[2]).__name__)
The example selects standard torch.nn.Linear layers. Custom linear classes or shared
weights may need model-specific handling. Omit skip to compress every ordinary linear
layer. If CompressedLinear cannot compress a layer’s weights, replacement raises an
error; use skip to keep that layer in its original form.
Combine compression with FP8 or INT8 computation
| Layer | Weight format | Computation |
|---|---|---|
CompressedLinear |
Input weight dtype | Standard linear operation, using the activation dtype |
CompressedFP8Linear |
FP8 E4M3FN codes | FP8 weights and activations, requires a CUDA GPU with SM8.9 or later |
CompressedINT8Linear |
INT8 codes | INT8 weights and activations, requires a CUDA GPU with SM8.0 or later |
For CompressedFP8Linear and CompressedINT8Linear, use config=None for FP8 or INT8
quantization alone, or pass LatticeRANSConfig(target_bpp=...) to apply further lossy
compression to the quantized weights. The target must be at least 0.001 bpp and less than 8 bpp.
import torch
import entropack as ep
linear = torch.nn.Linear(256, 256, dtype=torch.bfloat16, device="cuda")
layer = ep.CompressedINT8Linear.from_linear(
linear, config=ep.LatticeRANSConfig(target_bpp=4.0)
)
x = torch.randn(32, 256, dtype=torch.bfloat16, device="cuda")
with torch.inference_mode():
output = layer(x)
print(output.shape, layer.container_dtype, f"{layer.compressed_bits:.2f} bits per weight")
Use CompressedFP8Linear in the same pattern for FP8, on supported hardware.
To inspect the weights, codes() returns their FP8 or INT8 values, and dequantize()
returns their floating-point values after dequantization.
Measure storage
stored_nbytes reports the compressed weight size, including metadata and FP8 or INT8 quantization scales.
compressed_bits is 8 * stored_nbytes / (in_features * out_features).
For multiple layers, sum stored bytes and weight elements before computing the ratio.
Biases are separate from this weight-storage measure.
This measures weight storage, not peak inference memory.
Save and load a model
Save the model’s state_dict, then construct a model with the same architecture and
Compressed Linear classes before loading it. The following example compresses two layers
and restores their saved weights into a fresh model:
from pathlib import Path
import torch
import entropack as ep
config = ep.LatticeRANSConfig(target_bpp=4.0)
model = torch.nn.Sequential(
torch.nn.Linear(256, 256),
torch.nn.GELU(),
torch.nn.Linear(256, 64),
).to(device="cuda", dtype=torch.bfloat16)
for index in (0, 2):
model[index] = ep.CompressedLinear.from_linear(model[index], config=config)
model.eval()
x = torch.randn(8, 256, device="cuda", dtype=torch.bfloat16)
with torch.inference_mode():
expected = model(x)
path = Path("compressed_model.pt")
torch.save(model.state_dict(), path)
restored = torch.nn.Sequential(
ep.CompressedLinear(256, 256, config=config, device="cuda", dtype=torch.bfloat16),
torch.nn.GELU(),
ep.CompressedLinear(256, 64, config=config, device="cuda", dtype=torch.bfloat16),
).eval()
state = torch.load(path, map_location="cuda", weights_only=True)
restored.load_state_dict(state)
with torch.inference_mode():
actual = restored(x)
assert torch.allclose(actual, expected)
print(actual.shape)
The example saves compressed_model.pt in the current directory. Change the path as needed.
To load an existing checkpoint, construct the restored model, then call torch.load and
load_state_dict.
Keep the model architecture, layer names and classes, compression configurations, weight
dtypes, and library version with the checkpoint. Ordinary torch.nn.Linear layers
cannot load Compressed Linear checkpoints directly.