Benchmarking the utilization of standard compression methods to reduce size of SafeTensors blobs.
Safetensors has become, partly thanks to its utter simplicity, the leading format for storing and sharing DNN weights. The network and disk I/O associated with loading and moving the weights around is significant. A method to reduce its size, and therefore max(t_network, t_disk) would be beneficial.
To tackle the problem, I composed a quick experiment: How well do standard compression methods perform against real-world weights of leading open LLMs. Repository.
| Model | Scope | Method | Input MiB | Ratio | Saved | Compress (s) | Decompress (s) |
|---|---|---|---|---|---|---|---|
| Qwen/Qwen3-0.6B | all-safetensors | zstd-3 | 1433.7 | 1.280x | 21.9% | 2.37 | 1.44 |
| Qwen/Qwen3-0.6B | all-safetensors | zstd-9 | 1433.7 | 1.293x | 22.7% | 25.61 | 2.60 |
| Qwen/Qwen3-0.6B | all-safetensors | lz4 | 1433.7 | 1.002x | 0.2% | 0.91 | 0.51 |
| Qwen/Qwen3-0.6B | all-safetensors | deflate-6 | 1433.7 | 1.258x | 20.5% | 110.41 | 6.86 |
| Qwen/Qwen3-8B | all-safetensors | zstd-3 | 15622.6 | 1.275x | 21.6% | 30.24 | 20.00 |
| Qwen/Qwen3-8B | all-safetensors | zstd-9 | 15622.6 | 1.287x | 22.3% | 312.76 | 32.46 |
| Qwen/Qwen3-8B | all-safetensors | lz4 | 15622.6 | 1.001x | 0.1% | 11.40 | 6.94 |
| Qwen/Qwen3-8B | all-safetensors | deflate-6 | 15622.6 | 1.255x | 20.3% | 1266.42 | 80.22 |