Sapiens2 INT4-G128 Quantized Checkpoints
Collection
All 22 Sapiens2 INT4-G128 quants with fidelity scores. • 22 items • Updated • 1
Packed 4-bit derivative of facebook/sapiens2-pretrain-1b.
This artifact uses symmetric per-group INT4 packing with group size 128 for large floating-point weight tensors. Norms, biases, positional/rope tensors, and small tensors are kept in their source dtype. It is a storage/runtime-loader quant for the current official Sapiens2 code path, not an AWQ/GGUF/NVFP4 LLM artifact.
facebook__sapiens2-pretrain-1b-int4-g128.safetensors: packed INT4 safetensors artifact.load_sapiens2_int4.py: loader that reconstructs a PyTorch state dict for the official Sapiens2 model code.config.json and preprocessor_config.json: copied from the source repo.quantization_report.json: build report.fb127507e0b6c5de9b9516e60049b24bfae07de612858218967687538096647.7233x6862420.00560710The packed INT4 artifact was dequantized back to floating-point tensors and compared against the source checkpoint.
90.00%PASS99.332393%0.987518430from load_sapiens2_int4 import load_state_dict
state_dict = load_state_dict("facebook__sapiens2-pretrain-1b-int4-g128.safetensors", device="cpu")
# Then instantiate the matching official Sapiens2 architecture and load:
# model.load_state_dict(state_dict, strict=True)
This is a verified packed-weight artifact with a dequantizing loader. It does not claim native INT4 CUDA kernels for Sapiens2 yet. Runtime speedups require a Sapiens2-specific kernel/export path and should be benchmarked separately.
Base model
facebook/sapiens2-pretrain-1b