--- license: cc-by-nc-4.0 tags: - super-resolution - arbitrary-scale-super-resolution - image-restoration - implicit-neural-representation - gblsr pipeline_tag: image-to-image metrics: - psnr - lpips --- # GB-LSR: Local Spectral Decoding with a Learned Global Bandwidth [![arXiv](https://img.shields.io/badge/arXiv-2606.19617-b31b1b.svg)](https://arxiv.org/abs/2606.19617) [![DOI](https://img.shields.io/badge/DOI-10.48550%2FarXiv.2606.19617-blue.svg)](https://doi.org/10.48550/arXiv.2606.19617) Trained weights for GB-LSR, from [GB-LSR: Local Spectral Decoding with a Learned Global Bandwidth for Arbitrary-Scale Super-Resolution](https://arxiv.org/abs/2606.19617). GB-LSR partitions the image domain into a fixed grid of non-overlapping square patches. Each patch carries coefficients for a truncated Fourier basis, predicted by a single linear projection from shared convolutional-encoder features, and one trainable scalar bandwidth is shared across every patch and every image. As in earlier local spectral decoders, decoding at a continuous coordinate is a fixed-size basis contraction whose cost is set by the spectral cutoff; GB-LSR learns the bandwidth of that basis instead of fixing it. Code: . ## Models One checkpoint per configuration, from training seed 0. The paper's tables report means over three training seeds; the PSNR of the released checkpoints differs from those means by at most 0.021 dB. ### Arbitrary-scale super-resolution RDN encoder, trained on DIV2K at scales from x1 to x4. PSNR-Y (dB) and LPIPS at x4. Learned s: the bandwidth after training. | Folder | Model | Params (M) | Set5 | Set14 | B100 | Urban100 | DIV2Kval | Mean LPIPS (lower is better) | Learned s | |---|---|---|---|---|---|---|---|---|---| | `gblsr-scalar-asr` | GB-LSR-Scalar-ASR (base) | 22.02 | 32.19 | 28.72 | 27.66 | 26.43 | 30.67 | 0.2761 | 0.881 | | `gblsr-scalar-asr-noLE` | GB-LSR-Scalar-ASR-noLE | 22.02 | 32.19 | 28.71 | 27.67 | 26.45 | 30.68 | 0.2747 | 0.794 | | `gblsr-scalar-asr-nf96-noLE` | GB-LSR-Scalar-ASR-nf96+noLE | 24.93 | 32.23 | 28.74 | 27.68 | 26.47 | 30.70 | 0.2743 | 0.795 | | `gblsr-scalar-asr-nf48-noLE` | GB-LSR-Scalar-ASR-nf48+noLE | 20.61 | 32.18 | 28.72 | 27.67 | 26.40 | 30.66 | not scored | 0.784 | noLE: trained and evaluated without the local ensemble. nf96, nf48: RDN encoder with 96 or 48 base feature channels instead of 64. Scored as in the paper: PIL bicubic low-resolution inputs; PSNR-Y on the BT.601 luminance with a border as wide as the scale removed; Set5, Set14, B100, and Urban100 ground truth cropped at the bottom and right to a multiple of 12 pixels; DIV2Kval is the DIV2K validation split, scored on luminance and so not comparable with published DIV2K numbers; LPIPS-AlexNet on the RGB output. Latency at x4 (ms per image) and peak allocated GPU memory, from the paper's timing session on one NVIDIA H200 at batch size 1: median of 50 timed passes, averaged over the images and the three seeds' models. Speed: the base model's latency divided by the row's, geometric mean over the three test sets. | Folder | Set14 | B100 | Urban100 | Peak memory, Urban100 (MiB) | Speed vs base | |---|---|---|---|---|---| | `gblsr-scalar-asr` | 34.8 | 25.6 | 109.4 | 899 | 1.00x | | `gblsr-scalar-asr-noLE` | 17.7 | 14.6 | 52.5 | 900 | 1.93x | | `gblsr-scalar-asr-nf96-noLE` | 20.1 | 16.0 | 55.7 | 1,109 | 1.76x | | `gblsr-scalar-asr-nf48-noLE` | 18.8 | 14.4 | 52.9 | 730 | 1.90x | ### Design-study models (native reconstruction) GB-LSR-Scalar at patch size 32 and 16, spectral cutoff 16, feature width 128, trained on DTD and DIV2K. PSNR-RGB (dB) on images standardized to 256x256 as in the paper: a center crop, after upsampling any image with a side under 256. Latent: encoder feature values per image. Learned s is the implementation's value, as in the paper's design study (Appendix A.1); the bandwidth of the paper's Section 3.1 is this value times P/(P-1) for patch size P. Latency and memory as above, per 256x256 image. The design study compares these models with other GB-LSR variants only. | Folder | Patch size | Params | Latent | Kodak | Set14 | Urban100 | Learned s | Latency (ms) | Peak memory (MiB) | |---|---|---|---|---|---|---|---|---|---| | `gblsr-scalar-p32` | 32 | 989,955 | 8,192 | 26.32 | 25.34 | 22.55 | 0.758 | 1.367 | 117.8 | | `gblsr-scalar-p16` | 16 | 842,115 | 32,768 | 32.44 | 31.01 | 28.09 | 0.658 | 1.358 | 117.3 | ## Results in brief Against the authors' released LIIF, LTE, and SRNO checkpoints on the same RDN encoder, scored under one protocol and timed in one session per scale, the base model runs 1.25x faster than LIIF-RDN at x4 and as fast as SRNO-RDN, and trails the three by 0.07 to 0.79 dB PSNR-Y in distribution. Without the local ensemble, PSNR-Y stays within seed variation and the speedup over LIIF-RDN rises to 2.41x at x4 and 3.00x at x8, at the cost of value jumps at cell boundaries of 0.22 gray levels (of 255) on average at x4. The GB-LSR models scored with LPIPS (base, noLE, nf96+noLE) have the highest (worst) mean LPIPS at x4 of the methods compared. The paper gives the full comparison and its scope. ## Usage Install the package at the commit these files were checked with, and the Hub client: ```bash pip install "git+https://github.com/KempnerInstitute/gblsr@6a3cd604c2ba9e868daf01727551d70a76932e1c" huggingface_hub ``` Arbitrary-scale super-resolution: ```python import json, torch from huggingface_hub import snapshot_download from safetensors.torch import load_file from gblsr import GBLSRScalarASR root = snapshot_download("KempnerInstituteAI/gblsr", revision="arxiv-v2", allow_patterns="gblsr-scalar-asr/*") cfg = json.load(open(f"{root}/gblsr-scalar-asr/config.json")) model = GBLSRScalarASR(encoder_cfg=cfg["encoder_cfg"], decoder_cfg=cfg["decoder_cfg"]).eval() model.load_state_dict(load_file(f"{root}/gblsr-scalar-asr/model.safetensors")) lr = torch.rand(1, 3, 64, 64) # RGB in [0, 1] with torch.no_grad(): # any output size; tile_q chunks the decoder's queries hr = model.predict_full(lr, H_q=256, W_q=256, tile_q=30000).clamp(0, 1) ``` Design-study model: ```python import json, torch from huggingface_hub import snapshot_download from safetensors.torch import load_file from gblsr import BasisConfig, EncoderConfig, ModelConfig, build_model root = snapshot_download("KempnerInstituteAI/gblsr", revision="arxiv-v2", allow_patterns="gblsr-scalar-p32/*") b = json.load(open(f"{root}/gblsr-scalar-p32/config.json"))["build"] model = build_model( ModelConfig(arm=b["arm"], image_size=b["image_size"], patch_size=b["patch_size"], basis=BasisConfig(patch_size=b["patch_size"], p_max=b["p_max"], s_e_range=tuple(b["s_e_range"])), encoder=EncoderConfig(d_feat=b["d_feat"])), bandwidth_mode=b["bandwidth_mode"], adapt_order=b["adapt_order"]).eval() model.load_state_dict(load_file(f"{root}/gblsr-scalar-p32/model.safetensors")) with torch.no_grad(): recon = model(torch.rand(1, 3, 256, 256))["recon"].clamp(0, 1) # 256x256 RGB ``` `examples/infer.py` runs either kind of model on an image file. `SHA256SUMS` gives the sha256 of this card, the example, and every model file. ## Training Super-resolution: 1,000,000 steps on DIV2K with an L1 loss on 48x48 low-resolution patches; each example is a square crop of side round(48k) pixels with k drawn uniformly from [1, 4], downsampled to 48x48 with PIL bicubic, and the loss is taken on 2,304 of its pixels; batch size 16, random horizontal flips and 90 degree rotations, Adam at a learning rate of 1e-4 halved at 200,000, 400,000, 600,000, and 800,000 steps. The bandwidth starts at s = 1.0. Design study: 1,000,000 steps on 256x256 crops from a mixture of DTD and DIV2K, batch size 8, mean squared error loss, AdamW (beta = (0.9, 0.95), no weight decay) at a constant learning rate of 2e-4 with gradient-norm clipping at 1.0. The bandwidth passes through a log-space sigmoid onto [0.25, 2.0] and starts at s = 0.707. ## License and attribution Released under CC BY-NC 4.0. The super-resolution models are trained on [DIV2K](https://data.vision.ee.ethz.ch/cvl/DIV2K/), released "for academic research purpose only"; the design-study models are trained on [DTD](https://www.robots.ox.ac.uk/~vgg/data/dtd/), released for research purposes, and DIV2K. Use of the weights must respect those terms. The `gblsr` code is BSD-3-Clause. ## Versions - arXiv v2 (this version, tag `arxiv-v2`): the 2,000-step native model `gblsr-scalar` is replaced by the 1,000,000-step design-study models `gblsr-scalar-p32` and `gblsr-scalar-p16`; the four super-resolution weight files are unchanged. - arXiv v1: tag `arxiv-v1`. ## Citation ```bibtex @article{shad2026gblsr, title = {GB-LSR: Local Spectral Decoding with a Learned Global Bandwidth for Arbitrary-Scale Super-Resolution}, author = {Shad, Max and Khoshnevis, Naeem}, journal = {arXiv preprint arXiv:2606.19617}, year = {2026} } ```