Improving Depth Anything V2 Robustness to Video Compression

Community Article
Published May 7, 2026

Executive Summary

Managing the petabyte-scale AV real-time and synthetic video data requires compression. At the same time, applying compression to AV data pipelines requires confidence that downstream model accuracy is preserved. As a step forward in developing ML-safe video data processing, this study demonstrates that treating compression as a training strategy allows autonomous fleets to scale efficiently while preserving the perception accuracy AV systems depend on. This study focuses on depth estimation - the perception task that produces a depth map, and is among the most sensitive to any change in input.

Commonly used codecs are inherently lossy. While often visually lossless to human observers, they can distort spatial geometry, specifically for high-frequency information that machine learning models may rely on. In particular, dense, pixel-specific tasks in the AV perception stack - such as monocular depth estimation - are notoriously sensitive to any video transformation, including compression artifacts.

We propose a new approach: video compression should be utilized as data augmentation during model training to effectively improve the model learning and robustness. Just as models are trained with synthetic blur or noise to handle adverse weather, injecting compression artifacts into the training pipeline forces the network to learn geometric representations that are natively robust to telematics (i.e. transmission of computerized information) challenges, among which video compression is prominent.

Since lossy compression inherently discards fine spatial details regardless of the codec used, processing these compressed inputs degrades the model’s depth predictions compared to its clean-input baseline. To overcome this, we performed targeted model fine-tuning using the compressed videos, and achieved significant reduction in the validation errors using this approach.

1. Experiment Setup

In our prior testing, depth models showed measurable sensitivity to compressed input - even when the compression introduced no visible change to the video. To validate the “compression as augmentation” approach, we examined the impact of video compression on Depth Anything V2 (Base) [1], a state-of-the-art monocular relative depth estimation model. The model utilizes a pre-trained DINOv2 (ViT-B) [2] Vision Transformer as its backbone to extract rich, global semantic features from single 2D camera frames. These features are then systematically processed through a Dense Prediction Transformer (DPT) head, which fuses multi-scale representations to regress a continuous, high-resolution depth map.

To construct our training pipeline, we first fed the raw, uncompressed video frames into the baseline “teacher” model to generate reference depth maps. We then encoded these original video sequences using NVENC CABR HEVC hardware-accelerated encoding to produce compressed, optimized videos. Compared to standard NVENC HEVC encoding at equivalent quality, CABR achieves an overall 35.2% file size reduction across our validation set, which consists of 43 AV videos from the publicly available datasets Kitti[3], A2D2[4] and PandaSet[5]. Finally, we used these optimized, compressed frames as the primary input to fine-tune our “student” network forcing the model to learn how to recover the high-fidelity depth outputs directly from the compressed videos.

2. Finetuning for Improved Robustness

Since ground-truth depth is unavailable across diverse driving datasets, we employ a self-distillation methodology. A frozen "teacher" model processes uncompressed frames to generate pseudo ground-truth depth maps. A "student" model - initialized from the exact same weights learns to recover that high-fidelity depth geometry using the compressed frames as input.

2.1 Training Configuration

  • Base Model: Depth Anything V2 Base (~100M parameters).
  • Adaptation: LoRA[6] (rank 16, alpha 32) applied to all backbone attention and MLP layers, plus full fine-tuning of the neck and head. In total, this configuration updates only 13.5% of the network’s total parameters.
  • Loss Function: Masked L1. We ignore sky and extremely distant pixels (normalized depth < 0.02) to avoid penalizing the model on undefined geometry. While the main evaluation metric for this task is AbsRel, we optimize using Masked L1, since we want to penalize equally all parts of the depth image difference, to preserve the behaviour of the teacher model as much as possible.
  • Compression Augmentation: A multi-source dataset (A2D2, Kitti, PandaSet) split into 309 training videos (38,361 frames) and 43 held-out validation videos (6,513 frames), trained over 50 epochs. We utilize CABR compression as our primary spatial augmentation. By feeding the student model compressed inputs while targeting clean-frame pseudo-ground truth, the network learns to implicitly invert quantization artifacts.
  • Domain Retention: A 20% mix of clean frames is included during training to prevent catastrophic forgetting and ensure the model remains equally performant if uncompressed inputs are provided.
  • Zero-Overhead Inference: Post-training, the learnt LoRA weights are fully merged back into the base model. Consequently, the student model recovers compression-degraded depth with zero additional inference latency and zero additional VRAM overhead compared to the baseline teacher.

3. Distinguishing Artifacts from Natural ML Uncertainty

To accurately measure the true impact of video compression, we must first establish the teacher model's natural sensitivity to standard input variations. In real-world AV environments, even uncompressed camera feeds are subject to minor photometric variations (like ISO jitter or thermal noise) as well as mechanical calibration noise (such as subtle horizontal pixel shifts imitating physical camera movements on the vehicle). Since the baseline depth model is highly responsive to these small photometric and geometric shifts, its predictions naturally exhibit a baseline variance, even before any compression is applied.

Because of this inherent variance, measuring the student model against a strict zero-error baseline does not fully capture its practical performance. Instead, a more grounded approach is to evaluate whether the degradation introduced by compression actually exceeds the model's natural uncertainty under typical sensor and mechanical noise.

To answer this, we contextualized our assessment of the compression error by measuring the teacher model’s “noise floor” . For each clean validation frame, we generated 30 perturbation variants -combining photometric noise (Gaussian noise, brightness/contrast jitter) with minor geometric shifts (up to 2 horizontal pixels). These perturbations are designed to mimic natural, real-world sensor and calibration fluctuations. We then computed the per-pixel standard deviation (σ) of the teacher's depth predictions across these 30 noisy variants.

We evaluated whether the compression-induced error fell within 2σ (the teacher's natural uncertainty band). The P_valid metric denotes the percentage of pixels whose error falls within this natural variance band. If the student model's prediction on a compressed frame falls within this 2σ band, its output is statistically indistinguishable from the teacher's performance under natural sensor noise. This methodology mathematically separates actual compression damage from the model's baseline input sensitivity, proving whether the student has genuinely neutralized the compression artifacts to a safer deployment standard.

4. Results & Semantic Evaluation

4.1 Evaluation Metrics

Depth Anything V2 predicts relative (affine-invariant) depth rather than absolute metric distance (meters). Consequently, Absolute Relative Error (AbsRel) serves as the primary evaluation metric for this task, as it measures the percentage difference between the student's prediction and the teacher's baseline, normalizing for scale.

This is vital for AV pipelines: downstream perception stacks perform calibration priors to scale these relative camera depth maps into physical metric distances. If the relative depth geometry of a pedestrian is distorted by compression (resulting in a high AbsRel), that error propagates directly when scaled into real-world meters, leading to dangerous miscalculations in the perception pipeline. By minimizing AbsRel, the student model guarantees that the fundamental 3D geometry remains intact, ensuring that downstream metric scaling is safe and accurate.

4.2 Qualitative Visual Recovery

We visualize the direct impact of compression, and the student model's subsequent recovery, on individual driving frames. Reading vertically from top to bottom, these visualizations progress through: (1) the raw vs. CABR-compressed RGB inputs, (2) the Teacher on Clean (GT) depth map vs. the σ_Teacher noise floor, (3) the depth maps from the Teacher vs. the adapted Student when processing the compressed frame, (4) paired Absolute Relative Error (AbsRel) heatmaps, and finally (5) paired P_valid safety masks (where red pixels indicate geometric corruption exceeding the model's natural 2σ noise floor).

Blog image

Figure 1: Qualitative comparison of depth estimation on a scene with vehicles. The baseline teacher model (left-hand side) suffers global geometry distortion under video compression, particularly on vehicle surfaces and complex backgrounds (resulting in a heavily red P_valid mask of 13%). The student model (right-hand side) effectively neutralizes these artifacts, reducing the AbsRel error by over 3x and safely restoring 77.1% of the scene's pixels back into the natural 2σ operational noise floor.

For ETHAN

Figure 2: Qualitative comparison of depth estimation on a scene with Safety-Critical VRU. The baseline teacher model (left-hand side) suffers severe distortion on the pedestrian and building structures when subjected to video compression (evidenced by the bright AbsRel errors and red P_valid mask). The student model (right-hand side) successfully neutralizes these artifacts, restoring the safety-critical VRU geometry to fall safely within the model's natural 2σ noise floor.

4.3 Aggregated Depth Recovery

The student model was evaluated on 43 held-out validation videos (6,513 frames). All metrics are computed on valid (sky-masked) pixels only.

Metric Teacher (compressed) Student (compressed) Reduction
AbsRel 0.02516 0.02112 16.0%
RMSE 0.23832 0.21330 10.5%

Table 1: Teacher on compressed vs. Student on compressed

Note: When comparing the Student on compressed inputs directly against the Teacher on clean inputs, the student achieved metrics of, AbsRel = 0.0178, and RMSE = 0.1289

4.4 Per-Class Semantic Error Analysis

To understand which scene elements are most affected by compression, we ran semantic segmentation (SegFormer-B2[7], Cityscapes[8]) on the validation set. Reference: Teacher on clean inputs serves as ground truth.

Semantic Group Pixel % Teacher AbsRel Student AbsRel AbsRel Reduction Teacher RMSE Student RMSE RMSE Reduction
VRU (person/rider/bike) 0.8% 0.03 0.02 30.7% 0.42 0.3 29%
Building/wall/fence 30.6% 0.03 0.02 29.2% 0.23 0.15 34%
Sidewalk 6.2% 0.01 0.01 28% 0.15 0.11 26.5%
Vegetation/terrain 22.8% 0.05 0.03 26.6% 0.28 0.2 31.4%
Road surface 29.1% 0.01 0.01 23% 0.15 0.11 27.3%
Vehicles (car/truck/bus) 8.9% 0.02 0.01 22.8% 0.34 0.25 26%
Traffic infra (sign/light) 1.3% 0.04 0.04 12.1% 0.44 0.33 24.8%
Sky (masked) 0.2% 0.12 0.11 5.7% 0.36 0.31 13.2%

Table 2: Per-Class Semantic Error Analysis

The semantic breakdown highlights how compression degradation, and the student model's subsequent recovery, varies depending on the spatial frequency and geometric complexity of the scene elements. Most notably, Vulnerable Road Users (VRUs), which represent the most highly detailed and safety-critical objects in the AV environment, suffer heavily under baseline compression but see the greatest relative recovery from the student model (a 30.7% reduction in AbsRel and 29.0% in RMSE). High-frequency structural elements like buildings, fences, and vegetation also demonstrate massive geometric restoration, with RMSE reductions exceeding 30%. Conversely, large, homogeneous areas with low spatial frequency, such as road surfaces, inherently exhibit much lower baseline compression errors (a Teacher AbsRel of only 0.0101) and yet, the student model still successfully reduces this error by an additional 23.0%. This consistent improvement across both complex foreground objects and flat background planes confirms that the self-distillation process learned a generalized, robust understanding of 3D geometry rather than simply overfitting to specific object textures.

4.5 Noise-Floor & Uncertainty Analysis

To accurately contextualize the student model's recovery, we measured the baseline uncertainty of both the teacher and the student models. By applying specific perturbations to the uncompressed inputs, we establish natural "noise floors" (σ) for each network. We conducted three escalating experiments: Photometric noise only, followed by the addition of ±1 pixel and ±2 pixel horizontal shifts to simulate mechanical camera vibration.

4.5.1 Increased Intrinsic Stability (σ Comparison)

The initial analysis isolates photometric variance, simulating natural camera sensor fluctuations. We evaluated both models across 30 perturbation variants per frame, applying Gaussian noise (σ=0.015) and brightness/contrast jitter (±5%). We then introduced horizontal shifts of ±1 and ±2 pixels, which simulate natural calibration noise and natural camera movement during the drive.

noise_floor_sigma

Figure 3: Comparison of the natural noise floor (σ) between the Teacher and Student models across increasing levels of spatial perturbation. The x-axis represents the seven semantic classes, while the y-axis denotes the mean per-pixel standard deviation (the noise floor) in relative depth units.

The photometric and spatial noise analysis reveals a dual benefit of the compression-augmentation methodology. As illustrated across all three panels in Figure 3, the student model (blue) demonstrates a consistently tighter natural noise floor (σ_Student) than the baseline teacher (red). Across every semantic class, the student's baseline variance is 15% to 20% lower. This confirms that training with compression artifacts forced the network to become fundamentally more robust, learning a deterministic representation that resists both lighting variations and physical camera movements. As expected, both noise floors increase as spatial perturbation grows (moving from panel A to C), particularly for edge-heavy classes like Traffic Infrastructure, but the student maintains its superior stability throughout.

4.5.2 Validating Recovery Against Teacher Baseline (P_valid → σ_Teacher)

Having established the noise floors, we measured the percentage of pixels (P_valid) where compression-induced errors fall safely within the teacher's natural 2σ variance band.

pvalid_teacher_nf_noshift

Figure 4a: Pixel Validity within the Teacher's Natural Uncertainty Band - Photometric Noise Only

pvalid_teacher_nf_shift1

Figure 4b: Pixel Validity within the Teacher's Natural Uncertainty Band - Photometric Noise + ± 1px Shift

pvalid_teacher_nf_shift2

Figure 4c: Pixel Validity within the Teacher's Natural Uncertainty Band - Photometric Noise + ± 2px Shift

These charts evaluate performance against the baseline teacher's noise floor across (a) Photometric Only, (b) ±1px shift, and (c) ±2px shift experiments. In each chart, the x-axis categorizes the scene into the seven distinct semantic classes, while the y-axis measures the percentage of valid pixels P_valid safely within the target noise floor.

  • Red Bars (Baseline): Teacher model evaluated on compressed optimized video.
  • Blue Bars (Improvement): Student model evaluated on compressed optimized video, successfully pushing a higher percentage of artifacts back into the safe variance band.
  • Green Bars (Safety Check): Student model evaluated on clean uncompressed video.

These metrics highlight the two most critical technical achievements of the study:

  • Safety-Critical Artifact Recovery (Blue > Red): The baseline teacher struggles to maintain reliable geometry under compression. For example, under purely photometric noise, only 36.0% of compressed VRU pixels fall within the teacher's own natural 2σ uncertainty band. The student effectively repairs this damage, significantly outperforming the baseline by pushing 53.0% of those pixels back into the safe threshold. When evaluating under the more realistic ±2 pixel mechanical vibration envelope, this recovery surges: nearly three-quarters (74.5%) of the student's VRU predictions on heavily compressed frames are statistically indistinguishable from natural camera shake.
  • No Forgetting (Green > Blue): A critical requirement for AV deployment is that models must not degrade when uncompressed streams are available. The student model evaluated on clean inputs (green bars) boasts the highest validity across all classes. Under the ±2 pixel shift envelope, the student achieves 95.5% validity for both VRUs and Traffic Infrastructure. This mathematically proves the 20% clean-frame training mix successfully preserved high-bandwidth fidelity.

4.5.3 Self-Consistency Verification (P_valid → σ_Student)

Finally, we evaluated the student against its own, much tighter, 2σ noise floor.

pvalid_student_nf_noshift

Figure 5a: Pixel Validity within the Tighter Student Uncertainty Band- Photometric Noise Only

pvalid_student_nf_shift1

Figure 5b: Pixel Validity within the Tighter Student Uncertainty Band- Photometric Only Noise + ± 1 px shift

pvalid_student_nf_shift2

Figure 5c: Pixel Validity within the Tighter Student Uncertainty Band- Photometric Only Noise + ± 2 px shift

These charts evaluate the student model against its own internal noise floor across (a) Photometric Only, (b) ±1px shift, and (c) ±2px shift experiments. Similar to the previous evaluation, the x-axis represents the semantic classes and the y-axis measures the valid pixel percentage P_valid, but measured against the student's stricter internal baseline.

  • Blue Bars: Student performance on compressed optimized video.
  • Green Bars: Student performance on clean uncompressed video (Baseline consistency check).

Even against this stricter internal envelope, the model demonstrates remarkable self-consistency. As shown in the figures, 44% to 49% of the student's compressed predictions fall within its own tighter noise floor under purely photometric noise. When factoring in the ±2 pixel mechanical shift, this validity grows to 63%–78% on compressed frames, and reaches an exceptional 76%–94% on clean inputs. This confirms that the student is not merely memorizing and matching the teacher's outputs, but is generating genuinely more precise and internally consistent geometry across both compressed and uncompressed data.

4.6 Key Findings & Technical Impact

  1. Fine-tuning Increases Model Stability (Tighter Noise Floor): Beyond simply repairing spatial artifacts, utilizing compression as an augmentation mechanism fundamentally stabilized the overall model. Across all semantic classes, the student model's noise floor (σ_Student) is 15% to 20% lower than the teacher's. The LoRA fine-tuning forced the network to learn a more invariant, deterministic geometric representation that is highly robust to natural sensor noise.
  2. Safety-Critical Recovery (VRUs): Vulnerable road users (VRUs) benefit the most from this methodology, achieving a massive 30.7% AbsRel reduction and a 29.0% RMSE reduction. Under baseline compression, VRUs have the lowest valid pixel rate (36.0%), indicating compression degrades their high-frequency geometry. When accounting for realistic mechanical camera vibration, the student successfully pushes 74.5% of these compressed pixels safely back into the teacher's natural 2σ uncertainty band.
  3. No Forgetting on Clean Data: A critical requirement for AV deployment is that models must not degrade when uncompressed streams are available. The 20% clean-frame mixing during training was entirely successful. When evaluated on clean frames, the student model's AbsRel is 2x to 5x lower than its compressed error, and up to 95.5% of its pixels fall comfortably within the teacher's baseline noise floor.

5. Conclusion & Impact

This study proves that using video compression as an essential data augmentation step during model training improves the model robustness. By neutralizing the compression loss via this corresponding targeted augmentation, it becomes possible to meet the industry requirements and the set KPIs, AV pipelines can safely adopt compression and achieve drastically tighter, more accurate depth bounds for downstream perception tasks.

Across the validation set, fine-tuning produced consistent improvements at three levels: model stability, geometric restoration, and safety-critical recovery, without compromising clean-input performance:

  • Decoupling Cost from Safety: By pairing Beamr CABR encoding with this self-distillation methodology, AV pipelines can achieve an aggregate 35.2% reduction in video size over standard encoding, while simultaneously recovering up to 30.7% of the spatial geometry typically lost on safety-critical objects like VRUs.
  • Global Geometric Restoration: Across every foreground and background class, the student model successfully pushed a higher percentage of pixels back into the teacher's natural 2σ uncertainty band, proving the model neutralized compression artifacts globally rather than merely applying localized algorithmic smoothing.
  • Parameter-Efficient Proof of Concept: Notably, these recovery metrics were achieved through lightweight, highly parameter-efficient fine-tuning, updating only ~14% of the total model weights, rather than training the 100M parameter network from scratch. Because this simple, inexpensive adaptation so effectively neutralized severe compression damage, it strongly indicates that integrating compression augmentation natively into the pre-training phase of future AV foundation models will yield an even deeper, fundamental robustness to telematics bottlenecks.

Ultimately, this paradigm shift allows AV fleets to safely scale their data pipelines and aggressively cut infrastructure costs without compromising the high-confidence spatial KPIs required for safe vision-based perception.

References

  1. Yang, L., Kang, B., Huang, Z., et al. (2024). "Depth Anything V2". arXiv preprint arXiv:2406.09414. Link
  2. Oquab, M., Darcet, T., Moutakanni, T., et al. (2023). "DINOv2: Learning Robust Visual Features without Supervision". arXiv preprint arXiv:2304.07193. Link
  3. Geiger, A., Lenz, P., & Urtasun, R. (2012). "Are we ready for autonomous driving? The KITTI vision benchmark suite". Conference on Computer Vision and Pattern Recognition (CVPR). Link
  4. Geyer, W., Kassahun, Y., Mahmudi, M., et al. (2020). "A2D2: Audi Autonomous Driving Dataset". arXiv preprint arXiv:2004.06320. Link
  5. Xiao, P., Shao, Z., Hao, S., et al. (2021). "PandaSet: Advanced Sensor Suite Dataset for Autonomous Driving". IEEE International Intelligent Transportation Systems Conference (ITSC). Link
  6. Hu, E. J., Shen, Y., Wallis, P., et al. (2021). "LoRA: Low-Rank Adaptation of Large Language Models". International Conference on Learning Representations (ICLR). Link 7: Xie, E., Wang, W., Yu, Z., et al. (2021). "SegFormer: Simple and Efficient Design for Semantic Segmentation with Transformers". Advances in Neural Information Processing Systems (NeurIPS). Link 8: Cordts, M., Omran, M., Ramos, S., et al. (2016). "The Cityscapes Dataset for Semantic Urban Scene Understanding". Conference on Computer Vision and Pattern Recognition (CVPR). Link

Community

Sign up or log in to comment