A unified framework that reconstructs and renders 3D scenes directly from extreme low-light images — no SfM, no clean references needed.
Reconstructing 3D scenes under real-world low-light conditions remains challenging due to severe sensor noise, low signal-to-noise ratios, and degraded photometric consistency, which destabilize geometry estimation and novel view synthesis. Existing approaches often rely on well-lit reference data for reliable Structure-from-Motion (SfM) initialization, or apply per-view enhancement methods that introduce cross-view inconsistencies.
We propose NOVA-GS, a unified noise-aware framework for low-light 3D Gaussian Splatting that subsumes enhancement, denoising, and geometry optimization within a single process. Our method leverages VGGT-based feed-forward estimation to obtain robust camera poses and geometry directly from degraded inputs, eliminating the need for SfM. We further introduce noise-guided spherical harmonic regularization to suppress view-dependent artifacts in noisy regions. Extensive experiments on diverse real-world low-light datasets demonstrate improved geometric fidelity, color consistency, and robustness without requiring paired supervision or well-lit references.
Leverages VGGT feed-forward geometry estimation directly on degraded low-light inputs. No COLMAP, no well-lit references required.
Laplacian-guided masking focuses denoising effort on high-frequency, noise-prone regions via a Deep Attention-ResUNet.
Depth-guided reprojection with confidence-aware photometric loss enforces global geometric coherence across all input views.
Maps 2D noise estimates to 3D Gaussians, penalizing higher-order spherical harmonics in noisy regions to eliminate floaters.
NOVA-GS processes low-light inputs concurrently through three coupled stages — structure-aware enhancement, self-supervised denoising, and consistency-aware 3DGS optimization — all jointly trained end-to-end.
Drag the handles to compare the dark input images against the NOVA-GS reconstructed novel views.
NOVA-GS achieves best performance on LLNeRF (+1.58 dB over the next best) and consistently competitive results across all datasets, without using ground-truth supervision for any stage.
| Method | LOM | LLRS | LLNeRF | ||||||
|---|---|---|---|---|---|---|---|---|---|
| PSNR↑ | SSIM↑ | LPIPS↓ | PSNR↑ | SSIM↑ | LPIPS↓ | PSNR↑ | SSIM↑ | LPIPS↓ | |
| 2D Enhancement + GS | |||||||||
| MBLLEN + GS | 15.08 | 0.7008 | 0.3458 | 15.36 | 0.3958 | 0.6608 | 18.09 | 0.7011 | 0.3656 |
| URetinex-Net + GS | 20.29 | 0.8249 | 0.2991 | 15.10 | 0.4012 | 0.6670 | 20.08 | 0.8679 | 0.3946 |
| NeRF-Based Methods | |||||||||
| Aleth-NeRF † | 19.56 | 0.7822 | 0.3113 | 12.65 | 0.4054 | 0.8985 | 16.02 | 0.7558 | 0.6574 |
| LLNeRF | 17.60 | 0.7497 | 0.3654 | 14.93 | 0.3446 | 0.7292 | 18.82 | 0.8597 | 0.3377 |
| I2-NeRF † | 22.40 | 0.7877 | 0.2789 | 15.37 | 0.3763 | 0.6493 | 21.91 | 0.6048 | 0.6463 |
| 3D Gaussian Splatting Methods | |||||||||
| Luminance-GS † | 17.98 | 0.7893 | 0.3110 | 9.78 | 0.3334 | 0.7953 | 12.49 | 0.2178 | 0.5291 |
| LITA-GS | 19.99 | 0.7988 | 0.3058 | 14.87 | 0.4381 | 0.7234 | 15.50 | 0.8515 | 0.3973 |
| ✦ NOVA-GS (Ours) | 20.97 | 0.7904 | 0.2467 | 15.54 | 0.4088 | 0.6136 | 23.49 | 0.8978 | 0.3325 |
† uses ground-truth images for luminance / pose alignment. ■ Best ■ 2nd ■ 3rd