REVIEW 4 cited by
DiffV2IR: Visible-to-Infrared Diffusion Model via Vision-Language Understanding
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
The task of translating visible-to-infrared images (V2IR) is inherently challenging due to three main obstacles: 1) achieving semantic-aware translation, 2) managing the diverse wavelength spectrum in infrared imagery, and 3) the scarcity of comprehensive infrared datasets. Current leading methods tend to treat V2IR as a conventional image-to-image synthesis challenge, often overlooking these specific issues. To address this, we introduce DiffV2IR, a novel framework for image translation comprising two key elements: a Progressive Learning Module (PLM) and a Vision-Language Understanding Module (VLUM). PLM features an adaptive diffusion model architecture that leverages multi-stage knowledge learning to infrared transition from full-range to target wavelength. To improve V2IR translation, VLUM incorporates unified Vision-Language Understanding. We also collected a large infrared dataset, IR-500K, which includes 500,000 infrared images compiled by various scenes and objects under various environmental conditions. Through the combination of PLM, VLUM, and the extensive IR-500K dataset, DiffV2IR markedly improves the performance of V2IR. Experiments validate DiffV2IR's excellence in producing high-quality translations, establishing its efficacy and broad applicability. The code, dataset, and DiffV2IR model will be available at https://github.com/LidongWang-26/DiffV2IR.
Forward citations
Cited by 4 Pith papers
-
SpectraDINO: Modality-Conditioned Adaptation of RGB Vision Foundation Models Across Infrared Bands
SpectraDINO extends DINOv2 with lightweight per-modality adapters and staged distillation to handle NIR, SWIR, and LWIR in one backbone, but its SWIR gain is weakened by using the evaluation dataset for pretraining.
-
FusionRS: A Large-Scale RGB-Infrared-Style Remote Sensing Dataset for Cross-Modal Vision-Language Learning
FusionRS pairs 600,000 remote sensing images with synthetic infrared-style copies and captions, and shows tri-modal contrastive training plus infrared-aware captions improves retrieval and captioning on the synthetic ...
-
MTVDiff: Multimodal Conditional Latent Diffusion for Enhanced Thermal-to-Visible Face Translation
A latent-diffusion system that adds depth maps and text descriptions as extra conditions improves thermal-to-visible face translation, cutting FID by up to 48.3% and raising Rank-1 accuracy by up to 8.9 percentage points.
-
MonoIR-RS: Infrared Remote Sensing Vision-Language Learning with CLIP and VLM Adaptation
MonoIR-RS synthesizes 600K infrared remote-sensing images from visible sources, rewrites captions to be IR-aware, and shows that IR-aware fine-tuning improves CLIP retrieval by up to 12.8 points and drives VLM infrare...
Discussion (0). Continue with ORCID to comment.