Pith. sign in

REVIEW 4 cited by

DiffV2IR: Visible-to-Infrared Diffusion Model via Vision-Language Understanding

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2503.19012 v1 pith:AH7J3LXF submitted 2025-03-24 cs.CV

classification cs.CV
keywords diffv2irinfraredv2irdatasetmodeltranslationunderstandingvision-language
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The task of translating visible-to-infrared images (V2IR) is inherently challenging due to three main obstacles: 1) achieving semantic-aware translation, 2) managing the diverse wavelength spectrum in infrared imagery, and 3) the scarcity of comprehensive infrared datasets. Current leading methods tend to treat V2IR as a conventional image-to-image synthesis challenge, often overlooking these specific issues. To address this, we introduce DiffV2IR, a novel framework for image translation comprising two key elements: a Progressive Learning Module (PLM) and a Vision-Language Understanding Module (VLUM). PLM features an adaptive diffusion model architecture that leverages multi-stage knowledge learning to infrared transition from full-range to target wavelength. To improve V2IR translation, VLUM incorporates unified Vision-Language Understanding. We also collected a large infrared dataset, IR-500K, which includes 500,000 infrared images compiled by various scenes and objects under various environmental conditions. Through the combination of PLM, VLUM, and the extensive IR-500K dataset, DiffV2IR markedly improves the performance of V2IR. Experiments validate DiffV2IR's excellence in producing high-quality translations, establishing its efficacy and broad applicability. The code, dataset, and DiffV2IR model will be available at https://github.com/LidongWang-26/DiffV2IR.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. SpectraDINO: Modality-Conditioned Adaptation of RGB Vision Foundation Models Across Infrared Bands

    cs.CV 2026-05 unverdicted novelty 7.0 of 10

    SpectraDINO extends DINOv2 with lightweight per-modality adapters and staged distillation to handle NIR, SWIR, and LWIR in one backbone, but its SWIR gain is weakened by using the evaluation dataset for pretraining.

  2. FusionRS: A Large-Scale RGB-Infrared-Style Remote Sensing Dataset for Cross-Modal Vision-Language Learning

    cs.CV 2026-06 conditional novelty 6.0 of 10

    FusionRS pairs 600,000 remote sensing images with synthetic infrared-style copies and captions, and shows tri-modal contrastive training plus infrared-aware captions improves retrieval and captioning on the synthetic ...

  3. MTVDiff: Multimodal Conditional Latent Diffusion for Enhanced Thermal-to-Visible Face Translation

    cs.CV 2026-07 conditional novelty 5.0 of 10

    A latent-diffusion system that adds depth maps and text descriptions as extra conditions improves thermal-to-visible face translation, cutting FID by up to 48.3% and raising Rank-1 accuracy by up to 8.9 percentage points.

  4. MonoIR-RS: Infrared Remote Sensing Vision-Language Learning with CLIP and VLM Adaptation

    cs.CV 2026-07 conditional novelty 5.0 of 10

    MonoIR-RS synthesizes 600K infrared remote-sensing images from visible sources, rewrites captions to be IR-aware, and shows that IR-aware fine-tuning improves CLIP retrieval by up to 12.8 points and drives VLM infrare...

Pith tools