Pith. sign in

REVIEW 2 cited by

Semantic-CC: Boosting Remote Sensing Image Change Captioning via Foundational Knowledge and Semantic Guidance

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2407.14032 v1 pith:TCQFMXE4 submitted 2024-07-19 cs.CV

classification cs.CV
keywords changesemanticsemantic-cccaptioningdescriptionsguidanceknowledgeremote
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Remote sensing image change captioning (RSICC) aims to articulate the changes in objects of interest within bi-temporal remote sensing images using natural language. Given the limitations of current RSICC methods in expressing general features across multi-temporal and spatial scenarios, and their deficiency in providing granular, robust, and precise change descriptions, we introduce a novel change captioning (CC) method based on the foundational knowledge and semantic guidance, which we term Semantic-CC. Semantic-CC alleviates the dependency of high-generalization algorithms on extensive annotations by harnessing the latent knowledge of foundation models, and it generates more comprehensive and accurate change descriptions guided by pixel-level semantics from change detection (CD). Specifically, we propose a bi-temporal SAM-based encoder for dual-image feature extraction; a multi-task semantic aggregation neck for facilitating information interaction between heterogeneous tasks; a straightforward multi-scale change detection decoder to provide pixel-level semantic guidance; and a change caption decoder based on the large language model (LLM) to generate change description sentences. Moreover, to ensure the stability of the joint training of CD and CC, we propose a three-stage training strategy that supervises different tasks at various stages. We validate the proposed method on the LEVIR-CC and LEVIR-CD datasets. The experimental results corroborate the complementarity of CD and CC, demonstrating that Semantic-CC can generate more accurate change descriptions and achieve optimal performance across both tasks.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Robust Change Captioning in Remote Sensing: SECOND-CC Dataset and MModalCC Framework

    cs.CV 2025-01 conditional novelty 6.0 of 10

    A new dataset and a multimodal attention model improve remote sensing change captioning, but only when ground-truth semantic maps are supplied as input.

  2. Semantic-CD: Remote Sensing Image Semantic Change Detection towards Open-vocabulary Setting

    cs.CV 2025-01 conditional novelty 5.0 of 10

    Semantic-CD couples binary and semantic change detection in remote sensing imagery via CLIP text-guided cost volumes, reporting state-of-the-art scores on the SECOND dataset.

Pith tools