Abstract
Remote sensing image change captioning (RSICC) aims to generate natural-language descriptions of changes between bi-temporal satellite images. Existing studies, however, predominantly operate on RGB imagery and only partially exploit the richer cues available in multispectral data. Motivated by this gap, we focus on multispectral RSICC on the Sentinel-2-based MOSAIC-SEN2-CC dataset and investigate how explicit category guidance can be used to obtain more accurate and scene-consistent change descriptions. We propose a Category-Guided Change Captioning (CGCC) framework with a shared ResNet-101 backbone and a Transformer-based caption decoder. Within this framework, a lightweight category-guided encoder uses an auxiliary head to predict the scene category from fused multispectral features and maps this prediction to a learnable category embedding. An epoch-dependent interpolation between ground-truth and prediction-based embeddings is used to condition cross-temporal fusion on the scene type. Experiments on MOSAIC-SEN2-CC show that the proposed method consistently improves BLEU, METEOR, ROUGE-L, CIDEr-D, SPICE and the composite score $S_m^{\ast}$ over strong ResNet-101-based baselines, particularly on changed scenes, while achieving high category classification accuracy across all semantic classes.
Keywords
Subject Areas
OpenAlex SDG Match
SDGs auto-classified by OpenAlex (score ≥ 0.4 shown).