JL1-CC&QA: Extending the JL1-CD Benchmark with Change Captioning and Question Answering


Abstract

Remote sensing change detection (CD) traditionally focuses on pixel-level binary segmentation, which identifies where changes occur but neither what nor why. To bridge this semantic gap, we introduce JL1-CC&QA, a multi-task benchmark that extends the JL1-CD dataset with two complementary annotation layers: change captioning (CC) and change question answering (QA). Built upon 5,000 bi-temporal image pairs acquired by the Jilin-1 satellite at 0.5–0.75 m ground sample distance, the benchmark comprises: (i) JL1-CC, providing 17,021 quality-verified captions that describe diverse land-cover transformations; and (ii) JL1-QA, offering 20,060 question–answer pairs across eight question types, enabling fine-grained, interactive interrogation of surface changes. All annotations are produced via a three-stage pipeline consisting of multi-modal large language model (LLM) generation, vision-grounded LLM judging, and human expert verification. We hope that JL1-CC&QA, as a benchmark unifying binary change masks, change captions, and change-oriented QA over the same image set, will serve as a valuable resource for the community to advance multi-task change understanding in remote sensing. The dataset is available at https://github.com/circleLZY/JL1-CD.

Benchmark, change captioning, change question answering, change detection, remote sensing.

1 Introduction↩︎

Remote sensing change detection (CD) aims to identify surface changes from multi-temporal imagery acquired over the same geographic area, serving as an important technology for urban planning, environmental monitoring, disaster response, and resource management [1], [2]. Over the past decade, the rapid growth of both Earth observation data and deep learning methods has propelled CD into one of the most active research frontiers in remote sensing.

The predominant formulation of CD is binary change detection (BCD), which classifies each pixel as either changed or unchanged. A large number of BCD benchmarks have been established, spanning diverse geographic contexts, spatial resolutions, and change categories (see Table ¿tbl:tab:datasets? for a comprehensive summary) [3][19]. While the vast majority of these benchmarks adopt the bi-temporal setting and have driven extensive algorithmic development [10], [11], [20][24], the community has also explored single-temporal CD that exploits pretrained semantic representations to infer changes from unpaired imagery [25][27], as well as multi-temporal monitoring that tracks continuous change trajectories from dense satellite time series [28][31]. Despite these differences, all of the above approaches produce pixel-level masks that encode where change occurred but remain silent on what changed or why.

max width= Abbreviations: Opt.= Optical RGB; MS = Multispectral; HS = Hyperspectral; Cross = Cross-modal; Multi = Multi-modal fusion; GE = Google Earth; GF-2 = Gaofen-2; JL-1 = Jilin-1; S1/S2 = Sentinel-1/2; PV = Photovoltaic; Bi = Bi-temporal; Bldg.= Building(s); Veg.= Vegetation; VHR = Very High Resolution (\(<\)1 m). Dataset names are hyperlinked to download pages where available.

In parallel, the range of input modalities has expanded considerably. While optical RGB imagery remains the dominant data source for CD research [7], [10], [13], [15], [19], multispectral and hyperspectral sensors offer richer spectral discrimination for fine-grained change analysis [5], [8], [9], [32][35], and cross-modal fusion of optical and SAR data enables all-weather, day-and-night monitoring [14], [17], [36][40]. Although these modalities provide richer visual information, the output of BCD remains a pixel-level binary mask devoid of semantic content.

Semantic change detection (SCD) and building damage assessment (BDA) partially bridge this gap by assigning categorical labels to changed pixels. SCD introduces per-pixel land-cover annotations at both time phases, enabling identification of change types (e.g., farmland\(\to\)buildings) [30], [41][50]. BDA further introduces ordinal damage scales (e.g., minor/major/destroyed) for disaster response [51][57]. Nevertheless, the semantic content in these datasets is still encoded as discrete numerical class indices drawn from a closed taxonomy. This gap calls for a shift toward natural-language-grounded change understanding.

The advent of vision–language models (VLMs) has begun to close this gap. In the natural image domain, change captioning benchmarks first demonstrated the feasibility of describing visual differences in natural language [58], [59]. The remote sensing community quickly adopted this paradigm, producing a growing family of change captioning (CC) datasets that range from low-resolution multitemporal pairs to very-high-resolution urban and disaster scenes [60][66]. In parallel, change detection visual question answering (CDVQA) has emerged as a complementary task, enabling users to query specific aspects of surface changes through natural-language questions [67], [68], while instruction-tuning datasets have further extended the interaction to multi-turn conversational analysis [69][71]. Despite this rapid progress, rare benchmarks jointly provide binary change masks, change captions, and change-oriented question–answer pairs on the same set of image pairs, which limits multi-task learning and cross-task evaluation.

Figure 1: Timeline of the development of mainstream deep learning-based CD methods.

The evolution of datasets has been accompanied by corresponding advances in model architectures (see Fig. 1). Methods built on CNN backbones with change decoders [10], [11], [20], [21], [72][93] established the dominant Siamese paradigm, while Transformer [22], [23], [33], [94][104] introduced global cross-temporal attention, and Mamba-based methods [24], [40], [105][111] achieved comparable perception at linear computational cost. Foundation model adaptations have enabled zero-shot and few-shot CD by transferring pretrained CLIP and SAM representations [112][120]. Generative approaches based on GANs [37], [121][123] and diffusion models [124][130] have addressed data scarcity through synthetic bi-temporal image generation. Agent-based systems have integrated LLMs as reasoning engines for interactive change analysis [131][134]. A critical observation is that the text modality has been a key enabler at each stage: CLIP-based models use text prompts to guide change semantics, captioning models translate visual change into language, VQA models support interactive querying, and agents conduct multi-step reasoning in natural language. This trajectory underscores the need for language-grounded benchmarks that can support the training and evaluation of these increasingly capable architectures.

Motivated by the above analysis, we present JL1-CC&QA, extending the JL1-CD benchmark [19], which consists of 5,000 bi-temporal image pairs from the Jilin-1 satellite with a resolution of 0.5–0.75 m, with two new annotation layers: (i) JL1-CC, providing 17,021 quality-verified change captions spanning both anthropogenic and natural changes; and (ii) JL1-QA, offering 20,060 question–answer pairs across eight question types (existence, description, location, magnitude, temporal comparison, causation, relative comparison, and visual detail). All annotations are produced via a three-stage pipeline: multi-modal LLM generation, vision-grounded LLM judging, and human expert verification. Our main contributions are summarized as follows:

  1. We construct a change understanding benchmark that provides binary change masks, natural-language change descriptions, and diverse QA pairs over 5,000 bi-temporal satellite image pairs, facilitating joint training and evaluation across CD, CC, and QA tasks.

  2. We design a scalable, reproducible annotation pipeline that couples multi-modal LLM generation with vision-grounded LLM judging and human verification.

  3. We provide comprehensive dataset statistics, and release all data and code to support community research toward unified remote sensing change understanding.

2 Dataset Construction↩︎

This section describes the construction of JL1-CC&QA. We first introduce the source dataset JL1-CD (Section 2.1), then detail the change captioning pipeline (Section 2.2) and the question answering pipeline (Section 2.3).

2.1 JL1-CD↩︎

JL1-CC&QA is built upon JL1-CD [19], which comprises 5,000 bi-temporal image pairs captured by the Jilin-1 high-resolution optical satellite. The imagery was acquired across multiple provinces in China, including Shandong, Ningxia, Anhui, Hebei, and Hunan, between early 2022 and late 2023, and was carefully curated to exclude blur, cloud occlusion, and extreme illumination conditions. Each image pair consists of a pre-event image, a post-event image, and a pixel-level binary change mask annotated by professional interpreters. Key specifications are summarized in Table 1.

Table 1: Specifications of the JL1-CD Source Dataset
Attribute Value
Satellite Jilin-1 (JL1) Constellation
Spatial resolution 0.5–0.75 m GSD
Image size \(512 \times 512\) pixels
Spectral bands RGB (3 channels)
Total image pairs 5,000
Training / test split 4,000 / 1,000
Annotation Pixel-level binary mask
Temporal coverage 2022–2023
Geographic coverage Multiple provinces in China

Two properties of JL1-CD make it suitable for change captioning and question answering. First, the dataset is all-inclusive in terms of change types: it encompasses both anthropogenic changes (buildings, roads, hardened surfaces, photovoltaic panels, etc.) and natural changes (woodlands, grasslands, croplands, water bodies, etc.). This diversity ensures that the derived CC and QA annotations cover a broad spectrum of real-world surface dynamics. Second, the change area ratio (CAR)—defined as the proportion of changed pixels in each image pair—spans the full range from near-zero to 100%, with a mean of 9.7% and a median of 2.9% (Fig. 2). This long-tailed distribution mirrors the natural imbalance of real-world change scenarios and poses a meaningful challenge for language-grounded change understanding at varying scales.

Figure 2: Distribution of change area ratio (CAR) across 5,000 image pairs in JL1-CD.

2.2 JL1-CC↩︎

2.2.1 Task Definition↩︎

Given a bi-temporal image pair \((I_A, I_B)\) and its corresponding binary change mask \(M\), the change captioning task requires generating a set of natural-language sentences \(\{c_1, c_2, \ldots, c_k\}\) that accurately describe the semantic content of the observed surface changes, including the type, location, and extent of change.

2.2.2 Annotation Pipeline↩︎

As illustrated in Fig. 3, the JL1-CC annotation pipeline consists of three stages.

Figure 3: Overview of the JL1-CC annotation pipeline. Stage 1: a multi-modal LLM generates five candidate captions from the image pair and spatial metadata. Stage 2: a vision-grounded LLM judge scores each caption and retains the top three. Stage 3: human experts verify factual accuracy.

Stage 1: Multi-modal LLM Generation. For each image pair, we prompt a multi-modal LLM (Kimi-K2.6) with three visual inputs: the pre-event image \(I_A\), the post-event image \(I_B\), and the binary change mask \(M\), together with spatial metadata including the CAR and a textual description of the primary change region (e.g., “upper-left area of the image”). The model is instructed to generate five diverse captions, each describing the same change from a different perspective: change type, spatial location, visual appearance, scale, or implication.

Stage 2: Vision-Grounded LLM Judging. A second LLM call evaluates each caption by examining the original image pair alongside the generated text. Each caption is scored on a 1–10 integer scale across five criteria: (1) accuracy: whether the description matches the visible changes, with heavy penalties for hallucination; (2) specificity: whether concrete land-cover terms are used; (3) spatial correctness: whether location references are accurate; (4) naturalness: whether the English is fluent; and (5) informativeness: whether the caption conveys meaningful detail. The top-3 scoring captions are retained, and ties at the cutoff are preserved.

Stage 3: Human Expert Verification. A subset of the generated captions is reviewed by domain experts to verify factual accuracy and identify systematic errors, ensuring that the automated pipeline produces reliable annotations at scale.

2.2.3 Statistics↩︎

Table 2 summarizes the JL1-CC dataset statistics. The pipeline generates 25,000 candidate captions (5 per pair) and retains 17,021 after quality filtering, yielding a pass rate of 68.1%. The judge score distribution (Fig. 4) shows clear discrimination: 39.8% of captions score 9–10, 42.3% score 7–8, and 17.9% score below 7 and are rejected. The selected captions have a mean length of 26.2 words, with a vocabulary of 7,458 unique tokens. Fig. 5 presents a word cloud of the selected captions, where the most prominent terms—upper, lower, bare soil, agricultural, building, road, water—reflect both the spatial referencing style and the diversity of land-cover types in JL1-CD.

Figure 4: Judge score distribution for JL1-CC. Captions scoring below 7 (red) are rejected; those scoring 7–8 (green) and 9–10 (blue) are retained.
Figure 5: Word cloud of the 17,021 selected captions in JL1-CC.
Table 2: JL1-CC Dataset Statistics
Statistic Train Test Total
Image pairs 4,000 1,000 5,000
Generated captions 20,000 5,000 25,000
Selected captions 13,616 3,405 17,021
Avg.selected / pair 3.40 3.40 3.40
Avg.caption length (words) 26.2 26.4 26.2
Vocabulary size 6,837 4,064 7,458
Judge score (mean / median) 7.8 / 8.0 7.8 / 8.0

2.3 JL1-QA↩︎

2.3.1 Task Definition↩︎

Given a bi-temporal image pair \((I_A, I_B)\) and a natural-language question \(Q\), the change question answering task requires generating a textual answer \(A\) that accurately responds to the question based on the visible surface changes.

2.3.2 Question Taxonomy↩︎

We define eight question types adapted for open-ended answer generation, as summarized in Table 3. Each type targets a distinct aspect of change understanding, ranging from binary existence judgments to causal reasoning, ensuring comprehensive coverage of user information needs.

Table 3: Question Types in JL1-QA with Examples
Type Example Question
YES/NO Has the vegetation in the lower half been removed?
WHAT What happened to the farmland in the center?
WHERE Where did the most significant change occur?
HOW MUCH Is the change large-scale or localized?
BEFORE/AFTER What was present in the upper-left before the change?
CAUSE What type of development likely caused the changes?
DETAIL What do the new structures appear to be?
COMPARE Which area shows the most dramatic change?

2.3.3 Annotation Pipeline↩︎

As illustrated in Fig. 6, the JL1-QA annotation pipeline shares the same three-stage architecture as JL1-CC, with two key enhancements in the generation stage.

Figure 6: Overview of the JL1-QA annotation pipeline. The generation stage receives the image pair, change mask, and JL1-CC metadata (captions, CAR, change region) as context. The judge evaluates each QA pair for accuracy, quality, completeness, and redundancy.

Stage 1: Context-Enriched QA Generation. Beyond the three visual inputs (\(I_A\), \(I_B\), \(M\)), the LLM additionally receives contextual metadata from JL1-CC: the change area ratio, the spatial change region description, and up to three selected change captions. This context enrichment grounds the QA generation in verified change descriptions, yielding more accurate and diverse question–answer pairs. The model generates five QA pairs per image, each covering a different question type randomly sampled from the taxonomy. Crucially, the prompt explicitly prohibits copying exact numerical metadata (e.g., change percentages) into answers, as such precision cannot be derived from visual inspection alone.

Stage 2: Multi-Criteria Judging. Each QA pair is scored on a 1–10 integer scale across four criteria: (1) answer accuracy: whether the answer is factually correct given the images; (2) question quality: whether the question is clear, natural, and non-trivial; (3) answer completeness: whether the answer adequately addresses the question; and (4) redundancy: whether the QA pair is substantially different from the others for the same image. QA pairs scoring below 7 are discarded. The most common rejection reasons are hallucinated precise percentages (score 1–4), redundancy with other QA pairs (score 5–6), and vague answers (score 5–6).

Stage 3: Human Expert Verification. As with JL1-CC, domain experts review a subset of selected QA pairs to verify factual accuracy and check for systematic errors in question formulation or answer content.

2.3.4 Statistics↩︎

Table 4 summarizes the JL1-QA dataset statistics. From 24,995 generated QA pairs, 20,060 pass the quality threshold (score \(\geq\) 7), yielding a pass rate of 80.3%. The higher pass rate compared to JL1-CC (68.1%) is attributed to the context enrichment from change captions, which reduces hallucination in the generation stage. The judge score distribution (Fig. 7) confirms effective quality discrimination: 51.4% of QA pairs score 9–10, 28.4% score 7–8, and 20.2% are rejected. The question type distribution (Fig. 8) shows broad coverage across all eight categories, with YES/NO (21.2%), WHERE (18.3%), and WHAT (16.9%) being the most frequent. Questions average 11.1 words and answers 19.3 words in length.

Figure 7: Judge score distribution for JL1-QA. QA pairs scoring below 7 are rejected.
Figure 8: Distribution of question types in the 20,060 selected QA pairs of JL1-QA.
Table 4: JL1-QA Dataset Statistics
Statistic Train Test Total
Image pairs 3,999 1,000 4,999
Generated QA pairs 19,995 5,000 24,995
Selected QA pairs 16,055 4,005 20,060
Avg.selected / pair 4.01 4.01 4.01
Avg.question length (words) 11.1 11.1 11.1
Avg.answer length (words) 19.3 19.3 19.3
Judge score (mean / median) 7.9 / 9.0 7.9 / 9.0

3 Conclusion↩︎

In this paper, we presented JL1-CC&QA, a multi-task benchmark that extends the JL1-CD binary change detection dataset with two complementary annotation layers: change captioning (JL1-CC) and change question answering (JL1-QA). Both layers are produced through a three-stage pipeline—multi-modal LLM generation, vision-grounded LLM judging, and human expert verification—balancing scalability with factual reliability. We hope this benchmark serves as a valuable resource for the community to advance multi-task change understanding in remote sensing.

References↩︎

[1]
Z. Lv, T. Liu, J. A. Benediktsson, and N. Falco, “Land cover change detection techniques: Very-high-resolution optical images: A review,” IEEE Geoscience and Remote Sensing Magazine, vol. 10, no. 1, pp. 44–63, 2022, doi: 10.1109/MGRS.2021.3088865.
[2]
T. Bai, L. Wang, D. Yin, K. Sun, Y. Chen, and W. Li, “Deep learning for change detection in remote sensing: A review,” Geo-spatial Information Science, vol. 26, no. 3, pp. 262–288, 2023, doi: 10.1080/10095020.2022.2085633.
[3]
C. Benedek and T. Szirányi, “Change detection in optical aerial images by a multilayer conditional mixed Markov model,” IEEE Transactions on Geoscience and Remote Sensing, vol. 47, no. 10, pp. 3416–3430, 2009, doi: 10.1109/TGRS.2009.2022633.
[4]
N. Bourdis, D. Marraud, and H. Sahbi, “Constrained optical flow for aerial image change detection,” in 2011 IEEE international geoscience and remote sensing symposium, 2011, pp. 4176–4179, doi: 10.1109/IGARSS.2011.6050150.
[5]
R. C. Daudt, B. Le Saux, A. Boulch, and Y. Gousseau, “Urban change detection for multispectral earth observation using convolutional neural networks,” in IGARSS 2018 – 2018 IEEE international geoscience and remote sensing symposium, 2018, pp. 2115–2118, doi: 10.1109/IGARSS.2018.8518015.
[6]
M. A. Lebedev, Yu. V. Vizilter, O. V. Vygolov, V. A. Knyaz, and A. Yu. Rubis, “Change detection in remote sensing images using conditional adversarial networks,” The International Archives of the Photogrammetry, Remote Sensing and Spatial Information Sciences, vol. XLII–2, pp. 565–571, 2018, doi: 10.5194/isprs-archives-XLII-2-565-2018.
[7]
S. Ji, S. Wei, and M. Lu, “Fully convolutional networks for multisource building extraction from an open aerial and satellite imagery data set,” IEEE Transactions on Geoscience and Remote Sensing, vol. 57, no. 1, pp. 574–586, 2019, doi: 10.1109/TGRS.2018.2858817.
[8]
Q. Wang, Z. Yuan, Q. Du, and X. Li, GETNET: A general end-to-end 2-D CNN framework for hyperspectral image change detection,” IEEE Transactions on Geoscience and Remote Sensing, vol. 57, no. 1, pp. 3–13, 2019, doi: 10.1109/TGRS.2018.2849692.
[9]
J. López-Fandiño, D. B. Heras, F. Argüello, and M. Dalla Mura, GPU framework for change detection in multitemporal hyperspectral images,” International Journal of Parallel Programming, vol. 47, no. 2, pp. 272–292, 2019, doi: 10.1007/s10766-017-0547-5.
[10]
H. Chen and Z. Shi, “A spatial-temporal attention-based method and a new dataset for remote sensing image change detection,” Remote Sensing, vol. 12, no. 10, p. 1662, 2020, doi: 10.3390/rs12101662.
[11]
C. Zhang et al., “A deeply supervised image fusion network for change detection in high resolution bi-temporal remote sensing images,” ISPRS Journal of Photogrammetry and Remote Sensing, vol. 166, pp. 183–200, 2020, doi: 10.1016/j.isprsjprs.2020.06.003.
[12]
D. Peng, L. Bruzzone, Y. Zhang, H. Guan, H. Ding, and X. Huang, SemiCDNet: A semisupervised convolutional neural network for change detection in high resolution remote-sensing images,” IEEE Transactions on Geoscience and Remote Sensing, vol. 59, no. 7, pp. 5891–5906, 2021, doi: 10.1109/TGRS.2020.3011913.
[13]
L. Shen et al., S2Looking: A satellite side-looking dataset for building change detection,” Remote Sensing, vol. 13, no. 24, p. 5094, 2021, doi: 10.3390/rs13245094.
[14]
R. Shao, C. Du, H. Chen, and J. Li, SUNet: Change detection for heterogeneous remote sensing images from satellite and UAV using a dual-channel fully convolution network,” Remote Sensing, vol. 13, no. 18, p. 3750, 2021, doi: 10.3390/rs13183750.
[15]
Q. Shi, M. Liu, S. Li, X. Liu, F. Wang, and L. Zhang, “A deeply supervised attention metric-based network and an open aerial image dataset for remote sensing change detection,” IEEE Transactions on Geoscience and Remote Sensing, vol. 60, pp. 1–16, 2022, doi: 10.1109/TGRS.2021.3085870.
[16]
M. Liu, Z. Chai, H. Deng, and R. Liu, “A CNN-Transformer network with multiscale context aggregation for fine-grained cropland change detection,” IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing, vol. 15, pp. 4297–4306, 2022, doi: 10.1109/JSTARS.2022.3177235.
[17]
H. Li, F. Zhu, X. Zheng, M. Liu, and G. Chen, MSCDUNet: A deep learning framework for built-up area change detection integrating multispectral, SAR, and VHR data,” IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing, vol. 15, pp. 5163–5176, 2022, doi: 10.1109/JSTARS.2022.3181155.
[18]
S. Holail, T. Saleh, X. Xiao, and D. Li, AFDE-Net: Building change detection using attention-based feature differential enhancement for satellite imagery,” IEEE Geoscience and Remote Sensing Letters, vol. 20, pp. 1–5, 2023, doi: 10.1109/LGRS.2023.3283505.
[19]
Z. Liu, R. Zhu, L. Gao, Y. Zhou, J. Ma, and Y. Gu, JL1-CD: A new benchmark for remote sensing change detection and a robust multi-teacher knowledge distillation framework.” arXiv preprint arXiv:2502.13407, 2025, doi: 10.48550/arXiv.2502.13407.
[20]
R. C. Daudt, B. Le Saux, and A. Boulch, “Fully convolutional siamese networks for change detection,” in 2018 25th IEEE international conference on image processing (ICIP), 2018, pp. 4063–4067, doi: 10.1109/ICIP.2018.8451652.
[21]
S. Fang, K. Li, J. Shao, and Z. Li, SNUNet-CD: A densely connected siamese network for change detection of VHR images,” IEEE Geoscience and Remote Sensing Letters, vol. 19, pp. 1–5, 2022, doi: 10.1109/LGRS.2021.3056416.
[22]
H. Chen, Z. Qi, and Z. Shi, “Remote sensing image change detection with transformers,” IEEE Transactions on Geoscience and Remote Sensing, vol. 60, pp. 1–14, 2022, doi: 10.1109/TGRS.2021.3095166.
[23]
W. G. C. Bandara and V. M. Patel, “A transformer-based siamese network for change detection,” in IGARSS 2022 – 2022 IEEE international geoscience and remote sensing symposium, 2022, pp. 207–210, doi: 10.1109/IGARSS46834.2022.9883686.
[24]
H. Chen, J. Song, C. Han, J. Xia, and N. Yokoya, ChangeMamba: Remote sensing change detection with spatiotemporal state space model,” IEEE Transactions on Geoscience and Remote Sensing, vol. 62, pp. 1–20, 2024, doi: 10.1109/TGRS.2024.3417253.
[25]
Z. Zheng, A. Ma, L. Zhang, and Y. Zhong, “Change is everywhere: Single-temporal supervised object change detection in remote sensing imagery,” in IEEE/CVF international conference on computer vision (ICCV), 2021, pp. 15193–15202, doi: 10.1109/ICCV48922.2021.01491.
[26]
H. Chen, J. Song, C. Wu, B. Du, and N. Yokoya, “Exchange means change: An unsupervised single-temporal change detection framework based on intra- and inter-image patch exchange,” ISPRS Journal of Photogrammetry and Remote Sensing, vol. 206, pp. 87–105, 2023, doi: 10.1016/j.isprsjprs.2023.11.004.
[27]
T. Zhou, F. Luo, C. Fu, T. Guo, X. Wang, and B. Du, STMNet: Single-temporal mask-based network for self-supervised hyperspectral change detection,” IEEE Transactions on Geoscience and Remote Sensing, vol. 63, pp. 1–12, 2024, doi: 10.1109/TGRS.2024.3523541.
[28]
A. Toker, L. Kondmann, M. Weber, M. Eisenberger, A. Camero, and J. Hu, DynamicEarthNet: Daily multi-spectral satellite dataset for semantic change segmentation,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition (CVPR), 2022, pp. 21126–21135, doi: 10.1109/CVPR52688.2022.02048.
[29]
A. Van Etten, D. Hogan, J. M. Manso, J. Shermeyer, N. Weir, and R. Lewis, “The multi-temporal urban development SpaceNet dataset,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition (CVPR), 2021, pp. 6394–6403, doi: 10.1109/CVPR46437.2021.00633.
[30]
S. Shi et al., “Multi-temporal urban semantic understanding based on GF-2 remote sensing imagery: From tri-temporal datasets to multi-task mapping,” International Journal of Digital Earth, vol. 16, no. 1, pp. 3321–3347, 2023, doi: 10.1080/17538947.2023.2246445.
[31]
Y. Zhao, H.-C. Li, S. Lei, N. Liu, J. Pan, and T. Celik, COUD: Continual urbanization detector for time series building change detection,” IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing, vol. 17, pp. 19601–19615, 2024, doi: 10.1109/JSTARS.2024.3482559.
[32]
M. Hu, C. Wu, L. Zhang, and B. Du, “Hyperspectral anomaly change detection based on autoencoder,” IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing, vol. 14, pp. 3750–3762, 2021, doi: 10.1109/JSTARS.2021.3066508.
[33]
Y. Wang et al., “Spectral-spatial-temporal transformers for hyperspectral image change detection,” IEEE Transactions on Geoscience and Remote Sensing, vol. 60, pp. 1–14, 2022, doi: 10.1109/TGRS.2022.3203075.
[34]
W. Xie, X. Xu, and Y. Li, “Decentralized federated GAN for hyperspectral change detection in edge computing,” IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing, vol. 17, pp. 8863–8874, 2024, doi: 10.1109/JSTARS.2024.3389641.
[35]
W. Dong, J. Ren, S. Xiao, L. Fang, J. Qu, and Y. Li, “Cycle translation-based collaborative training for hyperspectral-RGB multimodal change detection,” IEEE Transactions on Image Processing, vol. 34, pp. 6347–6360, 2025, doi: 10.1109/TIP.2025.3607609.
[36]
X. He, S. Zhang, B. Xue, T. Zhao, and T. Wu, “Cross-modal change detection flood extraction based on convolutional neural network,” International Journal of Applied Earth Observation and Geoinformation, vol. 117, p. 103197, 2023, doi: 10.1016/j.jag.2023.103197.
[37]
X. Li, Z. Du, Y. Huang, and Z. Tan, “A deep translation (GAN) based change detection network for optical and SAR remote sensing images,” ISPRS Journal of Photogrammetry and Remote Sensing, vol. 179, pp. 14–34, 2021, doi: 10.1016/j.isprsjprs.2021.07.007.
[38]
J. Alatalo, T. Sipola, and M. Rantonen, “Improved difference images for change detection classifiers in SAR imagery using deep learning,” IEEE Transactions on Geoscience and Remote Sensing, vol. 61, pp. 1–14, 2023, doi: 10.1109/TGRS.2023.3324994.
[39]
Z. Liu, J. Zhang, W. Wang, and Y. Gu, M2CD: A unified MultiModal framework for optical-SAR change detection with mixture of experts and self-distillation,” IEEE Geoscience and Remote Sensing Letters, vol. 22, pp. 1–5, 2025, doi: 10.1109/LGRS.2025.3590959.
[40]
Y. Shen, S. Yao, Z. Qiang, and G. Pei, SD-Mamba: A lightweight synthetic-decompression network for cross-modal flood change detection,” International Journal of Applied Earth Observation and Geoinformation, vol. 136, p. 104409, 2025, doi: 10.1016/j.jag.2025.104409.
[41]
R. C. Daudt, B. Le Saux, A. Boulch, and Y. Gousseau, “Multitask learning for large-scale semantic change detection,” Computer Vision and Image Understanding, vol. 187, p. 102783, 2019, doi: 10.1016/j.cviu.2019.07.003.
[42]
K. Yang, G.-S. Xia, Z. Liu, B. Du, Y. Wen, and M. Pelillo, “Asymmetric siamese networks for semantic change detection in aerial images,” IEEE Transactions on Geoscience and Remote Sensing, vol. 60, pp. 1–18, 2022, doi: 10.1109/TGRS.2021.3113912.
[43]
S. Tian, A. Ma, Z. Zheng, Y. Zhong, X. Tan, and L. Zhang, “Large-scale deep learning based binary and semantic change detection in ultra high resolution remote sensing imagery: From benchmark datasets to urban application,” ISPRS Journal of Photogrammetry and Remote Sensing, vol. 193, pp. 164–186, 2022, doi: 10.1016/j.isprsjprs.2022.08.012.
[44]
L. Ding, H. Guo, S. Liu, L. Mou, J. Zhang, and L. Bruzzone, “Bi-temporal semantic reasoning for the semantic change detection in HR remote sensing images,” IEEE Transactions on Geoscience and Remote Sensing, vol. 60, pp. 1–14, 2022, doi: 10.1109/TGRS.2022.3154390.
[45]
P. Yuan, Q. Zhao, X. Zhao, X. Wang, X. Long, and Y. Zheng, “A transformer-based siamese network and an open optical dataset for semantic change detection of remote sensing images,” International Journal of Digital Earth, vol. 15, no. 1, pp. 1506–1525, 2022, doi: 10.1080/17538947.2022.2111470.
[46]
C. Pang, J. Wu, J. Ding, C. Song, and G.-S. Xia, “Detecting building changes with off-nadir aerial images,” Science China Information Sciences, vol. 66, no. 4, p. 140306, 2023, doi: 10.1007/s11432-022-3691-4.
[47]
K. Tang, F. Xu, X. Chen, Q. Dong, Y. Yuan, and J. Chen, “The ClearSCD model: Comprehensively leveraging semantics and change relationships for semantic change detection in high spatial resolution remote sensing imagery,” ISPRS Journal of Photogrammetry and Remote Sensing, vol. 211, pp. 299–317, 2024, doi: 10.1016/j.isprsjprs.2024.04.013.
[48]
M. Liu, S. Lin, Y. Zhong, Q. Shi, and J. Li, “A memory-guided network and a novel dataset for cropland semantic change detection,” IEEE Transactions on Geoscience and Remote Sensing, vol. 62, pp. 1–13, 2024, doi: 10.1109/TGRS.2024.3421654.
[49]
X. Tan et al., TripleS: Mitigating multi-task learning conflicts for semantic change detection in high-resolution remote sensing imagery,” ISPRS Journal of Photogrammetry and Remote Sensing, vol. 230, pp. 374–401, 2025, doi: 10.1016/j.isprsjprs.2025.09.019.
[50]
S. Fang, W. Li, Y. Song, Z. Li, and J. Zhao, “Rethinking semantic change detection from a semantic alignment perspective.” Authorea Preprints, 2025, doi: 10.36227/techrxiv.174016507.73028575/v1.
[51]
A. Fujita, K. Sakurada, T. Imaizumi, R. Ito, S. Hikosaka, and R. Nakamura, “Damage detection from aerial images via convolutional neural networks,” in 2017 fifteenth IAPR international conference on machine vision applications (MVA), 2017, pp. 5–8, doi: 10.23919/MVA.2017.7986759.
[52]
R. Gupta et al., “Creating xBD: A dataset for assessing building damage from satellite imagery,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition workshops (CVPRW), 2019, pp. 10–17, doi: 10.1109/CVPRW.2019.00047.
[53]
Z. Zheng, Y. Zhong, J. Wang, A. Ma, and L. Zhang, “Building damage assessment for rapid disaster response with a deep object-based semantic change detection framework: From natural disasters to man-made disasters,” Remote Sensing of Environment, vol. 265, p. 112636, 2021, doi: 10.1016/j.rse.2021.112636.
[54]
M. Rahnemoonfar, T. Chowdhury, and R. Murphy, RescueNet: A high resolution UAV semantic segmentation dataset for natural disaster damage assessment,” Scientific Data, vol. 10, no. 1, p. 913, 2023, doi: 10.1038/s41597-023-02799-4.
[55]
Y. Sun, Y. Wang, and M. Eineder, QuickQuakeBuildings: Post-earthquake SAR-optical dataset for quick damaged-building detection,” IEEE Geoscience and Remote Sensing Letters, vol. 21, pp. 1–5, 2024, doi: 10.1109/LGRS.2024.3406966.
[56]
H. Chen et al., BRIGHT: A globally distributed multimodal building damage assessment dataset with very-high-resolution for all-weather disaster response,” Earth System Science Data Discussions, pp. 1–51, 2025, doi: 10.5194/essd-2025-269.
[57]
H. Wang, W. He, Z. Li, and N. Yokoya, “Cross-scenario damaged building extraction network: Methodology, application, and efficiency using single-temporal HRRS imagery,” ISPRS Journal of Photogrammetry and Remote Sensing, vol. 228, pp. 228–248, 2025, doi: 10.1016/j.isprsjprs.2025.06.028.
[58]
H. Jhamtani and T. Berg-Kirkpatrick, “Learning to describe differences between pairs of similar images,” in Proceedings of the 2018 conference on empirical methods in natural language processing (EMNLP), 2018, pp. 4024–4034, doi: 10.18653/v1/D18-1436.
[59]
D. H. Park, T. Darrell, and A. Rohrbach, “Robust change captioning,” in Proceedings of the IEEE/CVF international conference on computer vision (ICCV), 2019, pp. 4624–4633, doi: 10.1109/ICCV.2019.00472.
[60]
G. Hoxha, S. Chouaf, F. Melgani, and Y. Smara, “Change captioning: A new paradigm for multitemporal remote sensing image analysis,” IEEE Transactions on Geoscience and Remote Sensing, vol. 60, pp. 1–14, 2022, doi: 10.1109/TGRS.2022.3163753.
[61]
C. Liu, R. Zhao, H. Chen, Z. Zou, and Z. Shi, “Remote sensing image change captioning with dual-branch transformers: A new method and a large scale dataset,” IEEE Transactions on Geoscience and Remote Sensing, vol. 60, pp. 1–20, 2022, doi: 10.1109/TGRS.2022.3218921.
[62]
C. Liu, K. Chen, B. Zhang, Z. Qi, Z. Zou, and Z. Shi, “Change detection as visual question answering,” IEEE Transactions on Geoscience and Remote Sensing, vol. 62, pp. 1–18, 2024, doi: 10.1109/TGRS.2024.3391584.
[63]
X. Li et al., SECOND-CC: A large-scale change captioning dataset with real-world challenges.” arXiv preprint arXiv:2501.10075, 2025, doi: 10.48550/arXiv.2501.10075.
[64]
Z. Chen et al., CCExpert: Advancing RSCC via an expert-level large vision-language model.” arXiv preprint arXiv:2411.11360, 2024, doi: 10.48550/arXiv.2411.11360.
[65]
Z. Li et al., “Towards real-world change captioning: A large-scale disaster-focused dataset,” in Advances in neural information processing systems (NeurIPS), 2025.
[66]
U. Türker et al., “Multispectral remote sensing image change captioning with Sentinel-2,” IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing, vol. 18, pp. 25410–25423, 2025, doi: 10.1109/JSTARS.2025.3566543.
[67]
Z. Yuan, L. Mou, Y. Hua, and X. X. Zhu, “Change detection meets visual question answering,” IEEE Transactions on Geoscience and Remote Sensing, vol. 60, pp. 1–13, 2022, doi: 10.1109/TGRS.2022.3160882.
[68]
K. Li et al., “Show me what and where has changed? Question answering and grounding for remote sensing change detection.” arXiv preprint arXiv:2410.23828, 2024, doi: 10.48550/arXiv.2410.23828.
[69]
P. Wu et al., ChangeChat: An interactive model for remote sensing change analysis via multimodal instruction tuning,” in IEEE international conference on acoustics, speech and signal processing (ICASSP), 2025, doi: 10.1109/ICASSP49660.2025.10888710.
[70]
J. Irvin, A. Gruca, et al., TEOChat: A large vision-language assistant for temporal earth observation data,” in International conference on learning representations (ICLR), 2025.
[71]
J. Wang et al., DisasterM3: A remote sensing vision-language dataset for disaster damage assessment and response.” arXiv preprint arXiv:2505.21089, 2025, doi: 10.48550/arXiv.2505.21089.
[72]
P. Chen, B. Zhang, D. Hong, Z. Chen, X. Yang, and B. Li, FCCDN: Feature constraint network for VHR image change detection,” ISPRS Journal of Photogrammetry and Remote Sensing, vol. 187, pp. 101–119, 2022, doi: 10.1016/j.isprsjprs.2022.02.021.
[73]
H. Chen, F. Pu, R. Yang, and X. Xu, RDP-Net: Region detail preserving network for change detection,” IEEE Transactions on Geoscience and Remote Sensing, vol. 60, pp. 1–10, 2022, doi: 10.1109/TGRS.2022.3227098.
[74]
H. Chen, W. Li, S. Chen, and Z. Shi, “Semantic-aware dense representation learning for remote sensing image change detection,” IEEE Transactions on Geoscience and Remote Sensing, vol. 60, pp. 1–18, 2022, doi: 10.1109/TGRS.2022.3203769.
[75]
H. Chen, W. Li, and Z. Shi, “Adversarial instance augmentation for building change detection in remote sensing images,” IEEE Transactions on Geoscience and Remote Sensing, vol. 60, pp. 1–16, 2021, doi: 10.1109/TGRS.2021.3066802.
[76]
Z. Li, C. Tang, X. Liu, W. Zhang, J. Dou, and L. Wang, “Lightweight remote sensing change detection with progressive feature aggregation and supervised attention,” IEEE Transactions on Geoscience and Remote Sensing, vol. 61, pp. 1–12, 2023, doi: 10.1109/TGRS.2023.3241436.
[77]
A. Codegoni, G. Lombardi, and A. Ferrari, TINYCD: A (not so) deep learning model for change detection,” Neural Computing and Applications, vol. 35, no. 11, pp. 8471–8486, 2023, doi: 10.1007/s00521-022-08122-3.
[78]
Y. Xing, J. Jiang, J. Xiang, E. Yan, Y. Song, and D. Mo, LightCDNet: Lightweight change detection network based on VHR images,” IEEE Geoscience and Remote Sensing Letters, vol. 20, pp. 1–5, 2023, doi: 10.1109/LGRS.2023.3304309.
[79]
T. Lei, X. Geng, H. Ning, Z. Lv, M. Gong, and Y. Jin, “Ultralightweight spatial-spectral feature cooperation network for change detection in remote sensing images,” IEEE Transactions on Geoscience and Remote Sensing, vol. 61, pp. 1–14, 2023, doi: 10.1109/TGRS.2023.3261273.
[80]
H. Zhang, M. Lin, G. Yang, and L. Zhang, ESCNet: An end-to-end superpixel-enhanced change detection network for very-high-resolution remote sensing images,” IEEE Transactions on Neural Networks and Learning Systems, vol. 34, no. 1, pp. 28–42, 2023, doi: 10.1109/TNNLS.2021.3089332.
[81]
Y. Feng, J. Jiang, H. Xu, and J. Zheng, “Change detection on remote sensing images using dual-branch multilevel intertemporal network,” IEEE Transactions on Geoscience and Remote Sensing, vol. 61, pp. 1–15, 2023, doi: 10.1109/TGRS.2023.3241257.
[82]
Y. Ye, M. Wang, L. Zhou, G. Lei, J. Fan, and Y. Qin, “Adjacent-level feature cross-fusion with 3-D CNN for remote sensing image change detection,” IEEE Transactions on Geoscience and Remote Sensing, vol. 61, pp. 1–14, 2023, doi: 10.1109/TGRS.2023.3305499.
[83]
C. Han, C. Wu, H. Guo, M. Hu, and H. Chen, HANet: A hierarchical attention network for change detection with bitemporal very-high-resolution remote sensing images,” IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing, vol. 16, pp. 3867–3878, 2023, doi: 10.1109/JSTARS.2023.3264802.
[84]
C. Han, C. Wu, H. Guo, M. Hu, J. Li, and H. Chen, “Change guiding network: Incorporating change prior to guide change detection in remote sensing imagery,” IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing, vol. 16, pp. 8395–8407, 2023, doi: 10.1109/JSTARS.2023.3310208.
[85]
Y. Cao and X. Huang, “A full-level fused cross-task transfer learning method for building change detection using noise-robust pretrained networks on crowdsourced labels,” Remote Sensing of Environment, vol. 284, p. 113371, 2023, doi: 10.1016/j.rse.2022.113371.
[86]
M. Cheng, W. He, Z. Li, G. Yang, and H. Zhang, “Harmony in diversity: Content cleansing change detection framework for very-high-resolution remote-sensing images,” ISPRS Journal of Photogrammetry and Remote Sensing, vol. 218, pp. 1–19, 2024, doi: 10.1016/j.isprsjprs.2024.09.002.
[87]
Y. Zhao, T. Celik, N. Liu, F. Gao, and H.-C. Li, SSLChange: A self-supervised change detection framework based on domain adaptation,” IEEE Transactions on Geoscience and Remote Sensing, vol. 62, pp. 1–14, 2024, doi: 10.1109/TGRS.2024.3489615.
[88]
Z. Zheng, Y. Zhong, J. Zhao, A. Ma, and L. Zhang, “Unifying remote sensing change detection via deep probabilistic change models: From principles, models to applications,” ISPRS Journal of Photogrammetry and Remote Sensing, vol. 215, pp. 239–255, 2024, doi: 10.1016/j.isprsjprs.2024.07.001.
[89]
Z. Liu, J. Zhang, W. Wang, and Y. Gu, “A novel multibranch self-distillation framework for optimizing remote sensing change detection,” IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing, vol. 18, pp. 22642–22655, 2025, doi: 10.1109/JSTARS.2025.3597817.
[90]
Y. Sun, L. Lei, Z. Li, G. Kuang, and Q. Yu, “Detecting changes without comparing images: Rules induced change detection in heterogeneous remote sensing images,” ISPRS Journal of Photogrammetry and Remote Sensing, vol. 230, pp. 241–257, 2025, doi: 10.1016/j.isprsjprs.2025.09.009.
[91]
L. Sun, M. Jin, J. Yan, (initial. missing). He, and L. Cao, Semantic-TemporalNet: A novel urban block change detection method based on semantic coherence analysis,” IEEE Transactions on Geoscience and Remote Sensing, vol. 63, pp. 1–16, 2025, doi: 10.1109/TGRS.2025.3611378.
[92]
Z. Liu, J. Zhang, W. Wang, and Y. Gu, “A novel multibranch self-distillation framework for optimizing remote sensing change detection,” IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing, vol. 18, pp. 22642–22655, 2025, doi: 10.1109/JSTARS.2025.3597817.
[93]
Z. Liu, J. Zhang, Q. Shi, and Y. Gu, “UnlearningCD: Distill to forget in change detection via perturbed knowledge,” IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing, pp. 1–13, 2026, doi: 10.1109/JSTARS.2026.3655788.
[94]
C. Zhang, L. Wang, S. Cheng, and Y. Li, SwinSUNet: Pure transformer network for remote sensing image change detection,” IEEE Transactions on Geoscience and Remote Sensing, vol. 60, pp. 1–13, 2022, doi: 10.1109/TGRS.2022.3160007.
[95]
Q. Li, R. Zhong, X. Du, and Y. Du, TransUNetCD: A hybrid transformer network for change detection in optical remote-sensing images,” IEEE Transactions on Geoscience and Remote Sensing, vol. 60, pp. 1–19, 2022, doi: 10.1109/TGRS.2022.3169479.
[96]
N. Shi, K. Chen, and G. Zhou, “A divided spatial and temporal context network for remote sensing change detection,” IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing, vol. 15, pp. 4897–4908, 2022, doi: 10.1109/JSTARS.2022.3176858.
[97]
Z. Zheng, Y. Zhong, S. Tian, A. Ma, and L. Zhang, ChangeMask: Deep multi-task encoder-transformer-decoder architecture for semantic change detection,” ISPRS Journal of Photogrammetry and Remote Sensing, vol. 183, pp. 228–239, 2022, doi: 10.1016/j.isprsjprs.2021.10.015.
[98]
M. Noman et al., “Remote sensing change detection with transformers trained from scratch,” IEEE Transactions on Geoscience and Remote Sensing, vol. 62, pp. 1–14, 2024, doi: 10.1109/TGRS.2024.3383800.
[99]
S. Fang, K. Li, and Z. Li, “Changer: Feature interaction is what you need for change detection,” IEEE Transactions on Geoscience and Remote Sensing, vol. 61, pp. 1–11, 2023, doi: 10.1109/TGRS.2023.3277496.
[100]
M. Bernhard, N. Strauß, and M. Schubert, MapFormer: Boosting change detection by using pre-change information,” in Proceedings of the IEEE/CVF international conference on computer vision (ICCV), 2023, pp. 16837–16846, doi: 10.1109/ICCV51070.2023.01544.
[101]
W. Yu, X. Zhang, S. Das, X. X. Zhu, and P. Ghamisi, MaskCD: A remote sensing change detection network based on mask classification,” IEEE Transactions on Geoscience and Remote Sensing, vol. 62, pp. 1–16, 2024, doi: 10.1109/TGRS.2024.3424300.
[102]
H. Zhang et al., BiFA: Remote sensing image change detection with bitemporal feature alignment,” IEEE Transactions on Geoscience and Remote Sensing, vol. 62, pp. 1–17, 2024, doi: 10.1109/TGRS.2024.3376673.
[103]
H. Zhang, H. Guo, K. Chen, H. Chen, Z. Zou, and Z. Shi, FoBa: A foreground-background co-guided method and new benchmark for remote sensing semantic change detection.” arXiv preprint arXiv:2509.15788, 2025, doi: 10.48550/arXiv.2509.15788.
[104]
D. Zhu, X. Huang, H. Huang, H. Zhou, and Z. Shao, Change3D: Revisiting change detection and captioning from a video modeling perspective,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition (CVPR), 2025, pp. 24011–24022, doi: 10.1109/CVPR52734.2025.02236.
[105]
H. Zhang, K. Chen, C. Liu, H. Chen, Z. Zou, and Z. Shi, CDMamba: Incorporating local clues into Mamba for remote sensing image binary change detection,” IEEE Transactions on Geoscience and Remote Sensing, vol. 63, pp. 1–16, 2025, doi: 10.1109/TGRS.2025.3545012.
[106]
S. Liu et al., CD-STMamba: Towards remote sensing image change detection with spatio-temporal interaction Mamba model,” IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing, vol. 18, pp. 10471–10485, 2025, doi: 10.1109/JSTARS.2025.3559085.
[107]
J. Zhao, J. Xie, Y. Zhou, W. Du, R. Yao, and A. El Saddik, ST-Mamba: Spatio-temporal synergistic model for remote sensing change detection,” IEEE Transactions on Geoscience and Remote Sensing, vol. 63, pp. 1–13, 2025, doi: 10.1109/TGRS.2025.3579617.
[108]
Y. Liu, G. Cheng, Q. Sun, C. Tian, and L. Wang, CWMamba: Leveraging CNN-Mamba fusion for enhanced change detection in remote sensing images,” IEEE Geoscience and Remote Sensing Letters, vol. 22, pp. 1–5, 2025, doi: 10.1109/LGRS.2025.3548145.
[109]
X. Liu et al., GSTM-SCD: Graph-enhanced spatio-temporal state space model for semantic change detection in multi-temporal remote sensing images,” ISPRS Journal of Photogrammetry and Remote Sensing, vol. 230, pp. 73–91, 2025, doi: 10.1016/j.isprsjprs.2025.09.003.
[110]
Y. Li, W. Liu, E. Li, L. Zhang, and X. Li, SAM-Mamba: A two-stage change detection network combining the adapting segment anything and Mamba models,” IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing, vol. 18, pp. 21607–21619, 2025, doi: 10.1109/JSTARS.2025.3601739.
[111]
J. N. Paranjape, C. De Melo, and V. M. Patel, “A Mamba-based siamese network for remote sensing change detection,” in 2025 IEEE/CVF winter conference on applications of computer vision (WACV), 2025, pp. 1186–1196, doi: 10.1109/WACV61041.2025.00123.
[112]
S. Dong, L. Wang, B. Du, and X. Meng, ChangeCLIP: Remote sensing change detection with multimodal vision-language representation learning,” ISPRS Journal of Photogrammetry and Remote Sensing, vol. 208, pp. 53–69, 2024, doi: 10.1016/j.isprsjprs.2024.01.004.
[113]
Z. Zheng, Y. Zhong, L. Zhang, and S. Ermon, “Segment any change,” in Advances in neural information processing systems, 2024, pp. 81204–81224, doi: 10.52202/079017-2581.
[114]
L. Ding, K. Zhu, D. Peng, H. Tang, K. Yang, and L. Bruzzone, “Adapting segment anything model for change detection in VHR remote sensing images,” IEEE Transactions on Geoscience and Remote Sensing, vol. 62, pp. 1–11, 2024, doi: 10.1109/TGRS.2024.3368168.
[115]
K. Li, X. Cao, and D. Meng, “A new learning paradigm for foundation model-based remote-sensing change detection,” IEEE Transactions on Geoscience and Remote Sensing, vol. 62, pp. 1–12, 2024, doi: 10.1109/TGRS.2024.3365825.
[116]
X. Tan, G. Chen, T. Wang, J. Wang, and X. Zhang, “Segment change model (SCM) for unsupervised change detection in VHR remote sensing images: A case study of buildings,” in IGARSS 2024 – 2024 IEEE international geoscience and remote sensing symposium, 2024, pp. 8577–8580, doi: 10.1109/IGARSS53475.2024.10642429.
[117]
K. Li et al., SemiCD-VL: Visual-language model guidance makes better semi-supervised change detector,” IEEE Transactions on Geoscience and Remote Sensing, vol. 63, pp. 1–13, 2024, doi: 10.1109/TGRS.2024.3512548.
[118]
J. Huang, J. Bao, M. Xia, and X. Yuan, SAM-based efficient feature integration network for remote sensing change detection: A case study on Macao sea reclamation,” IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing, vol. 18, pp. 16916–16928, 2025, doi: 10.1109/JSTARS.2025.3584145.
[119]
Y. Qin, C. Wang, Y. Fan, and C. Pan, SAM2-CD: Remote sensing image change detection with SAM2,” IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing, vol. 18, pp. 24575–24587, 2025, doi: 10.1109/JSTARS.2025.3610156.
[120]
K. Li et al., DynamicEarth: How far are we from open-vocabulary change detection?” arXiv preprint arXiv:2501.12931, 2025, doi: 10.48550/arXiv.2501.12931.
[121]
F. Jiang, M. Gong, T. Zhan, and X. Fan, “A semisupervised GAN-based multiple change detection framework in multi-spectral images,” IEEE Geoscience and Remote Sensing Letters, vol. 17, no. 7, pp. 1223–1227, 2020, doi: 10.1109/LGRS.2019.2941318.
[122]
P. Jian, K. Chen, and W. Cheng, GAN-based one-class classification for remote-sensing image change detection,” IEEE Geoscience and Remote Sensing Letters, vol. 19, pp. 1–5, 2022, doi: 10.1109/LGRS.2021.3066435.
[123]
Z. Zheng, S. Tian, A. Ma, L. Zhang, and Y. Zhong, “Scalable multi-temporal remote sensing change data generation via simulating stochastic change process,” in Proceedings of the IEEE/CVF international conference on computer vision (ICCV), 2023, pp. 21761–21770, doi: 10.1109/ICCV51070.2023.01994.
[124]
W. G. C. Bandara, N. G. Nair, and V. M. Patel, DDPM-CD: Denoising diffusion probabilistic models as feature extractors for change detection.” arXiv preprint arXiv:2206.11892, 2022, doi: 10.48550/arXiv.2206.11892.
[125]
Z. Zheng, S. Ermon, D. Kim, L. Zhang, and Y. Zhong, Changen2: Multi-temporal remote sensing generative change foundation model,” IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 47, no. 2, pp. 725–741, 2024, doi: 10.1109/TPAMI.2024.3475824.
[126]
K. Tang and J. Chen, ChangeAnywhere: Sample generation for remote sensing change detection via semantic latent diffusion model.” arXiv preprint arXiv:2404.08892, 2024, doi: 10.48550/arXiv.2404.08892.
[127]
J. Jia, G. Lee, Z. Wang, Z. Lyu, and Y. He, “Siamese meets diffusion network: SMDNet for enhanced change detection in high-resolution RS imagery,” IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing, vol. 17, pp. 8189–8202, 2024, doi: 10.1109/JSTARS.2024.3384545.
[128]
P. Han, Y. Gao, G. Chen, B. Zhao, and X. Li, “Hierarchical diffusion model for remote sensing image change detection,” IEEE Transactions on Geoscience and Remote Sensing, 2025, doi: 10.1109/TGRS.2025.3608791.
[129]
Q. Zang, J. Yang, S. Wang, D. Zhao, W. Yi, and Z. Zhong, ChangeDiff: A multi-temporal change detection data generator with flexible text prompts via diffusion model,” in Proceedings of the AAAI conference on artificial intelligence, 2025, pp. 9763–9771, doi: 10.1609/aaai.v39i9.33058.
[130]
Y. Benidir, N. Gonthier, and C. Mallet, “The change you want to detect: Semantic change detection in earth observation with hybrid data generation,” in IEEE/CVF conference on computer vision and pattern recognition (CVPR), 2025, pp. 2204–2214, doi: 10.1109/CVPR52734.2025.00211.
[131]
C. Liu, K. Chen, H. Zhang, Z. Qi, Z. Zou, and Z. Shi, “Change-agent: Towards interactive comprehensive remote sensing change interpretation and analysis,” IEEE Transactions on Geoscience and Remote Sensing, vol. 62, pp. 1–16, 2024, doi: 10.1109/TGRS.2024.3425815.
[132]
W. Xu et al., RS-Agent: Automating remote sensing tasks through intelligent agent.” arXiv preprint arXiv:2406.07089, 2024, doi: 10.48550/arXiv.2406.07089.
[133]
P. Feng et al., “Earth-agent: Unlocking the full landscape of earth observation with agents.” arXiv preprint arXiv:2509.23141, 2025, doi: 10.48550/arXiv.2509.23141.
[134]
H. Hu et al., RingMo-Agent: A unified remote sensing foundation model for multi-platform and multi-modal reasoning.” arXiv preprint arXiv:2507.20776, 2025, doi: 10.48550/arXiv.2507.20776.

  1. Ziyuan Liu and Yuantao Gu are with the Department of Electronic Engineering, Beijing National Research Center for Information Science and Technology, Tsinghua University, Beijing 100084, China (e-mail: liuziyua22@mails.tsinghua.edu.cn; gyt@tsinghua.edu.cn). Ruifei Zhu is with Chang Guang Satellite Technology Co., Ltd. (CGSTL) Changchun 130102, China (e-mail: zhuruifei@jl1.cn). Ouqiao Ma is with the College of Communications Engineering, Army Engineering University of PLA, Nanjing 210007, China (e-mail: moq@aeu.edu.cn). (Corresponding author: Yuantao Gu.)↩︎