Pith. sign in

Paper Citation Record · LEDGER

Multimodal Remote Sensing Scene Classification Using VLMs and Dual-Cross Attention Networks

As of 12 August 2026, this Paper Citation Record lists 45 of 45 outbound references and 0 inbound Pith citation observations for arXiv:2412.02531.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.02531 v1

Coverage vector

measured 45 of 45 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T23:25:10.376115Z

measured 45 of 45 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

45 of 45 outbound references displayed

  • verified exact1
  • verified fuzzy35
  • unresolved9
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 7071abf1-6c2f-486f-9014-4107d615da74 · outbound

This paper cites Deep learning for remote sensing image scene classification: A review and meta-analysis,.

Multimodal Remote Sensing Scene Classification Using VLMs and Dual-Cross Attention Networks Deep learning for remote sensing image scene classification: A review and meta-analysis,

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T23:25:10.229124Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:25:10.229124Z digest=sha256:5ad45ff657aae924f15d2fbfc8d7e60c4f762060c794be14e5fa7909265d3a06

Observation 527e1dd8-255b-4880-bccb-718267df7479 · outbound

This paper cites Scene and environment monitoring using aerial imagery and deep learning,.

Multimodal Remote Sensing Scene Classification Using VLMs and Dual-Cross Attention Networks Scene and environment monitoring using aerial imagery and deep learning,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:25:10.877827Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T23:25:10.233407Z digest=sha256:5b7521dc13c2830e5f7455a13dfce0a60bae7c4204b978934710dade47d6d416

Observation 7212442e-d9ac-4b65-abb3-523425eee87a · outbound

This paper cites Spatial context-aware method for urban land use classification using street view images,.

Multimodal Remote Sensing Scene Classification Using VLMs and Dual-Cross Attention Networks Spatial context-aware method for urban land use classification using street view images,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:25:10.868141Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T23:25:10.237070Z digest=sha256:8c8df90b10b1bd21379a3703033bafbf13d27c0a1f055110b582bf4272aaef94

Observation 05b9dae7-3a75-4c1a-8286-27dd85fd1598 · outbound

This paper cites Urban scene understanding based on semantic and socioeconomic features: From high-resolution remote sensing imagery to multi-source geographic datasets,.

Multimodal Remote Sensing Scene Classification Using VLMs and Dual-Cross Attention Networks Urban scene understanding based on semantic and socioeconomic features: From high-resolution remote sensing imagery to multi-source geographic datasets,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:25:10.857464Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T23:25:10.240960Z digest=sha256:6cc1aadd70bf52d2b52201ddfa5e515b0a9d01eced8c2549ed56b7370d08f83f

Observation 28541f4f-d411-41b0-b11e-a7d2d2c02d7e · outbound

This paper cites A survey on deep learning-driven remote sensing image scene understanding: Scene classification, scene retrieval and scene-guided object detection,.

Multimodal Remote Sensing Scene Classification Using VLMs and Dual-Cross Attention Networks A survey on deep learning-driven remote sensing image scene understanding: Scene classification, scene retrieval and scene-guided object detection,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:25:10.847730Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T23:25:10.244642Z digest=sha256:8878a96bd7297f069d559dd7bc1a60e88908691be40445ac3387b52b14ee06fb

Observation 2410f6e5-b6ba-4c04-ac03-4d9cfec9ad72 · outbound

This paper cites Evaluating the potential of texture and color descriptors for remote sensing image retrieval and classification,.

Multimodal Remote Sensing Scene Classification Using VLMs and Dual-Cross Attention Networks Evaluating the potential of texture and color descriptors for remote sensing image retrieval and classification,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:25:10.838031Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T23:25:10.248249Z digest=sha256:f7c19b3564df8f7f1a078d308f52b54972f5bc8f23940174f6f732ea0779336f

Observation 010b672d-a5bd-4dbf-9510-83834df5fac8 · outbound

This paper cites Indexing of remote sensing images with different resolutions by multiple features,.

Multimodal Remote Sensing Scene Classification Using VLMs and Dual-Cross Attention Networks Indexing of remote sensing images with different resolutions by multiple features,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:25:10.828162Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T23:25:10.252196Z digest=sha256:a6a162e79892ea9af7d89b1d79e20481a723ee2b8d3b2326560727b01085825c

Observation 2a2bf74e-c86c-4c47-86cd-481d6a9a096f · outbound

This paper cites High-resolution satellite scene classification using a sparse coding based multiple feature combination,.

Multimodal Remote Sensing Scene Classification Using VLMs and Dual-Cross Attention Networks High-resolution satellite scene classification using a sparse coding based multiple feature combination,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:25:10.818269Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T23:25:10.255697Z digest=sha256:ae573e2e7a2766548a72a1eca79ad670c5438c055a421b5545376b5a6c619acc

Observation ba22cb26-5dba-4745-ac5f-ac9a67248489 · outbound

This paper cites Remote sensing image scene classification meets deep learning: Challenges, methods, benchmarks, and opportunities,.

Multimodal Remote Sensing Scene Classification Using VLMs and Dual-Cross Attention Networks Remote sensing image scene classification meets deep learning: Challenges, methods, benchmarks, and opportunities,

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T23:25:10.259030Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:25:10.259030Z digest=sha256:ebaee3c8b964f71bca3b167dae1cefa89c773ccd220508a1891de7237378a1e8

Observation 49046a35-f649-43e9-8e76-6bfea94693c2 · outbound

This paper cites Scene classification based on multiscale convolutional neural network,.

Multimodal Remote Sensing Scene Classification Using VLMs and Dual-Cross Attention Networks Scene classification based on multiscale convolutional neural network,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:25:10.801839Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T23:25:10.262380Z digest=sha256:e2ff9c302d1c3b707cfe79b04211b53a643b561fac1f885af3bd4e8988bac358

Observation 0a0b01d1-c752-4d27-bf41-b941df70e6f3 · outbound

This paper cites Gan-based semisupervised scene classifi- cation of remote sensing image,.

Multimodal Remote Sensing Scene Classification Using VLMs and Dual-Cross Attention Networks Gan-based semisupervised scene classifi- cation of remote sensing image,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:25:10.792178Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T23:25:10.265554Z digest=sha256:39d5bec1ea27fd3c0268442f83992bc1f29211eb2baa7ee274f5061f25bc900d

Observation e7b264d0-4fd4-4d44-8d17-6190cd718c63 · outbound

This paper cites Scvit: A spatial-channel feature preserving vision transformer for remote sensing image scene classification,.

Multimodal Remote Sensing Scene Classification Using VLMs and Dual-Cross Attention Networks Scvit: A spatial-channel feature preserving vision transformer for remote sensing image scene classification,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:25:10.782374Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T23:25:10.269031Z digest=sha256:fda4abb7d07cc3f6dd2b2d52943252c33a4c7a68c671d49324b8897f0710fffb

Observation e7619c2e-677c-4029-a957-1d1e4a755059 · outbound

This paper cites Vision transformer with contrastive learning for remote sensing image scene classification,.

Multimodal Remote Sensing Scene Classification Using VLMs and Dual-Cross Attention Networks Vision transformer with contrastive learning for remote sensing image scene classification,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:25:10.772943Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T23:25:10.272302Z digest=sha256:b6593371f4c24ebb19bf5550320c6f5a2f6b3bc4e1a6b02ce56d8533748aa5bd

Observation 17e4a056-52c8-47bd-a3ed-22c767a01698 · outbound

This paper cites Deep learning-a new approach for multi-label scene classification in plan- etscope and sentinel-2 imagery,.

Multimodal Remote Sensing Scene Classification Using VLMs and Dual-Cross Attention Networks Deep learning-a new approach for multi-label scene classification in plan- etscope and sentinel-2 imagery,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:25:10.763313Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T23:25:10.275625Z digest=sha256:fa67ad4dad3a2dd9928a737ff0a5091ba91c565dd8948aa5b90d1060e0eab8d2

Observation 441c68e4-96bc-4e07-911f-356a7cf9cf4b · outbound

This paper cites A multi- level label-aware semi-supervised framework for remote sensing scene classification,.

Multimodal Remote Sensing Scene Classification Using VLMs and Dual-Cross Attention Networks A multi- level label-aware semi-supervised framework for remote sensing scene classification,

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T23:25:10.278716Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:25:10.278716Z digest=sha256:8adb39694fbfb1736dccb3ea34695c061c4be6f9abfe3d1bf575483167f8dee1

Observation 2a250829-f0c9-41c7-ac3e-bfda466284fc · outbound

This paper cites He and Q.

Multimodal Remote Sensing Scene Classification Using VLMs and Dual-Cross Attention Networks He and Q

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:25:10.747640Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T23:25:10.281890Z digest=sha256:212c5138ceed57131b714911644ee4fc5975de6706c17d71542e14b34b25c13a

Observation 711cc84a-e56f-4dbd-bfac-3865f06ea4dc · outbound

This paper cites A lightweight multi-scale crossmodal text-image retrieval method in remote sensing,.

Multimodal Remote Sensing Scene Classification Using VLMs and Dual-Cross Attention Networks A lightweight multi-scale crossmodal text-image retrieval method in remote sensing,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:25:10.737549Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T23:25:10.285092Z digest=sha256:b12f0644c9b3e23b8f0f734c55d4bcae1b17d69903f9efc886f62512d3ad157d

Observation 48b40467-48a4-495b-9d03-6b880c100835 · outbound

This paper cites Multi-label semantic feature fusion for remote sensing image captioning,.

Multimodal Remote Sensing Scene Classification Using VLMs and Dual-Cross Attention Networks Multi-label semantic feature fusion for remote sensing image captioning,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:25:10.727851Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T23:25:10.288191Z digest=sha256:76510d8c695c43e9d6de2863c78440978d2895e0fbdd2c87914f9d9a59d50e9d

Observation 58169b2b-4533-4cab-ac3d-e44861e1151f · outbound

This paper cites Teaw: Text-aware few-shot remote sensing image scene classification,.

Multimodal Remote Sensing Scene Classification Using VLMs and Dual-Cross Attention Networks Teaw: Text-aware few-shot remote sensing image scene classification,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:25:10.718106Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T23:25:10.291364Z digest=sha256:5c58cf87e1f78444c48e9ac559340205927080a2f3e87f620ec8eff6ac7f5a51

Observation 4e715f2b-7443-4163-9a93-995601a9a425 · outbound

This paper cites Remoteclip: A vision language foundation model for remote sensing,.

Multimodal Remote Sensing Scene Classification Using VLMs and Dual-Cross Attention Networks Remoteclip: A vision language foundation model for remote sensing,

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T23:25:10.294762Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:25:10.294762Z digest=sha256:3d2d5dba230d07aae1c280c93033af2edaf1d4593414ceeb8b296fd4296a4b67

Observation bbccb5ff-034d-402f-884d-77d00a688a07 · outbound

This paper cites Land-cover classification with high-resolution remote sensing images using transferable deep models,.

Multimodal Remote Sensing Scene Classification Using VLMs and Dual-Cross Attention Networks Land-cover classification with high-resolution remote sensing images using transferable deep models,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:25:10.702609Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T23:25:10.297929Z digest=sha256:628a4422365cb0b95ccedc69d37574716a952685b382767bb2b723046d60d3a4

Observation 4f9966a2-fa2b-4011-864c-199deac4b070 · outbound

This paper cites A dual- model architecture with grouping-attention-fusion for remote sensing scene classification,.

Multimodal Remote Sensing Scene Classification Using VLMs and Dual-Cross Attention Networks A dual- model architecture with grouping-attention-fusion for remote sensing scene classification,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:25:10.693529Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T23:25:10.301005Z digest=sha256:9c3711921bf983365d9ce8e70e94392f209641a6058c0d81d9ffbb553869877f

Observation bd3c34a4-8aa0-4bef-9741-1be01d262657 · outbound

This paper cites Scene classification with recurrent attention of vhr remote sensing images,.

Multimodal Remote Sensing Scene Classification Using VLMs and Dual-Cross Attention Networks Scene classification with recurrent attention of vhr remote sensing images,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:25:10.684148Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T23:25:10.304137Z digest=sha256:566e5849dd4c62d7169aaaac32b3ef7c4c9232e245e52932e1bba0af9f34fb46

Observation 2dc92c41-72be-498f-8556-d556561b435b · outbound

This paper cites Classification of remote sensing images using efficientnet-b3 cnn model with attention,.

Multimodal Remote Sensing Scene Classification Using VLMs and Dual-Cross Attention Networks Classification of remote sensing images using efficientnet-b3 cnn model with attention,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:25:10.675233Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T23:25:10.307138Z digest=sha256:de6aa0899b9fc4a68ec52b86184633c91da5fb017d0ce5d2a389c1a175c84617

Observation 7fb639c1-91dc-49ac-8e32-69e74a0c8722 · outbound

This paper cites Rs-deepsuperlearner: fusion of cnn ensemble for remote sensing scene classification,.

Multimodal Remote Sensing Scene Classification Using VLMs and Dual-Cross Attention Networks Rs-deepsuperlearner: fusion of cnn ensemble for remote sensing scene classification,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:25:10.666022Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T23:25:10.310320Z digest=sha256:f2c3c1f771964e35decc202e98889d42b2db8abcc3d313dbb17efe43bd2b7b94

Observation f157f96c-6b35-445a-b4c4-d9fb71c36280 · outbound

This paper cites Contextual spatial-channel attention network for remote sensing scene classification,.

Multimodal Remote Sensing Scene Classification Using VLMs and Dual-Cross Attention Networks Contextual spatial-channel attention network for remote sensing scene classification,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:25:10.656970Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T23:25:10.313382Z digest=sha256:0e8e46b1564a11502a39f6e394a6e3a48edc975da576721a7d1686ddbeef96e0

Observation 78461957-93eb-4f2d-a992-370049753e55 · outbound

This paper cites A local–global interactive vision trans- former for aerial scene classification,.

Multimodal Remote Sensing Scene Classification Using VLMs and Dual-Cross Attention Networks A local–global interactive vision trans- former for aerial scene classification,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:25:10.647236Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T23:25:10.316600Z digest=sha256:cc370a50c6b4b4dcd9ff71a89edfe66c95ecfc6c3ebe63ae9e597a0881cdec4c

Observation 753af8c8-49a9-410e-b950-ef4ead80cb6c · outbound

This paper cites Nir/rgb image fusion for scene classification using deep neural networks,.

Multimodal Remote Sensing Scene Classification Using VLMs and Dual-Cross Attention Networks Nir/rgb image fusion for scene classification using deep neural networks,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:25:10.637260Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T23:25:10.319967Z digest=sha256:0bb53772a59d600b6f30219bcced79a5030c3abf2a53a6012e9b0f3491ff8b36

Observation c746d677-d954-478a-a1d1-2aa7ff562e09 · outbound

This paper cites Mrssc: a benchmark dataset for multimodal remote sensing scene classification,.

Multimodal Remote Sensing Scene Classification Using VLMs and Dual-Cross Attention Networks Mrssc: a benchmark dataset for multimodal remote sensing scene classification,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:25:10.627429Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T23:25:10.323038Z digest=sha256:5948b059e29b29293ea5f9c5eb8bda86afda241284f44c5cb773fa11d253c6d6

Observation 63606dcf-7ef1-48ca-96c8-42c6ef88bc77 · outbound

This paper cites An optical image-aided approach for zero-shot sar image scene classifica- tion,.

Multimodal Remote Sensing Scene Classification Using VLMs and Dual-Cross Attention Networks An optical image-aided approach for zero-shot sar image scene classifica- tion,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:25:10.617723Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T23:25:10.326298Z digest=sha256:31de7eac5b3594729a0201331b8d4d40cf7a27ac1ddbccf70a379d627cb5561b

Observation a1a2c3d8-c511-4dae-a7ec-1a7bdb13eaf7 · outbound

This paper cites Alignment and fusion using distinct sensor data for multimodal aerial scene classification,.

Multimodal Remote Sensing Scene Classification Using VLMs and Dual-Cross Attention Networks Alignment and fusion using distinct sensor data for multimodal aerial scene classification,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:25:10.607260Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T23:25:10.329375Z digest=sha256:49abead28292632d4eb24a07ab31379227ee9dee931dd0bb050142a47b600c7f

Observation a75a814f-958d-4879-9e2a-90e740063b27 · outbound

This paper cites Text2seg: Remote sensing image semantic segmentation via text-guided visual foundation models,.

Multimodal Remote Sensing Scene Classification Using VLMs and Dual-Cross Attention Networks Text2seg: Remote sensing image semantic segmentation via text-guided visual foundation models,

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-11T23:25:10.332515Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:25:10.332515Z digest=sha256:cc2c40826f09389090a299f7317e49036e1a5b6431ee1a512c880a90acc2a7e4

Observation 4de625c1-270d-4dbe-b0df-17d64887ab11 · outbound

This paper cites Parameter-efficient transfer learning for remote sensing image-text retrieval,.

Multimodal Remote Sensing Scene Classification Using VLMs and Dual-Cross Attention Networks Parameter-efficient transfer learning for remote sensing image-text retrieval,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:25:10.597502Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T23:25:10.335765Z digest=sha256:8ce57001d92b45a72cda2ea0ad636d5f8841fa3d587df0b5f22d2ac8c8e97dd3

Observation f7b8332f-69e6-4264-b83f-52d170207ec3 · outbound

This paper cites Interacting- enhancing feature transformer for cross-modal remote-sensing image and text retrieval,.

Multimodal Remote Sensing Scene Classification Using VLMs and Dual-Cross Attention Networks Interacting- enhancing feature transformer for cross-modal remote-sensing image and text retrieval,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:25:10.587501Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T23:25:10.338847Z digest=sha256:21661f86cf7e36e9d5efadcf99966bb3ca1f8642e4e2b1efcdf55b208d876c7c

Observation 1f16899c-85ef-45dd-a398-b31376b6ccbb · outbound

This paper cites Few-shot remote sensing image scene classification: Recent advances, new baselines, and future trends,.

Multimodal Remote Sensing Scene Classification Using VLMs and Dual-Cross Attention Networks Few-shot remote sensing image scene classification: Recent advances, new baselines, and future trends,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:25:10.577824Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T23:25:10.341993Z digest=sha256:0184981601f323f5d8f4fa58fa5319cf2bb71d9731437fb8d0f4a003c1bd73ed

Observation 18f3fbe3-8b87-4d28-a26f-8ac5abeffb9c · outbound

This paper cites Few-shot medical image classification with simple shape and texture text descriptors using vision-language models.

Multimodal Remote Sensing Scene Classification Using VLMs and Dual-Cross Attention Networks Few-shot medical image classification with simple shape and texture text descriptors using vision-language models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-11T23:25:10.345426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:25:10.345426Z digest=sha256:cb09a5f59584ee9338afb2d5d5f7287a7f1b60d5ed32ef7cad63d14014156faf

Observation c94b1f13-b0d8-48a9-860f-f6d705d5e643 · outbound

This paper cites Vision-language model for generating textual descriptions from clinical images: model development and validation study,.

Multimodal Remote Sensing Scene Classification Using VLMs and Dual-Cross Attention Networks Vision-language model for generating textual descriptions from clinical images: model development and validation study,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:25:10.567331Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T23:25:10.349243Z digest=sha256:a764b41d55807eb2fef0fa457e27a5f4533bf4b4877818a8516cc6313153159d

Observation da38a411-eb5e-4bb6-9795-dbdbf2926c31 · outbound

This paper cites Medical image understanding with pretrained vision language models: A comprehensive study,.

Multimodal Remote Sensing Scene Classification Using VLMs and Dual-Cross Attention Networks Medical image understanding with pretrained vision language models: A comprehensive study,

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:25:10.557266Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T23:25:10.352529Z digest=sha256:bc5bae4096d5298dd026302121d41a656d465be5dd25f0c0112d3a7acb800ac2

Observation 64a15a57-9d45-4872-9e6a-b2d36ecd8e68 · outbound

This paper cites VLM-CPL: Consensus Pseudo Labels from Vision-Language Models for Annotation-Free Pathological Image Classification.

Multimodal Remote Sensing Scene Classification Using VLMs and Dual-Cross Attention Networks VLM-CPL: Consensus Pseudo Labels from Vision-Language Models for Annotation-Free Pathological Image Classification

Reference 39

Resolution
verified exact
local_arxiv, observed 2026-08-11T23:25:10.419574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T23:25:10.355724Z digest=sha256:e316201d310aae7cc12691c44ea1ed7b02dcc78bc5e902213c0c31c11d280959

Observation 42cf6449-0ff5-4167-a4f4-40965193b14d · outbound

This paper cites Improved zero-shot classification by adapting vlms with text descriptions,.

Multimodal Remote Sensing Scene Classification Using VLMs and Dual-Cross Attention Networks Improved zero-shot classification by adapting vlms with text descriptions,

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:25:10.547199Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T23:25:10.359331Z digest=sha256:ef0f96529b7cba28098a13c095ce1843744142230829cbac1eb7da5acad71681

Observation 33fd1be0-9798-49a7-b681-d2b250db52ac · outbound

This paper cites Exploiting lmm-based knowledge for image classification tasks,.

Multimodal Remote Sensing Scene Classification Using VLMs and Dual-Cross Attention Networks Exploiting lmm-based knowledge for image classification tasks,

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:25:10.536941Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T23:25:10.362556Z digest=sha256:89e7ef9ec24e395dfb2ece81320ca5a2809b69dffb2fdbd73a408a43aa643bd3

Observation 72d59773-5176-4f51-8408-cf63cff6aafa · outbound

This paper cites Visual instruction tuning,.

Multimodal Remote Sensing Scene Classification Using VLMs and Dual-Cross Attention Networks Visual instruction tuning,

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:25:10.526594Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T23:25:10.366014Z digest=sha256:c66f2b1a64dad0c32ed08daf5577f05ecbe207194bffc1cabc8aa1062eb9dbe1

Observation 80dc399b-2693-4174-ad2f-2f4381ea9fdf · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

Multimodal Remote Sensing Scene Classification Using VLMs and Dual-Cross Attention Networks An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-11T23:25:10.369305Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:25:10.369305Z digest=sha256:8901f029884ded318b43d8ab25d28f80fd47e17824c7973deca5e6290b6514f2

Observation 6a9a74f2-f280-466b-a6cb-68413ec2a2f5 · outbound

This paper cites Learning transferable visual models from natural language supervision,.

Multimodal Remote Sensing Scene Classification Using VLMs and Dual-Cross Attention Networks Learning transferable visual models from natural language supervision,

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-11T23:25:10.372826Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:25:10.372826Z digest=sha256:bd53c91fcb898aabb2741171f32e4aa7c6304f46abad441f88b6839e307fc736

Observation f43a904c-1605-4580-8050-66cb49da067f · outbound

This paper cites Segment anything,.

Multimodal Remote Sensing Scene Classification Using VLMs and Dual-Cross Attention Networks Segment anything,

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-11T23:25:10.376115Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:25:10.376115Z digest=sha256:d29ba96a7d8b32a3e47c5907321c14a8e94839fe359b349534845cf239076fe5

Pith citing papers

No inbound Pith citation observations are available.