Pith. sign in

Paper Citation Record · LEDGER

Vision-Language Modeling Meets Remote Sensing: Models, Datasets and Perspectives

As of 20 August 2026, this Paper Citation Record lists 100 of 299 outbound references and 1 inbound Pith citation observation for arXiv:2505.14361.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.14361 v1

Coverage vector

measured 100 of 299 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:38:55.619271Z

measured 101 of 101 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-01T05:11:58.947649Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-01T10:45:42.279098Z

Reference resolution

100 of 299 outbound references displayed

  • verified exact1
  • verified fuzzy0
  • unresolved99
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 3c13e333-f01a-4f75-a0a5-affae0186252 · outbound

This paper cites Rsgpt: A remote sensing vision language model and benchmark,.

Vision-Language Modeling Meets Remote Sensing: Models, Datasets and Perspectives Rsgpt: A remote sensing vision language model and benchmark,

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T15:38:43.064939Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:38:43.064939Z digest=sha256:9c8038622193fcbe5ec1a1a57fcd74dd8336c00d0c72f0c010f98776d41f9e3d

Observation 4d061b49-c907-4ddc-8d88-e98eff39cea3 · outbound

This paper cites Remoteclip: A vision language foundation model for remote sensing,.

Vision-Language Modeling Meets Remote Sensing: Models, Datasets and Perspectives Remoteclip: A vision language foundation model for remote sensing,

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T15:38:43.156146Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:38:43.156146Z digest=sha256:0597a7e0ed2fbf27f8b86910f8b1f9de7dc616020140c4e16415440f3951c811

Observation 28528bda-18a9-4da0-9c19-e9106bba1398 · outbound

This paper cites Geochat: Grounded large vision-language model for remote sensing,.

Vision-Language Modeling Meets Remote Sensing: Models, Datasets and Perspectives Geochat: Grounded large vision-language model for remote sensing,

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T15:38:43.283004Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:38:43.283004Z digest=sha256:47a82da135bf598315291128477ca0faf3545eb73461e8e1229285414db3266a

Observation 7edb7424-6f49-4b1a-844d-b3c0ca45e78f · outbound

This paper cites Remote sensing vision-language foundation models without annotations via ground remote alignment,.

Vision-Language Modeling Meets Remote Sensing: Models, Datasets and Perspectives Remote sensing vision-language foundation models without annotations via ground remote alignment,

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T15:38:43.452270Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:38:43.452270Z digest=sha256:78a20b7c89cc394efe42ebe610c41a61b5f78ac0781cb273c67d2577028178b0

Observation 80c7a8fa-64c7-4355-a09c-ba621b435a22 · outbound

This paper cites Skyeyegpt: Unifying remote sens- ing vision-language tasks via instruction tuning with large language model,.

Vision-Language Modeling Meets Remote Sensing: Models, Datasets and Perspectives Skyeyegpt: Unifying remote sens- ing vision-language tasks via instruction tuning with large language model,

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T15:38:43.608411Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:38:43.608411Z digest=sha256:61f5f886cd08236e36b37572992870a50300eb833f9ce67f93c325521b714311

Observation 8317d75c-4ee6-4c4c-aa23-a0b09e186326 · outbound

This paper cites Earthgpt: A universal multimodal large language model for multisensor image com- prehension in remote sensing domain,.

Vision-Language Modeling Meets Remote Sensing: Models, Datasets and Perspectives Earthgpt: A universal multimodal large language model for multisensor image com- prehension in remote sensing domain,

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T15:38:43.778323Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:38:43.778323Z digest=sha256:df7655d10ccd2b5ba58d9c0b15df84cae8c9049c4d3db97acd87002e44d025d2

Observation 9f968775-1c81-4dba-934d-ea48443625d5 · outbound

This paper cites Rs5m and georsclip: A large scale vision-language dataset and a large vision-language model for remote sensing,.

Vision-Language Modeling Meets Remote Sensing: Models, Datasets and Perspectives Rs5m and georsclip: A large scale vision-language dataset and a large vision-language model for remote sensing,

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T15:38:43.900412Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:38:43.900412Z digest=sha256:0a663d0e9a1c62a0d87b426fcdc6c3475fbc29798f9ae843b836dd6e0fff56c1

Observation 7991c0ca-874f-437a-bf88-f02088e1cb2d · outbound

This paper cites Skyscript: A large and semantically diverse vision-language dataset for remote sensing,.

Vision-Language Modeling Meets Remote Sensing: Models, Datasets and Perspectives Skyscript: A large and semantically diverse vision-language dataset for remote sensing,

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T15:38:44.024239Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:38:44.024239Z digest=sha256:6870c704266fbcb1238fdde1e9c6d97b31d4ba1d30126c581a89511e038b9e8c

Observation 85e592d4-8a4d-4a9b-8181-be5747a519a5 · outbound

This paper cites Lhrs-bot: Empowering remote sensing with vgi-enhanced large multimodal language model,.

Vision-Language Modeling Meets Remote Sensing: Models, Datasets and Perspectives Lhrs-bot: Empowering remote sensing with vgi-enhanced large multimodal language model,

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T15:38:44.181611Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:38:44.181611Z digest=sha256:b4754b9156f5c79197cc90439d9bd13d65c0a008e898a7725e9656ee2d21f0a7

Observation 6cf05be3-18b1-49f7-9990-60f465189842 · outbound

This paper cites Rs- llava: A large vision-language model for joint captioning and question answering in remote sensing imagery,.

Vision-Language Modeling Meets Remote Sensing: Models, Datasets and Perspectives Rs- llava: A large vision-language model for joint captioning and question answering in remote sensing imagery,

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T15:38:44.289917Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:38:44.289917Z digest=sha256:c4a557ba2cc8173cb77c7cac23932d4b759e0cb8577c63ef05005a09e5d92a97

Observation 9ae29f85-5d1f-4171-9a1e-bd4620c1b678 · outbound

This paper cites SkySenseGPT: A Fine-Grained Instruction Tuning Dataset and Model for Remote Sensing Vision-Language Understanding.

Vision-Language Modeling Meets Remote Sensing: Models, Datasets and Perspectives SkySenseGPT: A Fine-Grained Instruction Tuning Dataset and Model for Remote Sensing Vision-Language Understanding

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T15:38:44.401424Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:38:44.401424Z digest=sha256:f0778caf8deb58a9d19e271197b550a9bb0bd8a576e8f435d861de2c84f3ea94

Observation ff8d6127-e1ca-42d2-a67d-78004064ca42 · outbound

This paper cites Vrsbench: A versatile vision- language benchmark dataset for remote sensing image understanding,.

Vision-Language Modeling Meets Remote Sensing: Models, Datasets and Perspectives Vrsbench: A versatile vision- language benchmark dataset for remote sensing image understanding,

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T15:38:44.510141Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:38:44.510141Z digest=sha256:cae1d7e5e9bae92870e5c6dc7e267468fc34dbb2eaaeb3dfd636cbcc61c2b942

Observation 7e6441a9-928e-4ed8-ab1a-d9411385a790 · outbound

This paper cites Changeclip: Remote sens- ing change detection with multimodal vision-language representation learning,.

Vision-Language Modeling Meets Remote Sensing: Models, Datasets and Perspectives Changeclip: Remote sens- ing change detection with multimodal vision-language representation learning,

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T15:38:44.622345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:38:44.622345Z digest=sha256:b81e2317c9258533ee92f6ceb13468d888ea784ecab77db804a98abc09d09537

Observation 21ec66ba-65ee-4a29-8c3a-e81b9062f7dc · outbound

This paper cites Rs-clip: Zero shot remote sensing scene classification via contrastive vision-language supervision,.

Vision-Language Modeling Meets Remote Sensing: Models, Datasets and Perspectives Rs-clip: Zero shot remote sensing scene classification via contrastive vision-language supervision,

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T15:38:44.787635Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:38:44.787635Z digest=sha256:24e4aa09c07f62dd873a2172c7dd56faea806f07f304c54aca6832ba4a91bcec

Observation cb51dcff-43d7-4a42-8371-a7f4ee1c1383 · outbound

This paper cites Mind the Modality Gap: Towards a Remote Sensing Vision-Language Model via Cross-modal Alignment.

Vision-Language Modeling Meets Remote Sensing: Models, Datasets and Perspectives Mind the Modality Gap: Towards a Remote Sensing Vision-Language Model via Cross-modal Alignment

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T15:38:44.914015Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:38:44.914015Z digest=sha256:cb81bcda7e6ace1f326b73cebbbd9979846beff8894f86db9c591c6b3a32f53c

Observation c5713cf3-7683-4486-8685-24b06828bc24 · outbound

This paper cites Towards vision-language geo-foundation model: A survey,.

Vision-Language Modeling Meets Remote Sensing: Models, Datasets and Perspectives Towards vision-language geo-foundation model: A survey,

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T15:38:45.014750Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:38:45.014750Z digest=sha256:bbb093640936c18da4748beb8773b3af12a7d2694320d75707a1eee81f37fc6b

Observation 58a77bdf-717b-4601-bf47-2ba6fece9281 · outbound

This paper cites Vision-language models in remote sensing: Current progress and future trends,.

Vision-Language Modeling Meets Remote Sensing: Models, Datasets and Perspectives Vision-language models in remote sensing: Current progress and future trends,

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T15:38:45.170587Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:38:45.170587Z digest=sha256:e03095647065030d0f3ca630a5f99e6ba8d731e71d2861a517c9a78313095257

Observation 2ddba138-6449-4a73-a96c-46fe659847aa · outbound

This paper cites S-clip: Semi-supervised vision- language learning using few specialist captions,.

Vision-Language Modeling Meets Remote Sensing: Models, Datasets and Perspectives S-clip: Semi-supervised vision- language learning using few specialist captions,

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T15:38:45.235768Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:38:45.235768Z digest=sha256:b4cb09bfe91ca15553037bb310609ddc70e0e1140927802150b728b1c28c2996

Observation 962fc3d3-1b70-4987-8969-f0e56291152f · outbound

This paper cites Multi-view feature fusion and visual prompt for remote sensing image captioning,.

Vision-Language Modeling Meets Remote Sensing: Models, Datasets and Perspectives Multi-view feature fusion and visual prompt for remote sensing image captioning,

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T15:38:45.332966Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:38:45.332966Z digest=sha256:16b6a7505853a0c84ec683ddba5c5ce50d11b48011038e57c9f68a2816060195

Observation 34fe4cf5-db90-46af-aa20-a628f1c79d1a · outbound

This paper cites Vision- language models for zero-shot classification of remote sensing images,.

Vision-Language Modeling Meets Remote Sensing: Models, Datasets and Perspectives Vision- language models for zero-shot classification of remote sensing images,

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T15:38:45.468804Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:38:45.468804Z digest=sha256:473b6c118e66aaed63c491b9a06c1c11bdc563d4c1148ca99610d38d6e12a3c6

Observation 4591f1d6-c747-4afb-a06f-509911022cc7 · outbound

This paper cites Detecting cloud presence in satellite images using the rgb-based clip vision-language model,.

Vision-Language Modeling Meets Remote Sensing: Models, Datasets and Perspectives Detecting cloud presence in satellite images using the rgb-based clip vision-language model,

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T15:38:45.646020Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:38:45.646020Z digest=sha256:1193c8bb1e896920a21feb56ae446cafbfb3e6aa3e6e37deb09030653842cf3c

Observation 8fa01820-6df0-413f-8c19-17a775d5a98f · outbound

This paper cites Chatearthnet: a global- scale image-text dataset empowering vision-language geo-foundation models,.

Vision-Language Modeling Meets Remote Sensing: Models, Datasets and Perspectives Chatearthnet: a global- scale image-text dataset empowering vision-language geo-foundation models,

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T15:38:45.809267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:38:45.809267Z digest=sha256:507a4bf99d9902a3ac410e7533d9d86c5bc14c6e88c3853ff4ee8c6208d0dae8

Observation 9e13967b-f96d-4e1f-8a9b-3c50c95f9267 · outbound

This paper cites Bi-modal transformer-based approach for visual question answering in remote sensing imagery,.

Vision-Language Modeling Meets Remote Sensing: Models, Datasets and Perspectives Bi-modal transformer-based approach for visual question answering in remote sensing imagery,

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T15:38:45.933746Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:38:45.933746Z digest=sha256:955d59941f5f7e1badd32ad31c35accf53dc68147a4491eb7b5f0ff0b74fa8bc

Observation e010fd27-cd52-4b66-9022-b49a6f7b66c6 · outbound

This paper cites Popeye: A unified visual-language model for multi-source ship detection from remote sensing imagery,.

Vision-Language Modeling Meets Remote Sensing: Models, Datasets and Perspectives Popeye: A unified visual-language model for multi-source ship detection from remote sensing imagery,

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T15:38:46.052643Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:38:46.052643Z digest=sha256:9bb58fad9a1583409c8d4ccc64de79527308a0e004c3736f7f6f7e8c4f59c606

Observation f5d64ec1-2d08-49b0-9e2b-c884b2e27c38 · outbound

This paper cites Deep semantic-visual alignment for zero-shot remote sensing image scene classification,.

Vision-Language Modeling Meets Remote Sensing: Models, Datasets and Perspectives Deep semantic-visual alignment for zero-shot remote sensing image scene classification,

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T15:38:46.198276Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:38:46.198276Z digest=sha256:6dab81dcca14abde18edee74be2f7852be32d6d642aa0f2b85c9b3a2b3e13870

Observation 50a34e47-9fbc-4ba5-b65b-6eb24e18de1f · outbound

This paper cites Set-CLIP: Exploring Aligned Semantic From Low-Alignment Multimodal Data Through A Distribution View.

Vision-Language Modeling Meets Remote Sensing: Models, Datasets and Perspectives Set-CLIP: Exploring Aligned Semantic From Low-Alignment Multimodal Data Through A Distribution View

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-08-07T15:39:23.164372Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T15:38:46.336526Z digest=sha256:c759309dc9c5cf93d2b94fbd85c5bd61be6ad918b08b818da9c88f671de321bf

Observation 069b0439-49c8-441e-9d6e-9e43c7c134ad · outbound

This paper cites Segment change model (scm) for unsupervised change detection in vhr remote sensing images: a case study of buildings,.

Vision-Language Modeling Meets Remote Sensing: Models, Datasets and Perspectives Segment change model (scm) for unsupervised change detection in vhr remote sensing images: a case study of buildings,

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T15:38:46.502942Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:38:46.502942Z digest=sha256:c0f85b865fcbd7f57aee55a8b635b65ce11b161cb211267f148ae3a78b2f7f57

Observation dc1962f1-fe34-4f97-86b0-e2b6d62e2dde · outbound

This paper cites A new learning paradigm for foundation model-based remote-sensing change detection,.

Vision-Language Modeling Meets Remote Sensing: Models, Datasets and Perspectives A new learning paradigm for foundation model-based remote-sensing change detection,

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T15:38:46.641308Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:38:46.641308Z digest=sha256:7bba99cafb173e6b2908ab8b968be1564c476d4679fc77e46f6912060f46bfb2

Observation 2cc03d95-bcea-4853-aaac-ce83c706e8de · outbound

This paper cites Vlca: vision-language aligning model with cross-modal attention for bilingual remote sensing image captioning,.

Vision-Language Modeling Meets Remote Sensing: Models, Datasets and Perspectives Vlca: vision-language aligning model with cross-modal attention for bilingual remote sensing image captioning,

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T15:38:46.750482Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:38:46.750482Z digest=sha256:c4693a86b9fd7ec940ec6571ec47558c3ce66780d1d24a1ee449da5ae4e0f982

Observation 1c232cf3-5270-4340-81cf-8ad1dd03c2ae · outbound

This paper cites RS-GPT4V: A Unified Multimodal Instruction-Following Dataset for Remote Sensing Image Understanding.

Vision-Language Modeling Meets Remote Sensing: Models, Datasets and Perspectives RS-GPT4V: A Unified Multimodal Instruction-Following Dataset for Remote Sensing Image Understanding

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T15:38:46.888234Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:38:46.888234Z digest=sha256:b74919fdc646e328b9ddde01715697f5a22f285b8c86482c64e52f5695af1c65

Observation 0ddae0ae-3251-4bcb-a412-0779d4648ae9 · outbound

This paper cites Luojiahog: A hierarchy oriented geo-aware image caption dataset for remote sensing image–text retrieval,.

Vision-Language Modeling Meets Remote Sensing: Models, Datasets and Perspectives Luojiahog: A hierarchy oriented geo-aware image caption dataset for remote sensing image–text retrieval,

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T15:38:47.040530Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:38:47.040530Z digest=sha256:ef0823264cf3ed19511a889d309d7eb5a6d336a52fb3daa8854d29766d5a7700

Observation 2df3449b-f06b-4bfa-9f43-30aabb12c296 · outbound

This paper cites SATIN: A Multi-Task Metadataset for Classifying Satellite Imagery using Vision-Language Models.

Vision-Language Modeling Meets Remote Sensing: Models, Datasets and Perspectives SATIN: A Multi-Task Metadataset for Classifying Satellite Imagery using Vision-Language Models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T15:38:47.152096Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:38:47.152096Z digest=sha256:77058b975b8e696315f3b04bc4d4e25938c278854876447a5b9faf782b545d21

Observation 5f9f3b49-df9d-41cb-84d2-ad5c436dd404 · outbound

This paper cites Rsvg: Exploring data and models for visual grounding on remote sensing data,.

Vision-Language Modeling Meets Remote Sensing: Models, Datasets and Perspectives Rsvg: Exploring data and models for visual grounding on remote sensing data,

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T15:38:47.311652Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:38:47.311652Z digest=sha256:caa14f4fc854c6088ae09a6cf82ea20daa883c729af9c1c7c1b6c196008354c7

Observation 7fca31d7-7a71-41d2-8110-f68cd29bebeb · outbound

This paper cites Good at captioning bad at counting: Bench- marking gpt-4v on earth observation data,.

Vision-Language Modeling Meets Remote Sensing: Models, Datasets and Perspectives Good at captioning bad at counting: Bench- marking gpt-4v on earth observation data,

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T15:38:47.462892Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:38:47.462892Z digest=sha256:3598a654d4c3d72b00961a1e754f92a0181993d09ea12b0fea6a68877b92d216

Observation 240f4edb-64b2-4d99-8446-17152f569de6 · outbound

This paper cites Satclip: Global, general-purpose location embeddings with satellite imagery,.

Vision-Language Modeling Meets Remote Sensing: Models, Datasets and Perspectives Satclip: Global, general-purpose location embeddings with satellite imagery,

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T15:38:47.610998Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:38:47.610998Z digest=sha256:88132ac8269a30b4fb2628b3098b8b34f320ec52c7da3091e1f271b496d136a3

Observation c9e19baa-94ff-4442-b24d-75d4b146fd5b · outbound

This paper cites GeoCLIP: Clip-inspired alignment between locations and images for effective worldwide geo- localization,.

Vision-Language Modeling Meets Remote Sensing: Models, Datasets and Perspectives GeoCLIP: Clip-inspired alignment between locations and images for effective worldwide geo- localization,

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T15:38:47.742008Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:38:47.742008Z digest=sha256:cb208580edd1009911d58163a0b7397c7bf8dd3a90791faead9fee99f7141ca1

Observation b222e3fc-c3be-4be3-9492-1c3235c67cf6 · outbound

This paper cites Csp: Self-supervised contrastive spatial pre-training for geospatial-visual representations,.

Vision-Language Modeling Meets Remote Sensing: Models, Datasets and Perspectives Csp: Self-supervised contrastive spatial pre-training for geospatial-visual representations,

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T15:38:47.857469Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:38:47.857469Z digest=sha256:d2a0d0f0ca0242369ff38196317480d51453a98049eced97793c254a1e1b9dbc

Observation 5d3d29a9-925a-4789-a9e6-abb55d56fc23 · outbound

This paper cites Diffusionsat: A generative foundation model for satellite imagery,.

Vision-Language Modeling Meets Remote Sensing: Models, Datasets and Perspectives Diffusionsat: A generative foundation model for satellite imagery,

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T15:38:47.942387Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:38:47.942387Z digest=sha256:8d30dc13a865d8c9bc2cb9885bc245a3b77e5fa0029b684a48e45567943f195d

Observation d0624a66-64f5-40d7-a41d-597b684d9ab1 · outbound

This paper cites Crs-diff: Controllable remote sensing image generation with diffusion model,.

Vision-Language Modeling Meets Remote Sensing: Models, Datasets and Perspectives Crs-diff: Controllable remote sensing image generation with diffusion model,

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T15:38:48.144141Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:38:48.144141Z digest=sha256:4e9931ca9ec4bce7cf8744171925d8eee097f2b107d2ddea3fc8e45b55a9d6d8

Observation 431fd769-708a-4b4e-a89a-6f1e0aa122f9 · outbound

This paper cites Practical techniques for vision- language segmentation model in remote sensing,.

Vision-Language Modeling Meets Remote Sensing: Models, Datasets and Perspectives Practical techniques for vision- language segmentation model in remote sensing,

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T15:38:48.276207Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:38:48.276207Z digest=sha256:c90db2e5bbbbe6adfd0e19dc97ccca84954bbbf99b613b84265c843b68097cfd

Observation 023dc923-3566-4b26-b689-2de62c972118 · outbound

This paper cites Text2seg: Zero-shot remote sensing image semantic segmentation via text-guided visual foundation models,.

Vision-Language Modeling Meets Remote Sensing: Models, Datasets and Perspectives Text2seg: Zero-shot remote sensing image semantic segmentation via text-guided visual foundation models,

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T15:38:48.421970Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:38:48.421970Z digest=sha256:d005cf2f908954dbbcd451d0bc23a086032f8325c5c79f565000ddaf43ab6c66

Observation 261f59c7-267d-4a7e-96e2-a68940a51f15 · outbound

This paper cites A decoupling paradigm with prompt learning for remote sensing image change captioning,.

Vision-Language Modeling Meets Remote Sensing: Models, Datasets and Perspectives A decoupling paradigm with prompt learning for remote sensing image change captioning,

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T15:38:48.564888Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:38:48.564888Z digest=sha256:3cc8ae11cf5589a56b49ca7033815d743582ec8793b7b34e8941d5c33a326d33

Observation 97d41679-a69e-41ca-81bb-4d27abffb58a · outbound

This paper cites Bootstrapping interactive image-text alignment for remote sensing image captioning,.

Vision-Language Modeling Meets Remote Sensing: Models, Datasets and Perspectives Bootstrapping interactive image-text alignment for remote sensing image captioning,

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T15:38:48.738237Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:38:48.738237Z digest=sha256:abf304f133555e647d3c1dad8962efb62e737fb50f856be7854b9ee8ca713301

Observation 92996c50-923f-42e7-8516-296d00d0c1d0 · outbound

This paper cites Text- guided diverse image synthesis for long-tailed remote sensing object classification,.

Vision-Language Modeling Meets Remote Sensing: Models, Datasets and Perspectives Text- guided diverse image synthesis for long-tailed remote sensing object classification,

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T15:38:48.945772Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:38:48.945772Z digest=sha256:4bda3353835e018c98f2a5bdfbce2a9fedf3c5cfc9c41d301aaae3d2f599aee5

Observation 16f4bd2c-0ffd-4064-b3b6-752a3595ac07 · outbound

This paper cites Object detection in aerial images: A large-scale benchmark and challenges,.

Vision-Language Modeling Meets Remote Sensing: Models, Datasets and Perspectives Object detection in aerial images: A large-scale benchmark and challenges,

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T15:38:49.158228Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:38:49.158228Z digest=sha256:44eee36169d6f6dfe202c844ed31d503731b9941e5d05010bc93efab878ad030

Observation 977c3415-671e-43ac-87f8-48948426a08b · outbound

This paper cites InstructBLIP: Towards general-purpose vision-language models with instruction tuning,.

Vision-Language Modeling Meets Remote Sensing: Models, Datasets and Perspectives InstructBLIP: Towards general-purpose vision-language models with instruction tuning,

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T15:38:49.305788Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:38:49.305788Z digest=sha256:42197fa1ccf3c86a1e2194c20f697d1ff2918df6a26f8b63faa1bfe9f0c09592

Observation 2865e876-251a-47f7-8822-0dba7d917d7c · outbound

This paper cites Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality,.

Vision-Language Modeling Meets Remote Sensing: Models, Datasets and Perspectives Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality,

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T15:38:49.424549Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:38:49.424549Z digest=sha256:c9a7644205a1a1fc2fc9fb5c172fc4e24173dec38bf6ac0109f0600baa02682f

Observation dd525767-374d-41b4-90ae-70d878c23db1 · outbound

This paper cites Learning to rank question answer pairs with holographic dual lstm architecture,.

Vision-Language Modeling Meets Remote Sensing: Models, Datasets and Perspectives Learning to rank question answer pairs with holographic dual lstm architecture,

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T15:38:49.623916Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:38:49.623916Z digest=sha256:ae0313823daa3e2e358473e7dcf77f3a72b26fc4ab21ef6e48af6eaec5805ee5

Observation 83d2cb6a-951f-4d5b-a5ad-3acb6ec24634 · outbound

This paper cites Improved baselines with visual instruction tuning,.

Vision-Language Modeling Meets Remote Sensing: Models, Datasets and Perspectives Improved baselines with visual instruction tuning,

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T15:38:49.777690Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:38:49.777690Z digest=sha256:a53d132124544bdfe89fe5a65d9e15ce3d6eb7e9cc21c72e9f090bc7620951d9

Observation 9c5386c5-479e-4fd6-b4ee-5584bb0148ef · outbound

This paper cites Lora: Low-rank adaptation of large language models,.

Vision-Language Modeling Meets Remote Sensing: Models, Datasets and Perspectives Lora: Low-rank adaptation of large language models,

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T15:38:49.897162Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:38:49.897162Z digest=sha256:8bb5693562d12d1a7a32b1f2f0273bb7e01d546feb74a506ea8f6fd1f8515d4e

Observation e2d88973-a7c8-495d-926d-da6955896f03 · outbound

This paper cites Object detection in optical remote sensing images: A survey and a new benchmark,.

Vision-Language Modeling Meets Remote Sensing: Models, Datasets and Perspectives Object detection in optical remote sensing images: A survey and a new benchmark,

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T15:38:49.957361Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:38:49.957361Z digest=sha256:2699515e337b56140688c69990bc8f5f2e35e72e3e8fab1f7d649984a07a9b28

Observation c585a361-a717-4c94-a5f9-7885947955af · outbound

This paper cites Fair1m: A benchmark dataset for fine-grained object recognition in high-resolution remote sensing imagery,.

Vision-Language Modeling Meets Remote Sensing: Models, Datasets and Perspectives Fair1m: A benchmark dataset for fine-grained object recognition in high-resolution remote sensing imagery,

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T15:38:50.035586Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:38:50.035586Z digest=sha256:366f72b8eb8e92da9707aa576b7f25c0520968aadb27044a85270fb424f5b2e1

Observation 1929c793-e415-4b4d-9a32-587ad25b5520 · outbound

This paper cites Remote sensing image scene classifi- cation: Benchmark and state of the art,.

Vision-Language Modeling Meets Remote Sensing: Models, Datasets and Perspectives Remote sensing image scene classifi- cation: Benchmark and state of the art,

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T15:38:50.099116Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:38:50.099116Z digest=sha256:16dbe49bede1f6a41c19f3e91a8ee6a8705a7b2352d1d4f4b79d7abc84c42804

Observation 0a3af32e-b7a6-4bd1-a80a-7a6cdc3a5518 · outbound

This paper cites Rsvqa: Visual question answering for remote sensing data,.

Vision-Language Modeling Meets Remote Sensing: Models, Datasets and Perspectives Rsvqa: Visual question answering for remote sensing data,

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T15:38:50.178928Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:38:50.178928Z digest=sha256:850609484bf66fbb10aa925f3005b5df36106f17e92583db3a217fc1ed353870

Observation 7dc08abe-90cb-4c6b-b0d0-77e74f7a23d9 · outbound

This paper cites Floodnet: A high resolution aerial imagery dataset for post flood scene understanding,.

Vision-Language Modeling Meets Remote Sensing: Models, Datasets and Perspectives Floodnet: A high resolution aerial imagery dataset for post flood scene understanding,

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T15:38:50.251075Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:38:50.251075Z digest=sha256:d6fd4258773835ecf4c18711c5019c218a44ab547a52ce79b762f62d0d098d9d

Observation 5500eab6-e1e9-4c38-9631-29edbf09d0df · outbound

This paper cites Eva: Exploring the limits of masked visual representation learning at scale,.

Vision-Language Modeling Meets Remote Sensing: Models, Datasets and Perspectives Eva: Exploring the limits of masked visual representation learning at scale,

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T15:38:50.331731Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:38:50.331731Z digest=sha256:a5a3b1cf2aba84fa5c1a61cde01c943292d720c37ea8aed6f53d0943ec230392

Observation 1c4f9e10-9e25-444d-86ff-83b58f8829a4 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Vision-Language Modeling Meets Remote Sensing: Models, Datasets and Perspectives Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T15:38:50.389463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:38:50.389463Z digest=sha256:6f97905f16c0dac0f0b11243be3ab5a209ac72cc4409445771d5325c8b120bbd

Observation d4b0ffc7-3a98-4cb9-96ee-0dd05d2b16e0 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Vision-Language Modeling Meets Remote Sensing: Models, Datasets and Perspectives LLaMA: Open and Efficient Foundation Language Models

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-07T15:38:50.452910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:38:50.452910Z digest=sha256:8b9b4472c989a0f458d3105e6bc4e84e02cd6fee8d0a328e899334adb7652d1c

Observation c88bc01b-3d2c-4893-8763-3df3bf9063a1 · outbound

This paper cites MiniGPT-v2: large language model as a unified interface for vision-language multi-task learning.

Vision-Language Modeling Meets Remote Sensing: Models, Datasets and Perspectives MiniGPT-v2: large language model as a unified interface for vision-language multi-task learning

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-07T15:38:50.540598Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:38:50.540598Z digest=sha256:c62eac7885517b6937c9d4a59a2934a7af57d1b78db9cab8020f2f75d54be0a1

Observation 31a8632f-e268-4e83-93ab-9ab05bc6b133 · outbound

This paper cites Exploring models and data for remote sensing image caption generation,.

Vision-Language Modeling Meets Remote Sensing: Models, Datasets and Perspectives Exploring models and data for remote sensing image caption generation,

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T15:38:50.620624Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:38:50.620624Z digest=sha256:13eb4874b4e560fe94cb2bafc2c3b304a25ceea29ceff76ed04f2d3ed56d5ce1

Observation 1440f466-6029-4b44-b24c-01920da4a6d2 · outbound

This paper cites Exploring a fine-grained multiscale method for cross-modal remote sensing image retrieval,.

Vision-Language Modeling Meets Remote Sensing: Models, Datasets and Perspectives Exploring a fine-grained multiscale method for cross-modal remote sensing image retrieval,

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-07T15:38:50.727487Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:38:50.727487Z digest=sha256:cf8bb90aaf247f8a301d23a9a108271c020b1bd12b55a80fc6ab6cf79ef7c258

Observation 89e8dec0-2306-41a3-8e16-e1017e737596 · outbound

This paper cites Deep semantic understanding of high resolution remote sensing image,.

Vision-Language Modeling Meets Remote Sensing: Models, Datasets and Perspectives Deep semantic understanding of high resolution remote sensing image,

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-07T15:38:50.886595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:38:50.886595Z digest=sha256:faa23c9a2f8c1fe6b004cc85e20b0d9dd3129cebc46e269f9fd7e05a62b4f491

Observation d18e7c63-ea7e-4c00-8f97-e88ef95621ac · outbound

This paper cites Nwpu- captions dataset and mlca-net for remote sensing image captioning,.

Vision-Language Modeling Meets Remote Sensing: Models, Datasets and Perspectives Nwpu- captions dataset and mlca-net for remote sensing image captioning,

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-07T15:38:51.023865Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:38:51.023865Z digest=sha256:cca0d5aaa4af8fe3fb28a4fb316caf067d6bf50036e24b2ba41e01f4b04d7035

Observation 7f3341c5-52cf-46e0-8ea0-0256fed5e53e · outbound

This paper cites Capera: Captioning events in aerial videos,.

Vision-Language Modeling Meets Remote Sensing: Models, Datasets and Perspectives Capera: Captioning events in aerial videos,

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-07T15:38:51.137450Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:38:51.137450Z digest=sha256:9600d3e68d57cf648bd0916d2885890d83268e1e158c8cd892471cb0fad67579

Observation 25f376d8-5c9c-435d-bc12-7a7ed19908a3 · outbound

This paper cites Mutual attention inception network for remote sensing visual question answering,.

Vision-Language Modeling Meets Remote Sensing: Models, Datasets and Perspectives Mutual attention inception network for remote sensing visual question answering,

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-07T15:38:51.255800Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:38:51.255800Z digest=sha256:15c4d79d39c3885f72742e170907ed91705e3ff817d5230c805720075fd0652d

Observation b6989cdd-a7e0-46fe-a1e1-60f529be1bc9 · outbound

This paper cites Visual grounding in remote sensing images,.

Vision-Language Modeling Meets Remote Sensing: Models, Datasets and Perspectives Visual grounding in remote sensing images,

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-07T15:38:51.389367Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:38:51.389367Z digest=sha256:e16a3ac78d8beb6a2785e9915779806695d41536d213b1472cb652c052550c6f

Observation a2e78898-96b9-4583-a0a5-d2aba730a934 · outbound

This paper cites Dinov2: Learning robust visual features without supervision,.

Vision-Language Modeling Meets Remote Sensing: Models, Datasets and Perspectives Dinov2: Learning robust visual features without supervision,

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-07T15:38:51.485572Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:38:51.485572Z digest=sha256:326295bb1f1fa94c1c7e882b6e23023c43b90c8e40e551a0e0a626f0702d7129

Observation 9d34bc30-b6d2-4f1d-aebf-8a53983c8299 · outbound

This paper cites Multistep question-driven visual question answering for remote sensing,.

Vision-Language Modeling Meets Remote Sensing: Models, Datasets and Perspectives Multistep question-driven visual question answering for remote sensing,

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-07T15:38:51.603773Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:38:51.603773Z digest=sha256:4f0d7c916baf7d27d6c391399a4391aac9b1d8e13da0670bb5a6719aa76a9604

Observation ccf42811-0ac5-4cd3-8ef2-9250060d525c · outbound

This paper cites Eurosat: A novel dataset and deep learning benchmark for land use and land cover classification,.

Vision-Language Modeling Meets Remote Sensing: Models, Datasets and Perspectives Eurosat: A novel dataset and deep learning benchmark for land use and land cover classification,

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-07T15:38:51.743120Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:38:51.743120Z digest=sha256:0ad24a2b0ced30dc009b621bbc9cd51148ce407c235e3daa0c345306e56ff1d8

Observation 3ea23881-6c26-41c5-9df3-ad95e3888724 · outbound

This paper cites Bag-of-visual-words and spatial extensions for land-use classification,.

Vision-Language Modeling Meets Remote Sensing: Models, Datasets and Perspectives Bag-of-visual-words and spatial extensions for land-use classification,

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-07T15:38:51.805151Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:38:51.805151Z digest=sha256:ebd04e825a752d34a996c6a2022efa7109716a023ec610da218492880a4096d1

Observation 58fa19a6-1021-4da3-bed6-8d4e35b57de7 · outbound

This paper cites Satellite image classification via two-layer sparse coding with biased image representation,.

Vision-Language Modeling Meets Remote Sensing: Models, Datasets and Perspectives Satellite image classification via two-layer sparse coding with biased image representation,

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-07T15:38:51.888925Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:38:51.888925Z digest=sha256:6c3c8d0a82d2084fd6d42bfb16067fffb5295b1cf3b0cf149f8f60ff5d8a8e92

Observation 7c38ff30-6b01-4a80-b7e1-d6e732b0cb26 · outbound

This paper cites Deep learning based feature selection for remote sensing scene classification,.

Vision-Language Modeling Meets Remote Sensing: Models, Datasets and Perspectives Deep learning based feature selection for remote sensing scene classification,

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-07T15:38:52.018569Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:38:52.018569Z digest=sha256:b62bd44d5499b1cfdf39f3807ffc47b8183b7998660b8bbdaff6bc8267e2d1d2

Observation 10335210-ebb5-4d98-9f17-e7c17fa4500d · outbound

This paper cites A public dataset for ship classification in remote sensing images,.

Vision-Language Modeling Meets Remote Sensing: Models, Datasets and Perspectives A public dataset for ship classification in remote sensing images,

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-07T15:38:52.137785Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:38:52.137785Z digest=sha256:36ef1e5792870f0171ab40436a773b72bfb983bd42986ee3750f3b173f869db4

Observation feb5e806-9622-4061-ad42-36407c2f244d · outbound

This paper cites A public dataset for fine-grained ship classification in optical remote sensing images,.

Vision-Language Modeling Meets Remote Sensing: Models, Datasets and Perspectives A public dataset for fine-grained ship classification in optical remote sensing images,

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-07T15:38:52.252457Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:38:52.252457Z digest=sha256:b38f67efc27a00577394fe943161340adfac6f222b6d1b651d9d64e52100d5bd

Observation 0d191b95-75dc-4fbe-845b-24ce179a9a42 · outbound

This paper cites Rotation-insensitive and context- augmented object detection in remote sensing images,.

Vision-Language Modeling Meets Remote Sensing: Models, Datasets and Perspectives Rotation-insensitive and context- augmented object detection in remote sensing images,

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-07T15:38:52.381176Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:38:52.381176Z digest=sha256:f78367dea1f847e3e0ee1adaeab721b1ceabb13bd74884c876667e765c171a82

Observation 987be9d5-082a-450f-a7de-9a3295fa2199 · outbound

This paper cites Hierarchical and robust convolutional neural network for very high-resolution remote sensing object detection,.

Vision-Language Modeling Meets Remote Sensing: Models, Datasets and Perspectives Hierarchical and robust convolutional neural network for very high-resolution remote sensing object detection,

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-07T15:38:52.474659Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:38:52.474659Z digest=sha256:b41cfd2b3323d0275c273099b6b58f6ef4e7794c3cff8a8664fd026d1d52d8ef

Observation 38664c2c-4e0d-44bc-bc9f-1bfaab473bec · outbound

This paper cites Orientation robust object detection in aerial images using deep convolutional neural network,.

Vision-Language Modeling Meets Remote Sensing: Models, Datasets and Perspectives Orientation robust object detection in aerial images using deep convolutional neural network,

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-07T15:38:52.600351Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:38:52.600351Z digest=sha256:df4358c9fd38221dfdd7afcd216670a015e65319a7216ac26bbe8cec285f3675

Observation 81d5eaa9-da89-4d24-89c5-e57d9c5408c1 · outbound

This paper cites Detection and tracking meet drones challenge,.

Vision-Language Modeling Meets Remote Sensing: Models, Datasets and Perspectives Detection and tracking meet drones challenge,

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-07T15:38:52.772576Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:38:52.772576Z digest=sha256:3ae36559023293a02e6ffc68cd959e808c22248130fc35be5c817a92c445d011

Observation d30c5c91-1183-494a-82a4-8f749811cec9 · outbound

This paper cites Sar ship detection dataset (ssdd): Official release and comprehensive data analysis,.

Vision-Language Modeling Meets Remote Sensing: Models, Datasets and Perspectives Sar ship detection dataset (ssdd): Official release and comprehensive data analysis,

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-07T15:38:52.902513Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:38:52.902513Z digest=sha256:9bf3355233271ad959e1278ec7eee630d13f2c23fc64aa47daca9adad74de3ff

Observation d759d29c-a0ae-45d9-881c-8c582e48f16b · outbound

This paper cites Hit-uav: A high-altitude infrared thermal dataset for unmanned aerial vehicle- based object detection,.

Vision-Language Modeling Meets Remote Sensing: Models, Datasets and Perspectives Hit-uav: A high-altitude infrared thermal dataset for unmanned aerial vehicle- based object detection,

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-07T15:38:53.061271Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:38:53.061271Z digest=sha256:353531b69909590bcc079d3a8dbb50624724b0043e3805ea4d63feb802e3317c

Observation 208bc1da-2418-4272-9cf4-39f93be24714 · outbound

This paper cites Sea-shipping,.

Vision-Language Modeling Meets Remote Sensing: Models, Datasets and Perspectives Sea-shipping,

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-07T15:38:53.201727Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:38:53.201727Z digest=sha256:97e48b59c962d227d9bc09caa513fa94152bfca82a8c1c6af65ef31088d442f5

Observation 799f5ea0-37b9-4163-8f3e-e1e650e6c49d · outbound

This paper cites Infrared-security,.

Vision-Language Modeling Meets Remote Sensing: Models, Datasets and Perspectives Infrared-security,

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-07T15:38:53.334467Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:38:53.334467Z digest=sha256:72e55f34fb9bfd3862941be6a879dac00206dafab594d467b62483959f34d5f3

Observation 22d82886-fbba-4fa1-aa18-76e1ab140edd · outbound

This paper cites Aerial-mancar,.

Vision-Language Modeling Meets Remote Sensing: Models, Datasets and Perspectives Aerial-mancar,

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-07T15:38:53.458005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:38:53.458005Z digest=sha256:0497e871fc6a7f13e9792011074d242d2f0dabef7dc51c9a624eca5210bfc930

Observation 934f876c-0fa5-459f-9829-7b47cf2efbb0 · outbound

This paper cites Double-light-vehicle,.

Vision-Language Modeling Meets Remote Sensing: Models, Datasets and Perspectives Double-light-vehicle,

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-07T15:38:53.615236Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:38:53.615236Z digest=sha256:4fdb620bdfa8708e2e83af9c0af0a444a000cf27f30aa96629f7b08ff4634553

Observation c9252e02-6cb8-4857-a81d-dabf6f28d5f6 · outbound

This paper cites Oceanic-ship,.

Vision-Language Modeling Meets Remote Sensing: Models, Datasets and Perspectives Oceanic-ship,

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-07T15:38:53.767902Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:38:53.767902Z digest=sha256:864f09fb4be407a408fa649ab48c6b3883f87f305b764fb62dd804aa867de5e9

Observation bd39e153-1de2-4451-b5bf-85dd1427fe7e · outbound

This paper cites Learning transferable visual models from natural language supervision,.

Vision-Language Modeling Meets Remote Sensing: Models, Datasets and Perspectives Learning transferable visual models from natural language supervision,

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-07T15:38:53.856194Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:38:53.856194Z digest=sha256:a7afa62891b13fe26cc673d93c90a8b6bcc2f6952257fc0b3924ad550061abc3

Observation 8d669b70-2684-4702-98cd-2f6004dbf9cc · outbound

This paper cites A novel svm-based decoder for remote sensing image captioning,.

Vision-Language Modeling Meets Remote Sensing: Models, Datasets and Perspectives A novel svm-based decoder for remote sensing image captioning,

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-07T15:38:53.940975Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:38:53.940975Z digest=sha256:cdfc779d3ed78cfe061c42481cde689915780b7cff44976e02278df0a6f16adf

Observation f94161d5-5c71-4a9e-a1f3-0b29eebc6008 · outbound

This paper cites Star: A first- ever dataset and a large-scale benchmark for scene graph generation in large-size satellite imagery,.

Vision-Language Modeling Meets Remote Sensing: Models, Datasets and Perspectives Star: A first- ever dataset and a large-scale benchmark for scene graph generation in large-size satellite imagery,

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-07T15:38:54.042866Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:38:54.042866Z digest=sha256:5de834c9fd92174d29b2da5c3eb696699eb265dcf820bea8e2d35eb4fe8cf9ba

Observation 0f6ab84d-fb29-4562-8e98-555a0e36d4b6 · outbound

This paper cites Fine-grained recognition for oriented ship against complex scenes in optical remote sensing images,.

Vision-Language Modeling Meets Remote Sensing: Models, Datasets and Perspectives Fine-grained recognition for oriented ship against complex scenes in optical remote sensing images,

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-07T15:38:54.193358Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:38:54.193358Z digest=sha256:202b3b23ff88b5fffd7184a5bf3f6a1b4d34a8756d9130e511c1db1c4f981ed0

Observation 12e4f184-fb94-4dd7-a226-7c059ac8fdb4 · outbound

This paper cites Hrsid: A high-resolution sar images dataset for ship detection and instance segmentation,.

Vision-Language Modeling Meets Remote Sensing: Models, Datasets and Perspectives Hrsid: A high-resolution sar images dataset for ship detection and instance segmentation,

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-07T15:38:54.375287Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:38:54.375287Z digest=sha256:74e5be721bcf5bd666f19422438c9c83838f1202639015e142000b3671c6098c

Observation 9767e98b-45c8-4f16-bb92-16656e945f49 · outbound

This paper cites Deepsat: a learning framework for satellite imagery,.

Vision-Language Modeling Meets Remote Sensing: Models, Datasets and Perspectives Deepsat: a learning framework for satellite imagery,

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-07T15:38:54.541073Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:38:54.541073Z digest=sha256:00aba9c5ba1c96def93ec1b72cb39f0cb3d4009a148973eec8a4b7c6fbfc0882

Observation 092404e3-5ee1-4800-9439-a696adec9209 · outbound

This paper cites Nasc-tg2: Natural scene classification with tiangong-2 remotely sensed imagery,.

Vision-Language Modeling Meets Remote Sensing: Models, Datasets and Perspectives Nasc-tg2: Natural scene classification with tiangong-2 remotely sensed imagery,

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-07T15:38:54.677808Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:38:54.677808Z digest=sha256:d9ea93dc021b70f0cdc053741c7be7a240211aa05edad1299899ae13f2cfbf15

Observation 24ddc95e-49e5-4702-afdd-33cfdf9e1ce2 · outbound

This paper cites Feature significance-based multibag- of-visual-words model for remote sensing image scene classification,.

Vision-Language Modeling Meets Remote Sensing: Models, Datasets and Perspectives Feature significance-based multibag- of-visual-words model for remote sensing image scene classification,

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-07T15:38:54.847291Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:38:54.847291Z digest=sha256:0f0e977b92e99610677c14167acc928f2b947e2fb6a4000d01f06c39d2e8d333

Observation 6b49a768-8919-4179-afc5-6c7c478e92c9 · outbound

This paper cites Dirichlet-derived multiple topic scene classification model for high spatial resolution remote sensing imagery,.

Vision-Language Modeling Meets Remote Sensing: Models, Datasets and Perspectives Dirichlet-derived multiple topic scene classification model for high spatial resolution remote sensing imagery,

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-07T15:38:54.995650Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:38:54.995650Z digest=sha256:79ac459f38502467832187b051e3bd687e227cbd1cec472308dae121c965730e

Observation 25deeedd-14ab-498b-a3d1-7a4f4784a03f · outbound

This paper cites Patternnet: A benchmark dataset for performance evaluation of remote sensing image retrieval,.

Vision-Language Modeling Meets Remote Sensing: Models, Datasets and Perspectives Patternnet: A benchmark dataset for performance evaluation of remote sensing image retrieval,

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-07T15:38:55.164443Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:38:55.164443Z digest=sha256:342d3bfed8399a893e8818978aca5c4b7618af6b458867308b8cdc64d04cc08e

Observation 3a7e90a7-a976-4e2f-ba1d-b185f8b532f6 · outbound

This paper cites Accurate object localization in remote sensing images based on convolutional neural networks,.

Vision-Language Modeling Meets Remote Sensing: Models, Datasets and Perspectives Accurate object localization in remote sensing images based on convolutional neural networks,

Reference 96

Resolution
unresolved
no resolver link, observed 2026-08-07T15:38:55.256524Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:38:55.256524Z digest=sha256:57348d6b8cfcdef153b97ca430df8a0d2b993d8decbfbcdbcb61ded692f61e70

Observation d7d6f5c8-e6fa-446a-b1cc-fb89c5014327 · outbound

This paper cites Land-cover classification with high-resolution remote sensing images using transferable deep models,.

Vision-Language Modeling Meets Remote Sensing: Models, Datasets and Perspectives Land-cover classification with high-resolution remote sensing images using transferable deep models,

Reference 97

Resolution
unresolved
no resolver link, observed 2026-08-07T15:38:55.326585Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:38:55.326585Z digest=sha256:6ac664e9c01323b13bf3f8a623de51f1db6897e60c152ca73b33fb8c0fb0e4f4

Observation 101371cc-8551-4285-99eb-cec4523d8062 · outbound

This paper cites Scene classification with recurrent attention of vhr remote sensing images,.

Vision-Language Modeling Meets Remote Sensing: Models, Datasets and Perspectives Scene classification with recurrent attention of vhr remote sensing images,

Reference 98

Resolution
unresolved
no resolver link, observed 2026-08-07T15:38:55.387112Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:38:55.387112Z digest=sha256:4dc3d365a91b3edd90093a023d488f708e34a9d0a5857b99ac61c53d9d47db56

Observation 3e4482f4-aff9-45e9-b94c-ed82e06fa7af · outbound

This paper cites Clrs: Continual learning benchmark for remote sensing image scene classification,.

Vision-Language Modeling Meets Remote Sensing: Models, Datasets and Perspectives Clrs: Continual learning benchmark for remote sensing image scene classification,

Reference 99

Resolution
unresolved
no resolver link, observed 2026-08-07T15:38:55.503910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:38:55.503910Z digest=sha256:5c6ce3d62e07767505adaefa324f12b958917b32e0634cd423c182467aaba8b4

Observation 29b7051a-932f-48f1-8806-dfd06eb16a7c · outbound

This paper cites On creating benchmark dataset for aerial image interpreta- tion: Reviews, guidances, and million-aid,.

Vision-Language Modeling Meets Remote Sensing: Models, Datasets and Perspectives On creating benchmark dataset for aerial image interpreta- tion: Reviews, guidances, and million-aid,

Reference 100

Resolution
unresolved
no resolver link, observed 2026-08-07T15:38:55.619271Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:38:55.619271Z digest=sha256:855567c23740232327983f802896f300c33352adf9eea324e972dc12cb20a15f

Pith citing papers

Observation f489b34b-47d3-4a3c-8fb1-24e0b4a0f13a · inbound

TreeAgent: A Generalizable Multi-Agent Framework for Automated Bias Labeling in Forestry via Compiled Expert Rules and Vision-Language Models cites this paper.

TreeAgent: A Generalizable Multi-Agent Framework for Automated Bias Labeling in Forestry via Compiled Expert Rules and Vision-Language Models Vision-Language Modeling Meets Remote Sensing: Models, Datasets and Perspectives

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-07-01T10:45:42.280738Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-07-01T05:11:58.947649Z digest=sha256:ca324cb0035323277bd1847df4694934083d10e58d63f0778a8c77b3221b4eee