Pith. sign in

Paper Citation Record · LEDGER

Remote Sensing Large Vision-Language Model: Semantic-augmented Multi-level Alignment and Semantic-aware Expert Modeling

As of 17 August 2026, this Paper Citation Record lists 77 of 77 outbound references and 1 inbound Pith citation observation for arXiv:2506.21863.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.21863 v1

Coverage vector

measured 77 of 77 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T22:22:41.831349Z

measured 78 of 78 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-10T13:48:08.135538Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-10T14:10:29.814089Z

Reference resolution

77 of 77 outbound references displayed

  • verified exact0
  • verified fuzzy30
  • unresolved47
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 26b619e7-1fd1-4cff-a1b7-3a42e454f4de · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Remote Sensing Large Vision-Language Model: Semantic-augmented Multi-level Alignment and Semantic-aware Expert Modeling LLaMA: Open and Efficient Foundation Language Models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T22:22:35.223058Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:22:35.223058Z digest=sha256:e046611bfb03f17d5c234958b5f22325a8b646b7d0331939120aaade36698560

Observation a367dd25-5608-4ce3-91e5-54c2a34bcb63 · outbound

This paper cites Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality,.

Remote Sensing Large Vision-Language Model: Semantic-augmented Multi-level Alignment and Semantic-aware Expert Modeling Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality,

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T22:22:36.850304Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:22:36.850304Z digest=sha256:7d9c6f8a69b19bf090a7539481c8d23fe7969de4d8f096141ef288227824a94e

Observation 011fc25c-32c0-4045-9e49-0680ba162714 · outbound

This paper cites Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone.

Remote Sensing Large Vision-Language Model: Semantic-augmented Multi-level Alignment and Semantic-aware Expert Modeling Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T22:22:36.905458Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:22:36.905458Z digest=sha256:164b433407e005985d27040666150703fbe8949658a5c7618040529401dce173

Observation b9800eb2-efdb-4547-a677-6792b4e84941 · outbound

This paper cites Stanford alpaca: An instruction-following llama model,.

Remote Sensing Large Vision-Language Model: Semantic-augmented Multi-level Alignment and Semantic-aware Expert Modeling Stanford alpaca: An instruction-following llama model,

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T22:22:36.997126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:22:36.997126Z digest=sha256:e2e395d3545206cfa2e20e077d3d427449df31e51ae234708126e6a9d1c25865

Observation 8ed410b6-ec2c-4dfe-9cb6-af69b0c31ba2 · outbound

This paper cites GPT-4 Technical Report.

Remote Sensing Large Vision-Language Model: Semantic-augmented Multi-level Alignment and Semantic-aware Expert Modeling GPT-4 Technical Report

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T22:22:37.102606Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:22:37.102606Z digest=sha256:a7283b7c4600689920fc80757e02a66772b8bb69796a9dfbf9e73edb0d21bcdf

Observation ceffb3a3-2a06-46e5-8057-8f3c954513fc · outbound

This paper cites Learning transferable visual models from natural language supervision,.

Remote Sensing Large Vision-Language Model: Semantic-augmented Multi-level Alignment and Semantic-aware Expert Modeling Learning transferable visual models from natural language supervision,

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T22:22:37.201063Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:22:37.201063Z digest=sha256:92fbfdd83733bf867155d5b31bac90f0cdd20be94d9c2610c751160112a42aa8

Observation 0f3c7262-fc54-4fe3-9321-b2dd64f1f0d1 · outbound

This paper cites Blip-2: Bootstrapping language- image pre-training with frozen image encoders and large language models,.

Remote Sensing Large Vision-Language Model: Semantic-augmented Multi-level Alignment and Semantic-aware Expert Modeling Blip-2: Bootstrapping language- image pre-training with frozen image encoders and large language models,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:23:08.137893Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T22:22:37.282011Z digest=sha256:c376f2d95e80c4b97dfbf6143bc19fceb7dff6058f4d16e3f1c4103c7ccbef08

Observation d7472608-b0c8-4bb4-97d0-3f532735ae94 · outbound

This paper cites Instructblip: towards general-purpose vision-language mod- els with instruction tuning,.

Remote Sensing Large Vision-Language Model: Semantic-augmented Multi-level Alignment and Semantic-aware Expert Modeling Instructblip: towards general-purpose vision-language mod- els with instruction tuning,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:23:08.025541Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T22:22:37.307798Z digest=sha256:657704fe72b4888ce9258365c666ead04a895c83ff92f9cee53ffc5b80321fa3

Observation be8d938d-911a-477e-be1d-78920bae6ad8 · outbound

This paper cites Visual instruction tuning,.

Remote Sensing Large Vision-Language Model: Semantic-augmented Multi-level Alignment and Semantic-aware Expert Modeling Visual instruction tuning,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:23:07.904765Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T22:22:37.340244Z digest=sha256:76239d2ad97ff987998f6af0ec9adc2bd34317ecd95148adc42ecd17dd82a116

Observation 35dd2474-e173-4e13-abf0-5519d5d892b4 · outbound

This paper cites Cogvlm: Visual expert for pretrained language models,.

Remote Sensing Large Vision-Language Model: Semantic-augmented Multi-level Alignment and Semantic-aware Expert Modeling Cogvlm: Visual expert for pretrained language models,

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T22:22:37.379971Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:22:37.379971Z digest=sha256:1e9214259d6ada88003d638fea5795d693fe0802233669b330ed6cc7e09914d3

Observation ba83cc8c-8833-4ca0-8872-f73e3468f0f7 · outbound

This paper cites mplug-owl2: Revolutionizing multi-modal large language model with modality collaboration,.

Remote Sensing Large Vision-Language Model: Semantic-augmented Multi-level Alignment and Semantic-aware Expert Modeling mplug-owl2: Revolutionizing multi-modal large language model with modality collaboration,

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T22:22:37.425036Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:22:37.425036Z digest=sha256:bda1d7d03618db3a49186dc14b20d08d4e5052f086890878475b92a403e95571

Observation bbf9b0fa-3c69-4555-ab15-6338cf0f33ae · outbound

This paper cites Flamingo: a visual language model for few-shot learning,.

Remote Sensing Large Vision-Language Model: Semantic-augmented Multi-level Alignment and Semantic-aware Expert Modeling Flamingo: a visual language model for few-shot learning,

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T22:22:37.458153Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:22:37.458153Z digest=sha256:f9139da159467ac0446427e21d465b18c5fb60598b6294906546dfab1023a5e3

Observation fa6b6200-8df9-4657-bc92-5e9670131043 · outbound

This paper cites Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks,.

Remote Sensing Large Vision-Language Model: Semantic-augmented Multi-level Alignment and Semantic-aware Expert Modeling Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks,

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T22:22:37.512451Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:22:37.512451Z digest=sha256:12eeefd6fe9315cfb42771473da1614443ac3f2a15fc4237180c2d659d3c4da2

Observation dd754c83-1607-4e06-badf-284d63151ec8 · outbound

This paper cites How far are we to gpt-4v? closing the gap to commercial multimodal models with open-source suites,.

Remote Sensing Large Vision-Language Model: Semantic-augmented Multi-level Alignment and Semantic-aware Expert Modeling How far are we to gpt-4v? closing the gap to commercial multimodal models with open-source suites,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:23:07.740051Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T22:22:37.555854Z digest=sha256:f11c38942220e6ecf5bb75d5208ba35312c49fdb615a32f6f0acf85ec5afee74

Observation f94173a1-c875-42c0-b7a9-a514fe7be0c8 · outbound

This paper cites Qwen2.5-VL Technical Report.

Remote Sensing Large Vision-Language Model: Semantic-augmented Multi-level Alignment and Semantic-aware Expert Modeling Qwen2.5-VL Technical Report

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T22:22:37.581702Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:22:37.581702Z digest=sha256:c51e61cbdfbcf71caf157cb37929272867c90a815864922f022c4f9fb817c012

Observation 3b8029ff-bebd-4799-878e-2bc272449bc2 · outbound

This paper cites Multimodal Chain-of-Thought Reasoning in Language Models.

Remote Sensing Large Vision-Language Model: Semantic-augmented Multi-level Alignment and Semantic-aware Expert Modeling Multimodal Chain-of-Thought Reasoning in Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T22:22:37.620896Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:22:37.620896Z digest=sha256:cc9e823c01cd783d6ca97a99ab5e5c7820ac603bca9fd811da6fa615245a6d13

Observation 38cd8281-1b22-4352-8eff-46b7aecea09b · outbound

This paper cites DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding.

Remote Sensing Large Vision-Language Model: Semantic-augmented Multi-level Alignment and Semantic-aware Expert Modeling DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T22:22:37.650879Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:22:37.650879Z digest=sha256:b0a2bb775ccd4c66795801ab1b958512234525c632a00ea749d5f74a58e95bf5

Observation 6a00b048-1af5-49e0-a168-9ac0d95eebb4 · outbound

This paper cites Visionllm: Large language model is also an open-ended decoder for vision-centric tasks,.

Remote Sensing Large Vision-Language Model: Semantic-augmented Multi-level Alignment and Semantic-aware Expert Modeling Visionllm: Large language model is also an open-ended decoder for vision-centric tasks,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:23:07.626489Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T22:22:37.685120Z digest=sha256:a2c3aa4712055fb6a50b9df348dbae30ddb6569d67487b41357c297527b0022d

Observation 006f5fa3-1fc7-4cbe-9933-4fa979b7adf1 · outbound

This paper cites Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language Models.

Remote Sensing Large Vision-Language Model: Semantic-augmented Multi-level Alignment and Semantic-aware Expert Modeling Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T22:22:37.709672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:22:37.709672Z digest=sha256:5a69f781af6cc1b0fefd8f42d8cacaa0533fd3a1b57c499be94b5cfcabca1145

Observation f6a5387e-36b4-4fbd-b673-82fb7b301473 · outbound

This paper cites Learning source-invariant deep hashing convolutional neural networks for cross-source remote sensing image retrieval,.

Remote Sensing Large Vision-Language Model: Semantic-augmented Multi-level Alignment and Semantic-aware Expert Modeling Learning source-invariant deep hashing convolutional neural networks for cross-source remote sensing image retrieval,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:23:07.384742Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T22:22:37.735468Z digest=sha256:c9b3cf8b6c7c40e028370462ad1ff4bf470a69712e429f0c811e6190607c20ed

Observation 4825f794-4740-45eb-b570-3a66cf21449b · outbound

This paper cites Automatic radiometric normalization for multitemporal remote sensing imagery with iterative slow feature anal- 12 ysis,.

Remote Sensing Large Vision-Language Model: Semantic-augmented Multi-level Alignment and Semantic-aware Expert Modeling Automatic radiometric normalization for multitemporal remote sensing imagery with iterative slow feature anal- 12 ysis,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:23:07.237268Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T22:22:37.759575Z digest=sha256:d225d8092fc7b8899e00d140a5a933184bad0a56edecfd61d1fa66859e496f2c

Observation 8cd2adb0-69d2-4b66-90b4-5bf616d57c2b · outbound

This paper cites A supervised progressive growing generative adversarial network for remote sensing image scene classification,.

Remote Sensing Large Vision-Language Model: Semantic-augmented Multi-level Alignment and Semantic-aware Expert Modeling A supervised progressive growing generative adversarial network for remote sensing image scene classification,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:23:07.054740Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T22:22:37.795105Z digest=sha256:bc59b4b3c034c7166d800bfe1d39a6335d91f767957d60e9f800c6bd833083dc

Observation b237e4cc-abc1-4e2d-bcec-3e145eadc888 · outbound

This paper cites Multimodal remote sensing image matching combining learning features and delaunay triangula- tion,.

Remote Sensing Large Vision-Language Model: Semantic-augmented Multi-level Alignment and Semantic-aware Expert Modeling Multimodal remote sensing image matching combining learning features and delaunay triangula- tion,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:23:06.906216Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T22:22:37.825702Z digest=sha256:7a62886c4748afddfc927f42c0559f509219475c0c66c70510e56d5c49962276

Observation aa33687a-f580-4333-97d4-2a08a59da360 · outbound

This paper cites Urban flood-related remote sensing: research trends, gaps and opportunities,.

Remote Sensing Large Vision-Language Model: Semantic-augmented Multi-level Alignment and Semantic-aware Expert Modeling Urban flood-related remote sensing: research trends, gaps and opportunities,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:23:06.680198Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T22:22:37.864695Z digest=sha256:cac9f9e64810dbc50482944b0d9a1d24a042440603857eee109d6618cb5b6d9b

Observation 107a23c2-1ed6-4912-8df3-21a2d23c0e9e · outbound

This paper cites Remote sensing-based proxies for urban disaster risk management and resilience: A review,.

Remote Sensing Large Vision-Language Model: Semantic-augmented Multi-level Alignment and Semantic-aware Expert Modeling Remote sensing-based proxies for urban disaster risk management and resilience: A review,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:23:06.535824Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T22:22:37.917813Z digest=sha256:603049cee85ba0d66e37567a9f00cfbb2a442013d2c9ef1052183459e3450b91

Observation 091bfab1-7671-4e9d-a78a-12980f751795 · outbound

This paper cites Remote sensing in multirisk assess- ment: Improving disaster preparedness,.

Remote Sensing Large Vision-Language Model: Semantic-augmented Multi-level Alignment and Semantic-aware Expert Modeling Remote sensing in multirisk assess- ment: Improving disaster preparedness,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:23:06.468657Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T22:22:37.951695Z digest=sha256:8cddeb28be80d2c0cf30bed75f494929a9791b091ef5b156658b2a1f5645862b

Observation 48d1a99b-3fff-4879-86fb-83e40371dc0c · outbound

This paper cites Geochat: Grounded large vision-language model for remote sensing,.

Remote Sensing Large Vision-Language Model: Semantic-augmented Multi-level Alignment and Semantic-aware Expert Modeling Geochat: Grounded large vision-language model for remote sensing,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:23:06.359877Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T22:22:37.992766Z digest=sha256:4bc22d4c695f61807bf8325a08dcb28d75c80d18e5aba58e68ff3b29fee9b244

Observation bf95f217-69ae-4b4d-8275-716b44b4473f · outbound

This paper cites Lhrs-bot: Empowering remote sensing with vgi-enhanced large multimodal language model,.

Remote Sensing Large Vision-Language Model: Semantic-augmented Multi-level Alignment and Semantic-aware Expert Modeling Lhrs-bot: Empowering remote sensing with vgi-enhanced large multimodal language model,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:23:06.234735Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T22:22:38.024677Z digest=sha256:a0286451a9ef4c69ab3fa36aa94ca190741b443083157180f1bb868f5b8645c6

Observation 8ed055c6-7b5c-4429-ae66-7caf29855f62 · outbound

This paper cites Skyeyegpt: Unifying remote sensing vision-language tasks via instruction tuning with large language model,.

Remote Sensing Large Vision-Language Model: Semantic-augmented Multi-level Alignment and Semantic-aware Expert Modeling Skyeyegpt: Unifying remote sensing vision-language tasks via instruction tuning with large language model,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:23:06.061858Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T22:22:38.095736Z digest=sha256:55ecf72ae8e311f0c0e67c8a3a65755ec9102cfa74d1bc52104a46abd5a3da21

Observation 1790d3fa-4750-40bd-8f65-f765a54d9134 · outbound

This paper cites H2rsvlm: Towards helpful and honest remote sensing large vision language model,.

Remote Sensing Large Vision-Language Model: Semantic-augmented Multi-level Alignment and Semantic-aware Expert Modeling H2rsvlm: Towards helpful and honest remote sensing large vision language model,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:23:05.609068Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T22:22:38.155117Z digest=sha256:8e23401753557f52068693a26ecef02521c86b0a6153aa438a32ffbb4838a5aa

Observation aec9247a-2ce7-4106-ba4c-d5a89bbc388e · outbound

This paper cites Rs- llava: A large vision-language model for joint captioning and question answering in remote sensing imagery,.

Remote Sensing Large Vision-Language Model: Semantic-augmented Multi-level Alignment and Semantic-aware Expert Modeling Rs- llava: A large vision-language model for joint captioning and question answering in remote sensing imagery,

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T22:22:38.225898Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:22:38.225898Z digest=sha256:2524088c9b34cb7a012b5b5a6084d05027975ebc1bd414b25716c9678ac6f8ac

Observation f16d1259-155d-4449-8054-a9a99b154a5a · outbound

This paper cites SkySenseGPT: A Fine-Grained Instruction Tuning Dataset and Model for Remote Sensing Vision-Language Understanding.

Remote Sensing Large Vision-Language Model: Semantic-augmented Multi-level Alignment and Semantic-aware Expert Modeling SkySenseGPT: A Fine-Grained Instruction Tuning Dataset and Model for Remote Sensing Vision-Language Understanding

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T22:22:38.288515Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:22:38.288515Z digest=sha256:10a332514d1922b3fe0a43f01d1454f1fe21cadf1a3a9c28fc8ee1fe88abf9e3

Observation 0f5c73c2-eed5-4256-95f8-765f8e89c25f · outbound

This paper cites Rsgpt: A remote sensing vision language model and benchmark,.

Remote Sensing Large Vision-Language Model: Semantic-augmented Multi-level Alignment and Semantic-aware Expert Modeling Rsgpt: A remote sensing vision language model and benchmark,

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T22:22:38.369262Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:22:38.369262Z digest=sha256:22604a035aec053850952469ec382a8ecb491bbb524dd85adb6b35dcbc5d9418

Observation f83e90f2-a829-40bc-b0ea-8b0f4cb61f5f · outbound

This paper cites RSUniVLM: A Unified Vision Language Model for Remote Sensing via Granularity-oriented Mixture of Experts.

Remote Sensing Large Vision-Language Model: Semantic-augmented Multi-level Alignment and Semantic-aware Expert Modeling RSUniVLM: A Unified Vision Language Model for Remote Sensing via Granularity-oriented Mixture of Experts

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T22:22:38.461173Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:22:38.461173Z digest=sha256:902eddcab7c9ca599dfad7e320c827a1729f5032546edd1d7c60fc393104d874

Observation ec569c52-8643-4cc7-83d8-c0bd36107b6f · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Remote Sensing Large Vision-Language Model: Semantic-augmented Multi-level Alignment and Semantic-aware Expert Modeling Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T22:22:38.535644Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:22:38.535644Z digest=sha256:15b743fd1e5d06007beb2780a4792a599faaaaa2a93723ff9bdb4502951911a8

Observation bb233f38-f6ac-42d0-844a-370aa015a0e8 · outbound

This paper cites The Llama 3 Herd of Models.

Remote Sensing Large Vision-Language Model: Semantic-augmented Multi-level Alignment and Semantic-aware Expert Modeling The Llama 3 Herd of Models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T22:22:38.617226Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:22:38.617226Z digest=sha256:19c4600679a42b7a11f2dc989685825a5b5c97d3638f76dcb4d3b86b1af36091

Observation 77549bf0-b983-43f4-84e4-b843211d9a8e · outbound

This paper cites Qwen Technical Report.

Remote Sensing Large Vision-Language Model: Semantic-augmented Multi-level Alignment and Semantic-aware Expert Modeling Qwen Technical Report

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T22:22:38.685861Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:22:38.685861Z digest=sha256:5857c3790cdba025d9eb6eda6e8a815dfe4bfc28ef4a3df4b64da0f523e350da

Observation b20196fd-9769-4162-830a-65bd6f5a89c8 · outbound

This paper cites Internlm: A multilingual language model with progressively enhanced capabilities,.

Remote Sensing Large Vision-Language Model: Semantic-augmented Multi-level Alignment and Semantic-aware Expert Modeling Internlm: A multilingual language model with progressively enhanced capabilities,

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T22:22:38.777895Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:22:38.777895Z digest=sha256:fc9b9048a8035b0bb84fddcb595b8524fd2c79e2ef27cb498a37478ee3383cf3

Observation 253ce06b-4d3e-4d1b-bb05-902024b5477d · outbound

This paper cites Llava-next: Improved reasoning, ocr, and world knowledge,.

Remote Sensing Large Vision-Language Model: Semantic-augmented Multi-level Alignment and Semantic-aware Expert Modeling Llava-next: Improved reasoning, ocr, and world knowledge,

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T22:22:38.870626Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:22:38.870626Z digest=sha256:52ed1a43fedabefed9ff8eadb8b5187fdaaf36e8f461d23c784a2cfb7be069bb

Observation d9dcd1a2-1e1c-42ab-934b-5e83343b41f0 · outbound

This paper cites InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models.

Remote Sensing Large Vision-Language Model: Semantic-augmented Multi-level Alignment and Semantic-aware Expert Modeling InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T22:22:38.978885Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:22:38.978885Z digest=sha256:8ec74d3056dfaa555ad17fc41b356b2c26771af8a11f70d4b303dfdef2da2277

Observation e4d4b9f3-1b07-4cb6-b2d9-5521e39fca66 · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

Remote Sensing Large Vision-Language Model: Semantic-augmented Multi-level Alignment and Semantic-aware Expert Modeling MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T22:22:39.049798Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:22:39.049798Z digest=sha256:fc501c04aa35ddc8072dbfee1b498c4bffb29a9377f8e2132d793bc2f69532ca

Observation 2a7d4b5d-d066-4e70-bbd3-4208638aae9f · outbound

This paper cites Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models.

Remote Sensing Large Vision-Language Model: Semantic-augmented Multi-level Alignment and Semantic-aware Expert Modeling Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T22:22:39.123538Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:22:39.123538Z digest=sha256:7af546bba6a1f15bd75b8af8bcf76aee73f40e6ce6e5b407134db9feb0b8d5b1

Observation a62c908d-af4f-4995-9b14-dc0f73021938 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

Remote Sensing Large Vision-Language Model: Semantic-augmented Multi-level Alignment and Semantic-aware Expert Modeling Gemini: A Family of Highly Capable Multimodal Models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T22:22:39.212756Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:22:39.212756Z digest=sha256:55e881940cb10c06fd0af89a25b30bb415c9241c595975f80cd9daba440368fa

Observation e759238f-dab3-4c82-8923-2bcb3f9e021a · outbound

This paper cites KOSMOS-2.5: A Multimodal Literate Model.

Remote Sensing Large Vision-Language Model: Semantic-augmented Multi-level Alignment and Semantic-aware Expert Modeling KOSMOS-2.5: A Multimodal Literate Model

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T22:22:39.299853Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:22:39.299853Z digest=sha256:81b7206978089aa15637f91850e5872a2b87b47ef6efc74c584662b361e6663f

Observation ff274aec-2bf1-4ed8-8007-f2c30dc020d1 · outbound

This paper cites Scaling up visual and vision-language representation learning with noisy text supervision,.

Remote Sensing Large Vision-Language Model: Semantic-augmented Multi-level Alignment and Semantic-aware Expert Modeling Scaling up visual and vision-language representation learning with noisy text supervision,

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T22:22:39.393855Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:22:39.393855Z digest=sha256:431b018171e2face8c9f63c1b741a82ed92d8a3155870389bd3a529a40a60981

Observation ef79ffac-a7ca-4ea5-9e95-168cee76a890 · outbound

This paper cites Sharegpt4v: Improving large multi-modal models with better captions,.

Remote Sensing Large Vision-Language Model: Semantic-augmented Multi-level Alignment and Semantic-aware Expert Modeling Sharegpt4v: Improving large multi-modal models with better captions,

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:23:03.827568Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T22:22:39.491263Z digest=sha256:ead1a188f09173eb4e212ad8f29d58e730f5bdf4e3c3068e3c170217fb082425

Observation e83556a7-25f4-4a34-845c-d81d4b8bfba9 · outbound

This paper cites MiniGPT-v2: large language model as a unified interface for vision-language multi-task learning.

Remote Sensing Large Vision-Language Model: Semantic-augmented Multi-level Alignment and Semantic-aware Expert Modeling MiniGPT-v2: large language model as a unified interface for vision-language multi-task learning

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-06T22:22:39.583072Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:22:39.583072Z digest=sha256:623ab7206daa3927cf371d340ecf38da0fb72db03e62bbd54d4676d3cca020e1

Observation 6634d2c4-535b-4d8f-889a-8032b1f00373 · outbound

This paper cites InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition.

Remote Sensing Large Vision-Language Model: Semantic-augmented Multi-level Alignment and Semantic-aware Expert Modeling InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-06T22:22:39.652994Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:22:39.652994Z digest=sha256:311bb84c138ef502108179af59e93681876d5cb169f43766952374294b7932d7

Observation b281da44-fbdf-4dc3-9560-9aa1687ef19c · outbound

This paper cites Gpt4roi: Instruction tuning large language model on regionof-interest, 2024,.

Remote Sensing Large Vision-Language Model: Semantic-augmented Multi-level Alignment and Semantic-aware Expert Modeling Gpt4roi: Instruction tuning large language model on regionof-interest, 2024,

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:23:03.548855Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T22:22:39.761352Z digest=sha256:da7b8acac73821120501d23573018a2bfe571a9c751fab2b7033a28745329dc5

Observation 0409c39d-37a3-487d-9e11-a04f0cd333e9 · outbound

This paper cites Glamm: Pixel grounding large multimodal model,.

Remote Sensing Large Vision-Language Model: Semantic-augmented Multi-level Alignment and Semantic-aware Expert Modeling Glamm: Pixel grounding large multimodal model,

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-06T22:22:39.844011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:22:39.844011Z digest=sha256:8545d4a5f1946dc7478d0065ed23cad7f1d6f1a1e8b0e53e76cac8f186b9114d

Observation d4964b9e-ef44-4714-bbbb-b500462ca1fb · outbound

This paper cites LHRS-Bot-Nova: Improved Multimodal Large Language Model for Remote Sensing Vision-Language Interpretation.

Remote Sensing Large Vision-Language Model: Semantic-augmented Multi-level Alignment and Semantic-aware Expert Modeling LHRS-Bot-Nova: Improved Multimodal Large Language Model for Remote Sensing Vision-Language Interpretation

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-06T22:22:39.933256Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:22:39.933256Z digest=sha256:47d9073ec9dbf4ab57adaf6f553197bcbd13ffb6872a8b54e16383986336b262

Observation e5b11529-a876-47bc-afe7-58fcfbb95c0d · outbound

This paper cites EarthMarker: A Visual Prompting Multi-modal Large Language Model for Remote Sensing.

Remote Sensing Large Vision-Language Model: Semantic-augmented Multi-level Alignment and Semantic-aware Expert Modeling EarthMarker: A Visual Prompting Multi-modal Large Language Model for Remote Sensing

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-06T22:22:40.050681Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:22:40.050681Z digest=sha256:21a289ab0f08fb39f772b00f5dd48a016aa04da5af6a4a7d1d2f15343b35d595

Observation cddd690e-fa14-49b9-a2bc-321b147da3a4 · outbound

This paper cites GeoLLaVA: Efficient Fine-Tuned Vision-Language Models for Temporal Change Detection in Remote Sensing.

Remote Sensing Large Vision-Language Model: Semantic-augmented Multi-level Alignment and Semantic-aware Expert Modeling GeoLLaVA: Efficient Fine-Tuned Vision-Language Models for Temporal Change Detection in Remote Sensing

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-06T22:22:40.114061Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:22:40.114061Z digest=sha256:4eb2b9434cf9cad84d81c5f842bec4d1db3d3763aae7dc02988f757779b0238c

Observation b837ee10-fd68-4969-a7c7-5ae3e609ebc5 · outbound

This paper cites TEOChat: A Large Vision-Language Assistant for Temporal Earth Observation Data.

Remote Sensing Large Vision-Language Model: Semantic-augmented Multi-level Alignment and Semantic-aware Expert Modeling TEOChat: A Large Vision-Language Assistant for Temporal Earth Observation Data

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-06T22:22:40.225645Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:22:40.225645Z digest=sha256:993899f0ccc85ea42c65740375314033dff2eb55ed925bfde48ace412231b64d

Observation ee836c34-910a-4761-b513-d54b9060d421 · outbound

This paper cites Earthgpt: A universal multi-modal large language model for multi-sensor image comprehension in remote sensing domain,.

Remote Sensing Large Vision-Language Model: Semantic-augmented Multi-level Alignment and Semantic-aware Expert Modeling Earthgpt: A universal multi-modal large language model for multi-sensor image comprehension in remote sensing domain,

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:23:03.371380Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T22:22:40.319247Z digest=sha256:61e08ce196546575e8843a6dcf970c04dca4c7d49669e8606130b5a25d72b031

Observation 76d8790e-b373-4c74-a4c0-9e8653b0605b · outbound

This paper cites Ringmogpt: A unified remote sensing foundation model for vision, language, and grounded tasks,.

Remote Sensing Large Vision-Language Model: Semantic-augmented Multi-level Alignment and Semantic-aware Expert Modeling Ringmogpt: A unified remote sensing foundation model for vision, language, and grounded tasks,

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:23:03.204907Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T22:22:40.405051Z digest=sha256:3d2b298240d3c94e951370bf01c094a3411c63e959481177b5885fbc14af7fa0

Observation 1e2553fc-aa01-4bcf-9126-187abe71100c · outbound

This paper cites Skyscript: A large and semantically diverse vision-language dataset for remote sens- ing,.

Remote Sensing Large Vision-Language Model: Semantic-augmented Multi-level Alignment and Semantic-aware Expert Modeling Skyscript: A large and semantically diverse vision-language dataset for remote sens- ing,

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-06T22:22:40.494200Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:22:40.494200Z digest=sha256:39f76079ef1f65013e56a0101934d740249f21c5cbb508195cf890cd573030bc

Observation f30ee937-9a04-41b4-b979-557da59ec5a0 · outbound

This paper cites Deep semantic understanding of high resolution remote sensing image,.

Remote Sensing Large Vision-Language Model: Semantic-augmented Multi-level Alignment and Semantic-aware Expert Modeling Deep semantic understanding of high resolution remote sensing image,

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:22:45.081454Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T22:22:40.564185Z digest=sha256:b2f9325c6ea7ed4533cb9b7a0848ca00912c581bf8123800db0773f04d6bcf61

Observation 9e93813d-4eaa-4ec6-8557-c15fbd3fc5cf · outbound

This paper cites Remote sensing image scene classifi- cation: Benchmark and state of the art,.

Remote Sensing Large Vision-Language Model: Semantic-augmented Multi-level Alignment and Semantic-aware Expert Modeling Remote sensing image scene classifi- cation: Benchmark and state of the art,

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-06T22:22:40.639987Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:22:40.639987Z digest=sha256:4c4a828c8a1e40f61d9b123f27d24d3562aab0df82360a14464149550f3175b6

Observation 1b3dcfdb-32b9-4c65-a9d6-911c92551ff7 · outbound

This paper cites Nwpu- captions dataset and mlca-net for remote sensing image captioning,.

Remote Sensing Large Vision-Language Model: Semantic-augmented Multi-level Alignment and Semantic-aware Expert Modeling Nwpu- captions dataset and mlca-net for remote sensing image captioning,

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:22:44.694995Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T22:22:40.714308Z digest=sha256:1233e0892245ad5b596a501f1c98995f5c1f85cd1f10cbe8a7040ede6e41b491

Observation 4911cca2-b81f-45e1-9d61-c6f362d666ac · outbound

This paper cites Exploring a Fine-Grained Multiscale Method for Cross-Modal Remote Sensing Image Retrieval.

Remote Sensing Large Vision-Language Model: Semantic-augmented Multi-level Alignment and Semantic-aware Expert Modeling Exploring a Fine-Grained Multiscale Method for Cross-Modal Remote Sensing Image Retrieval

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-06T22:22:40.798399Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:22:40.798399Z digest=sha256:6c4d339049814818bbe818857e4932e2573c5656811a4836f97667d884673231

Observation 51d48c8a-48b8-4d32-99bb-97de27ae9e6e · outbound

This paper cites METER-ML: A Multi-Sensor Earth Observation Benchmark for Automated Methane Source Mapping.

Remote Sensing Large Vision-Language Model: Semantic-augmented Multi-level Alignment and Semantic-aware Expert Modeling METER-ML: A Multi-Sensor Earth Observation Benchmark for Automated Methane Source Mapping

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-06T22:22:40.908652Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:22:40.908652Z digest=sha256:d7852fe4f23d5c157cd8b1f78b9b6445f25b7508f39bd7683c87a677609fdec5

Observation bca7d714-1582-4d43-bce5-a7e9262d9514 · outbound

This paper cites Functional map of the world,.

Remote Sensing Large Vision-Language Model: Semantic-augmented Multi-level Alignment and Semantic-aware Expert Modeling Functional map of the world,

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-06T22:22:41.031030Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:22:41.031030Z digest=sha256:0406c3f4e77f080b68d8e8877cce16971294e482a20b8621d8661195c3ccae95

Observation a2dcb4e5-4bc9-4eb3-a623-418848a1c8e2 · outbound

This paper cites Exploring models and data for remote sensing image caption generation,.

Remote Sensing Large Vision-Language Model: Semantic-augmented Multi-level Alignment and Semantic-aware Expert Modeling Exploring models and data for remote sensing image caption generation,

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:22:44.460150Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T22:22:41.078684Z digest=sha256:0a707dfdd9e802fa6ff27e67fe43bb607352eb92bb654471d92c6995976439b7

Observation ff9def5d-e6e7-4689-9bc6-520ecde441ec · outbound

This paper cites Rsvqa: Visual question answering for remote sensing data,.

Remote Sensing Large Vision-Language Model: Semantic-augmented Multi-level Alignment and Semantic-aware Expert Modeling Rsvqa: Visual question answering for remote sensing data,

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:22:44.146974Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T22:22:41.132016Z digest=sha256:ed510cb23b9f89c9de89e2871328d44dd06153bef0a2e814398cf09755fd40f3

Observation e0692b84-d5bb-411a-8ae5-72431ce9f3c4 · outbound

This paper cites Visual grounding in remote sensing images,.

Remote Sensing Large Vision-Language Model: Semantic-augmented Multi-level Alignment and Semantic-aware Expert Modeling Visual grounding in remote sensing images,

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:22:43.880645Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T22:22:41.183387Z digest=sha256:c2d572192006a0985b87f10649ec0584a84cfacae55e26f1509453d4b3941ac1

Observation ad0bd306-6f06-44b9-a4c6-77ded1d15432 · outbound

This paper cites Rsvg: Exploring data and models for visual grounding on remote sensing data,.

Remote Sensing Large Vision-Language Model: Semantic-augmented Multi-level Alignment and Semantic-aware Expert Modeling Rsvg: Exploring data and models for visual grounding on remote sensing data,

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:22:43.585787Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T22:22:41.237485Z digest=sha256:635d997df6807a26efde2b77a8ac1d4cffd1bd52f866818956694f85a151337f

Observation 0ace819e-8c81-45a1-8691-900ac20deb2a · outbound

This paper cites Object detection in optical remote sensing images: A survey and a new benchmark,.

Remote Sensing Large Vision-Language Model: Semantic-augmented Multi-level Alignment and Semantic-aware Expert Modeling Object detection in optical remote sensing images: A survey and a new benchmark,

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-06T22:22:41.294447Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:22:41.294447Z digest=sha256:d2d0c9eb3ef8080ce507c852f5c507554044db1635a519f81a9db00878a8456b

Observation 74381c94-fca9-478f-8ce9-fefc3ea57b87 · outbound

This paper cites Dota: A large-scale dataset for object detection in aerial images,.

Remote Sensing Large Vision-Language Model: Semantic-augmented Multi-level Alignment and Semantic-aware Expert Modeling Dota: A large-scale dataset for object detection in aerial images,

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-06T22:22:41.350812Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:22:41.350812Z digest=sha256:74d99cfadfa18e9e2a8d6d1ea17024875f6de5ce3b2e2e5d07f369198d5f05c6

Observation d5a745c7-c55c-4c8a-8348-2d69ffe7ae2b · outbound

This paper cites Aid: A benchmark data set for performance evaluation of aerial scene classification,.

Remote Sensing Large Vision-Language Model: Semantic-augmented Multi-level Alignment and Semantic-aware Expert Modeling Aid: A benchmark data set for performance evaluation of aerial scene classification,

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:22:43.248276Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T22:22:41.398190Z digest=sha256:8afacda6bc038f527db89069ba199144ca99132071ba50bb33603d14be36c7d4

Observation f0ac9d96-da54-4527-9b83-f21b3531fddd · outbound

This paper cites Satellite image classification via two-layer sparse coding with biased image representation,.

Remote Sensing Large Vision-Language Model: Semantic-augmented Multi-level Alignment and Semantic-aware Expert Modeling Satellite image classification via two-layer sparse coding with biased image representation,

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:22:43.002910Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T22:22:41.467391Z digest=sha256:90ae658fac6df61794b3cdec7109369733df3f40f1aa6e18b8df018dcfba96ee

Observation 16b87da5-b324-4c2a-aa0a-a7711128364f · outbound

This paper cites Bag-of-visual- words scene classifier with local and global features for high spatial resolution remote sensing imagery,.

Remote Sensing Large Vision-Language Model: Semantic-augmented Multi-level Alignment and Semantic-aware Expert Modeling Bag-of-visual- words scene classifier with local and global features for high spatial resolution remote sensing imagery,

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:22:42.750777Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T22:22:41.529916Z digest=sha256:c0ceef70418bea6e0fd3078bfff3af12df440c17ef2e356431e215925224ddb8

Observation 7c171966-67c3-453a-8619-10b7c97a4124 · outbound

This paper cites Improved baselines with visual instruction tuning,.

Remote Sensing Large Vision-Language Model: Semantic-augmented Multi-level Alignment and Semantic-aware Expert Modeling Improved baselines with visual instruction tuning,

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-06T22:22:41.583771Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:22:41.583771Z digest=sha256:a7b13a674d4fdad3d021dd4d4f03940c3364928f620c7461d5fa1d77b6a829fe

Observation ab0645a4-d411-4393-a611-ed09529fa31e · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

Remote Sensing Large Vision-Language Model: Semantic-augmented Multi-level Alignment and Semantic-aware Expert Modeling Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-06T22:22:41.662531Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:22:41.662531Z digest=sha256:94fe475f218df751497e4e5687ba5185e843ef8740c78826ee6705611e028df0

Observation b3dcb610-4927-4266-88be-ce3d69c6f970 · outbound

This paper cites SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features.

Remote Sensing Large Vision-Language Model: Semantic-augmented Multi-level Alignment and Semantic-aware Expert Modeling SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-06T22:22:41.716132Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:22:41.716132Z digest=sha256:e70e549b3a50dd25fd1c9976a0f2387a91e8b2b6135e56e663d5bbcd74a9f9b1

Observation f5c432bd-ec86-4cd4-8ee2-4da9858061f4 · outbound

This paper cites Qwen2.5: A party of foundation models,.

Remote Sensing Large Vision-Language Model: Semantic-augmented Multi-level Alignment and Semantic-aware Expert Modeling Qwen2.5: A party of foundation models,

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-06T22:22:41.767782Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:22:41.767782Z digest=sha256:e064a6ab53f46e6e02218fdfb8d913175b30ca725a035adb4fa51fbcae3554c4

Observation 59bd219b-acaf-4158-9483-a56f0ca825b3 · outbound

This paper cites Vhm: Versatile and honest vision language model for remote sensing image analysis,.

Remote Sensing Large Vision-Language Model: Semantic-augmented Multi-level Alignment and Semantic-aware Expert Modeling Vhm: Versatile and honest vision language model for remote sensing image analysis,

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:22:42.582486Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T22:22:41.831349Z digest=sha256:fcd7634fbe77cb02b8265a735b4ee5d7d4706e5d2023492f48bf6355e92381ab

Pith citing papers

Observation 522cebc6-71b4-418f-9961-afaf5b036e74 · inbound

Vision-and-Language Navigation for UAVs: Progress, Challenges, and a Research Roadmap cites this paper.

Vision-and-Language Navigation for UAVs: Progress, Challenges, and a Research Roadmap Remote Sensing Large Vision-Language Model: Semantic-augmented Multi-level Alignment and Semantic-aware Expert Modeling

Reference 148

Resolution
verified exact
arxiv_id, observed 2026-05-10T14:10:29.816125Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-10T13:48:08.135538Z digest=sha256:9f3ca066d64c41ab274615a75982f5e346fe9f37e6968e36f8557401d6435773