Pith. sign in

Paper Citation Record · LEDGER

UniRS: Unifying Multi-temporal Remote Sensing Tasks through Vision Language Models

As of 14 August 2026, this Paper Citation Record lists 56 of 56 outbound references and 2 inbound Pith citation observations for arXiv:2412.20742.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.20742 v1

Coverage vector

measured 56 of 56 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T23:18:21.438754Z

measured 58 of 58 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T00:14:05.499836Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-07T00:14:11.438315Z

Reference resolution

56 of 56 outbound references displayed

  • verified exact0
  • verified fuzzy44
  • unresolved12
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation cfdd125b-a09c-4491-883a-9cbe2377ad50 · outbound

This paper cites Visual instruction tuning,.

UniRS: Unifying Multi-temporal Remote Sensing Tasks through Vision Language Models Visual instruction tuning,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:18:23.458753Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T23:18:21.129990Z digest=sha256:5c82e3f5a0cb8551957838bccdadd6081529a5419dbe012e718d5b813c9c520f

Observation ce25b4aa-3cbd-43d6-8ce7-7b336e7685b2 · outbound

This paper cites Flamingo: a visual language model for few-shot learning,.

UniRS: Unifying Multi-temporal Remote Sensing Tasks through Vision Language Models Flamingo: a visual language model for few-shot learning,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:18:23.367346Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T23:18:21.136231Z digest=sha256:7491b834918b00a1420efd1a70a0f70a435dd7551316b72e46b43ab8ae8d941a

Observation 530008c7-1235-4052-950f-d8ec06735c83 · outbound

This paper cites Vila: On pre-training for visual language models,.

UniRS: Unifying Multi-temporal Remote Sensing Tasks through Vision Language Models Vila: On pre-training for visual language models,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:18:23.350477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T23:18:21.141609Z digest=sha256:dc4d9f26bdc4675eface941036b99264731a3d09f0b708a93b368352dc7caf5c

Observation fbe617b2-21fb-44de-8a92-073732e6ca3c · outbound

This paper cites MiniGPT-v2: large language model as a unified interface for vision-language multi-task learning.

UniRS: Unifying Multi-temporal Remote Sensing Tasks through Vision Language Models MiniGPT-v2: large language model as a unified interface for vision-language multi-task learning

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-10T23:18:21.147250Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:18:21.147250Z digest=sha256:1279a37c0dc28a1939f70a83e0aaa1460912691ff40ab50ba2128a0ef5f30924

Observation 06d2b9f4-d435-4e2c-9e8b-d09ecacff320 · outbound

This paper cites Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation,.

UniRS: Unifying Multi-temporal Remote Sensing Tasks through Vision Language Models Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:18:23.333186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T23:18:21.153597Z digest=sha256:3d77f15a503eb2e4c479445b0eda9c8467bf0c99708e2208351648fda7efae89

Observation 799829fc-f48b-40d5-9955-a60296e47ea9 · outbound

This paper cites Blip-2: Bootstrapping language- image pre-training with frozen image encoders and large language models,.

UniRS: Unifying Multi-temporal Remote Sensing Tasks through Vision Language Models Blip-2: Bootstrapping language- image pre-training with frozen image encoders and large language models,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:18:23.315300Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T23:18:21.158791Z digest=sha256:d2689aa3ef2fa5e0b4d1b0fd0181e1b307db387c7b3108f4058152cb0d553d6e

Observation 21c7731b-5099-4296-a867-84ceff71c929 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

UniRS: Unifying Multi-temporal Remote Sensing Tasks through Vision Language Models LLaMA: Open and Efficient Foundation Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-10T23:18:21.164175Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:18:21.164175Z digest=sha256:250d6c3206d8a9fbbc75954e2bedaf41d9f0ee1a3665bc65fc41ed90a99be59d

Observation df83d8ad-0e80-44f5-a8c3-905779f10e98 · outbound

This paper cites Training language models to follow instructions with human feedback,.

UniRS: Unifying Multi-temporal Remote Sensing Tasks through Vision Language Models Training language models to follow instructions with human feedback,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:18:23.296497Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T23:18:21.170197Z digest=sha256:5333b1a760b91d59bcade60adc6fbb01bc2ea330c345e3dc3c9e4ec69462718e

Observation c88ce50f-906b-4c8f-8b5d-37fc8e4da3d3 · outbound

This paper cites Qwen Technical Report.

UniRS: Unifying Multi-temporal Remote Sensing Tasks through Vision Language Models Qwen Technical Report

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-10T23:18:21.175296Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:18:21.175296Z digest=sha256:044c3f9c6981dc604a23b409dc7de73d40241ddf8e1454859a7564d4fd60c710

Observation 5b31fda4-bd17-4c4d-9c1f-41a87607715a · outbound

This paper cites Language mod- els are few-shot learners,.

UniRS: Unifying Multi-temporal Remote Sensing Tasks through Vision Language Models Language mod- els are few-shot learners,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:18:23.278960Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T23:18:21.180320Z digest=sha256:d4a85498806e38262e29f522730800933ff78eb6497738b18c3a687842dc3de2

Observation 09aa9e4d-7db1-406f-b233-051a3d217e69 · outbound

This paper cites Laion- 5b: An open large-scale dataset for training next generation image-text models,.

UniRS: Unifying Multi-temporal Remote Sensing Tasks through Vision Language Models Laion- 5b: An open large-scale dataset for training next generation image-text models,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:18:23.261896Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T23:18:21.185583Z digest=sha256:0773a97ddb600bd30f25aefe75bba0188928efb54e853aff1f7f3734415d7c47

Observation f3555b1b-c8ba-41f1-9012-807276725dd0 · outbound

This paper cites Llava-med: Training a large language-and-vision assistant for biomedicine in one day,.

UniRS: Unifying Multi-temporal Remote Sensing Tasks through Vision Language Models Llava-med: Training a large language-and-vision assistant for biomedicine in one day,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:18:23.245810Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T23:18:21.190988Z digest=sha256:5f0e7f92844ab52989a3d366205593d716b502a648c0cca0dd7d046607972203

Observation f0d34255-e33e-49db-837d-38721ecf3295 · outbound

This paper cites Mobility VLA: Multimodal Instruction Navigation with Long-Context VLMs and Topological Graphs.

UniRS: Unifying Multi-temporal Remote Sensing Tasks through Vision Language Models Mobility VLA: Multimodal Instruction Navigation with Long-Context VLMs and Topological Graphs

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-10T23:18:21.196144Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:18:21.196144Z digest=sha256:875f41d40053584fd645ce9d2f06b44b36ebb964e32ecf117ea27706ad634039

Observation 4632d040-0eac-4350-aef7-8a6f7e96d63c · outbound

This paper cites Rsvqa: Visual question answering for remote sensing data,.

UniRS: Unifying Multi-temporal Remote Sensing Tasks through Vision Language Models Rsvqa: Visual question answering for remote sensing data,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:18:23.229120Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T23:18:21.201942Z digest=sha256:076353ee5410f0b6b79889f45fe92e893f427fbab3fe809fe4999c48f579270f

Observation f0c8b997-0e45-4fcd-b1dd-eba28ef142ce · outbound

This paper cites Answer-type prediction for visual question answering,.

UniRS: Unifying Multi-temporal Remote Sensing Tasks through Vision Language Models Answer-type prediction for visual question answering,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:18:23.212597Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T23:18:21.208078Z digest=sha256:41c9d50973c6215cda77a295f56a35d8475ca48330ee11021005bdbf5d5a2552

Observation ea401cdc-6899-4110-a644-8fcca327b6c4 · outbound

This paper cites Multi-step question-driven visual ques- tion answering for remote sensing,.

UniRS: Unifying Multi-temporal Remote Sensing Tasks through Vision Language Models Multi-step question-driven visual ques- tion answering for remote sensing,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:18:23.160301Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T23:18:21.213271Z digest=sha256:d8865256625feeb793d017b0c13748ab318060308e283892459fb973faa7ec54

Observation 7652e053-07a3-41f4-8123-ab8624cd9906 · outbound

This paper cites Bi-modal transformer-based approach for visual question answering in remote sensing imagery,.

UniRS: Unifying Multi-temporal Remote Sensing Tasks through Vision Language Models Bi-modal transformer-based approach for visual question answering in remote sensing imagery,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:18:23.051315Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T23:18:21.218321Z digest=sha256:bcd1299bbcbad93bfacda080e8e9c1ad12261fe679bd8554f1c4f7cbda9b4a34

Observation d79f23aa-1b6c-4153-b41b-a9432905d186 · outbound

This paper cites Changes to captions: An attentive network for remote sensing change captioning,.

UniRS: Unifying Multi-temporal Remote Sensing Tasks through Vision Language Models Changes to captions: An attentive network for remote sensing change captioning,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:18:22.907158Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T23:18:21.224247Z digest=sha256:39cf9cfb16a6b61202b8b67371cb5bde2d981b8bc3486d3b14da693e6254eb64

Observation 9ddb6af0-974c-41dc-a04d-b6a1edf19bfc · outbound

This paper cites Progressive scale-aware network for remote sensing image change captioning,.

UniRS: Unifying Multi-temporal Remote Sensing Tasks through Vision Language Models Progressive scale-aware network for remote sensing image change captioning,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:18:22.805680Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T23:18:21.229092Z digest=sha256:9d1ea49ef3505e7d0d9ede5862e9a4a6b5b9254dd43a0a67946a1a95afede88b

Observation 12399aa5-b068-4bfb-beb3-a61dc53e7f88 · outbound

This paper cites Futh-net: fusing temporal relations and holistic features for aerial video classification,.

UniRS: Unifying Multi-temporal Remote Sensing Tasks through Vision Language Models Futh-net: fusing temporal relations and holistic features for aerial video classification,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:18:22.779334Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T23:18:21.234463Z digest=sha256:b9c89e58ceb41e38545204c5044535bea398e67a33ed108859ed9ffda4599f2d

Observation 9362b406-f560-4720-b98e-23397670ffe5 · outbound

This paper cites Ok-vqa: A visual question answering benchmark requiring external knowledge,.

UniRS: Unifying Multi-temporal Remote Sensing Tasks through Vision Language Models Ok-vqa: A visual question answering benchmark requiring external knowledge,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:18:22.762710Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T23:18:21.240700Z digest=sha256:2f9abcba126d3e89cfada5e162b290e496f5bed4cddf2846050cc18736af51be

Observation 76fd99e3-cf0a-4563-a1de-b0420b434339 · outbound

This paper cites Deep learning based event recognition in aerial imagery,.

UniRS: Unifying Multi-temporal Remote Sensing Tasks through Vision Language Models Deep learning based event recognition in aerial imagery,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:18:22.744621Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T23:18:21.247367Z digest=sha256:ab1d8afb4a3aae24be44af4a4cacf55b0b3be9eb024aa70260209f925e68b1eb

Observation 4ad2a81b-9a8d-4704-904f-80df22718a0b · outbound

This paper cites A decoupling paradigm with prompt learning for remote sensing image change cap- tioning,.

UniRS: Unifying Multi-temporal Remote Sensing Tasks through Vision Language Models A decoupling paradigm with prompt learning for remote sensing image change cap- tioning,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:18:22.727137Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T23:18:21.254060Z digest=sha256:863b60b1570c27811c3e7f28d100adc67eb43098dec6f18ddcea9cc44f92d99d

Observation 1892636e-0c9b-4294-a9e2-aee3f275b794 · outbound

This paper cites LHRS-Bot: Empowering Remote Sensing with VGI-Enhanced Large Multimodal Language Model.

UniRS: Unifying Multi-temporal Remote Sensing Tasks through Vision Language Models LHRS-Bot: Empowering Remote Sensing with VGI-Enhanced Large Multimodal Language Model

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-10T23:18:21.260642Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:18:21.260642Z digest=sha256:ae873258ea1363bafb0e76578857437c368665acc5300a0dbb92f90729fd2eb3

Observation 806d599f-8e76-4721-a274-ab799ca5cd0e · outbound

This paper cites Geochat: Grounded large vision-language model for remote sensing,.

UniRS: Unifying Multi-temporal Remote Sensing Tasks through Vision Language Models Geochat: Grounded large vision-language model for remote sensing,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:18:22.709468Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T23:18:21.266215Z digest=sha256:ecdffe59ee5eaea0d94bf2815279ad0d2c6c1fb7ca3b74bb855e31c9feda044b

Observation 1a1863ba-bd5f-4407-a0d8-7a0c620301cd · outbound

This paper cites SkyEyeGPT: Unifying Remote Sensing Vision-Language Tasks via Instruction Tuning with Large Language Model.

UniRS: Unifying Multi-temporal Remote Sensing Tasks through Vision Language Models SkyEyeGPT: Unifying Remote Sensing Vision-Language Tasks via Instruction Tuning with Large Language Model

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-10T23:18:21.272731Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:18:21.272731Z digest=sha256:1fac7306d44e2ed09b9bd2cace07f97f3fe35f6a829c106aa147ae35fda951b4

Observation 3e974fef-266d-4b36-9be6-62a27b8b8c7d · outbound

This paper cites TEOChat: A Large Vision-Language Assistant for Temporal Earth Observation Data.

UniRS: Unifying Multi-temporal Remote Sensing Tasks through Vision Language Models TEOChat: A Large Vision-Language Assistant for Temporal Earth Observation Data

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-10T23:18:21.279759Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:18:21.279759Z digest=sha256:eb8c2901a0c0225f872f6f3b96d930fbc6a62f864421310a21fc5f8875273973

Observation 4bab5494-172b-4a59-b513-4cb577f94817 · outbound

This paper cites Rs- llava: A large vision-language model for joint captioning and question answering in remote sensing imagery,.

UniRS: Unifying Multi-temporal Remote Sensing Tasks through Vision Language Models Rs- llava: A large vision-language model for joint captioning and question answering in remote sensing imagery,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:18:22.690464Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T23:18:21.288265Z digest=sha256:e8e4db4189f9127e29d52a595e55967c7223f488efb6b425ca6a2d6dabe23fd2

Observation 0221ae97-370c-4d17-9822-1dfbd82cf01e · outbound

This paper cites RSGPT: A Remote Sensing Vision Language Model and Benchmark.

UniRS: Unifying Multi-temporal Remote Sensing Tasks through Vision Language Models RSGPT: A Remote Sensing Vision Language Model and Benchmark

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-10T23:18:21.294085Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:18:21.294085Z digest=sha256:b0e10ebdf37e7f1b6fac4d8f4d2ca2043b5323a019e2f76c174a449af9c91c18

Observation 1362db8c-f70c-4af0-a999-173bbdf1a48f · outbound

This paper cites Earthgpt: A universal multi-modal large language model for multi-sensor image comprehension in remote sensing domain,.

UniRS: Unifying Multi-temporal Remote Sensing Tasks through Vision Language Models Earthgpt: A universal multi-modal large language model for multi-sensor image comprehension in remote sensing domain,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:18:22.636011Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T23:18:21.301355Z digest=sha256:652100818176a665c0267f59bc5c883faa5a6a71e1cc9e0cefb31e2e0ae748db

Observation 07501f7c-3e75-40cb-878c-73b447bb457e · outbound

This paper cites Remote sensing image change captioning with dual-branch transformers: A new method and a large scale dataset,.

UniRS: Unifying Multi-temporal Remote Sensing Tasks through Vision Language Models Remote sensing image change captioning with dual-branch transformers: A new method and a large scale dataset,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:18:22.539895Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T23:18:21.307209Z digest=sha256:497e0ed4fa2b4dda9be2d44f832b021a14c7774bd9afce18797650cbe0f3d4fc

Observation 7f315012-a026-4801-812b-b468fbf60e3c · outbound

This paper cites Era: A data set and deep learning benchmark for event recognition in aerial videos [software and data sets],.

UniRS: Unifying Multi-temporal Remote Sensing Tasks through Vision Language Models Era: A data set and deep learning benchmark for event recognition in aerial videos [software and data sets],

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:18:22.357789Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T23:18:21.314462Z digest=sha256:33ca693ff4b7c21e559bfb291856bd644b24283955e4497bd2fb29ac50ee6b27

Observation 453961bf-5b70-47b7-8e70-968a953bf34f · outbound

This paper cites GPT-4 Technical Report.

UniRS: Unifying Multi-temporal Remote Sensing Tasks through Vision Language Models GPT-4 Technical Report

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-10T23:18:21.320759Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:18:21.320759Z digest=sha256:31098746280a37477b223a42ffdc21ce55b641e7c3de2f4add06794c4f6c7322

Observation ad6d56c1-74cf-41cb-bd94-3a6849e34b0f · outbound

This paper cites Zero-shot video moment retrieval from frozen vision-language models,.

UniRS: Unifying Multi-temporal Remote Sensing Tasks through Vision Language Models Zero-shot video moment retrieval from frozen vision-language models,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:18:22.280819Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T23:18:21.327714Z digest=sha256:eb65c8e3c53c04edb657c0c12b5a533c15a1ee6d2417a1767948920d21f5e6fb

Observation 9c85abab-b506-4dbb-ac5f-f504453a2c63 · outbound

This paper cites Visual narratives: Large-scale hi- erarchical classification of art-historical images,.

UniRS: Unifying Multi-temporal Remote Sensing Tasks through Vision Language Models Visual narratives: Large-scale hi- erarchical classification of art-historical images,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:18:22.255200Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T23:18:21.334898Z digest=sha256:75a99d96f097f9674cd623b7ac6737f007c895b5ec5a7c440937a8cb23d1f148

Observation 184674d7-b23c-4a64-ad60-e2dc6ce7cb30 · outbound

This paper cites Learning transferable visual models from natural language supervision,.

UniRS: Unifying Multi-temporal Remote Sensing Tasks through Vision Language Models Learning transferable visual models from natural language supervision,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:18:22.162406Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T23:18:21.341042Z digest=sha256:1ba8fd545b6bc578658ecf3b345b3266bd7571b563ba8d72971b160a26cabddf

Observation ea49b646-97eb-4168-8986-1375fd0d1da5 · outbound

This paper cites Sigmoid loss for language image pre-training,.

UniRS: Unifying Multi-temporal Remote Sensing Tasks through Vision Language Models Sigmoid loss for language image pre-training,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:18:22.028018Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T23:18:21.346292Z digest=sha256:bad1b35fee95641c7fb47cd71317ac3b704c45ef338d6cfef8c21a7176a61701

Observation fddb3fee-08e5-48ad-b128-9367fe93cd6c · outbound

This paper cites Improved baselines with visual instruction tuning,.

UniRS: Unifying Multi-temporal Remote Sensing Tasks through Vision Language Models Improved baselines with visual instruction tuning,

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:18:21.979100Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T23:18:21.350968Z digest=sha256:27316ca002276466da4f708ea4837a4ae5ffc9629f4c237fa61503ecdd8f2145

Observation f3867446-ec3d-4bc5-b5ed-6bd479515b0e · outbound

This paper cites SkySenseGPT: A Fine-Grained Instruction Tuning Dataset and Model for Remote Sensing Vision-Language Understanding.

UniRS: Unifying Multi-temporal Remote Sensing Tasks through Vision Language Models SkySenseGPT: A Fine-Grained Instruction Tuning Dataset and Model for Remote Sensing Vision-Language Understanding

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-10T23:18:21.355874Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:18:21.355874Z digest=sha256:8af11767e43d33f9d83e39cc2a2d07bee2056e69bdfb1531cfeaebd1b20b5c9a

Observation 210f211f-57c8-45f8-82e3-8aea311fc432 · outbound

This paper cites Nwpu- captions dataset and mlca-net for remote sensing image captioning,.

UniRS: Unifying Multi-temporal Remote Sensing Tasks through Vision Language Models Nwpu- captions dataset and mlca-net for remote sensing image captioning,

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:18:21.962185Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T23:18:21.360954Z digest=sha256:e4d39a0ca2853785c738c654407fddee5c3e3bccf87caf08d3d20200d3199e0e

Observation f960d65f-2fc5-4c05-b7da-203a8ee61fcd · outbound

This paper cites Multiscale spatio- temporal network for aerial video event recognition,.

UniRS: Unifying Multi-temporal Remote Sensing Tasks through Vision Language Models Multiscale spatio- temporal network for aerial video event recognition,

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:18:21.945407Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T23:18:21.365922Z digest=sha256:335a39cd21b5b5da9f464934fc74cf3f0ba3b19ea341c5caca5f67608b292004

Observation 884a5ea8-b3bf-4a5c-9ed4-4c0d35801b68 · outbound

This paper cites Alleviating spatial misalignment and motion interference for uav-based video recognition,.

UniRS: Unifying Multi-temporal Remote Sensing Tasks through Vision Language Models Alleviating spatial misalignment and motion interference for uav-based video recognition,

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:18:21.926977Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T23:18:21.370699Z digest=sha256:d19b73409239538f153f6a787b05ee707a58339902c82016006880e62c381c3e

Observation d9ef3b67-f80d-40f8-a649-9074db3cdc69 · outbound

This paper cites Multitask learning: A knowledge-based source of inductive bias,.

UniRS: Unifying Multi-temporal Remote Sensing Tasks through Vision Language Models Multitask learning: A knowledge-based source of inductive bias,

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:18:21.909314Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T23:18:21.375349Z digest=sha256:b2738cd22dab070f415ecf53872e8d203593fdf7c1506c739a29bdd776dfb2bf

Observation d7233d7c-ecbd-49bf-992d-7220d1f77ccf · outbound

This paper cites Multitask learning,.

UniRS: Unifying Multi-temporal Remote Sensing Tasks through Vision Language Models Multitask learning,

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:18:21.890641Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T23:18:21.380080Z digest=sha256:85a9a8b519c08fd477f23b81c6d3e4e69019f0d2d4a9a5e1b838dc748eb6d64e

Observation 73f794b3-9073-475d-a394-c26c4b5a34de · outbound

This paper cites Multitask learning for crash analysis: A fine-tuned llm framework using twitter data,.

UniRS: Unifying Multi-temporal Remote Sensing Tasks through Vision Language Models Multitask learning for crash analysis: A fine-tuned llm framework using twitter data,

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:18:21.873166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T23:18:21.384800Z digest=sha256:1254742dd1fe22936c350520fd1f2d667b69788082142f8b52b6ab36706c6652

Observation 2199e715-5976-4748-bfc8-cc11eb5bce13 · outbound

This paper cites TigerBot: An Open Multilingual Multitask LLM.

UniRS: Unifying Multi-temporal Remote Sensing Tasks through Vision Language Models TigerBot: An Open Multilingual Multitask LLM

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-10T23:18:21.389831Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:18:21.389831Z digest=sha256:54185c41314784bd47dc0f107753012abfa1ea844444499740fb112c26679d37

Observation ce2dd9f4-c4d0-4e63-9ac2-b06eb3a7eb98 · outbound

This paper cites Mftcoder: Boosting code llms with multitask fine-tuning,.

UniRS: Unifying Multi-temporal Remote Sensing Tasks through Vision Language Models Mftcoder: Boosting code llms with multitask fine-tuning,

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:18:21.855590Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T23:18:21.394659Z digest=sha256:68d42e52dc50f7c9b45432744312236266a4923a5daf9f785c262ad2d19bf80d

Observation e7289dbe-7768-43e2-a089-033eeb79ff8e · outbound

This paper cites Sheared LLaMA: Accelerating Language Model Pre-training via Structured Pruning.

UniRS: Unifying Multi-temporal Remote Sensing Tasks through Vision Language Models Sheared LLaMA: Accelerating Language Model Pre-training via Structured Pruning

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-10T23:18:21.398995Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:18:21.398995Z digest=sha256:0c86fa4d8424ae1c24c2e260e12a9dd1ce9ffc3249b05ecbeedcdb54401f4890

Observation bb21a836-9d5c-4f3c-9b37-2b0d49b7ab2c · outbound

This paper cites Dota: A large-scale dataset for object detection in aerial images,.

UniRS: Unifying Multi-temporal Remote Sensing Tasks through Vision Language Models Dota: A large-scale dataset for object detection in aerial images,

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:18:21.837573Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T23:18:21.404590Z digest=sha256:9748a46ce2c85101420e82b5d33d4b0df08d53d038ff3163b6d95789bce04886

Observation cd9dd48d-bcf2-4dcb-a9ee-dc8d7bbafe6c · outbound

This paper cites Anchor- free oriented proposal generator for object detection,.

UniRS: Unifying Multi-temporal Remote Sensing Tasks through Vision Language Models Anchor- free oriented proposal generator for object detection,

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:18:21.820206Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T23:18:21.409442Z digest=sha256:78e581340bcd66165321b1715bf2dd04950075f0a9c60e86df5dca635a70309d

Observation a556e187-3f3e-4375-bf3f-6d507618b135 · outbound

This paper cites Fair1m: A benchmark dataset for fine-grained object recognition in high-resolution remote sensing imagery,.

UniRS: Unifying Multi-temporal Remote Sensing Tasks through Vision Language Models Fair1m: A benchmark dataset for fine-grained object recognition in high-resolution remote sensing imagery,

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:18:21.800383Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T23:18:21.413999Z digest=sha256:7cbcc08d2a27fb63696a9cd0ad8f1922d38463d4b810c337e61f5310a10538cc

Observation 0baf7aa0-97f2-4948-9f25-6ba66454b5b6 · outbound

This paper cites Remote sensing image scene classifi- cation: Benchmark and state of the art,.

UniRS: Unifying Multi-temporal Remote Sensing Tasks through Vision Language Models Remote sensing image scene classifi- cation: Benchmark and state of the art,

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:18:21.782752Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T23:18:21.419113Z digest=sha256:37575f382535d66b60ac4faf55700ebbe5a67bbbed3f86472ccf88453a6836bd

Observation c013b11a-333a-453e-9990-53ae176d121d · outbound

This paper cites Floodnet: A high resolution aerial imagery dataset for post flood scene understanding,.

UniRS: Unifying Multi-temporal Remote Sensing Tasks through Vision Language Models Floodnet: A high resolution aerial imagery dataset for post flood scene understanding,

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:18:21.764505Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T23:18:21.424505Z digest=sha256:ac58200f4062deae3d0b6fa3ee5480a9c0d1563e4da94ef4f5d6a086450adce8

Observation 16960b0b-0ce9-47a7-84bf-a5a805e7ea54 · outbound

This paper cites A spatial-temporal attention-based method and a new dataset for remote sensing image change detection,.

UniRS: Unifying Multi-temporal Remote Sensing Tasks through Vision Language Models A spatial-temporal attention-based method and a new dataset for remote sensing image change detection,

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:18:21.744693Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T23:18:21.429256Z digest=sha256:8a380484654a6572230242b27105492e70caa7edbf524547365ec29220c4fd90

Observation d3fff566-c8d6-4717-ba45-78a27bdbc18a · outbound

This paper cites A spatial hierarchical reasoning network for remote sensing visual question answering,.

UniRS: Unifying Multi-temporal Remote Sensing Tasks through Vision Language Models A spatial hierarchical reasoning network for remote sensing visual question answering,

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:18:21.726638Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T23:18:21.434050Z digest=sha256:6803a9a2b8f2a5971cdc825db1b25afe732605291f63935a21d5f0dc2b3a8802

Observation f4ccce72-88d5-4d46-9aa5-cc0389180fae · outbound

This paper cites Temporal rela- tions matter: A two-pathway network for aerial video recognition,.

UniRS: Unifying Multi-temporal Remote Sensing Tasks through Vision Language Models Temporal rela- tions matter: A two-pathway network for aerial video recognition,

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:18:21.706541Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T23:18:21.438754Z digest=sha256:41a87d01b9d3149fd956bea5842daa0759013ec46f2087c03c74cb20f8639c75

Pith citing papers

Observation 7285c014-15ee-4005-8ea9-9c153e94f6a6 · inbound

RSVideo: Are Your Vision-Language Models Ready for Remote Sensing Videos? cites this paper.

RSVideo: Are Your Vision-Language Models Ready for Remote Sensing Videos? UniRS: Unifying Multi-temporal Remote Sensing Tasks through Vision Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-04T16:16:50.580454Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:16:50.580454Z digest=sha256:a0498e4cdaa6fda4f780c426bd912268fb02ee7e5f9514ae743c0021b1c4b6a3

Observation 0df9cd09-f311-4e22-8d2e-b4f62b846751 · inbound

RSVideo: Are Your Vision-Language Models Ready for Remote Sensing Videos? cites this paper.

RSVideo: Are Your Vision-Language Models Ready for Remote Sensing Videos? UniRS: Unifying Multi-temporal Remote Sensing Tasks through Vision Language Models

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-08-07T00:14:11.514686Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T00:14:05.499836Z digest=sha256:d346b4845ad698e67677721a88b866dbc01015056572ee2e8e7b70bf99024198