Pith. sign in

Paper Citation Record · LEDGER

UniRS: Unifying Multi-temporal Remote Sensing Tasks through Vision Language Models

As of 14 August 2026, this Paper Citation Record lists 56 of 56 outbound references and 2 inbound Pith citation observations for arXiv:2412.20742.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.20742 v1

Coverage vector

measured 56 of 56 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T23:18:21.438754Z

measured 58 of 58 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T00:14:05.499836Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-07T00:14:11.438315Z

Reference resolution

56 of 56 outbound references displayed

  • verified exact0
  • verified fuzzy44
  • unresolved12
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation cfdd125b-a09c-4491-883a-9cbe2377ad50 · outbound

This paper cites Visual instruction tuning,.

UniRS: Unifying Multi-temporal Remote Sensing Tasks through Vision Language Models Visual instruction tuning,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:18:23.458753Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T23:18:21.129990Z digest=sha256:46db9c19b199ac8369bdface76efa1554cac61bf24778660c804984f015d237f

Observation ce25b4aa-3cbd-43d6-8ce7-7b336e7685b2 · outbound

This paper cites Flamingo: a visual language model for few-shot learning,.

UniRS: Unifying Multi-temporal Remote Sensing Tasks through Vision Language Models Flamingo: a visual language model for few-shot learning,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:18:23.367346Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T23:18:21.136231Z digest=sha256:9bf0db78d9ab5d2f8a417f4e469bb2b64a2519f637ae954527af4ffd8d68d358

Observation 530008c7-1235-4052-950f-d8ec06735c83 · outbound

This paper cites Vila: On pre-training for visual language models,.

UniRS: Unifying Multi-temporal Remote Sensing Tasks through Vision Language Models Vila: On pre-training for visual language models,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:18:23.350477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T23:18:21.141609Z digest=sha256:d4363869eb66d11d9abfa5418e9c9bed51292d907f3e8c7d91b9a07b25e5a62d

Observation fbe617b2-21fb-44de-8a92-073732e6ca3c · outbound

This paper cites MiniGPT-v2: large language model as a unified interface for vision-language multi-task learning.

UniRS: Unifying Multi-temporal Remote Sensing Tasks through Vision Language Models MiniGPT-v2: large language model as a unified interface for vision-language multi-task learning

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-10T23:18:21.147250Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:18:21.147250Z digest=sha256:1279a37c0dc28a1939f70a83e0aaa1460912691ff40ab50ba2128a0ef5f30924

Observation 06d2b9f4-d435-4e2c-9e8b-d09ecacff320 · outbound

This paper cites Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation,.

UniRS: Unifying Multi-temporal Remote Sensing Tasks through Vision Language Models Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:18:23.333186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T23:18:21.153597Z digest=sha256:78ee465cbaa034f8bdab9ad2b775167ee46610456d9e24aaefafe42d51bfa074

Observation 799829fc-f48b-40d5-9955-a60296e47ea9 · outbound

This paper cites Blip-2: Bootstrapping language- image pre-training with frozen image encoders and large language models,.

UniRS: Unifying Multi-temporal Remote Sensing Tasks through Vision Language Models Blip-2: Bootstrapping language- image pre-training with frozen image encoders and large language models,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:18:23.315300Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T23:18:21.158791Z digest=sha256:03be88cc2641a9ccc56403ee13e96e2e8dfa506befec09e2931a2906d2362927

Observation 21c7731b-5099-4296-a867-84ceff71c929 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

UniRS: Unifying Multi-temporal Remote Sensing Tasks through Vision Language Models LLaMA: Open and Efficient Foundation Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-10T23:18:21.164175Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:18:21.164175Z digest=sha256:250d6c3206d8a9fbbc75954e2bedaf41d9f0ee1a3665bc65fc41ed90a99be59d

Observation df83d8ad-0e80-44f5-a8c3-905779f10e98 · outbound

This paper cites Training language models to follow instructions with human feedback,.

UniRS: Unifying Multi-temporal Remote Sensing Tasks through Vision Language Models Training language models to follow instructions with human feedback,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:18:23.296497Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T23:18:21.170197Z digest=sha256:edf8f5d1bf7c3722dc2b4237f703cf58672a541e8076efa72901435f6bdb8dec

Observation c88ce50f-906b-4c8f-8b5d-37fc8e4da3d3 · outbound

This paper cites Qwen Technical Report.

UniRS: Unifying Multi-temporal Remote Sensing Tasks through Vision Language Models Qwen Technical Report

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-10T23:18:21.175296Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:18:21.175296Z digest=sha256:044c3f9c6981dc604a23b409dc7de73d40241ddf8e1454859a7564d4fd60c710

Observation 5b31fda4-bd17-4c4d-9c1f-41a87607715a · outbound

This paper cites Language mod- els are few-shot learners,.

UniRS: Unifying Multi-temporal Remote Sensing Tasks through Vision Language Models Language mod- els are few-shot learners,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:18:23.278960Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T23:18:21.180320Z digest=sha256:3699f09b72156cefd9c2e65d8c2015fb7b4777605d50f9e3111147aa286000ff

Observation 09aa9e4d-7db1-406f-b233-051a3d217e69 · outbound

This paper cites Laion- 5b: An open large-scale dataset for training next generation image-text models,.

UniRS: Unifying Multi-temporal Remote Sensing Tasks through Vision Language Models Laion- 5b: An open large-scale dataset for training next generation image-text models,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:18:23.261896Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T23:18:21.185583Z digest=sha256:870f9efbcd8f1372ef64f34307664776fda70c92b9b6262f5ecd0baebfcaa40f

Observation f3555b1b-c8ba-41f1-9012-807276725dd0 · outbound

This paper cites Llava-med: Training a large language-and-vision assistant for biomedicine in one day,.

UniRS: Unifying Multi-temporal Remote Sensing Tasks through Vision Language Models Llava-med: Training a large language-and-vision assistant for biomedicine in one day,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:18:23.245810Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T23:18:21.190988Z digest=sha256:d81f4ad2160f459b38b5b4679411c77245fd4659dc335ee133b78357a2f03cbc

Observation f0d34255-e33e-49db-837d-38721ecf3295 · outbound

This paper cites Mobility VLA: Multimodal Instruction Navigation with Long-Context VLMs and Topological Graphs.

UniRS: Unifying Multi-temporal Remote Sensing Tasks through Vision Language Models Mobility VLA: Multimodal Instruction Navigation with Long-Context VLMs and Topological Graphs

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-10T23:18:21.196144Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:18:21.196144Z digest=sha256:875f41d40053584fd645ce9d2f06b44b36ebb964e32ecf117ea27706ad634039

Observation 4632d040-0eac-4350-aef7-8a6f7e96d63c · outbound

This paper cites Rsvqa: Visual question answering for remote sensing data,.

UniRS: Unifying Multi-temporal Remote Sensing Tasks through Vision Language Models Rsvqa: Visual question answering for remote sensing data,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:18:23.229120Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T23:18:21.201942Z digest=sha256:46c180994647fdd44fae1bbef0b3f2f7748299de2f446fae89561c6ecf6ec895

Observation f0c8b997-0e45-4fcd-b1dd-eba28ef142ce · outbound

This paper cites Answer-type prediction for visual question answering,.

UniRS: Unifying Multi-temporal Remote Sensing Tasks through Vision Language Models Answer-type prediction for visual question answering,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:18:23.212597Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T23:18:21.208078Z digest=sha256:f41d7b9b99e6384d9e96a6cb4a2f56ab2036a7e87a39030618f1825ebec62ee2

Observation ea401cdc-6899-4110-a644-8fcca327b6c4 · outbound

This paper cites Multi-step question-driven visual ques- tion answering for remote sensing,.

UniRS: Unifying Multi-temporal Remote Sensing Tasks through Vision Language Models Multi-step question-driven visual ques- tion answering for remote sensing,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:18:23.160301Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T23:18:21.213271Z digest=sha256:df992c93f400d068834cfa8b06c6706e91524c324022040ec49b8ece045d6a12

Observation 7652e053-07a3-41f4-8123-ab8624cd9906 · outbound

This paper cites Bi-modal transformer-based approach for visual question answering in remote sensing imagery,.

UniRS: Unifying Multi-temporal Remote Sensing Tasks through Vision Language Models Bi-modal transformer-based approach for visual question answering in remote sensing imagery,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:18:23.051315Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T23:18:21.218321Z digest=sha256:948244f809d26be276e084991712b186f93defe84b67eaf54523776b91dd99f3

Observation d79f23aa-1b6c-4153-b41b-a9432905d186 · outbound

This paper cites Changes to captions: An attentive network for remote sensing change captioning,.

UniRS: Unifying Multi-temporal Remote Sensing Tasks through Vision Language Models Changes to captions: An attentive network for remote sensing change captioning,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:18:22.907158Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T23:18:21.224247Z digest=sha256:9d13fa3fda4baa368404fd0e971d0a19d705fe857a7e4d67666cd49cc2b1c37f

Observation 9ddb6af0-974c-41dc-a04d-b6a1edf19bfc · outbound

This paper cites Progressive scale-aware network for remote sensing image change captioning,.

UniRS: Unifying Multi-temporal Remote Sensing Tasks through Vision Language Models Progressive scale-aware network for remote sensing image change captioning,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:18:22.805680Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T23:18:21.229092Z digest=sha256:429d56631a522758bb3554825a86c6383c3d7ea5da90b258301a6b45785d8075

Observation 12399aa5-b068-4bfb-beb3-a61dc53e7f88 · outbound

This paper cites Futh-net: fusing temporal relations and holistic features for aerial video classification,.

UniRS: Unifying Multi-temporal Remote Sensing Tasks through Vision Language Models Futh-net: fusing temporal relations and holistic features for aerial video classification,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:18:22.779334Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T23:18:21.234463Z digest=sha256:0a641f7a48a97e16fa924246937cdc8ef83ae9881ea68e4f1c87bb295b44df18

Observation 9362b406-f560-4720-b98e-23397670ffe5 · outbound

This paper cites Ok-vqa: A visual question answering benchmark requiring external knowledge,.

UniRS: Unifying Multi-temporal Remote Sensing Tasks through Vision Language Models Ok-vqa: A visual question answering benchmark requiring external knowledge,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:18:22.762710Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T23:18:21.240700Z digest=sha256:fa31fd4bf87e22fbd832c8456d96687a56297358815f4b58b268e1d937f05b02

Observation 76fd99e3-cf0a-4563-a1de-b0420b434339 · outbound

This paper cites Deep learning based event recognition in aerial imagery,.

UniRS: Unifying Multi-temporal Remote Sensing Tasks through Vision Language Models Deep learning based event recognition in aerial imagery,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:18:22.744621Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T23:18:21.247367Z digest=sha256:5c53e55b8f89e343626818b87c656dbff4fb7d0309d33d88972f28b2a908f875

Observation 4ad2a81b-9a8d-4704-904f-80df22718a0b · outbound

This paper cites A decoupling paradigm with prompt learning for remote sensing image change cap- tioning,.

UniRS: Unifying Multi-temporal Remote Sensing Tasks through Vision Language Models A decoupling paradigm with prompt learning for remote sensing image change cap- tioning,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:18:22.727137Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T23:18:21.254060Z digest=sha256:2166a89351d155b7476f206577a10867501b625ac8521c0a5329c11e20fc6a1e

Observation 1892636e-0c9b-4294-a9e2-aee3f275b794 · outbound

This paper cites LHRS-Bot: Empowering Remote Sensing with VGI-Enhanced Large Multimodal Language Model.

UniRS: Unifying Multi-temporal Remote Sensing Tasks through Vision Language Models LHRS-Bot: Empowering Remote Sensing with VGI-Enhanced Large Multimodal Language Model

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-10T23:18:21.260642Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:18:21.260642Z digest=sha256:ae873258ea1363bafb0e76578857437c368665acc5300a0dbb92f90729fd2eb3

Observation 806d599f-8e76-4721-a274-ab799ca5cd0e · outbound

This paper cites Geochat: Grounded large vision-language model for remote sensing,.

UniRS: Unifying Multi-temporal Remote Sensing Tasks through Vision Language Models Geochat: Grounded large vision-language model for remote sensing,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:18:22.709468Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T23:18:21.266215Z digest=sha256:20b5ba3444838c34456073ca6a9e8ea8e5cd9f24723671b73519194c946db274

Observation 1a1863ba-bd5f-4407-a0d8-7a0c620301cd · outbound

This paper cites SkyEyeGPT: Unifying Remote Sensing Vision-Language Tasks via Instruction Tuning with Large Language Model.

UniRS: Unifying Multi-temporal Remote Sensing Tasks through Vision Language Models SkyEyeGPT: Unifying Remote Sensing Vision-Language Tasks via Instruction Tuning with Large Language Model

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-10T23:18:21.272731Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:18:21.272731Z digest=sha256:b8c29c04c18b44999a294060fe6b69b9f5661b1b9873f8ff6e27c702e2d9c268

Observation 3e974fef-266d-4b36-9be6-62a27b8b8c7d · outbound

This paper cites TEOChat: A Large Vision-Language Assistant for Temporal Earth Observation Data.

UniRS: Unifying Multi-temporal Remote Sensing Tasks through Vision Language Models TEOChat: A Large Vision-Language Assistant for Temporal Earth Observation Data

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-10T23:18:21.279759Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:18:21.279759Z digest=sha256:eb8c2901a0c0225f872f6f3b96d930fbc6a62f864421310a21fc5f8875273973

Observation 4bab5494-172b-4a59-b513-4cb577f94817 · outbound

This paper cites Rs- llava: A large vision-language model for joint captioning and question answering in remote sensing imagery,.

UniRS: Unifying Multi-temporal Remote Sensing Tasks through Vision Language Models Rs- llava: A large vision-language model for joint captioning and question answering in remote sensing imagery,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:18:22.690464Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T23:18:21.288265Z digest=sha256:95b10654c9b08a83964aa52fe61c468ddc70935efb580fffb93828fc4c945efe

Observation 0221ae97-370c-4d17-9822-1dfbd82cf01e · outbound

This paper cites RSGPT: A Remote Sensing Vision Language Model and Benchmark.

UniRS: Unifying Multi-temporal Remote Sensing Tasks through Vision Language Models RSGPT: A Remote Sensing Vision Language Model and Benchmark

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-10T23:18:21.294085Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:18:21.294085Z digest=sha256:b0e10ebdf37e7f1b6fac4d8f4d2ca2043b5323a019e2f76c174a449af9c91c18

Observation 1362db8c-f70c-4af0-a999-173bbdf1a48f · outbound

This paper cites Earthgpt: A universal multi-modal large language model for multi-sensor image comprehension in remote sensing domain,.

UniRS: Unifying Multi-temporal Remote Sensing Tasks through Vision Language Models Earthgpt: A universal multi-modal large language model for multi-sensor image comprehension in remote sensing domain,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:18:22.636011Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T23:18:21.301355Z digest=sha256:3befea64f7a531b04ad57ce14f0453ca41039dff5956f7d5aa8ef131adf5f28e

Observation 07501f7c-3e75-40cb-878c-73b447bb457e · outbound

This paper cites Remote sensing image change captioning with dual-branch transformers: A new method and a large scale dataset,.

UniRS: Unifying Multi-temporal Remote Sensing Tasks through Vision Language Models Remote sensing image change captioning with dual-branch transformers: A new method and a large scale dataset,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:18:22.539895Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T23:18:21.307209Z digest=sha256:acca24d770572dd3e879d45403032193ba83c59916f2cfba0b4fffc750ebdfcf

Observation 7f315012-a026-4801-812b-b468fbf60e3c · outbound

This paper cites Era: A data set and deep learning benchmark for event recognition in aerial videos [software and data sets],.

UniRS: Unifying Multi-temporal Remote Sensing Tasks through Vision Language Models Era: A data set and deep learning benchmark for event recognition in aerial videos [software and data sets],

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:18:22.357789Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T23:18:21.314462Z digest=sha256:b9256179468d56c729d6aa46459fc78213207292b0978c154725760817cd672d

Observation 453961bf-5b70-47b7-8e70-968a953bf34f · outbound

This paper cites GPT-4 Technical Report.

UniRS: Unifying Multi-temporal Remote Sensing Tasks through Vision Language Models GPT-4 Technical Report

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-10T23:18:21.320759Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:18:21.320759Z digest=sha256:31098746280a37477b223a42ffdc21ce55b641e7c3de2f4add06794c4f6c7322

Observation ad6d56c1-74cf-41cb-bd94-3a6849e34b0f · outbound

This paper cites Zero-shot video moment retrieval from frozen vision-language models,.

UniRS: Unifying Multi-temporal Remote Sensing Tasks through Vision Language Models Zero-shot video moment retrieval from frozen vision-language models,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:18:22.280819Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T23:18:21.327714Z digest=sha256:854ce334c87ad847c4b4594c97ba0fbd81ea71ffed26f6e3922edf124c31ce8a

Observation 9c85abab-b506-4dbb-ac5f-f504453a2c63 · outbound

This paper cites Visual narratives: Large-scale hi- erarchical classification of art-historical images,.

UniRS: Unifying Multi-temporal Remote Sensing Tasks through Vision Language Models Visual narratives: Large-scale hi- erarchical classification of art-historical images,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:18:22.255200Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T23:18:21.334898Z digest=sha256:76fd66002d10d6b5f5d8c0fc6494916256e20c36f34322fa0b8bef82e0cf308c

Observation 184674d7-b23c-4a64-ad60-e2dc6ce7cb30 · outbound

This paper cites Learning transferable visual models from natural language supervision,.

UniRS: Unifying Multi-temporal Remote Sensing Tasks through Vision Language Models Learning transferable visual models from natural language supervision,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:18:22.162406Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T23:18:21.341042Z digest=sha256:0519f7af3e25f00feddcf40b11b2eb8eac4e01acbbb44accfccda9d84050d8ce

Observation ea49b646-97eb-4168-8986-1375fd0d1da5 · outbound

This paper cites Sigmoid loss for language image pre-training,.

UniRS: Unifying Multi-temporal Remote Sensing Tasks through Vision Language Models Sigmoid loss for language image pre-training,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:18:22.028018Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T23:18:21.346292Z digest=sha256:1b76a2217d191b037eb8619f2b0284402c445fbeb0c30a80f498d03e4c404adc

Observation fddb3fee-08e5-48ad-b128-9367fe93cd6c · outbound

This paper cites Improved baselines with visual instruction tuning,.

UniRS: Unifying Multi-temporal Remote Sensing Tasks through Vision Language Models Improved baselines with visual instruction tuning,

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:18:21.979100Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T23:18:21.350968Z digest=sha256:5a496969da05fac320a23e885e5e1cee3ebc5f84799fdd66c857cf94d83d9a1b

Observation f3867446-ec3d-4bc5-b5ed-6bd479515b0e · outbound

This paper cites SkySenseGPT: A Fine-Grained Instruction Tuning Dataset and Model for Remote Sensing Vision-Language Understanding.

UniRS: Unifying Multi-temporal Remote Sensing Tasks through Vision Language Models SkySenseGPT: A Fine-Grained Instruction Tuning Dataset and Model for Remote Sensing Vision-Language Understanding

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-10T23:18:21.355874Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:18:21.355874Z digest=sha256:8af11767e43d33f9d83e39cc2a2d07bee2056e69bdfb1531cfeaebd1b20b5c9a

Observation 210f211f-57c8-45f8-82e3-8aea311fc432 · outbound

This paper cites Nwpu- captions dataset and mlca-net for remote sensing image captioning,.

UniRS: Unifying Multi-temporal Remote Sensing Tasks through Vision Language Models Nwpu- captions dataset and mlca-net for remote sensing image captioning,

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:18:21.962185Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T23:18:21.360954Z digest=sha256:a7c4068c1a608590f33f9fd2477f1a0485a3529f7669aefde106f7afebc3fa15

Observation f960d65f-2fc5-4c05-b7da-203a8ee61fcd · outbound

This paper cites Multiscale spatio- temporal network for aerial video event recognition,.

UniRS: Unifying Multi-temporal Remote Sensing Tasks through Vision Language Models Multiscale spatio- temporal network for aerial video event recognition,

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:18:21.945407Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T23:18:21.365922Z digest=sha256:8a253205ebd6734f13b110e385ab729043782887db6507889df4f4249f191eed

Observation 884a5ea8-b3bf-4a5c-9ed4-4c0d35801b68 · outbound

This paper cites Alleviating spatial misalignment and motion interference for uav-based video recognition,.

UniRS: Unifying Multi-temporal Remote Sensing Tasks through Vision Language Models Alleviating spatial misalignment and motion interference for uav-based video recognition,

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:18:21.926977Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T23:18:21.370699Z digest=sha256:72be5e6ad8b1cd26c85b61063799d0b9c291c8a103a7c0a54224e8bcb0baae26

Observation d9ef3b67-f80d-40f8-a649-9074db3cdc69 · outbound

This paper cites Multitask learning: A knowledge-based source of inductive bias,.

UniRS: Unifying Multi-temporal Remote Sensing Tasks through Vision Language Models Multitask learning: A knowledge-based source of inductive bias,

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:18:21.909314Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T23:18:21.375349Z digest=sha256:07ca9f3e4af8e7a4a051ba2420065892a03f99922c4145fd8590d25bf47dc874

Observation d7233d7c-ecbd-49bf-992d-7220d1f77ccf · outbound

This paper cites Multitask learning,.

UniRS: Unifying Multi-temporal Remote Sensing Tasks through Vision Language Models Multitask learning,

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:18:21.890641Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T23:18:21.380080Z digest=sha256:8c8efb28f1606e4f6e327c70bc1ad94eaf3c85c02b4802bf1ca95dc4d3796090

Observation 73f794b3-9073-475d-a394-c26c4b5a34de · outbound

This paper cites Multitask learning for crash analysis: A fine-tuned llm framework using twitter data,.

UniRS: Unifying Multi-temporal Remote Sensing Tasks through Vision Language Models Multitask learning for crash analysis: A fine-tuned llm framework using twitter data,

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:18:21.873166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T23:18:21.384800Z digest=sha256:3b0a1e48053bdce207c492e377ec51c21fd17586654f4fa70dcbe147305de82d

Observation 2199e715-5976-4748-bfc8-cc11eb5bce13 · outbound

This paper cites TigerBot: An Open Multilingual Multitask LLM.

UniRS: Unifying Multi-temporal Remote Sensing Tasks through Vision Language Models TigerBot: An Open Multilingual Multitask LLM

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-10T23:18:21.389831Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:18:21.389831Z digest=sha256:54185c41314784bd47dc0f107753012abfa1ea844444499740fb112c26679d37

Observation ce2dd9f4-c4d0-4e63-9ac2-b06eb3a7eb98 · outbound

This paper cites Mftcoder: Boosting code llms with multitask fine-tuning,.

UniRS: Unifying Multi-temporal Remote Sensing Tasks through Vision Language Models Mftcoder: Boosting code llms with multitask fine-tuning,

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:18:21.855590Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T23:18:21.394659Z digest=sha256:367dfd8a7ac7bc3c38c8441435c922e87feb746b723215cff690eb17b8dc2922

Observation e7289dbe-7768-43e2-a089-033eeb79ff8e · outbound

This paper cites Sheared LLaMA: Accelerating Language Model Pre-training via Structured Pruning.

UniRS: Unifying Multi-temporal Remote Sensing Tasks through Vision Language Models Sheared LLaMA: Accelerating Language Model Pre-training via Structured Pruning

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-10T23:18:21.398995Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:18:21.398995Z digest=sha256:0c86fa4d8424ae1c24c2e260e12a9dd1ce9ffc3249b05ecbeedcdb54401f4890

Observation bb21a836-9d5c-4f3c-9b37-2b0d49b7ab2c · outbound

This paper cites Dota: A large-scale dataset for object detection in aerial images,.

UniRS: Unifying Multi-temporal Remote Sensing Tasks through Vision Language Models Dota: A large-scale dataset for object detection in aerial images,

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:18:21.837573Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T23:18:21.404590Z digest=sha256:a8ce9972b7c46700910489ed6a1b8f30f800136d9c4f953d34f8e4d8200b7b0e

Observation cd9dd48d-bcf2-4dcb-a9ee-dc8d7bbafe6c · outbound

This paper cites Anchor- free oriented proposal generator for object detection,.

UniRS: Unifying Multi-temporal Remote Sensing Tasks through Vision Language Models Anchor- free oriented proposal generator for object detection,

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:18:21.820206Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T23:18:21.409442Z digest=sha256:a8ea7dc9c59d30b0a71addf8b69ed94ad740217eee756083d4a9eed36521d6b1

Observation a556e187-3f3e-4375-bf3f-6d507618b135 · outbound

This paper cites Fair1m: A benchmark dataset for fine-grained object recognition in high-resolution remote sensing imagery,.

UniRS: Unifying Multi-temporal Remote Sensing Tasks through Vision Language Models Fair1m: A benchmark dataset for fine-grained object recognition in high-resolution remote sensing imagery,

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:18:21.800383Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T23:18:21.413999Z digest=sha256:9b2f8b65fb300f338eaf859ac705e6d69f014f4af946c935e1ae2c447666c1ce

Observation 0baf7aa0-97f2-4948-9f25-6ba66454b5b6 · outbound

This paper cites Remote sensing image scene classifi- cation: Benchmark and state of the art,.

UniRS: Unifying Multi-temporal Remote Sensing Tasks through Vision Language Models Remote sensing image scene classifi- cation: Benchmark and state of the art,

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:18:21.782752Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T23:18:21.419113Z digest=sha256:ad8f03eec1f187d0d9d8241c1c074ef07082701ac86cd2350a04d81f1c10b7b5

Observation c013b11a-333a-453e-9990-53ae176d121d · outbound

This paper cites Floodnet: A high resolution aerial imagery dataset for post flood scene understanding,.

UniRS: Unifying Multi-temporal Remote Sensing Tasks through Vision Language Models Floodnet: A high resolution aerial imagery dataset for post flood scene understanding,

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:18:21.764505Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T23:18:21.424505Z digest=sha256:24114fc8fc79c1afe2fa1c8bdb44b641e6384d639afe3d86617575216e970a9d

Observation 16960b0b-0ce9-47a7-84bf-a5a805e7ea54 · outbound

This paper cites A spatial-temporal attention-based method and a new dataset for remote sensing image change detection,.

UniRS: Unifying Multi-temporal Remote Sensing Tasks through Vision Language Models A spatial-temporal attention-based method and a new dataset for remote sensing image change detection,

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:18:21.744693Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T23:18:21.429256Z digest=sha256:47307ed74d3e8811aa756b6bdb2d8646b79fe2f376dda35fec73faad4dd73211

Observation d3fff566-c8d6-4717-ba45-78a27bdbc18a · outbound

This paper cites A spatial hierarchical reasoning network for remote sensing visual question answering,.

UniRS: Unifying Multi-temporal Remote Sensing Tasks through Vision Language Models A spatial hierarchical reasoning network for remote sensing visual question answering,

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:18:21.726638Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T23:18:21.434050Z digest=sha256:83be2220495e600c1622b483e9823f57d9a8e18982db5bc9965d62d578c70f5a

Observation f4ccce72-88d5-4d46-9aa5-cc0389180fae · outbound

This paper cites Temporal rela- tions matter: A two-pathway network for aerial video recognition,.

UniRS: Unifying Multi-temporal Remote Sensing Tasks through Vision Language Models Temporal rela- tions matter: A two-pathway network for aerial video recognition,

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:18:21.706541Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T23:18:21.438754Z digest=sha256:c8c0c18b91650d9593a169eca07c582abb76e2699506f52a57763e27678e7ab6

Pith citing papers

Observation 7285c014-15ee-4005-8ea9-9c153e94f6a6 · inbound

RSVideo: Are Your Vision-Language Models Ready for Remote Sensing Videos? cites this paper.

RSVideo: Are Your Vision-Language Models Ready for Remote Sensing Videos? UniRS: Unifying Multi-temporal Remote Sensing Tasks through Vision Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-04T16:16:50.580454Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:16:50.580454Z digest=sha256:a0498e4cdaa6fda4f780c426bd912268fb02ee7e5f9514ae743c0021b1c4b6a3

Observation 0df9cd09-f311-4e22-8d2e-b4f62b846751 · inbound

RSVideo: Are Your Vision-Language Models Ready for Remote Sensing Videos? cites this paper.

RSVideo: Are Your Vision-Language Models Ready for Remote Sensing Videos? UniRS: Unifying Multi-temporal Remote Sensing Tasks through Vision Language Models

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-08-07T00:14:11.514686Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T00:14:05.499836Z digest=sha256:31413bdfab4b27e6e671abc78478786c57860f476fe261c1f01c4d337fa05df7