Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-10T23:18:21.438754Z
Paper Citation Record · LEDGER
As of 14 August 2026, this Paper Citation Record lists 56 of 56 outbound references and 2 inbound Pith citation observations for arXiv:2412.20742.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-10T23:18:21.438754Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T00:14:05.499836Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-07T00:14:11.438315Z
56 of 56 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation cfdd125b-a09c-4491-883a-9cbe2377ad50 · outbound
UniRS: Unifying Multi-temporal Remote Sensing Tasks through Vision Language Models Visual instruction tuning,
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation ce25b4aa-3cbd-43d6-8ce7-7b336e7685b2 · outbound
UniRS: Unifying Multi-temporal Remote Sensing Tasks through Vision Language Models Flamingo: a visual language model for few-shot learning,
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 530008c7-1235-4052-950f-d8ec06735c83 · outbound
UniRS: Unifying Multi-temporal Remote Sensing Tasks through Vision Language Models Vila: On pre-training for visual language models,
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation fbe617b2-21fb-44de-8a92-073732e6ca3c · outbound
UniRS: Unifying Multi-temporal Remote Sensing Tasks through Vision Language Models MiniGPT-v2: large language model as a unified interface for vision-language multi-task learning
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 06d2b9f4-d435-4e2c-9e8b-d09ecacff320 · outbound
UniRS: Unifying Multi-temporal Remote Sensing Tasks through Vision Language Models Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation,
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 799829fc-f48b-40d5-9955-a60296e47ea9 · outbound
UniRS: Unifying Multi-temporal Remote Sensing Tasks through Vision Language Models Blip-2: Bootstrapping language- image pre-training with frozen image encoders and large language models,
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 21c7731b-5099-4296-a867-84ceff71c929 · outbound
UniRS: Unifying Multi-temporal Remote Sensing Tasks through Vision Language Models LLaMA: Open and Efficient Foundation Language Models
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation df83d8ad-0e80-44f5-a8c3-905779f10e98 · outbound
UniRS: Unifying Multi-temporal Remote Sensing Tasks through Vision Language Models Training language models to follow instructions with human feedback,
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation c88ce50f-906b-4c8f-8b5d-37fc8e4da3d3 · outbound
UniRS: Unifying Multi-temporal Remote Sensing Tasks through Vision Language Models Qwen Technical Report
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5b31fda4-bd17-4c4d-9c1f-41a87607715a · outbound
UniRS: Unifying Multi-temporal Remote Sensing Tasks through Vision Language Models Language mod- els are few-shot learners,
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 09aa9e4d-7db1-406f-b233-051a3d217e69 · outbound
UniRS: Unifying Multi-temporal Remote Sensing Tasks through Vision Language Models Laion- 5b: An open large-scale dataset for training next generation image-text models,
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation f3555b1b-c8ba-41f1-9012-807276725dd0 · outbound
UniRS: Unifying Multi-temporal Remote Sensing Tasks through Vision Language Models Llava-med: Training a large language-and-vision assistant for biomedicine in one day,
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation f0d34255-e33e-49db-837d-38721ecf3295 · outbound
UniRS: Unifying Multi-temporal Remote Sensing Tasks through Vision Language Models Mobility VLA: Multimodal Instruction Navigation with Long-Context VLMs and Topological Graphs
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4632d040-0eac-4350-aef7-8a6f7e96d63c · outbound
UniRS: Unifying Multi-temporal Remote Sensing Tasks through Vision Language Models Rsvqa: Visual question answering for remote sensing data,
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation f0c8b997-0e45-4fcd-b1dd-eba28ef142ce · outbound
UniRS: Unifying Multi-temporal Remote Sensing Tasks through Vision Language Models Answer-type prediction for visual question answering,
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation ea401cdc-6899-4110-a644-8fcca327b6c4 · outbound
UniRS: Unifying Multi-temporal Remote Sensing Tasks through Vision Language Models Multi-step question-driven visual ques- tion answering for remote sensing,
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 7652e053-07a3-41f4-8123-ab8624cd9906 · outbound
UniRS: Unifying Multi-temporal Remote Sensing Tasks through Vision Language Models Bi-modal transformer-based approach for visual question answering in remote sensing imagery,
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation d79f23aa-1b6c-4153-b41b-a9432905d186 · outbound
UniRS: Unifying Multi-temporal Remote Sensing Tasks through Vision Language Models Changes to captions: An attentive network for remote sensing change captioning,
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 9ddb6af0-974c-41dc-a04d-b6a1edf19bfc · outbound
UniRS: Unifying Multi-temporal Remote Sensing Tasks through Vision Language Models Progressive scale-aware network for remote sensing image change captioning,
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 12399aa5-b068-4bfb-beb3-a61dc53e7f88 · outbound
UniRS: Unifying Multi-temporal Remote Sensing Tasks through Vision Language Models Futh-net: fusing temporal relations and holistic features for aerial video classification,
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 9362b406-f560-4720-b98e-23397670ffe5 · outbound
UniRS: Unifying Multi-temporal Remote Sensing Tasks through Vision Language Models Ok-vqa: A visual question answering benchmark requiring external knowledge,
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 76fd99e3-cf0a-4563-a1de-b0420b434339 · outbound
UniRS: Unifying Multi-temporal Remote Sensing Tasks through Vision Language Models Deep learning based event recognition in aerial imagery,
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 4ad2a81b-9a8d-4704-904f-80df22718a0b · outbound
UniRS: Unifying Multi-temporal Remote Sensing Tasks through Vision Language Models A decoupling paradigm with prompt learning for remote sensing image change cap- tioning,
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 1892636e-0c9b-4294-a9e2-aee3f275b794 · outbound
UniRS: Unifying Multi-temporal Remote Sensing Tasks through Vision Language Models LHRS-Bot: Empowering Remote Sensing with VGI-Enhanced Large Multimodal Language Model
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 806d599f-8e76-4721-a274-ab799ca5cd0e · outbound
UniRS: Unifying Multi-temporal Remote Sensing Tasks through Vision Language Models Geochat: Grounded large vision-language model for remote sensing,
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 1a1863ba-bd5f-4407-a0d8-7a0c620301cd · outbound
UniRS: Unifying Multi-temporal Remote Sensing Tasks through Vision Language Models SkyEyeGPT: Unifying Remote Sensing Vision-Language Tasks via Instruction Tuning with Large Language Model
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3e974fef-266d-4b36-9be6-62a27b8b8c7d · outbound
UniRS: Unifying Multi-temporal Remote Sensing Tasks through Vision Language Models TEOChat: A Large Vision-Language Assistant for Temporal Earth Observation Data
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4bab5494-172b-4a59-b513-4cb577f94817 · outbound
UniRS: Unifying Multi-temporal Remote Sensing Tasks through Vision Language Models Rs- llava: A large vision-language model for joint captioning and question answering in remote sensing imagery,
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 0221ae97-370c-4d17-9822-1dfbd82cf01e · outbound
UniRS: Unifying Multi-temporal Remote Sensing Tasks through Vision Language Models RSGPT: A Remote Sensing Vision Language Model and Benchmark
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1362db8c-f70c-4af0-a999-173bbdf1a48f · outbound
UniRS: Unifying Multi-temporal Remote Sensing Tasks through Vision Language Models Earthgpt: A universal multi-modal large language model for multi-sensor image comprehension in remote sensing domain,
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 07501f7c-3e75-40cb-878c-73b447bb457e · outbound
UniRS: Unifying Multi-temporal Remote Sensing Tasks through Vision Language Models Remote sensing image change captioning with dual-branch transformers: A new method and a large scale dataset,
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 7f315012-a026-4801-812b-b468fbf60e3c · outbound
UniRS: Unifying Multi-temporal Remote Sensing Tasks through Vision Language Models Era: A data set and deep learning benchmark for event recognition in aerial videos [software and data sets],
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 453961bf-5b70-47b7-8e70-968a953bf34f · outbound
UniRS: Unifying Multi-temporal Remote Sensing Tasks through Vision Language Models GPT-4 Technical Report
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ad6d56c1-74cf-41cb-bd94-3a6849e34b0f · outbound
UniRS: Unifying Multi-temporal Remote Sensing Tasks through Vision Language Models Zero-shot video moment retrieval from frozen vision-language models,
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 9c85abab-b506-4dbb-ac5f-f504453a2c63 · outbound
UniRS: Unifying Multi-temporal Remote Sensing Tasks through Vision Language Models Visual narratives: Large-scale hi- erarchical classification of art-historical images,
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 184674d7-b23c-4a64-ad60-e2dc6ce7cb30 · outbound
UniRS: Unifying Multi-temporal Remote Sensing Tasks through Vision Language Models Learning transferable visual models from natural language supervision,
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation ea49b646-97eb-4168-8986-1375fd0d1da5 · outbound
UniRS: Unifying Multi-temporal Remote Sensing Tasks through Vision Language Models Sigmoid loss for language image pre-training,
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation fddb3fee-08e5-48ad-b128-9367fe93cd6c · outbound
UniRS: Unifying Multi-temporal Remote Sensing Tasks through Vision Language Models Improved baselines with visual instruction tuning,
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation f3867446-ec3d-4bc5-b5ed-6bd479515b0e · outbound
UniRS: Unifying Multi-temporal Remote Sensing Tasks through Vision Language Models SkySenseGPT: A Fine-Grained Instruction Tuning Dataset and Model for Remote Sensing Vision-Language Understanding
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 210f211f-57c8-45f8-82e3-8aea311fc432 · outbound
UniRS: Unifying Multi-temporal Remote Sensing Tasks through Vision Language Models Nwpu- captions dataset and mlca-net for remote sensing image captioning,
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation f960d65f-2fc5-4c05-b7da-203a8ee61fcd · outbound
UniRS: Unifying Multi-temporal Remote Sensing Tasks through Vision Language Models Multiscale spatio- temporal network for aerial video event recognition,
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 884a5ea8-b3bf-4a5c-9ed4-4c0d35801b68 · outbound
UniRS: Unifying Multi-temporal Remote Sensing Tasks through Vision Language Models Alleviating spatial misalignment and motion interference for uav-based video recognition,
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation d9ef3b67-f80d-40f8-a649-9074db3cdc69 · outbound
UniRS: Unifying Multi-temporal Remote Sensing Tasks through Vision Language Models Multitask learning: A knowledge-based source of inductive bias,
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation d7233d7c-ecbd-49bf-992d-7220d1f77ccf · outbound
UniRS: Unifying Multi-temporal Remote Sensing Tasks through Vision Language Models Multitask learning,
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 73f794b3-9073-475d-a394-c26c4b5a34de · outbound
UniRS: Unifying Multi-temporal Remote Sensing Tasks through Vision Language Models Multitask learning for crash analysis: A fine-tuned llm framework using twitter data,
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 2199e715-5976-4748-bfc8-cc11eb5bce13 · outbound
UniRS: Unifying Multi-temporal Remote Sensing Tasks through Vision Language Models TigerBot: An Open Multilingual Multitask LLM
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ce2dd9f4-c4d0-4e63-9ac2-b06eb3a7eb98 · outbound
UniRS: Unifying Multi-temporal Remote Sensing Tasks through Vision Language Models Mftcoder: Boosting code llms with multitask fine-tuning,
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation e7289dbe-7768-43e2-a089-033eeb79ff8e · outbound
UniRS: Unifying Multi-temporal Remote Sensing Tasks through Vision Language Models Sheared LLaMA: Accelerating Language Model Pre-training via Structured Pruning
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bb21a836-9d5c-4f3c-9b37-2b0d49b7ab2c · outbound
UniRS: Unifying Multi-temporal Remote Sensing Tasks through Vision Language Models Dota: A large-scale dataset for object detection in aerial images,
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation cd9dd48d-bcf2-4dcb-a9ee-dc8d7bbafe6c · outbound
UniRS: Unifying Multi-temporal Remote Sensing Tasks through Vision Language Models Anchor- free oriented proposal generator for object detection,
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation a556e187-3f3e-4375-bf3f-6d507618b135 · outbound
UniRS: Unifying Multi-temporal Remote Sensing Tasks through Vision Language Models Fair1m: A benchmark dataset for fine-grained object recognition in high-resolution remote sensing imagery,
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 0baf7aa0-97f2-4948-9f25-6ba66454b5b6 · outbound
UniRS: Unifying Multi-temporal Remote Sensing Tasks through Vision Language Models Remote sensing image scene classifi- cation: Benchmark and state of the art,
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation c013b11a-333a-453e-9990-53ae176d121d · outbound
UniRS: Unifying Multi-temporal Remote Sensing Tasks through Vision Language Models Floodnet: A high resolution aerial imagery dataset for post flood scene understanding,
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 16960b0b-0ce9-47a7-84bf-a5a805e7ea54 · outbound
UniRS: Unifying Multi-temporal Remote Sensing Tasks through Vision Language Models A spatial-temporal attention-based method and a new dataset for remote sensing image change detection,
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation d3fff566-c8d6-4717-ba45-78a27bdbc18a · outbound
UniRS: Unifying Multi-temporal Remote Sensing Tasks through Vision Language Models A spatial hierarchical reasoning network for remote sensing visual question answering,
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation f4ccce72-88d5-4d46-9aa5-cc0389180fae · outbound
UniRS: Unifying Multi-temporal Remote Sensing Tasks through Vision Language Models Temporal rela- tions matter: A two-pathway network for aerial video recognition,
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 7285c014-15ee-4005-8ea9-9c153e94f6a6 · inbound
RSVideo: Are Your Vision-Language Models Ready for Remote Sensing Videos? UniRS: Unifying Multi-temporal Remote Sensing Tasks through Vision Language Models
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0df9cd09-f311-4e22-8d2e-b4f62b846751 · inbound
RSVideo: Are Your Vision-Language Models Ready for Remote Sensing Videos? UniRS: Unifying Multi-temporal Remote Sensing Tasks through Vision Language Models
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.