Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-01T18:11:04.192167Z
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 51 of 51 outbound references and 0 inbound Pith citation observations for arXiv:2607.17386.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-01T18:11:04.192167Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
51 of 51 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 7c911503-5255-4be9-9c32-4a06b43b0c4a · outbound
SkyVLaM: Multimodal Large Language Model for UAV Video Understanding in Remote Sensing One token to seg them all: Language instructed reasoning seg- mentation in videos.Advances in Neural Information Processing Systems, 37:6833–6859, 2024
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e726ec18-5a18-4e01-a66d-acf19c22c573 · outbound
SkyVLaM: Multimodal Large Language Model for UAV Video Understanding in Remote Sensing Meteor: An automatic metric for mt evaluation with improved correlation with human judgments
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bc7c76ca-1343-4e91-9975-31bbadbdd027 · outbound
SkyVLaM: Multimodal Large Language Model for UAV Video Understanding in Remote Sensing End-to-end referring video object segmentation with multimodal transformers
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ebeeb569-da85-4acb-8d79-5b96c294852e · outbound
SkyVLaM: Multimodal Large Language Model for UAV Video Understanding in Remote Sensing Streamingtom: Streaming token compres- sion for efficient video understanding.arXiv preprint arXiv:2510.18269, 2025
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aac304ac-f50d-4d06-89f3-224d7201ef3f · outbound
SkyVLaM: Multimodal Large Language Model for UAV Video Understanding in Remote Sensing Mevis: A large-scale benchmark for video segmentation with motion expressions
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f0495812-3451-4c19-85ed-ad679685c4c3 · outbound
SkyVLaM: Multimodal Large Language Model for UAV Video Understanding in Remote Sensing The unmanned aerial vehicle benchmark: Object detection and tracking
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f6ace37c-9484-43c0-8418-585971a9f62c · outbound
SkyVLaM: Multimodal Large Language Model for UAV Video Understanding in Remote Sensing Framefusion: Combining similarity and importance for video token reduction on large vision language models
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e7486384-4444-47a1-b64c-6ab04edcd40b · outbound
SkyVLaM: Multimodal Large Language Model for UAV Video Understanding in Remote Sensing Chain-of-Frames: Advancing Video Understanding in Multimodal LLMs via Frame-Aware Reasoning
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 954877f9-8f77-44ff-bdd2-48c34a27fa1d · outbound
SkyVLaM: Multimodal Large Language Model for UAV Video Understanding in Remote Sensing The devil is in temporal token: High quality video reasoning segmentation
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c2dae05a-a68c-4d7b-9543-b872f27f70de · outbound
SkyVLaM: Multimodal Large Language Model for UAV Video Understanding in Remote Sensing Rsgpt: A remote sensing vision language model and benchmark.ISPRS Journal of Photogrammetry and Remote Sensing, 224:272–286, 2025
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 69ff8e28-3474-4e55-8196-1d2bf37d04b2 · outbound
SkyVLaM: Multimodal Large Language Model for UAV Video Understanding in Remote Sensing Prunevid: Visual token pruning for efficient video large language models
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dfecb740-b81a-49cf-a800-cfeac3bca07f · outbound
SkyVLaM: Multimodal Large Language Model for UAV Video Understanding in Remote Sensing Multi-granular spatio-temporal token merging for training-free acceleration of video llms
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c57608d3-d131-4995-8b1c-eff11246edac · outbound
SkyVLaM: Multimodal Large Language Model for UAV Video Understanding in Remote Sensing TEOChat: A Large Vision-Language Assistant for Temporal Earth Observation Data
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a485416f-d454-4b76-9ef5-7fc4ede0b8cf · outbound
SkyVLaM: Multimodal Large Language Model for UAV Video Understanding in Remote Sensing Lisa: Reasoning segmentation via large language model
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 83cd5b5c-dd0b-47d7-a5e5-ee9ef889736f · outbound
SkyVLaM: Multimodal Large Language Model for UAV Video Understanding in Remote Sensing Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6731052e-dbfa-41ad-814f-7b040ed21374 · outbound
SkyVLaM: Multimodal Large Language Model for UAV Video Understanding in Remote Sensing Referdino: Referring video object segmentation with visual grounding foundations
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4f45cead-57fa-4193-934d-3aad75ef807a · outbound
SkyVLaM: Multimodal Large Language Model for UAV Video Understanding in Remote Sensing Video-llava: Learning united visual representation by alignment before projection
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 27c684a3-7358-48c6-886a-e007211e6294 · outbound
SkyVLaM: Multimodal Large Language Model for UAV Video Understanding in Remote Sensing Glus: Global-local reasoning unified into a single large language model for video segmentation
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1e1e3f91-2614-4465-aaeb-bde5dc3df7a4 · outbound
SkyVLaM: Multimodal Large Language Model for UAV Video Understanding in Remote Sensing Visual instruction tuning.Advances in neural information processing systems, 36:34892–34916, 2023
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bf66e690-c950-4d14-b84c-12030a074936 · outbound
SkyVLaM: Multimodal Large Language Model for UAV Video Understanding in Remote Sensing RSUniVLM: A Unified Vision Language Model for Remote Sensing via Granularity-oriented Mixture of Experts
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7da24c0e-c78c-49c0-915a-672bacdb4bf6 · outbound
SkyVLaM: Multimodal Large Language Model for UAV Video Understanding in Remote Sensing SkySenseGPT: A Fine-Grained Instruction Tuning Dataset and Model for Remote Sensing Vision-Language Understanding
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0ebe0656-7020-49d2-88e4-e00b58c04323 · outbound
SkyVLaM: Multimodal Large Language Model for UAV Video Understanding in Remote Sensing Video-chatgpt: Towards detailed video understanding via large vision and language models
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b7b0d143-8b00-4c13-8bfd-e131ad13230a · outbound
SkyVLaM: Multimodal Large Language Model for UAV Video Understanding in Remote Sensing Videoglamm: A large multimodal model for pixel-level visual grounding in videos
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 113d5682-d67d-40cb-b88c-a2e3a48dad5a · outbound
SkyVLaM: Multimodal Large Language Model for UAV Video Understanding in Remote Sensing Geopix: A multimodal large language model for pixel-level image understanding in remote sensing.IEEE Geoscience and Remote Sensing Magazine, 2025
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 65940e87-c7ca-4c12-b52c-4d4054028358 · outbound
SkyVLaM: Multimodal Large Language Model for UAV Video Understanding in Remote Sensing Vhm: Versatile and honest vision language model for remote sensing image analysis
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9ddf7593-9019-48ce-b84c-fa5f2f6664a9 · outbound
SkyVLaM: Multimodal Large Language Model for UAV Video Understanding in Remote Sensing Llava++: extending visual capabilities with llama-3 and phi-3 (2024).URL https://github.com/mbzuai-oryx/LLaVA- pp, 2024
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 086ae9bd-1812-4e22-acf0-874e4e16621b · outbound
SkyVLaM: Multimodal Large Language Model for UAV Video Understanding in Remote Sensing Glamm: Pixel grounding large multimodal model
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fb476229-664f-4803-89e2-2fd277f44817 · outbound
SkyVLaM: Multimodal Large Language Model for UAV Video Understanding in Remote Sensing SAM 2: Segment Anything in Images and Videos
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eec54fbe-cb03-45c0-b031-d8fcb329e4d4 · outbound
SkyVLaM: Multimodal Large Language Model for UAV Video Understanding in Remote Sensing Pixellm: Pixel reasoning with large multimodal model
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 867c4e28-917a-4a41-93be-31ea3459c95a · outbound
SkyVLaM: Multimodal Large Language Model for UAV Video Understanding in Remote Sensing Moviechat: From dense token to sparse memory for long video understanding
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d83d7eb4-6c63-4448-94a0-952dadbdb274 · outbound
SkyVLaM: Multimodal Large Language Model for UAV Video Understanding in Remote Sensing Earthdial: Turning multi-sensory earth observations to interactive dialogues
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 52daab6d-b37e-48b3-9d2e-d8222f199e69 · outbound
SkyVLaM: Multimodal Large Language Model for UAV Video Understanding in Remote Sensing Drone-based rgb-infrared cross-modality vehicle detection via uncertainty-aware learning.IEEE Transactions on Circuits and Systems for Video Technology, 32(10):6700–6713, 2022
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 17863fb7-2119-4768-8b96-53666e87d913 · outbound
SkyVLaM: Multimodal Large Language Model for UAV Video Understanding in Remote Sensing Adaptive keyframe sampling for long video understanding
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4e899a43-c9ad-4aeb-9ac4-0626688eb142 · outbound
SkyVLaM: Multimodal Large Language Model for UAV Video Understanding in Remote Sensing Cider: Consensus-based image description evaluation
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a24a7245-aef3-4bb0-aa6b-10394c3e1801 · outbound
SkyVLaM: Multimodal Large Language Model for UAV Video Understanding in Remote Sensing Geollava-8k: scaling remote-sensing multimodal large language models to 8k resolution.arXiv preprint arXiv:2505.21375, 2025
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2dd4ab56-621a-491c-bbf2-07dda40d7475 · outbound
SkyVLaM: Multimodal Large Language Model for UAV Video Understanding in Remote Sensing Instructseg: Unifying instructed visual segmentation with multi-modal large language models
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aa94ce04-88b5-493b-9d0e-b29f1600024d · outbound
SkyVLaM: Multimodal Large Language Model for UAV Video Understanding in Remote Sensing Longvlm: Efficient long video understanding via large language models
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1a711df3-dda3-4d74-9bfd-c7b3a4daab95 · outbound
SkyVLaM: Multimodal Large Language Model for UAV Video Understanding in Remote Sensing Language as queries for referring video object segmentation
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9ce06e3c-e938-4b4b-ab75-acf4c575595c · outbound
SkyVLaM: Multimodal Large Language Model for UAV Video Understanding in Remote Sensing SlowFast-LLaVA: A Strong Training-Free Baseline for Video Large Language Models
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 32849a1b-d5f7-4a20-b197-fc29806e3be2 · outbound
SkyVLaM: Multimodal Large Language Model for UAV Video Understanding in Remote Sensing Visa: Reasoning video object segmentation via large language models
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 126070c8-ffb6-4120-bcfd-8e74f9716739 · outbound
SkyVLaM: Multimodal Large Language Model for UAV Video Understanding in Remote Sensing Referred by multi-modality: A unified temporal transformer for video object segmentation
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f4431f6f-4b71-442c-8f56-39b9b7774506 · outbound
SkyVLaM: Multimodal Large Language Model for UAV Video Understanding in Remote Sensing Self-chained image-language model for video localization and question answering.Advances in Neural Information Processing Systems, 36:76749–76771, 2023
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4a2ee8db-e65b-4904-8f56-7f723e0d407d · outbound
SkyVLaM: Multimodal Large Language Model for UAV Video Understanding in Remote Sensing Frame-Voyager: Learning to Query Frames for Video Large Language Models
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4817844c-a271-4236-a712-ef3db908dedc · outbound
SkyVLaM: Multimodal Large Language Model for UAV Video Understanding in Remote Sensing Sa2VA: Marrying SAM2 with LLaVA for Dense Grounded Understanding of Images and Videos
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bb233953-67a4-4e45-9c3c-ed97252e8ef0 · outbound
SkyVLaM: Multimodal Large Language Model for UAV Video Understanding in Remote Sensing Skyeyegpt: Unifying remote sensing vision- language tasks via instruction tuning with large language model.ISPRS Journal of Photogram- metry and Remote Sensing, 221:64–77, 2025
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bd172da8-7b7c-4156-9dcf-d6053158fda9 · outbound
SkyVLaM: Multimodal Large Language Model for UAV Video Understanding in Remote Sensing Video-llama: An instruction-tuned audio-visual language model for video understanding
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0cd98b9a-d26e-4ea7-90ad-819a85d924ee · outbound
SkyVLaM: Multimodal Large Language Model for UAV Video Understanding in Remote Sensing GeoRSMLLM: A Multimodal Large Language Model for Vision-Language Tasks in Geoscience and Remote Sensing
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a1b883f2-6312-4707-a5a4-da08e0549566 · outbound
SkyVLaM: Multimodal Large Language Model for UAV Video Understanding in Remote Sensing Tifre: Text-guided video frame reduction for efficient video multi-modal large language models.arXiv preprint arXiv:2602.08861, 2026
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9860a9dd-5836-408c-b37e-fe9a0a0d5ded · outbound
SkyVLaM: Multimodal Large Language Model for UAV Video Understanding in Remote Sensing Reason: Reinforced causal search with information bottleneck for video understanding
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 890e5b91-6ce5-49a7-9e45-11915eb46598 · outbound
SkyVLaM: Multimodal Large Language Model for UAV Video Understanding in Remote Sensing Detection and tracking meet drones challenge.IEEE transactions on pattern analysis and machine intelligence, 44(11):7380–7399, 2021
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 42c5d8b8-44ec-4295-ac63-564a1914705d · outbound
SkyVLaM: Multimodal Large Language Model for UAV Video Understanding in Remote Sensing Focus: Efficient keyframe selection for long video understanding.arXiv preprint arXiv:2510.27280, 2025
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.