Pith. sign in

Paper Citation Record · LEDGER

SkyVLaM: Multimodal Large Language Model for UAV Video Understanding in Remote Sensing

As of 8 August 2026, this Paper Citation Record lists 51 of 51 outbound references and 0 inbound Pith citation observations for arXiv:2607.17386.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.17386 v1

Coverage vector

measured 51 of 51 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-01T18:11:04.192167Z

measured 51 of 51 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

51 of 51 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved51
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 7c911503-5255-4be9-9c32-4a06b43b0c4a · outbound

This paper cites One token to seg them all: Language instructed reasoning seg- mentation in videos.Advances in Neural Information Processing Systems, 37:6833–6859, 2024.

SkyVLaM: Multimodal Large Language Model for UAV Video Understanding in Remote Sensing One token to seg them all: Language instructed reasoning seg- mentation in videos.Advances in Neural Information Processing Systems, 37:6833–6859, 2024

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-01T18:11:00.691639Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:11:00.691639Z digest=sha256:3c086ff47e94783b4088d15cb1d3fc03d332d74624f1cf570a62cb306ba4d1ac

Observation e726ec18-5a18-4e01-a66d-acf19c22c573 · outbound

This paper cites Meteor: An automatic metric for mt evaluation with improved correlation with human judgments.

SkyVLaM: Multimodal Large Language Model for UAV Video Understanding in Remote Sensing Meteor: An automatic metric for mt evaluation with improved correlation with human judgments

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-01T18:11:00.744246Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:11:00.744246Z digest=sha256:8aacb2edc016587de544a4f511da6952e9a0100b0ddcdda3c5e172151e286d62

Observation bc7c76ca-1343-4e91-9975-31bbadbdd027 · outbound

This paper cites End-to-end referring video object segmentation with multimodal transformers.

SkyVLaM: Multimodal Large Language Model for UAV Video Understanding in Remote Sensing End-to-end referring video object segmentation with multimodal transformers

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-01T18:11:00.790110Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:11:00.790110Z digest=sha256:6011a4a2f683f7268716ea85b10780d985c964fa0491090271377491abcc355e

Observation ebeeb569-da85-4acb-8d79-5b96c294852e · outbound

This paper cites Streamingtom: Streaming token compres- sion for efficient video understanding.arXiv preprint arXiv:2510.18269, 2025.

SkyVLaM: Multimodal Large Language Model for UAV Video Understanding in Remote Sensing Streamingtom: Streaming token compres- sion for efficient video understanding.arXiv preprint arXiv:2510.18269, 2025

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-01T18:11:00.854409Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:11:00.854409Z digest=sha256:3c7024625d1a0e14a88cb6b0a461feaffdd4030991b75ca9b6b958dcbf120aa3

Observation aac304ac-f50d-4d06-89f3-224d7201ef3f · outbound

This paper cites Mevis: A large-scale benchmark for video segmentation with motion expressions.

SkyVLaM: Multimodal Large Language Model for UAV Video Understanding in Remote Sensing Mevis: A large-scale benchmark for video segmentation with motion expressions

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-01T18:11:00.919449Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:11:00.919449Z digest=sha256:f461312482744612d75b40db44b7f4135f0f6d426f9b7935e0e1cd1990ac9d01

Observation f0495812-3451-4c19-85ed-ad679685c4c3 · outbound

This paper cites The unmanned aerial vehicle benchmark: Object detection and tracking.

SkyVLaM: Multimodal Large Language Model for UAV Video Understanding in Remote Sensing The unmanned aerial vehicle benchmark: Object detection and tracking

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-01T18:11:00.984738Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:11:00.984738Z digest=sha256:dfba8ac5b4fc13a680c70c31bfb701a3b40cb28a2e554a588d70d6bfb37acb7a

Observation f6ace37c-9484-43c0-8418-585971a9f62c · outbound

This paper cites Framefusion: Combining similarity and importance for video token reduction on large vision language models.

SkyVLaM: Multimodal Large Language Model for UAV Video Understanding in Remote Sensing Framefusion: Combining similarity and importance for video token reduction on large vision language models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-01T18:11:01.082341Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:11:01.082341Z digest=sha256:47711c5493f0ba6257e87a56b0d06bf186b4c8769c128de273ea34ebc58e100b

Observation e7486384-4444-47a1-b64c-6ab04edcd40b · outbound

This paper cites Chain-of-Frames: Advancing Video Understanding in Multimodal LLMs via Frame-Aware Reasoning.

SkyVLaM: Multimodal Large Language Model for UAV Video Understanding in Remote Sensing Chain-of-Frames: Advancing Video Understanding in Multimodal LLMs via Frame-Aware Reasoning

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-01T18:11:01.204751Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:11:01.204751Z digest=sha256:9c4b6b2883b704e1eb1dd7d38a0fce371945743cce5fc96920554d42aa06aae7

Observation 954877f9-8f77-44ff-bdd2-48c34a27fa1d · outbound

This paper cites The devil is in temporal token: High quality video reasoning segmentation.

SkyVLaM: Multimodal Large Language Model for UAV Video Understanding in Remote Sensing The devil is in temporal token: High quality video reasoning segmentation

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-01T18:11:01.296336Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:11:01.296336Z digest=sha256:f87e156e458c786ae078ef818f17e722db910341ef89fd145330f868b7baf605

Observation c2dae05a-a68c-4d7b-9543-b872f27f70de · outbound

This paper cites Rsgpt: A remote sensing vision language model and benchmark.ISPRS Journal of Photogrammetry and Remote Sensing, 224:272–286, 2025.

SkyVLaM: Multimodal Large Language Model for UAV Video Understanding in Remote Sensing Rsgpt: A remote sensing vision language model and benchmark.ISPRS Journal of Photogrammetry and Remote Sensing, 224:272–286, 2025

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-01T18:11:01.367418Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:11:01.367418Z digest=sha256:283d7ed460c443179551a227b4fb501e69695dc11e895164d4f3ff3730ec770c

Observation 69ff8e28-3474-4e55-8196-1d2bf37d04b2 · outbound

This paper cites Prunevid: Visual token pruning for efficient video large language models.

SkyVLaM: Multimodal Large Language Model for UAV Video Understanding in Remote Sensing Prunevid: Visual token pruning for efficient video large language models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-01T18:11:01.457054Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:11:01.457054Z digest=sha256:f3b90c871adc1647dcc732f27b1002274d6937ce229d7f1a52d021df16e262f7

Observation dfecb740-b81a-49cf-a800-cfeac3bca07f · outbound

This paper cites Multi-granular spatio-temporal token merging for training-free acceleration of video llms.

SkyVLaM: Multimodal Large Language Model for UAV Video Understanding in Remote Sensing Multi-granular spatio-temporal token merging for training-free acceleration of video llms

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-01T18:11:01.521166Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:11:01.521166Z digest=sha256:36a600a2b513349850889c7b29468e2d809c12535d815aa48661f22a935f3aba

Observation c57608d3-d131-4995-8b1c-eff11246edac · outbound

This paper cites TEOChat: A Large Vision-Language Assistant for Temporal Earth Observation Data.

SkyVLaM: Multimodal Large Language Model for UAV Video Understanding in Remote Sensing TEOChat: A Large Vision-Language Assistant for Temporal Earth Observation Data

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-01T18:11:01.580630Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:11:01.580630Z digest=sha256:f183bce9290e7f16c562a05ce9e270d6c4d434a9cfc476f054ebb324032fb042

Observation a485416f-d454-4b76-9ef5-7fc4ede0b8cf · outbound

This paper cites Lisa: Reasoning segmentation via large language model.

SkyVLaM: Multimodal Large Language Model for UAV Video Understanding in Remote Sensing Lisa: Reasoning segmentation via large language model

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-01T18:11:01.696273Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:11:01.696273Z digest=sha256:223ee89b0f460dcd3af36ecae52d2f98981dfefdbe65ec1eea557b275fa725e5

Observation 83cd5b5c-dd0b-47d7-a5e5-ee9ef889736f · outbound

This paper cites Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models.

SkyVLaM: Multimodal Large Language Model for UAV Video Understanding in Remote Sensing Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-01T18:11:01.761048Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:11:01.761048Z digest=sha256:af5a6041544b4de3a32717d796b040e489fb7581c5526f2bdc2b378a5969fe77

Observation 6731052e-dbfa-41ad-814f-7b040ed21374 · outbound

This paper cites Referdino: Referring video object segmentation with visual grounding foundations.

SkyVLaM: Multimodal Large Language Model for UAV Video Understanding in Remote Sensing Referdino: Referring video object segmentation with visual grounding foundations

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-01T18:11:01.849069Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:11:01.849069Z digest=sha256:4b362cda9b93a949bdde119a5d83b1290ed73bb32138036d4f104abf393fb7ad

Observation 4f45cead-57fa-4193-934d-3aad75ef807a · outbound

This paper cites Video-llava: Learning united visual representation by alignment before projection.

SkyVLaM: Multimodal Large Language Model for UAV Video Understanding in Remote Sensing Video-llava: Learning united visual representation by alignment before projection

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-01T18:11:01.971650Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:11:01.971650Z digest=sha256:ef9d451637145b9802a4540b320bb09cc35e297418e8ae1f3902d38732e602d7

Observation 27c684a3-7358-48c6-886a-e007211e6294 · outbound

This paper cites Glus: Global-local reasoning unified into a single large language model for video segmentation.

SkyVLaM: Multimodal Large Language Model for UAV Video Understanding in Remote Sensing Glus: Global-local reasoning unified into a single large language model for video segmentation

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-01T18:11:02.034375Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:11:02.034375Z digest=sha256:bb0c705c243ecd7993395094774fccd524c45d1d322b6000d0a1be3bed2cca86

Observation 1e1e3f91-2614-4465-aaeb-bde5dc3df7a4 · outbound

This paper cites Visual instruction tuning.Advances in neural information processing systems, 36:34892–34916, 2023.

SkyVLaM: Multimodal Large Language Model for UAV Video Understanding in Remote Sensing Visual instruction tuning.Advances in neural information processing systems, 36:34892–34916, 2023

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-01T18:11:02.108301Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:11:02.108301Z digest=sha256:7ba0baf1b857d5c5d7df461aec833c98e0882136e4a201f5643840c04fd1a24b

Observation bf66e690-c950-4d14-b84c-12030a074936 · outbound

This paper cites RSUniVLM: A Unified Vision Language Model for Remote Sensing via Granularity-oriented Mixture of Experts.

SkyVLaM: Multimodal Large Language Model for UAV Video Understanding in Remote Sensing RSUniVLM: A Unified Vision Language Model for Remote Sensing via Granularity-oriented Mixture of Experts

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-01T18:11:02.265647Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:11:02.265647Z digest=sha256:d6e009c1996a126d5a4f1d7dbe86692201464d62b9ff4a40239269b0cddc2f6f

Observation 7da24c0e-c78c-49c0-915a-672bacdb4bf6 · outbound

This paper cites SkySenseGPT: A Fine-Grained Instruction Tuning Dataset and Model for Remote Sensing Vision-Language Understanding.

SkyVLaM: Multimodal Large Language Model for UAV Video Understanding in Remote Sensing SkySenseGPT: A Fine-Grained Instruction Tuning Dataset and Model for Remote Sensing Vision-Language Understanding

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-01T18:11:02.304478Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:11:02.304478Z digest=sha256:d643e7ab2aa4861f2e0c73e1ec395dad78242077c5e78095f2bb8fdbcf1de4bf

Observation 0ebe0656-7020-49d2-88e4-e00b58c04323 · outbound

This paper cites Video-chatgpt: Towards detailed video understanding via large vision and language models.

SkyVLaM: Multimodal Large Language Model for UAV Video Understanding in Remote Sensing Video-chatgpt: Towards detailed video understanding via large vision and language models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-01T18:11:02.369009Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:11:02.369009Z digest=sha256:9eae14242a344e9f3307edb041631df71b6a99a5b8880b1cf2dde58d01bcc5ec

Observation b7b0d143-8b00-4c13-8bfd-e131ad13230a · outbound

This paper cites Videoglamm: A large multimodal model for pixel-level visual grounding in videos.

SkyVLaM: Multimodal Large Language Model for UAV Video Understanding in Remote Sensing Videoglamm: A large multimodal model for pixel-level visual grounding in videos

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-01T18:11:02.428042Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:11:02.428042Z digest=sha256:822873c08f7bc49dc8f82f8f03d43d1003c8751b283c70f90e4c7264381a6fd5

Observation 113d5682-d67d-40cb-b88c-a2e3a48dad5a · outbound

This paper cites Geopix: A multimodal large language model for pixel-level image understanding in remote sensing.IEEE Geoscience and Remote Sensing Magazine, 2025.

SkyVLaM: Multimodal Large Language Model for UAV Video Understanding in Remote Sensing Geopix: A multimodal large language model for pixel-level image understanding in remote sensing.IEEE Geoscience and Remote Sensing Magazine, 2025

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-01T18:11:02.501383Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:11:02.501383Z digest=sha256:dba0b14a10e9c16e0c20db823efecf44040d9d3e3962ba750d262f6bf4deb39e

Observation 65940e87-c7ca-4c12-b52c-4d4054028358 · outbound

This paper cites Vhm: Versatile and honest vision language model for remote sensing image analysis.

SkyVLaM: Multimodal Large Language Model for UAV Video Understanding in Remote Sensing Vhm: Versatile and honest vision language model for remote sensing image analysis

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-01T18:11:02.552720Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:11:02.552720Z digest=sha256:14103f058a3876cbe3a3988f2f36dc4d59436f716a9ee68f906117906dfbcb4e

Observation 9ddf7593-9019-48ce-b84c-fa5f2f6664a9 · outbound

This paper cites Llava++: extending visual capabilities with llama-3 and phi-3 (2024).URL https://github.com/mbzuai-oryx/LLaVA- pp, 2024.

SkyVLaM: Multimodal Large Language Model for UAV Video Understanding in Remote Sensing Llava++: extending visual capabilities with llama-3 and phi-3 (2024).URL https://github.com/mbzuai-oryx/LLaVA- pp, 2024

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-01T18:11:02.593053Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:11:02.593053Z digest=sha256:6f32dd4e59ecbc959716571942e5f53d3f85de5b6621953fb616126904d3e2eb

Observation 086ae9bd-1812-4e22-acf0-874e4e16621b · outbound

This paper cites Glamm: Pixel grounding large multimodal model.

SkyVLaM: Multimodal Large Language Model for UAV Video Understanding in Remote Sensing Glamm: Pixel grounding large multimodal model

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-01T18:11:02.644608Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:11:02.644608Z digest=sha256:4797a946f771a83ebc50d8aa86786e4e15b55f636f383f241c06ffe5fc10268e

Observation fb476229-664f-4803-89e2-2fd277f44817 · outbound

This paper cites SAM 2: Segment Anything in Images and Videos.

SkyVLaM: Multimodal Large Language Model for UAV Video Understanding in Remote Sensing SAM 2: Segment Anything in Images and Videos

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-01T18:11:02.721970Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:11:02.721970Z digest=sha256:c0043e969e0ec4b93403a7c814658326bac6ebf1615de672b1c10fd199047a82

Observation eec54fbe-cb03-45c0-b031-d8fcb329e4d4 · outbound

This paper cites Pixellm: Pixel reasoning with large multimodal model.

SkyVLaM: Multimodal Large Language Model for UAV Video Understanding in Remote Sensing Pixellm: Pixel reasoning with large multimodal model

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-01T18:11:02.803839Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:11:02.803839Z digest=sha256:622ff1c16614fb26f8be3a81156119bfb46b01889f1aa02485035880fd1bff1f

Observation 867c4e28-917a-4a41-93be-31ea3459c95a · outbound

This paper cites Moviechat: From dense token to sparse memory for long video understanding.

SkyVLaM: Multimodal Large Language Model for UAV Video Understanding in Remote Sensing Moviechat: From dense token to sparse memory for long video understanding

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-01T18:11:02.860093Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:11:02.860093Z digest=sha256:91dac5efdb1025dc7ab7832245293db761f7463b0677d94cc0a0cb8012f06fe6

Observation d83d7eb4-6c63-4448-94a0-952dadbdb274 · outbound

This paper cites Earthdial: Turning multi-sensory earth observations to interactive dialogues.

SkyVLaM: Multimodal Large Language Model for UAV Video Understanding in Remote Sensing Earthdial: Turning multi-sensory earth observations to interactive dialogues

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-01T18:11:02.915593Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:11:02.915593Z digest=sha256:406225b92ae4bccff34023bb24adc744f559e22ee0586fccec76652dd339050f

Observation 52daab6d-b37e-48b3-9d2e-d8222f199e69 · outbound

This paper cites Drone-based rgb-infrared cross-modality vehicle detection via uncertainty-aware learning.IEEE Transactions on Circuits and Systems for Video Technology, 32(10):6700–6713, 2022.

SkyVLaM: Multimodal Large Language Model for UAV Video Understanding in Remote Sensing Drone-based rgb-infrared cross-modality vehicle detection via uncertainty-aware learning.IEEE Transactions on Circuits and Systems for Video Technology, 32(10):6700–6713, 2022

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-01T18:11:02.973608Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:11:02.973608Z digest=sha256:398f2595b31b58b8ee04f4ffed3f068709a6b4e4729137146a85e6292d14c523

Observation 17863fb7-2119-4768-8b96-53666e87d913 · outbound

This paper cites Adaptive keyframe sampling for long video understanding.

SkyVLaM: Multimodal Large Language Model for UAV Video Understanding in Remote Sensing Adaptive keyframe sampling for long video understanding

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-01T18:11:03.030943Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:11:03.030943Z digest=sha256:c540e8332d4ea3d9141ce39ca35bc91c9f76f9c386ca64cb5da5857e1de51422

Observation 4e899a43-c9ad-4aeb-9ac4-0626688eb142 · outbound

This paper cites Cider: Consensus-based image description evaluation.

SkyVLaM: Multimodal Large Language Model for UAV Video Understanding in Remote Sensing Cider: Consensus-based image description evaluation

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-01T18:11:03.074265Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:11:03.074265Z digest=sha256:7142297036833892bf5b2f61287bc34dfd4b387f0c3435f23404aabff7e84d5f

Observation a24a7245-aef3-4bb0-aa6b-10394c3e1801 · outbound

This paper cites Geollava-8k: scaling remote-sensing multimodal large language models to 8k resolution.arXiv preprint arXiv:2505.21375, 2025.

SkyVLaM: Multimodal Large Language Model for UAV Video Understanding in Remote Sensing Geollava-8k: scaling remote-sensing multimodal large language models to 8k resolution.arXiv preprint arXiv:2505.21375, 2025

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-01T18:11:03.124523Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:11:03.124523Z digest=sha256:e59401e5fdd6feeeb7dd7fa760eff4f71f8c344d27c7980b628fc98c86a9aa5a

Observation 2dd4ab56-621a-491c-bbf2-07dda40d7475 · outbound

This paper cites Instructseg: Unifying instructed visual segmentation with multi-modal large language models.

SkyVLaM: Multimodal Large Language Model for UAV Video Understanding in Remote Sensing Instructseg: Unifying instructed visual segmentation with multi-modal large language models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-01T18:11:03.221579Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:11:03.221579Z digest=sha256:83727057dd748377d7f85e05d58cdba97bdca8329b59d0c56514a570434806a2

Observation aa94ce04-88b5-493b-9d0e-b29f1600024d · outbound

This paper cites Longvlm: Efficient long video understanding via large language models.

SkyVLaM: Multimodal Large Language Model for UAV Video Understanding in Remote Sensing Longvlm: Efficient long video understanding via large language models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-01T18:11:03.350493Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:11:03.350493Z digest=sha256:4e4c70f862d1cc3b096100fc79374e624784d099d63a69a1da425361ea78e0d0

Observation 1a711df3-dda3-4d74-9bfd-c7b3a4daab95 · outbound

This paper cites Language as queries for referring video object segmentation.

SkyVLaM: Multimodal Large Language Model for UAV Video Understanding in Remote Sensing Language as queries for referring video object segmentation

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-01T18:11:03.422889Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:11:03.422889Z digest=sha256:d29ce238a956a13e1f450e86a7efb2d5c31b7a19361b0914f248335e19056987

Observation 9ce06e3c-e938-4b4b-ab75-acf4c575595c · outbound

This paper cites SlowFast-LLaVA: A Strong Training-Free Baseline for Video Large Language Models.

SkyVLaM: Multimodal Large Language Model for UAV Video Understanding in Remote Sensing SlowFast-LLaVA: A Strong Training-Free Baseline for Video Large Language Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-01T18:11:03.486399Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:11:03.486399Z digest=sha256:f29e7812f5f6134672673e87807b33e9bd52151501db798b04eff1ffd19cd8ad

Observation 32849a1b-d5f7-4a20-b197-fc29806e3be2 · outbound

This paper cites Visa: Reasoning video object segmentation via large language models.

SkyVLaM: Multimodal Large Language Model for UAV Video Understanding in Remote Sensing Visa: Reasoning video object segmentation via large language models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-01T18:11:03.576528Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:11:03.576528Z digest=sha256:d64e5bde71b0418c5abd9d7e6f97dfb96232fef191de650b4dfe4e2b850e71a9

Observation 126070c8-ffb6-4120-bcfd-8e74f9716739 · outbound

This paper cites Referred by multi-modality: A unified temporal transformer for video object segmentation.

SkyVLaM: Multimodal Large Language Model for UAV Video Understanding in Remote Sensing Referred by multi-modality: A unified temporal transformer for video object segmentation

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-01T18:11:03.579616Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:11:03.579616Z digest=sha256:3660eee1374b6ae0e8e32740aaf03f248dffe9086238e3f4a9604be729a0dd1f

Observation f4431f6f-4b71-442c-8f56-39b9b7774506 · outbound

This paper cites Self-chained image-language model for video localization and question answering.Advances in Neural Information Processing Systems, 36:76749–76771, 2023.

SkyVLaM: Multimodal Large Language Model for UAV Video Understanding in Remote Sensing Self-chained image-language model for video localization and question answering.Advances in Neural Information Processing Systems, 36:76749–76771, 2023

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-01T18:11:03.612401Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:11:03.612401Z digest=sha256:0c6bee80c6fb627628c7985ebf05c9db784ea9541f9ca2831900933001f6356f

Observation 4a2ee8db-e65b-4904-8f56-7f723e0d407d · outbound

This paper cites Frame-Voyager: Learning to Query Frames for Video Large Language Models.

SkyVLaM: Multimodal Large Language Model for UAV Video Understanding in Remote Sensing Frame-Voyager: Learning to Query Frames for Video Large Language Models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-01T18:11:03.661242Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:11:03.661242Z digest=sha256:fa6280c1fd7e3db4472ff2a47de1faab985203464162aafa789c3bb9a149aaf0

Observation 4817844c-a271-4236-a712-ef3db908dedc · outbound

This paper cites Sa2VA: Marrying SAM2 with LLaVA for Dense Grounded Understanding of Images and Videos.

SkyVLaM: Multimodal Large Language Model for UAV Video Understanding in Remote Sensing Sa2VA: Marrying SAM2 with LLaVA for Dense Grounded Understanding of Images and Videos

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-01T18:11:03.743258Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:11:03.743258Z digest=sha256:eed1b2659dd325d0c19a2801f173d2cdd3c5ab627d0b687dab9e620114a16703

Observation bb233953-67a4-4e45-9c3c-ed97252e8ef0 · outbound

This paper cites Skyeyegpt: Unifying remote sensing vision- language tasks via instruction tuning with large language model.ISPRS Journal of Photogram- metry and Remote Sensing, 221:64–77, 2025.

SkyVLaM: Multimodal Large Language Model for UAV Video Understanding in Remote Sensing Skyeyegpt: Unifying remote sensing vision- language tasks via instruction tuning with large language model.ISPRS Journal of Photogram- metry and Remote Sensing, 221:64–77, 2025

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-01T18:11:03.814064Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:11:03.814064Z digest=sha256:97a742db39df9de2a0288d558c3cb3495843e48bbc3559a6009353e928081617

Observation bd172da8-7b7c-4156-9dcf-d6053158fda9 · outbound

This paper cites Video-llama: An instruction-tuned audio-visual language model for video understanding.

SkyVLaM: Multimodal Large Language Model for UAV Video Understanding in Remote Sensing Video-llama: An instruction-tuned audio-visual language model for video understanding

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-01T18:11:03.885115Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:11:03.885115Z digest=sha256:ef3e5f322b318d0db05722ba25281fe0e11eb4a4165e27229dc12c94b4e7c7b9

Observation 0cd98b9a-d26e-4ea7-90ad-819a85d924ee · outbound

This paper cites GeoRSMLLM: A Multimodal Large Language Model for Vision-Language Tasks in Geoscience and Remote Sensing.

SkyVLaM: Multimodal Large Language Model for UAV Video Understanding in Remote Sensing GeoRSMLLM: A Multimodal Large Language Model for Vision-Language Tasks in Geoscience and Remote Sensing

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-01T18:11:03.944636Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:11:03.944636Z digest=sha256:df0981c6e1edc82272c64fd850222d1f3e7eb32a683eefbfef87bbe2efe0ee52

Observation a1b883f2-6312-4707-a5a4-da08e0549566 · outbound

This paper cites Tifre: Text-guided video frame reduction for efficient video multi-modal large language models.arXiv preprint arXiv:2602.08861, 2026.

SkyVLaM: Multimodal Large Language Model for UAV Video Understanding in Remote Sensing Tifre: Text-guided video frame reduction for efficient video multi-modal large language models.arXiv preprint arXiv:2602.08861, 2026

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-01T18:11:04.007847Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:11:04.007847Z digest=sha256:0cdeff6688df4b38f4bf908e399e16c17f148ce05818f62ee7fb04d77fcf4ef9

Observation 9860a9dd-5836-408c-b37e-fe9a0a0d5ded · outbound

This paper cites Reason: Reinforced causal search with information bottleneck for video understanding.

SkyVLaM: Multimodal Large Language Model for UAV Video Understanding in Remote Sensing Reason: Reinforced causal search with information bottleneck for video understanding

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-01T18:11:04.079355Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:11:04.079355Z digest=sha256:d9642a1fab67a4d82fb3d44aee8662510122284084b97ec728d184ed3fe03f9a

Observation 890e5b91-6ce5-49a7-9e45-11915eb46598 · outbound

This paper cites Detection and tracking meet drones challenge.IEEE transactions on pattern analysis and machine intelligence, 44(11):7380–7399, 2021.

SkyVLaM: Multimodal Large Language Model for UAV Video Understanding in Remote Sensing Detection and tracking meet drones challenge.IEEE transactions on pattern analysis and machine intelligence, 44(11):7380–7399, 2021

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-01T18:11:04.130946Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:11:04.130946Z digest=sha256:a29bcbe19f725ea571499b517f23890d563c5fd0ad71b16bad32684de4630017

Observation 42c5d8b8-44ec-4295-ac63-564a1914705d · outbound

This paper cites Focus: Efficient keyframe selection for long video understanding.arXiv preprint arXiv:2510.27280, 2025.

SkyVLaM: Multimodal Large Language Model for UAV Video Understanding in Remote Sensing Focus: Efficient keyframe selection for long video understanding.arXiv preprint arXiv:2510.27280, 2025

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-01T18:11:04.192167Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:11:04.192167Z digest=sha256:7cac1100e55121a66a450c877aae81cecf2529f56766996978d3d902634f093f

Pith citing papers

No inbound Pith citation observations are available.