Pith. sign in

Paper Citation Record · LEDGER

Memory-Augmented Multimodal Large Language Models for Small Object Understanding in Streaming Aerial Videos

As of 10 August 2026, this Paper Citation Record lists 47 of 47 outbound references and 0 inbound Pith citation observations for arXiv:2607.19857.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.19857 v1

Coverage vector

measured 47 of 47 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-01T11:33:44.814543Z

measured 47 of 47 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

47 of 47 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved47
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 9b995d4e-2191-4b15-b83f-b8fc2466296d · outbound

This paper cites The small-drone revolution is coming—scientists need to ensure it will be safe,.

Memory-Augmented Multimodal Large Language Models for Small Object Understanding in Streaming Aerial Videos The small-drone revolution is coming—scientists need to ensure it will be safe,

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-01T11:33:41.010000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:33:41.010000Z digest=sha256:b924bd4fcdba467a6e4532565435a831a5327e664397be7c340142f57c335944

Observation 49b07a8c-e707-4b7c-9a55-2c743b824f6b · outbound

This paper cites Champion-level drone racing using deep reinforcement learning,.

Memory-Augmented Multimodal Large Language Models for Small Object Understanding in Streaming Aerial Videos Champion-level drone racing using deep reinforcement learning,

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-01T11:33:41.077507Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:33:41.077507Z digest=sha256:b5dc002083c6550103c2c617dc946fe67f25b6fc8b27abb0c207218858afdf0b

Observation 761e62ab-1b1c-4192-b68f-2dad832a8038 · outbound

This paper cites Video object segmentation without tem- poral information,.

Memory-Augmented Multimodal Large Language Models for Small Object Understanding in Streaming Aerial Videos Video object segmentation without tem- poral information,

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-01T11:33:41.210959Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:33:41.210959Z digest=sha256:b4650adff20d00834b5c87e612760db01e599a72db1a6a50ae44c0fbc6697a16

Observation 5c39917d-b99c-406e-b8f3-61d24b26056c · outbound

This paper cites 3d question answering for city scene understanding,.

Memory-Augmented Multimodal Large Language Models for Small Object Understanding in Streaming Aerial Videos 3d question answering for city scene understanding,

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-01T11:33:41.290989Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:33:41.290989Z digest=sha256:33f349a4cbb90f655fea2a21e9e86a39918a1077320d4ebcb030b5f0503337ea

Observation 43efa3e9-286a-4e1b-be86-a3197c84754d · outbound

This paper cites Phenobench: A large dataset and benchmarks for semantic image interpretation in the agricultural do- main,.

Memory-Augmented Multimodal Large Language Models for Small Object Understanding in Streaming Aerial Videos Phenobench: A large dataset and benchmarks for semantic image interpretation in the agricultural do- main,

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-01T11:33:41.391352Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:33:41.391352Z digest=sha256:65bd27b28dd6408c29845bbb086d102420f30dbe763a73091ca7b7ccc218eb18

Observation 01ecd4ed-e87b-4b06-8dec-7ad48ad22ef2 · outbound

This paper cites Detecting flying objects using a single moving camera,.

Memory-Augmented Multimodal Large Language Models for Small Object Understanding in Streaming Aerial Videos Detecting flying objects using a single moving camera,

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-01T11:33:41.479024Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:33:41.479024Z digest=sha256:e57da6da69b3707a2faa739a5fee7c99dbfadaa2fcc051585b66d00f2c1b696a

Observation 4f4d291c-5418-4571-8780-b59da9a6df49 · outbound

This paper cites Revisiting image-language networks for open-ended phrase detection,.

Memory-Augmented Multimodal Large Language Models for Small Object Understanding in Streaming Aerial Videos Revisiting image-language networks for open-ended phrase detection,

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-01T11:33:41.574070Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:33:41.574070Z digest=sha256:1d0b9d0c874add21a23830e60a69042e3e20054f2c003cf0f688810083c98880

Observation d2b93844-3363-4b7a-aff3-0037efbdcd01 · outbound

This paper cites Mevis: A large- scale benchmark for video segmentation with motion expressions,.

Memory-Augmented Multimodal Large Language Models for Small Object Understanding in Streaming Aerial Videos Mevis: A large- scale benchmark for video segmentation with motion expressions,

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-01T11:33:41.647764Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:33:41.647764Z digest=sha256:167b19a44831165c3b3a2b4e168ba1f93969717f7ea683718a012822d51715d8

Observation d5431214-f946-448e-974c-5886b8e16458 · outbound

This paper cites Lamot: Language- guided multi-object tracking,.

Memory-Augmented Multimodal Large Language Models for Small Object Understanding in Streaming Aerial Videos Lamot: Language- guided multi-object tracking,

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-01T11:33:41.750581Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:33:41.750581Z digest=sha256:d772ff2cc0b1a135e039e563644964c447e8acec40126863166595e31463f05d

Observation a6822768-0774-409a-b3c4-642633a0a9f6 · outbound

This paper cites Skyfind: A large-scale benchmark unveiling referring expres- sion comprehension for uav.

Memory-Augmented Multimodal Large Language Models for Small Object Understanding in Streaming Aerial Videos Skyfind: A large-scale benchmark unveiling referring expres- sion comprehension for uav

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-01T11:33:41.793599Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:33:41.793599Z digest=sha256:b6f42554aa6aea71390a80d9ce324a95cae838ffe539cccc49419507ef1fbaaa

Observation 50c4b8f7-f820-4c1e-96ac-d23aae4cef2d · outbound

This paper cites Aerialmind: Towards referring multi-object tracking in uav scenarios,.

Memory-Augmented Multimodal Large Language Models for Small Object Understanding in Streaming Aerial Videos Aerialmind: Towards referring multi-object tracking in uav scenarios,

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-01T11:33:41.860865Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:33:41.860865Z digest=sha256:a9fee240eb690764b306ba76d6382d9268a0099b33f68d6587bbd8820d690cea

Observation 6da0cf3b-dd9b-4492-acd3-8afc41de44ce · outbound

This paper cites Event-aware instructed assistant for referring video segmentation,.

Memory-Augmented Multimodal Large Language Models for Small Object Understanding in Streaming Aerial Videos Event-aware instructed assistant for referring video segmentation,

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-01T11:33:41.918717Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:33:41.918717Z digest=sha256:416b733d751892ad6cfc2877a16d3fcc0f56923d4ef6fcfbce6258629ffa36b3

Observation cd07ae9b-fc66-4d12-97e5-44f84799e2fb · outbound

This paper cites City-vlm: Towards multidomain perception scene understanding via multimodal incomplete learning,.

Memory-Augmented Multimodal Large Language Models for Small Object Understanding in Streaming Aerial Videos City-vlm: Towards multidomain perception scene understanding via multimodal incomplete learning,

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-01T11:33:42.000557Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:33:42.000557Z digest=sha256:6c033557de1d71f0020f791d6287e9545ce8eb1a5da7de5d1dbe6376d48cae19

Observation 71600a21-e282-4709-8a33-42faab77c5fb · outbound

This paper cites Qwen2.5 technical report,.

Memory-Augmented Multimodal Large Language Models for Small Object Understanding in Streaming Aerial Videos Qwen2.5 technical report,

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-01T11:33:42.085352Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:33:42.085352Z digest=sha256:5155eebdd6069197648c1f738a7e9c842f30ee349e7eb1bfc63636ac973bd174

Observation 366e8549-1cbc-4afa-984b-b330d1f213b2 · outbound

This paper cites SAM 3: Segment Anything with Concepts.

Memory-Augmented Multimodal Large Language Models for Small Object Understanding in Streaming Aerial Videos SAM 3: Segment Anything with Concepts

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-01T11:33:42.259681Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:33:42.259681Z digest=sha256:b9c92ee7cc5d2ab8a89d61dc13216c9f576b844d394ddbe06048850ca06502bc

Observation 305d387f-576c-47d0-99f8-484771a6ffed · outbound

This paper cites React: Streaming video analytics on the edge with asynchronous cloud support,.

Memory-Augmented Multimodal Large Language Models for Small Object Understanding in Streaming Aerial Videos React: Streaming video analytics on the edge with asynchronous cloud support,

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-01T11:33:42.350558Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:33:42.350558Z digest=sha256:2d31836bec4df403fc0c9c612678098e2aad4224b4bb432586d8907ead8ceada

Observation 5b29eab7-6a4d-4271-8344-f237d491cea0 · outbound

This paper cites Multiple-object-tracking algorithm based on dense trajectory voting in aerial videos,.

Memory-Augmented Multimodal Large Language Models for Small Object Understanding in Streaming Aerial Videos Multiple-object-tracking algorithm based on dense trajectory voting in aerial videos,

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-01T11:33:42.439765Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:33:42.439765Z digest=sha256:ab0fffac0cad5b772ce88f2d418db2046ffdcd081689e65c9b37d19f89fc8764

Observation 08fa439c-bdb6-4acf-8ed9-9f15f82ab321 · outbound

This paper cites Urvos: Unified referring video object segmentation network with a large-scale benchmark,.

Memory-Augmented Multimodal Large Language Models for Small Object Understanding in Streaming Aerial Videos Urvos: Unified referring video object segmentation network with a large-scale benchmark,

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-01T11:33:42.515394Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:33:42.515394Z digest=sha256:0d8e1f56e7cb7dd892f2b6de47f953c869cabf11a85cae55e2d0992b11e7816f

Observation b14fc791-78fa-4110-bcd4-66ae0c8a6cf8 · outbound

This paper cites Sharegpt4video: Improving video understanding and generation with better captions,.

Memory-Augmented Multimodal Large Language Models for Small Object Understanding in Streaming Aerial Videos Sharegpt4video: Improving video understanding and generation with better captions,

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-01T11:33:42.604426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:33:42.604426Z digest=sha256:0362b6c5b5b748c76fdb03b42ab443adfafb248a950a3d085ebcdad40fa9b5de

Observation 79b21db3-821c-4694-8687-61f39cbdae15 · outbound

This paper cites Mevis: A multi-modal dataset for referring motion expression video segmentation,.

Memory-Augmented Multimodal Large Language Models for Small Object Understanding in Streaming Aerial Videos Mevis: A multi-modal dataset for referring motion expression video segmentation,

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-01T11:33:42.675678Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:33:42.675678Z digest=sha256:c348d4e94f28edac6ec623f37de0d4f7f928d3d1a916a0479e471de7adaf5080

Observation 844c35ca-a119-472a-8294-da0f114f5e6f · outbound

This paper cites Jtd-uav: Mllm-enhanced joint tracking and description framework for anti-uav systems,.

Memory-Augmented Multimodal Large Language Models for Small Object Understanding in Streaming Aerial Videos Jtd-uav: Mllm-enhanced joint tracking and description framework for anti-uav systems,

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-01T11:33:42.778744Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:33:42.778744Z digest=sha256:efdb57c324f3d6e3bbdd5c5865bc68997034243198932701741bf02268cc0adb

Observation 0e892caa-5ce7-4f56-bdb7-f50fa6c94cd4 · outbound

This paper cites Visa: Reasoning video object segmentation via large lan- guage models,.

Memory-Augmented Multimodal Large Language Models for Small Object Understanding in Streaming Aerial Videos Visa: Reasoning video object segmentation via large lan- guage models,

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-01T11:33:42.829375Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:33:42.829375Z digest=sha256:d048246dd4094474677de70da36dd2c6ab3eea152fd258c11738a39137afa8b5

Observation eed37ea5-009d-4755-b182-9cc0b34b6b93 · outbound

This paper cites Sa2VA: Marrying SAM2 with LLaVA for Dense Grounded Understanding of Images and Videos.

Memory-Augmented Multimodal Large Language Models for Small Object Understanding in Streaming Aerial Videos Sa2VA: Marrying SAM2 with LLaVA for Dense Grounded Understanding of Images and Videos

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-01T11:33:42.889990Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:33:42.889990Z digest=sha256:b08dd6a5a8993ebb7de419fe689e3dbd710c2af671b6de765a5f41a45a0816da

Observation 26c53c5e-2219-4323-a01b-152283af7035 · outbound

This paper cites Unipixel: Unified object referring and segmentation for pixel-level visual reason- ing,.

Memory-Augmented Multimodal Large Language Models for Small Object Understanding in Streaming Aerial Videos Unipixel: Unified object referring and segmentation for pixel-level visual reason- ing,

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-01T11:33:42.985730Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:33:42.985730Z digest=sha256:849666b80cb70f111f9088f75a61742e560ad947b40bfead46803acd06b1fec9

Observation 166f5840-37de-4732-b725-fca92a964b90 · outbound

This paper cites SAM 2: Segment Anything in Images and Videos.

Memory-Augmented Multimodal Large Language Models for Small Object Understanding in Streaming Aerial Videos SAM 2: Segment Anything in Images and Videos

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-01T11:33:43.065979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:33:43.065979Z digest=sha256:667d488e9f73f6ba7bcf167121cb6648850e1d22ca8d7dac3863176aa87c1ff5

Observation 4af657c0-610a-4099-a718-52299e96ebc4 · outbound

This paper cites Videoglamm: A large multimodal model for pixel-level visual JOURNAL OF LATEX CLASS FILES, VOL. 14, NO. 8, AUGUST 2021 14 grounding in videos,.

Memory-Augmented Multimodal Large Language Models for Small Object Understanding in Streaming Aerial Videos Videoglamm: A large multimodal model for pixel-level visual JOURNAL OF LATEX CLASS FILES, VOL. 14, NO. 8, AUGUST 2021 14 grounding in videos,

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-01T11:33:43.135856Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:33:43.135856Z digest=sha256:c90e1dd433f074b4f31e598dd483d99eaabc25d42ddbf2db6d8277c1d921d54c

Observation 7a0db379-3b2a-4df5-b8ca-796f9e2f9a06 · outbound

This paper cites Glus: Global-local reasoning unified into a single large language model for video segmentation,.

Memory-Augmented Multimodal Large Language Models for Small Object Understanding in Streaming Aerial Videos Glus: Global-local reasoning unified into a single large language model for video segmentation,

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-01T11:33:43.233507Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:33:43.233507Z digest=sha256:6dd1f21d420d9a6ec8b2e9a5506fb003403617a592205fa176de50c801173103

Observation 5c1cf94f-b519-42a1-b153-da195082896f · outbound

This paper cites The devil is in temporal token: High quality video reasoning segmentation,.

Memory-Augmented Multimodal Large Language Models for Small Object Understanding in Streaming Aerial Videos The devil is in temporal token: High quality video reasoning segmentation,

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-01T11:33:43.328278Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:33:43.328278Z digest=sha256:9eb13c135564c77d53b21b8ebd999955662285bbc0c1ee2b961fdd2036218f1d

Observation 1fc75169-9d8b-4957-8153-a284702ec0b5 · outbound

This paper cites Instructseg: Unifying instructed visual segmentation with multi-modal large language models,.

Memory-Augmented Multimodal Large Language Models for Small Object Understanding in Streaming Aerial Videos Instructseg: Unifying instructed visual segmentation with multi-modal large language models,

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-01T11:33:43.435658Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:33:43.435658Z digest=sha256:8ca86faeb62427c258b892143ff1adfefb28a9b865f510effae38126ef8049c5

Observation df21c549-3a59-43ec-94da-f40429c71d43 · outbound

This paper cites Geochat: Grounded large vision-language model for remote sensing,.

Memory-Augmented Multimodal Large Language Models for Small Object Understanding in Streaming Aerial Videos Geochat: Grounded large vision-language model for remote sensing,

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-01T11:33:43.532666Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:33:43.532666Z digest=sha256:0abd436e447e06ec80e9ff191846419b631cb032c1bd1bae3a8e263f8630f515

Observation aa2add77-efb9-4878-be89-fdcd1b629294 · outbound

This paper cites SkySenseGPT: A Fine-Grained Instruction Tuning Dataset and Model for Remote Sensing Vision-Language Understanding.

Memory-Augmented Multimodal Large Language Models for Small Object Understanding in Streaming Aerial Videos SkySenseGPT: A Fine-Grained Instruction Tuning Dataset and Model for Remote Sensing Vision-Language Understanding

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-01T11:33:43.613930Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:33:43.613930Z digest=sha256:2eece73503ec5c9a66f47f34f47f7e0b407290ba283916e9796f895e8af85a9f

Observation 2ef00eab-20ef-407f-b9a5-a1423abd28ee · outbound

This paper cites Earthgpt: A universal multimodal large language model for multisensor image comprehension in remote sensing domain,.

Memory-Augmented Multimodal Large Language Models for Small Object Understanding in Streaming Aerial Videos Earthgpt: A universal multimodal large language model for multisensor image comprehension in remote sensing domain,

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-01T11:33:43.690981Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:33:43.690981Z digest=sha256:9119cbe2e1d7b796f1dab25ccf9a049f7016b307760f146c28f4c316c1dd27b9

Observation 1b7777f8-6659-4ffb-9d4d-397c763d51cd · outbound

This paper cites Kimi K2.5: Visual Agentic Intelligence.

Memory-Augmented Multimodal Large Language Models for Small Object Understanding in Streaming Aerial Videos Kimi K2.5: Visual Agentic Intelligence

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-01T11:33:43.758425Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:33:43.758425Z digest=sha256:0844eb6cd36263c5ae364005d4a2186c5c5e535b64e9e71587b04a633ff5fde4

Observation beedfb6f-dfe7-4f2b-9f22-bf52ce0350a9 · outbound

This paper cites Qwen3.6-Plus: Towards real world agents,.

Memory-Augmented Multimodal Large Language Models for Small Object Understanding in Streaming Aerial Videos Qwen3.6-Plus: Towards real world agents,

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-01T11:33:43.831961Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:33:43.831961Z digest=sha256:96baeb94e76cd82cb8847fec6b116d9ac6362ad8c4935c53a29da411ce17b753

Observation 211f0754-f661-4860-b3d8-3eb118d94675 · outbound

This paper cites Visual instruction tuning,.

Memory-Augmented Multimodal Large Language Models for Small Object Understanding in Streaming Aerial Videos Visual instruction tuning,

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-01T11:33:43.921225Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:33:43.921225Z digest=sha256:eef518d0d0af277725df99e105d8df32b409f704152704861bd16e786bb595fa

Observation 49e040f5-7cb0-4289-9175-9dbffc0da9c2 · outbound

This paper cites Qwen2.5-VL Technical Report.

Memory-Augmented Multimodal Large Language Models for Small Object Understanding in Streaming Aerial Videos Qwen2.5-VL Technical Report

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-01T11:33:43.990599Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:33:43.990599Z digest=sha256:992be9f4ccdf661f03e0f77af90a56e230163369988e127d31d625009589c67d

Observation e7b2fe0a-3f59-41ca-9521-652d554e1f04 · outbound

This paper cites StreamingVLM: Real-Time Understanding for Infinite Video Streams.

Memory-Augmented Multimodal Large Language Models for Small Object Understanding in Streaming Aerial Videos StreamingVLM: Real-Time Understanding for Infinite Video Streams

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-01T11:33:44.071908Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:33:44.071908Z digest=sha256:0211c5e387eae8ec27f2bbe2fcdd3ea9037c4c677150c22fa6824a939570461a

Observation 9f512ed1-8155-4f28-bd60-74bb222a0f54 · outbound

This paper cites A fast and accurate one-stage approach to visual grounding,.

Memory-Augmented Multimodal Large Language Models for Small Object Understanding in Streaming Aerial Videos A fast and accurate one-stage approach to visual grounding,

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-01T11:33:44.133296Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:33:44.133296Z digest=sha256:166c349aa0c95ae0ecbb9c3a94f3f10fbb9d075c2949e7996131d58d72678fbd

Observation dfd9d33c-78e5-40bc-b1bf-c29bba6295cf · outbound

This paper cites Improving one-stage visual grounding by recursive sub-query construction,.

Memory-Augmented Multimodal Large Language Models for Small Object Understanding in Streaming Aerial Videos Improving one-stage visual grounding by recursive sub-query construction,

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-01T11:33:44.252921Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:33:44.252921Z digest=sha256:5e4328213e26116f098d6cad3ad3a4b0c25b3824946768cb7ba1994206aabddd

Observation 994a86ff-27c6-4364-b2da-a43282072bfd · outbound

This paper cites Referring transformer: A one-step approach to multi-task visual grounding,.

Memory-Augmented Multimodal Large Language Models for Small Object Understanding in Streaming Aerial Videos Referring transformer: A one-step approach to multi-task visual grounding,

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-01T11:33:44.355563Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:33:44.355563Z digest=sha256:ef4bf3597172941401e33077c4a4916d2807a562456a6b03c2a4fe5e9538b90a

Observation ed070ee0-2232-42c8-b209-4f4b6c51f151 · outbound

This paper cites Transvg: End-to-end visual grounding with transformers,.

Memory-Augmented Multimodal Large Language Models for Small Object Understanding in Streaming Aerial Videos Transvg: End-to-end visual grounding with transformers,

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-01T11:33:44.438359Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:33:44.438359Z digest=sha256:81517673b0b3d07169cb7f16f4b46187845990d0f14a6a8dcb8653fee1fb810a

Observation a3617427-5543-4c34-81cf-13884fbc05a6 · outbound

This paper cites Improving visual grounding with visual-linguistic verification and iterative reasoning,.

Memory-Augmented Multimodal Large Language Models for Small Object Understanding in Streaming Aerial Videos Improving visual grounding with visual-linguistic verification and iterative reasoning,

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-01T11:33:44.518896Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:33:44.518896Z digest=sha256:a8247108b6eb59bb5ea7c87e13350ae459fe95b30044a8738896e1fb6d093143

Observation 15da3c5a-576d-4eba-b291-7e1adf5e5b66 · outbound

This paper cites Seqtr: A simple yet universal network for visual grounding,.

Memory-Augmented Multimodal Large Language Models for Small Object Understanding in Streaming Aerial Videos Seqtr: A simple yet universal network for visual grounding,

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-01T11:33:44.607850Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:33:44.607850Z digest=sha256:f3fc4d0edd91f42257aa7d1b6c4b6c31f98184b72941d99d913eea0dc733e45f

Observation 2c406c7c-c960-4212-b552-37e5b05d6790 · outbound

This paper cites Shifting more attention to visual backbone: Query-modulated refinement networks for end-to-end visual grounding,.

Memory-Augmented Multimodal Large Language Models for Small Object Understanding in Streaming Aerial Videos Shifting more attention to visual backbone: Query-modulated refinement networks for end-to-end visual grounding,

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-01T11:33:44.682676Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:33:44.682676Z digest=sha256:84adac8a3ac858cc157a9c4c52a12ccfa0c71e2f622e7597da3878b1a0dd4272

Observation 8f6255a7-0317-45e2-98b1-3e35f30b41d7 · outbound

This paper cites A survivor in the era of large- scale pretraining: An empirical study of one-stage referring expression comprehension,.

Memory-Augmented Multimodal Large Language Models for Small Object Understanding in Streaming Aerial Videos A survivor in the era of large- scale pretraining: An empirical study of one-stage referring expression comprehension,

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-01T11:33:44.745993Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:33:44.745993Z digest=sha256:10446388985dc2401fdec2519f07bca73e7168feb7027687fbb50190e6a1942e

Observation 2bc27e84-0bb9-4c6c-9f13-e9a704e31956 · outbound

This paper cites Polyformer: Referring image segmentation as sequential polygon generation,.

Memory-Augmented Multimodal Large Language Models for Small Object Understanding in Streaming Aerial Videos Polyformer: Referring image segmentation as sequential polygon generation,

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-01T11:33:44.814543Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:33:44.814543Z digest=sha256:2b806d80ed09d62a2411006777a9b238f754ea2d3b7ccf44453a510fe9e044a6

Observation c2c19032-b299-4504-a5f6-cad0810c8554 · outbound

This paper cites Qwen2.5 Technical Report.

Memory-Augmented Multimodal Large Language Models for Small Object Understanding in Streaming Aerial Videos Qwen2.5 Technical Report

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-01T11:33:42.188553Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:33:42.188553Z digest=sha256:9c8211e4304e518b9f54d55e244e464e0518f51da15b438f6e6eb4413780783c

Pith citing papers

No inbound Pith citation observations are available.