Pith. sign in

Paper Citation Record · LEDGER

Memory-Augmented Multimodal Large Language Models for Small Object Understanding in Streaming Aerial Videos

As of 13 August 2026, this Paper Citation Record lists 47 of 47 outbound references and 0 inbound Pith citation observations for arXiv:2607.19857.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.19857 v1

Coverage vector

measured 47 of 47 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-01T11:33:44.814543Z

measured 47 of 47 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

47 of 47 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved47
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 9b995d4e-2191-4b15-b83f-b8fc2466296d · outbound

This paper cites The small-drone revolution is coming—scientists need to ensure it will be safe,.

Memory-Augmented Multimodal Large Language Models for Small Object Understanding in Streaming Aerial Videos The small-drone revolution is coming—scientists need to ensure it will be safe,

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-01T11:33:41.010000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:33:41.010000Z digest=sha256:ae61a404898e09127232d99371248377175f466a5785aa5e97a0096c66f407b3

Observation 49b07a8c-e707-4b7c-9a55-2c743b824f6b · outbound

This paper cites Champion-level drone racing using deep reinforcement learning,.

Memory-Augmented Multimodal Large Language Models for Small Object Understanding in Streaming Aerial Videos Champion-level drone racing using deep reinforcement learning,

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-01T11:33:41.077507Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:33:41.077507Z digest=sha256:b854b23c57a2ae304cddc866d36873f7a1be5078009719c5773d294d28147292

Observation 761e62ab-1b1c-4192-b68f-2dad832a8038 · outbound

This paper cites Video object segmentation without tem- poral information,.

Memory-Augmented Multimodal Large Language Models for Small Object Understanding in Streaming Aerial Videos Video object segmentation without tem- poral information,

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-01T11:33:41.210959Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:33:41.210959Z digest=sha256:10c4ff5595367fc94a845ce0570fc1dc276e1a8dd8131aa512629de9daf1c4ce

Observation 5c39917d-b99c-406e-b8f3-61d24b26056c · outbound

This paper cites 3d question answering for city scene understanding,.

Memory-Augmented Multimodal Large Language Models for Small Object Understanding in Streaming Aerial Videos 3d question answering for city scene understanding,

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-01T11:33:41.290989Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:33:41.290989Z digest=sha256:76f16efbb6cb249c909e5fd854acbae776d073130d6fcb70ba7e7bd7f28fbd16

Observation 43efa3e9-286a-4e1b-be86-a3197c84754d · outbound

This paper cites Phenobench: A large dataset and benchmarks for semantic image interpretation in the agricultural do- main,.

Memory-Augmented Multimodal Large Language Models for Small Object Understanding in Streaming Aerial Videos Phenobench: A large dataset and benchmarks for semantic image interpretation in the agricultural do- main,

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-01T11:33:41.391352Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:33:41.391352Z digest=sha256:51b070836c6b187ef0e09019b097230e317ef3afc079c19e6a6ebaff815bea1e

Observation 01ecd4ed-e87b-4b06-8dec-7ad48ad22ef2 · outbound

This paper cites Detecting flying objects using a single moving camera,.

Memory-Augmented Multimodal Large Language Models for Small Object Understanding in Streaming Aerial Videos Detecting flying objects using a single moving camera,

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-01T11:33:41.479024Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:33:41.479024Z digest=sha256:607f59acd3b965661c689c152478f202ee6e34132ce5f5dc4102f6c56054048c

Observation 4f4d291c-5418-4571-8780-b59da9a6df49 · outbound

This paper cites Revisiting image-language networks for open-ended phrase detection,.

Memory-Augmented Multimodal Large Language Models for Small Object Understanding in Streaming Aerial Videos Revisiting image-language networks for open-ended phrase detection,

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-01T11:33:41.574070Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:33:41.574070Z digest=sha256:3483dda9811fb0a40da23b46ea12de52421edb2c20944bb81859af6a85af27fc

Observation d2b93844-3363-4b7a-aff3-0037efbdcd01 · outbound

This paper cites Mevis: A large- scale benchmark for video segmentation with motion expressions,.

Memory-Augmented Multimodal Large Language Models for Small Object Understanding in Streaming Aerial Videos Mevis: A large- scale benchmark for video segmentation with motion expressions,

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-01T11:33:41.647764Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:33:41.647764Z digest=sha256:cc76f4a9bcee558cc2639a335d6b511c59be9ce6e1af6c11b5dbd00a468a2217

Observation d5431214-f946-448e-974c-5886b8e16458 · outbound

This paper cites Lamot: Language- guided multi-object tracking,.

Memory-Augmented Multimodal Large Language Models for Small Object Understanding in Streaming Aerial Videos Lamot: Language- guided multi-object tracking,

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-01T11:33:41.750581Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:33:41.750581Z digest=sha256:a9004bf426f9babd3a3bb87e3270829863944d3c34281cd5201189c1e391b987

Observation a6822768-0774-409a-b3c4-642633a0a9f6 · outbound

This paper cites Skyfind: A large-scale benchmark unveiling referring expres- sion comprehension for uav.

Memory-Augmented Multimodal Large Language Models for Small Object Understanding in Streaming Aerial Videos Skyfind: A large-scale benchmark unveiling referring expres- sion comprehension for uav

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-01T11:33:41.793599Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:33:41.793599Z digest=sha256:955a98e019b7306fda49d8eb554233d989dd7f25eabd1d5fb82b5936aa571f2d

Observation 50c4b8f7-f820-4c1e-96ac-d23aae4cef2d · outbound

This paper cites Aerialmind: Towards referring multi-object tracking in uav scenarios,.

Memory-Augmented Multimodal Large Language Models for Small Object Understanding in Streaming Aerial Videos Aerialmind: Towards referring multi-object tracking in uav scenarios,

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-01T11:33:41.860865Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:33:41.860865Z digest=sha256:ec86f0c25d8e2b00601f2f32c8c2ee980b21d476bdc3c8979ed44ae0ba3ad2c9

Observation 6da0cf3b-dd9b-4492-acd3-8afc41de44ce · outbound

This paper cites Event-aware instructed assistant for referring video segmentation,.

Memory-Augmented Multimodal Large Language Models for Small Object Understanding in Streaming Aerial Videos Event-aware instructed assistant for referring video segmentation,

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-01T11:33:41.918717Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:33:41.918717Z digest=sha256:792ca1364eea35eb2f1438b8eba2f7969a446ff4de853da7df476fcc0e18d636

Observation cd07ae9b-fc66-4d12-97e5-44f84799e2fb · outbound

This paper cites City-vlm: Towards multidomain perception scene understanding via multimodal incomplete learning,.

Memory-Augmented Multimodal Large Language Models for Small Object Understanding in Streaming Aerial Videos City-vlm: Towards multidomain perception scene understanding via multimodal incomplete learning,

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-01T11:33:42.000557Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:33:42.000557Z digest=sha256:c1e8a0c5e1d3468eaedc4d66f61a20a3dc3ced92a5faac74787dea4717550b3e

Observation 71600a21-e282-4709-8a33-42faab77c5fb · outbound

This paper cites Qwen2.5 technical report,.

Memory-Augmented Multimodal Large Language Models for Small Object Understanding in Streaming Aerial Videos Qwen2.5 technical report,

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-01T11:33:42.085352Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:33:42.085352Z digest=sha256:25c82bdd745e14a9a4c87eeef798d79d1666627166fe91099150d9b268f438a0

Observation 366e8549-1cbc-4afa-984b-b330d1f213b2 · outbound

This paper cites SAM 3: Segment Anything with Concepts.

Memory-Augmented Multimodal Large Language Models for Small Object Understanding in Streaming Aerial Videos SAM 3: Segment Anything with Concepts

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-01T11:33:42.259681Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:33:42.259681Z digest=sha256:abb18bfbf6c64b9b8ba3efe9c9bed66ff0fd878da3b8aec1224a811f659edde7

Observation 305d387f-576c-47d0-99f8-484771a6ffed · outbound

This paper cites React: Streaming video analytics on the edge with asynchronous cloud support,.

Memory-Augmented Multimodal Large Language Models for Small Object Understanding in Streaming Aerial Videos React: Streaming video analytics on the edge with asynchronous cloud support,

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-01T11:33:42.350558Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:33:42.350558Z digest=sha256:b83085cf5e658fb254faf66f2ae737501e91cbbfc303a9fd92d51da0ab295b1d

Observation 5b29eab7-6a4d-4271-8344-f237d491cea0 · outbound

This paper cites Multiple-object-tracking algorithm based on dense trajectory voting in aerial videos,.

Memory-Augmented Multimodal Large Language Models for Small Object Understanding in Streaming Aerial Videos Multiple-object-tracking algorithm based on dense trajectory voting in aerial videos,

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-01T11:33:42.439765Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:33:42.439765Z digest=sha256:d7411159b63475cb6e9d386f9123d8bf9c51751876601faded6d4fcfca95e8f8

Observation 08fa439c-bdb6-4acf-8ed9-9f15f82ab321 · outbound

This paper cites Urvos: Unified referring video object segmentation network with a large-scale benchmark,.

Memory-Augmented Multimodal Large Language Models for Small Object Understanding in Streaming Aerial Videos Urvos: Unified referring video object segmentation network with a large-scale benchmark,

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-01T11:33:42.515394Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:33:42.515394Z digest=sha256:4fdecf371fa8010420cfd3b22849cba626124e74f3e7a099095cbaac35a113eb

Observation b14fc791-78fa-4110-bcd4-66ae0c8a6cf8 · outbound

This paper cites Sharegpt4video: Improving video understanding and generation with better captions,.

Memory-Augmented Multimodal Large Language Models for Small Object Understanding in Streaming Aerial Videos Sharegpt4video: Improving video understanding and generation with better captions,

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-01T11:33:42.604426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:33:42.604426Z digest=sha256:df218ec170ec91178c7ae07dbf8ce19b6b44c92a027206efa5c3253c9d7236cc

Observation 79b21db3-821c-4694-8687-61f39cbdae15 · outbound

This paper cites Mevis: A multi-modal dataset for referring motion expression video segmentation,.

Memory-Augmented Multimodal Large Language Models for Small Object Understanding in Streaming Aerial Videos Mevis: A multi-modal dataset for referring motion expression video segmentation,

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-01T11:33:42.675678Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:33:42.675678Z digest=sha256:5cbc0d33dd2b1b7ba00498e46be21f4fb3e9d51cf105f2ccc4345fee7a79eae8

Observation 844c35ca-a119-472a-8294-da0f114f5e6f · outbound

This paper cites Jtd-uav: Mllm-enhanced joint tracking and description framework for anti-uav systems,.

Memory-Augmented Multimodal Large Language Models for Small Object Understanding in Streaming Aerial Videos Jtd-uav: Mllm-enhanced joint tracking and description framework for anti-uav systems,

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-01T11:33:42.778744Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:33:42.778744Z digest=sha256:53db4b81ddcc62fd668cebc09625fa884d2410d7b413e635dddaa06bb00f6995

Observation 0e892caa-5ce7-4f56-bdb7-f50fa6c94cd4 · outbound

This paper cites Visa: Reasoning video object segmentation via large lan- guage models,.

Memory-Augmented Multimodal Large Language Models for Small Object Understanding in Streaming Aerial Videos Visa: Reasoning video object segmentation via large lan- guage models,

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-01T11:33:42.829375Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:33:42.829375Z digest=sha256:2ae592028662dde2ffeaae45a6cc6aab0ee6dd554dcd9a331f22b6f2b21e147c

Observation eed37ea5-009d-4755-b182-9cc0b34b6b93 · outbound

This paper cites Sa2VA: Marrying SAM2 with LLaVA for Dense Grounded Understanding of Images and Videos.

Memory-Augmented Multimodal Large Language Models for Small Object Understanding in Streaming Aerial Videos Sa2VA: Marrying SAM2 with LLaVA for Dense Grounded Understanding of Images and Videos

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-01T11:33:42.889990Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:33:42.889990Z digest=sha256:3ef08985cff4ba46e4d5bc756a9425925f5341220dc5f1d78311e51abe0c43cc

Observation 26c53c5e-2219-4323-a01b-152283af7035 · outbound

This paper cites Unipixel: Unified object referring and segmentation for pixel-level visual reason- ing,.

Memory-Augmented Multimodal Large Language Models for Small Object Understanding in Streaming Aerial Videos Unipixel: Unified object referring and segmentation for pixel-level visual reason- ing,

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-01T11:33:42.985730Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:33:42.985730Z digest=sha256:6761af2ffa2d798dcc53560af84a1dad8ab21cc23c52140cdaeb17a4e26a2e9a

Observation 166f5840-37de-4732-b725-fca92a964b90 · outbound

This paper cites SAM 2: Segment Anything in Images and Videos.

Memory-Augmented Multimodal Large Language Models for Small Object Understanding in Streaming Aerial Videos SAM 2: Segment Anything in Images and Videos

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-01T11:33:43.065979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:33:43.065979Z digest=sha256:75041f61e0f80a1f13a084c5d6a7a7e85ec201e2921804183ffc79fa667d12e1

Observation 4af657c0-610a-4099-a718-52299e96ebc4 · outbound

This paper cites Videoglamm: A large multimodal model for pixel-level visual JOURNAL OF LATEX CLASS FILES, VOL. 14, NO. 8, AUGUST 2021 14 grounding in videos,.

Memory-Augmented Multimodal Large Language Models for Small Object Understanding in Streaming Aerial Videos Videoglamm: A large multimodal model for pixel-level visual JOURNAL OF LATEX CLASS FILES, VOL. 14, NO. 8, AUGUST 2021 14 grounding in videos,

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-01T11:33:43.135856Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:33:43.135856Z digest=sha256:67e64c05023a14c99ed07f215cbe56d40e5b89f09141267717764fda2f78a4d1

Observation 7a0db379-3b2a-4df5-b8ca-796f9e2f9a06 · outbound

This paper cites Glus: Global-local reasoning unified into a single large language model for video segmentation,.

Memory-Augmented Multimodal Large Language Models for Small Object Understanding in Streaming Aerial Videos Glus: Global-local reasoning unified into a single large language model for video segmentation,

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-01T11:33:43.233507Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:33:43.233507Z digest=sha256:65398f23fe34b7f1df6fe1f6147249fd5554ed1bd2d958a34f47589d6d201fd7

Observation 5c1cf94f-b519-42a1-b153-da195082896f · outbound

This paper cites The devil is in temporal token: High quality video reasoning segmentation,.

Memory-Augmented Multimodal Large Language Models for Small Object Understanding in Streaming Aerial Videos The devil is in temporal token: High quality video reasoning segmentation,

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-01T11:33:43.328278Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:33:43.328278Z digest=sha256:b88c78a8af0223de060c33876e76cb6bc8de154a227ed5b897f555c93e3b5754

Observation 1fc75169-9d8b-4957-8153-a284702ec0b5 · outbound

This paper cites Instructseg: Unifying instructed visual segmentation with multi-modal large language models,.

Memory-Augmented Multimodal Large Language Models for Small Object Understanding in Streaming Aerial Videos Instructseg: Unifying instructed visual segmentation with multi-modal large language models,

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-01T11:33:43.435658Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:33:43.435658Z digest=sha256:1bf07d42073709744650437fd14b4eb93069728064f8093817d10aeb7518c001

Observation df21c549-3a59-43ec-94da-f40429c71d43 · outbound

This paper cites Geochat: Grounded large vision-language model for remote sensing,.

Memory-Augmented Multimodal Large Language Models for Small Object Understanding in Streaming Aerial Videos Geochat: Grounded large vision-language model for remote sensing,

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-01T11:33:43.532666Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:33:43.532666Z digest=sha256:82d7d04f9619b565e9c36e57b7c91e7e74b67db81b311cabec9145df08c24fb1

Observation aa2add77-efb9-4878-be89-fdcd1b629294 · outbound

This paper cites SkySenseGPT: A Fine-Grained Instruction Tuning Dataset and Model for Remote Sensing Vision-Language Understanding.

Memory-Augmented Multimodal Large Language Models for Small Object Understanding in Streaming Aerial Videos SkySenseGPT: A Fine-Grained Instruction Tuning Dataset and Model for Remote Sensing Vision-Language Understanding

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-01T11:33:43.613930Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:33:43.613930Z digest=sha256:a64dc9dd44a84ddf43ed86d1df1cd78cd191b422a9634c04c18efbe95337c2eb

Observation 2ef00eab-20ef-407f-b9a5-a1423abd28ee · outbound

This paper cites Earthgpt: A universal multimodal large language model for multisensor image comprehension in remote sensing domain,.

Memory-Augmented Multimodal Large Language Models for Small Object Understanding in Streaming Aerial Videos Earthgpt: A universal multimodal large language model for multisensor image comprehension in remote sensing domain,

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-01T11:33:43.690981Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:33:43.690981Z digest=sha256:7b4e5d3bd007bd1d63f7d4291272d7afe93e1762db410a7b47c091228b3e027d

Observation 1b7777f8-6659-4ffb-9d4d-397c763d51cd · outbound

This paper cites Kimi K2.5: Visual Agentic Intelligence.

Memory-Augmented Multimodal Large Language Models for Small Object Understanding in Streaming Aerial Videos Kimi K2.5: Visual Agentic Intelligence

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-01T11:33:43.758425Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:33:43.758425Z digest=sha256:bc007dfe0439adc031ad704edd550a3fe54d9fec5861f62d5540c62a4744e64c

Observation beedfb6f-dfe7-4f2b-9f22-bf52ce0350a9 · outbound

This paper cites Qwen3.6-Plus: Towards real world agents,.

Memory-Augmented Multimodal Large Language Models for Small Object Understanding in Streaming Aerial Videos Qwen3.6-Plus: Towards real world agents,

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-01T11:33:43.831961Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:33:43.831961Z digest=sha256:5842f9df0e805cb4d60db321e7dbb5774c791c8f5d267852a6a414263ed23525

Observation 211f0754-f661-4860-b3d8-3eb118d94675 · outbound

This paper cites Visual instruction tuning,.

Memory-Augmented Multimodal Large Language Models for Small Object Understanding in Streaming Aerial Videos Visual instruction tuning,

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-01T11:33:43.921225Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:33:43.921225Z digest=sha256:645932bda8ce9e0d738f607d477134e0b53ef7946f0e3e444d70741483b10c10

Observation 49e040f5-7cb0-4289-9175-9dbffc0da9c2 · outbound

This paper cites Qwen2.5-VL Technical Report.

Memory-Augmented Multimodal Large Language Models for Small Object Understanding in Streaming Aerial Videos Qwen2.5-VL Technical Report

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-01T11:33:43.990599Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:33:43.990599Z digest=sha256:821a52e5c0748daee3c004da303ae86b41a8339956292b3873f4c2238a0915b1

Observation e7b2fe0a-3f59-41ca-9521-652d554e1f04 · outbound

This paper cites StreamingVLM: Real-Time Understanding for Infinite Video Streams.

Memory-Augmented Multimodal Large Language Models for Small Object Understanding in Streaming Aerial Videos StreamingVLM: Real-Time Understanding for Infinite Video Streams

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-01T11:33:44.071908Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:33:44.071908Z digest=sha256:c2aedb5b1cc94288ec2141e5b3fa89a567b28d5890d5771850d064c078df80e0

Observation 9f512ed1-8155-4f28-bd60-74bb222a0f54 · outbound

This paper cites A fast and accurate one-stage approach to visual grounding,.

Memory-Augmented Multimodal Large Language Models for Small Object Understanding in Streaming Aerial Videos A fast and accurate one-stage approach to visual grounding,

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-01T11:33:44.133296Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:33:44.133296Z digest=sha256:bdfa04729628d452925993864fa3c7c3a32fadb057232855610d16de4342a879

Observation dfd9d33c-78e5-40bc-b1bf-c29bba6295cf · outbound

This paper cites Improving one-stage visual grounding by recursive sub-query construction,.

Memory-Augmented Multimodal Large Language Models for Small Object Understanding in Streaming Aerial Videos Improving one-stage visual grounding by recursive sub-query construction,

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-01T11:33:44.252921Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:33:44.252921Z digest=sha256:ddfdf4659fe188e5c8204fdd87e7e2d15d31b3fac0cc89eaf2acdea542ed2eda

Observation 994a86ff-27c6-4364-b2da-a43282072bfd · outbound

This paper cites Referring transformer: A one-step approach to multi-task visual grounding,.

Memory-Augmented Multimodal Large Language Models for Small Object Understanding in Streaming Aerial Videos Referring transformer: A one-step approach to multi-task visual grounding,

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-01T11:33:44.355563Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:33:44.355563Z digest=sha256:725e4d818db0ddb2a5faee216da361cf868a61f05b680a36489966fbb2baf469

Observation ed070ee0-2232-42c8-b209-4f4b6c51f151 · outbound

This paper cites Transvg: End-to-end visual grounding with transformers,.

Memory-Augmented Multimodal Large Language Models for Small Object Understanding in Streaming Aerial Videos Transvg: End-to-end visual grounding with transformers,

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-01T11:33:44.438359Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:33:44.438359Z digest=sha256:292c6681c1639efee6ea3225693a0f885e17daa5b0b3fc2ec4a1742d21225c06

Observation a3617427-5543-4c34-81cf-13884fbc05a6 · outbound

This paper cites Improving visual grounding with visual-linguistic verification and iterative reasoning,.

Memory-Augmented Multimodal Large Language Models for Small Object Understanding in Streaming Aerial Videos Improving visual grounding with visual-linguistic verification and iterative reasoning,

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-01T11:33:44.518896Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:33:44.518896Z digest=sha256:151d786da48b00dfae444d15e7c0098e4f47fb561270998933806cff84fa64a7

Observation 15da3c5a-576d-4eba-b291-7e1adf5e5b66 · outbound

This paper cites Seqtr: A simple yet universal network for visual grounding,.

Memory-Augmented Multimodal Large Language Models for Small Object Understanding in Streaming Aerial Videos Seqtr: A simple yet universal network for visual grounding,

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-01T11:33:44.607850Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:33:44.607850Z digest=sha256:f32f95dafcda505eb8c7d08b5034c7ac47223878df14e4961c19df96af092726

Observation 2c406c7c-c960-4212-b552-37e5b05d6790 · outbound

This paper cites Shifting more attention to visual backbone: Query-modulated refinement networks for end-to-end visual grounding,.

Memory-Augmented Multimodal Large Language Models for Small Object Understanding in Streaming Aerial Videos Shifting more attention to visual backbone: Query-modulated refinement networks for end-to-end visual grounding,

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-01T11:33:44.682676Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:33:44.682676Z digest=sha256:95007330635251e111fb05c49dead73817482dcd3c56f9a14d6fb126361c610a

Observation 8f6255a7-0317-45e2-98b1-3e35f30b41d7 · outbound

This paper cites A survivor in the era of large- scale pretraining: An empirical study of one-stage referring expression comprehension,.

Memory-Augmented Multimodal Large Language Models for Small Object Understanding in Streaming Aerial Videos A survivor in the era of large- scale pretraining: An empirical study of one-stage referring expression comprehension,

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-01T11:33:44.745993Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:33:44.745993Z digest=sha256:4b7b5cafe4656c4d7a0f867c8fbb285bf464181846a952bd4a5fe90880f5a23f

Observation 2bc27e84-0bb9-4c6c-9f13-e9a704e31956 · outbound

This paper cites Polyformer: Referring image segmentation as sequential polygon generation,.

Memory-Augmented Multimodal Large Language Models for Small Object Understanding in Streaming Aerial Videos Polyformer: Referring image segmentation as sequential polygon generation,

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-01T11:33:44.814543Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:33:44.814543Z digest=sha256:4683910d80c55ebe3be6bc1f6070d7f021ffe3ea33d37b8e415bae940776946c

Observation c2c19032-b299-4504-a5f6-cad0810c8554 · outbound

This paper cites Qwen2.5 Technical Report.

Memory-Augmented Multimodal Large Language Models for Small Object Understanding in Streaming Aerial Videos Qwen2.5 Technical Report

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-01T11:33:42.188553Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:33:42.188553Z digest=sha256:ed8821136f90c433cd868a7c749561ec91c2815a06cd828481fece92d573b22f

Pith citing papers

No inbound Pith citation observations are available.