Pith. sign in

Paper Citation Record · LEDGER

SALOVA: Segment-Augmented Long Video Assistant for Targeted Retrieval and Routing in Long-Form Video Analysis

As of 20 August 2026, this Paper Citation Record lists 73 of 73 outbound references and 3 inbound Pith citation observations for arXiv:2411.16173.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.16173 v2

Coverage vector

measured 73 of 73 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T13:31:04.642937Z

measured 76 of 76 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-12T10:18:48.890691Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-20T18:53:38.996149Z

Reference resolution

73 of 73 outbound references displayed

  • verified exact0
  • verified fuzzy24
  • unresolved48
  • parse uncertain1
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 57e231a7-ea8f-496b-a33d-a6ae56e950b3 · outbound

This paper cites Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone.

SALOVA: Segment-Augmented Long Video Assistant for Targeted Retrieval and Routing in Long-Form Video Analysis Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-12T13:31:03.955094Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:31:03.955094Z digest=sha256:bd9efa672a115acc71570ca9cffadcd78f317feaf1541045a44bcd3f88285af6

Observation f30c46be-108a-4b7c-a23a-05c66c022612 · outbound

This paper cites Flamingo: a visual language model for few-shot learning.

SALOVA: Segment-Augmented Long Video Assistant for Targeted Retrieval and Routing in Long-Form Video Analysis Flamingo: a visual language model for few-shot learning

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-12T13:31:03.961916Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:31:03.961916Z digest=sha256:e7812c9931b30315a507c9560f0b0cacd846e418e7e7933ac2456576faa9031b

Observation 35b0cc40-492e-4f56-b313-504ba88fe11d · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

SALOVA: Segment-Augmented Long Video Assistant for Targeted Retrieval and Routing in Long-Form Video Analysis Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-12T13:31:03.970336Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:31:03.970336Z digest=sha256:a0b69d83b7710ccde9f09aadfd59e5ca63ba33500fabdb7605327006be0be189

Observation 58727375-4b88-4fce-87d1-500447a09a19 · outbound

This paper cites Frozen in time: A joint video and image encoder for end-to-end retrieval.

SALOVA: Segment-Augmented Long Video Assistant for Targeted Retrieval and Routing in Long-Form Video Analysis Frozen in time: A joint video and image encoder for end-to-end retrieval

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-12T13:31:03.975755Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:31:03.975755Z digest=sha256:aacebc2a1e3b9563491a52ebc7ee473ccb74d199dc92a61cbf0604fd38411949

Observation 70d5b1df-bbd2-4b14-ad30-864834dd6e21 · outbound

This paper cites Lan- guage models are few-shot learners.

SALOVA: Segment-Augmented Long Video Assistant for Targeted Retrieval and Routing in Long-Form Video Analysis Lan- guage models are few-shot learners

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-12T13:31:03.987744Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:31:03.987744Z digest=sha256:e6849a1779932236d1fc3866ec8e8f29c11eeb50479550f92da4a2601a82c48b

Observation aeb4c7d7-ef0a-43a8-9e10-a9bd27aab885 · outbound

This paper cites Activitynet: A large-scale video benchmark for human activity understanding.

SALOVA: Segment-Augmented Long Video Assistant for Targeted Retrieval and Routing in Long-Form Video Analysis Activitynet: A large-scale video benchmark for human activity understanding

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:31:06.495865Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T13:31:03.997771Z digest=sha256:a4e7c6f114cb9fa92bb4e04d5ada7f09af90ab63daef8c2273ee33276eb4ee7b

Observation 9b4856bd-ad5b-49d8-9f3b-c6ebb171409e · outbound

This paper cites Honeybee: Locality-enhanced Projector for Multimodal LLM.

SALOVA: Segment-Augmented Long Video Assistant for Targeted Retrieval and Routing in Long-Form Video Analysis Honeybee: Locality-enhanced Projector for Multimodal LLM

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-12T13:31:04.008071Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:31:04.008071Z digest=sha256:5769fe2c52ba309a4be2ea288c32d77bb3aec79121378e986b6da69b3e2f2282

Observation d115cd80-01c9-43a4-be4d-04a3b491d210 · outbound

This paper cites ShareGPT4Video: Improving Video Understanding and Generation with Better Captions.

SALOVA: Segment-Augmented Long Video Assistant for Targeted Retrieval and Routing in Long-Form Video Analysis ShareGPT4Video: Improving Video Understanding and Generation with Better Captions

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-12T13:31:04.016872Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:31:04.016872Z digest=sha256:95ea697747ec389bdb6624f71126c004c17d02992897265d8442c8d00924f70d

Observation a7d7b95f-b852-4a5c-9dd7-ac6c129150fa · outbound

This paper cites Training Deep Nets with Sublinear Memory Cost.

SALOVA: Segment-Augmented Long Video Assistant for Targeted Retrieval and Routing in Long-Form Video Analysis Training Deep Nets with Sublinear Memory Cost

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-12T13:31:04.025102Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:31:04.025102Z digest=sha256:2a02db0d279aa7bd196ea003e17a29468e7731a1aa33649206cf7ac58915f9ca

Observation f5ebbed0-bea7-42cd-9ccf-14966dd6a3f3 · outbound

This paper cites Panda-70m: Captioning 70m videos with multiple cross-modality teachers.

SALOVA: Segment-Augmented Long Video Assistant for Targeted Retrieval and Routing in Long-Form Video Analysis Panda-70m: Captioning 70m videos with multiple cross-modality teachers

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-12T13:31:04.036608Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:31:04.036608Z digest=sha256:590cec8e942525ae1284f79d91d43c6f53eb2caadb6a74fca08046d9c7559d07

Observation c2f1a3ba-b6cd-4faa-9ae6-b73478f4aff1 · outbound

This paper cites InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks.

SALOVA: Segment-Augmented Long Video Assistant for Targeted Retrieval and Routing in Long-Form Video Analysis InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-12T13:31:04.042655Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:31:04.042655Z digest=sha256:85c1f406932584e676786ba24b6508ef7bfd22d07343fc717e0b5363894acd14

Observation e3a92d58-ee71-4499-ac88-fd22ed3be7b1 · outbound

This paper cites VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs.

SALOVA: Segment-Augmented Long Video Assistant for Targeted Retrieval and Routing in Long-Form Video Analysis VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-12T13:31:04.051362Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:31:04.051362Z digest=sha256:8ce78a69667fd2967dac3b396b8c28a9074136bb21ef9dc9f87ca0e94ab18e6e

Observation 7a423278-109c-4505-a66e-7b863c51a2cd · outbound

This paper cites Gonzalez, Ion Stoica, and Eric P.

SALOVA: Segment-Augmented Long Video Assistant for Targeted Retrieval and Routing in Long-Form Video Analysis Gonzalez, Ion Stoica, and Eric P

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-12T13:31:04.058867Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:31:04.058867Z digest=sha256:aaaa0ad3f0088323cc305bbbeb3a84543a2b6f32290e3b6507a8435a9eed459d

Observation 1531c6b0-7512-4441-b98c-a26d17260198 · outbound

This paper cites InstructBLIP: Towards general-purpose vision- language models with instruction tuning.

SALOVA: Segment-Augmented Long Video Assistant for Targeted Retrieval and Routing in Long-Form Video Analysis InstructBLIP: Towards general-purpose vision- language models with instruction tuning

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:31:06.464072Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T13:31:04.067305Z digest=sha256:321cab16725058ceccd92f513e677290a0817d9488022ff868076f768d18f7e3

Observation e98ad8a7-fb69-4b87-8731-3af104786d0a · outbound

This paper cites FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning.

SALOVA: Segment-Augmented Long Video Assistant for Targeted Retrieval and Routing in Long-Form Video Analysis FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-12T13:31:04.077714Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:31:04.077714Z digest=sha256:a578575199d2ebbc90007f06915161ef624309320fc129086b38d4d057d0ba9a

Observation d2c8bc61-368b-4447-b94a-6ff9e3e29571 · outbound

This paper cites Patch n’pack: Navit, a vision transformer for any aspect ratio and resolution.

SALOVA: Segment-Augmented Long Video Assistant for Targeted Retrieval and Routing in Long-Form Video Analysis Patch n’pack: Navit, a vision transformer for any aspect ratio and resolution

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:31:06.449361Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T13:31:04.086355Z digest=sha256:982446ce1b967b83c21d3fbc12a6fa51b4a67323bb269e447879f3551b53f4b5

Observation 9855ebe5-7360-4b12-bc9a-c8eefe3d9b8d · outbound

This paper cites InternLM-XComposer2: Mastering Free-form Text-Image Composition and Comprehension in Vision-Language Large Model.

SALOVA: Segment-Augmented Long Video Assistant for Targeted Retrieval and Routing in Long-Form Video Analysis InternLM-XComposer2: Mastering Free-form Text-Image Composition and Comprehension in Vision-Language Large Model

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-12T13:31:04.096194Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:31:04.096194Z digest=sha256:43533bb7fce53851ec81648f06969b1f5411da1e586e3284a66641d2969aef01

Observation a8d85cd7-4d49-49e3-859d-04e545fedc8c · outbound

This paper cites The Llama 3 Herd of Models.

SALOVA: Segment-Augmented Long Video Assistant for Targeted Retrieval and Routing in Long-Form Video Analysis The Llama 3 Herd of Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-12T13:31:04.102463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:31:04.102463Z digest=sha256:eef56271f89034520bbce5f2744671fb882985fa3b39b987f586a83b4b08c1c9

Observation a9b923ba-1570-4c45-a686-8919fb0580a9 · outbound

This paper cites Slowfast networks for video recognition.

SALOVA: Segment-Augmented Long Video Assistant for Targeted Retrieval and Routing in Long-Form Video Analysis Slowfast networks for video recognition

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:31:06.435705Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T13:31:04.113593Z digest=sha256:c304c8c9f522d0d89dd1d4f81d2baf1d9b4df8ea890c42dfb0bc8b829ae784db

Observation a483f1a2-da3e-4714-808b-17345e50c169 · outbound

This paper cites Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis.

SALOVA: Segment-Augmented Long Video Assistant for Targeted Retrieval and Routing in Long-Form Video Analysis Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-12T13:31:04.119531Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:31:04.119531Z digest=sha256:b420d14658eea13d5b3ef2b8df07d24607a1a988621395c7a4f6173cc1817b25

Observation 31425c59-ec50-4615-a346-5b13beed029b · outbound

This paper cites Long Story Short: Story-level Video Understanding from 20K Short Films.

SALOVA: Segment-Augmented Long Video Assistant for Targeted Retrieval and Routing in Long-Form Video Analysis Long Story Short: Story-level Video Understanding from 20K Short Films

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-12T13:31:04.127586Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:31:04.127586Z digest=sha256:baee2c6be7511a6bd4f7837b4a515f15e1c489546f120da7d64fe2d8dc45b5ed

Observation 593182dd-316b-4b18-a780-637c6680b74a · outbound

This paper cites Gemini, 2023.

SALOVA: Segment-Augmented Long Video Assistant for Targeted Retrieval and Routing in Long-Form Video Analysis Gemini, 2023

Reference 22

Resolution
parse uncertain
no resolver link, observed 2026-08-12T13:31:04.134099Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:31:04.134099Z digest=sha256:372d8c7dd58f7fa90559eed0dbf767389ea118beec5d57598c3d680cd9bf6e5f

Observation 67b4aa85-c730-4367-b10c-e3009ae6c505 · outbound

This paper cites Ma-lmm: Memory-augmented large multimodal model for long-term video understanding.

SALOVA: Segment-Augmented Long Video Assistant for Targeted Retrieval and Routing in Long-Form Video Analysis Ma-lmm: Memory-augmented large multimodal model for long-term video understanding

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:31:06.410889Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T13:31:04.140328Z digest=sha256:d33cac7276a1d78923fd9da0d5c19b0f95b45dd0fe924aa528672d0282a1f038

Observation 955a6b36-8383-4a74-8f11-c3f770b2fc6a · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

SALOVA: Segment-Augmented Long Video Assistant for Targeted Retrieval and Routing in Long-Form Video Analysis LoRA: Low-Rank Adaptation of Large Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-12T13:31:04.150807Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:31:04.150807Z digest=sha256:ce767cfbd141cc17e8b6f831aec0864c6db54b261d32c06b894588c258741444

Observation b61891ba-2acd-45cf-ba8d-e0846530d678 · outbound

This paper cites Language is not all you need: Aligning perception with language mod- els.

SALOVA: Segment-Augmented Long Video Assistant for Targeted Retrieval and Routing in Long-Form Video Analysis Language is not all you need: Aligning perception with language mod- els

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-12T13:31:04.156057Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:31:04.156057Z digest=sha256:b31eda9d626dafedffeb692d98172dcf613af4f0e63a104bf1622ac79ed38b8f

Observation 3b625dcc-6d42-4325-9f5b-550aa38baf81 · outbound

This paper cites Chat-univi: Unified visual representation em- powers large language models with image and video un- derstanding.

SALOVA: Segment-Augmented Long Video Assistant for Targeted Retrieval and Routing in Long-Form Video Analysis Chat-univi: Unified visual representation em- powers large language models with image and video un- derstanding

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:31:06.380546Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T13:31:04.161739Z digest=sha256:1323fc918eef4149a6e17c4396d9bfdde3271435b0dbfba4a9b4c2461b5fc3b8

Observation ecb673a5-47a4-4bd5-815c-1231256f4f9b · outbound

This paper cites MiraData: A Large-Scale Video Dataset with Long Durations and Structured Captions.

SALOVA: Segment-Augmented Long Video Assistant for Targeted Retrieval and Routing in Long-Form Video Analysis MiraData: A Large-Scale Video Dataset with Long Durations and Structured Captions

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-12T13:31:04.168046Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:31:04.168046Z digest=sha256:882482621fe6be6b35818cd8df0435e1b763d95df3162d3df33188aab4c8902a

Observation af23ba6b-26fe-4df3-b4be-2cc75786ba3f · outbound

This paper cites Retrieval-augmented generation for knowledge-intensive nlp tasks.

SALOVA: Segment-Augmented Long Video Assistant for Targeted Retrieval and Routing in Long-Form Video Analysis Retrieval-augmented generation for knowledge-intensive nlp tasks

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-12T13:31:04.176091Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:31:04.176091Z digest=sha256:86c83c5ebe27e09e75ede3fa4e64272a9be8e4f17f37785edf91f526865eac26

Observation 4200f90e-e588-4b41-9218-d25a475f9c82 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

SALOVA: Segment-Augmented Long Video Assistant for Targeted Retrieval and Routing in Long-Form Video Analysis LLaVA-OneVision: Easy Visual Task Transfer

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-12T13:31:04.185703Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:31:04.185703Z digest=sha256:9544bdf404e1a7fc210f93fea2807c8fe9ebc69203f04a35bf87374edcf7cb7c

Observation afd45411-2347-471e-979c-258f4d30e4cb · outbound

This paper cites LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models.

SALOVA: Segment-Augmented Long Video Assistant for Targeted Retrieval and Routing in Long-Form Video Analysis LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-12T13:31:04.194172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:31:04.194172Z digest=sha256:e63f620ee82114ab4851883f90ac5443ccb47de6c1a24d03def357d695cd28f0

Observation ce9f6c74-54a6-4763-9dc7-d899d4e50287 · outbound

This paper cites Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models.

SALOVA: Segment-Augmented Long Video Assistant for Targeted Retrieval and Routing in Long-Form Video Analysis Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:31:06.356650Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T13:31:04.200577Z digest=sha256:585672088bdbf3190e37b26e5315a63b499a1a63035ab9685460d47e8341e169

Observation b6cf23d7-e0a9-4138-8760-5c6d1b56e995 · outbound

This paper cites Mvbench: A comprehensive multi-modal video understand- ing benchmark.

SALOVA: Segment-Augmented Long Video Assistant for Targeted Retrieval and Routing in Long-Form Video Analysis Mvbench: A comprehensive multi-modal video understand- ing benchmark

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:31:06.342554Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T13:31:04.206745Z digest=sha256:70c628f8fcbd81f841ef3879b4e97db5505a242d777950ae623a0b8496f55079

Observation 7617c986-c22f-4574-8bad-e999925e9145 · outbound

This paper cites Mo- mentdiff: Generative video moment retrieval from random to real.

SALOVA: Segment-Augmented Long Video Assistant for Targeted Retrieval and Routing in Long-Form Video Analysis Mo- mentdiff: Generative video moment retrieval from random to real

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:31:06.327010Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T13:31:04.223008Z digest=sha256:e4a3c33935061e43f14b0df171522fcfcea8035f46767903e7303beef9a63c63

Observation d3a3f9f2-fd55-4d69-886b-072e49de2732 · outbound

This paper cites Llama-vid: An image is worth 2 tokens in large language models.

SALOVA: Segment-Augmented Long Video Assistant for Targeted Retrieval and Routing in Long-Form Video Analysis Llama-vid: An image is worth 2 tokens in large language models

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:31:06.309061Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T13:31:04.230679Z digest=sha256:bc18414d985d0d9b08194847305426cac5f180f3195043a51605dd136c1381fd

Observation 98bd8e06-81c9-436b-8443-e79d3005ca83 · outbound

This paper cites Video-LLaVA: Learning United Visual Representation by Alignment Before Projection.

SALOVA: Segment-Augmented Long Video Assistant for Targeted Retrieval and Routing in Long-Form Video Analysis Video-LLaVA: Learning United Visual Representation by Alignment Before Projection

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-12T13:31:04.241041Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:31:04.241041Z digest=sha256:7021925325fdeedb3459391c15f98851963c9c6301fb329c921ff421e2287e10

Observation 92c44ada-7a89-4e0e-a43d-58809150ce82 · outbound

This paper cites Vila: On pre-training for vi- sual language models.

SALOVA: Segment-Augmented Long Video Assistant for Targeted Retrieval and Routing in Long-Form Video Analysis Vila: On pre-training for vi- sual language models

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:31:06.272272Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T13:31:04.246495Z digest=sha256:ed1ce318786332570eef37d183185f3c3dff8a74520df1af205f57382a732780

Observation c9180e20-dc9b-4d94-a9c9-7b2a3274d05c · outbound

This paper cites Improved Baselines with Visual Instruction Tuning.

SALOVA: Segment-Augmented Long Video Assistant for Targeted Retrieval and Routing in Long-Form Video Analysis Improved Baselines with Visual Instruction Tuning

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-12T13:31:04.252680Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:31:04.252680Z digest=sha256:a4b596a4ba5c8b17d02b902383213e3be72628535d8eeeafb5f4f770b1ac3700

Observation 187f7b3f-9c74-4ddc-ac85-d5c472fc4290 · outbound

This paper cites Visual instruction tuning.

SALOVA: Segment-Augmented Long Video Assistant for Targeted Retrieval and Routing in Long-Form Video Analysis Visual instruction tuning

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:31:06.253897Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T13:31:04.259329Z digest=sha256:1693c425afdcbf6b252baaa3c9d4a42788f4a2f483014d3ed4bf18608a01e90a

Observation 76d12409-f3aa-4272-bbf8-01b22871879c · outbound

This paper cites St-llm: Large language models are effective tem- poral learners.

SALOVA: Segment-Augmented Long Video Assistant for Targeted Retrieval and Routing in Long-Form Video Analysis St-llm: Large language models are effective tem- poral learners

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:31:06.231627Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T13:31:04.264431Z digest=sha256:bc4e0506139c16db26967cf01043247edc169c75aafbdd107a4bd8ab5d839b0f

Observation 72545038-e80b-4558-ac72-88e87644066f · outbound

This paper cites Sora: A Review on Background, Technology, Limitations, and Opportunities of Large Vision Models.

SALOVA: Segment-Augmented Long Video Assistant for Targeted Retrieval and Routing in Long-Form Video Analysis Sora: A Review on Background, Technology, Limitations, and Opportunities of Large Vision Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-12T13:31:04.276218Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:31:04.276218Z digest=sha256:139c9f69c0b8cb9ce1f336c6b24c62c78b95d895a227f6b1892d72d26d8a1b1b

Observation 5f7cf55e-fe58-4508-8046-b1e46a8e1da8 · outbound

This paper cites Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models.

SALOVA: Segment-Augmented Long Video Assistant for Targeted Retrieval and Routing in Long-Form Video Analysis Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-12T13:31:04.286311Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:31:04.286311Z digest=sha256:bf578a1e53aa210e3206c88f8c4ba791dc7c648dde52349298ed2d6b31a31c75

Observation 6c6c3ab3-4ad6-405e-8bfb-4be734a5e7a6 · outbound

This paper cites MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training.

SALOVA: Segment-Augmented Long Video Assistant for Targeted Retrieval and Routing in Long-Form Video Analysis MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-12T13:31:04.298444Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:31:04.298444Z digest=sha256:7c9e97d67e966dc932f5e926b8e06dfe52542d669b91256ca9ca568a2600b6d4

Observation c11bf21f-bd96-4bb1-8a9b-1f1e41bdafbb · outbound

This paper cites an unresolved cited work.

SALOVA: Segment-Augmented Long Video Assistant for Targeted Retrieval and Routing in Long-Form Video Analysis Unresolved cited work

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-12T13:31:04.314096Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:31:04.314096Z digest=sha256:1ff41606fbf25e37e35ee239308787d02bf0ad4c8478e9b062438485fc09bc46

Observation e1bff855-85a9-404b-9f39-3caf5e0eba79 · outbound

This paper cites Gpt-4 technical report, 2023.

SALOVA: Segment-Augmented Long Video Assistant for Targeted Retrieval and Routing in Long-Form Video Analysis Gpt-4 technical report, 2023

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:31:06.199637Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T13:31:04.321808Z digest=sha256:a4ba2ff93f19fd4408471638c220e324a09a5eef3e28635338d4b20cb3b23365

Observation 20d9bed8-ea03-41be-8302-c3beb6c69331 · outbound

This paper cites GPT-4V(ision) System Card, 2023.

SALOVA: Segment-Augmented Long Video Assistant for Targeted Retrieval and Routing in Long-Form Video Analysis GPT-4V(ision) System Card, 2023

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-12T13:31:04.330818Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:31:04.330818Z digest=sha256:b0f24b05a992c2a759d40f97bdb078f55b8cf3daf178c0712f591a47cd397363

Observation 78b39bb4-a728-49bb-966f-22c22a91afcd · outbound

This paper cites Hello gpt-4o, 2024.

SALOVA: Segment-Augmented Long Video Assistant for Targeted Retrieval and Routing in Long-Form Video Analysis Hello gpt-4o, 2024

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:31:06.175294Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T13:31:04.339055Z digest=sha256:a6999c58967e89405c277072d1306bcbedbe74dd34b3656daadf650342095c21

Observation dedd7203-000c-475d-b2ea-113194113098 · outbound

This paper cites Per- ception test: A diagnostic benchmark for multimodal video models.

SALOVA: Segment-Augmented Long Video Assistant for Targeted Retrieval and Routing in Long-Form Video Analysis Per- ception test: A diagnostic benchmark for multimodal video models

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:31:06.160148Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T13:31:04.348248Z digest=sha256:59dafedb7eb060e1059fd87beb51fc0ff911ef7ef81d40c714b72eea498a7924

Observation 8f7b9ade-22f0-46ad-af56-680ccbc8cbae · outbound

This paper cites Learning transferable visual models from natural language supervi- sion.

SALOVA: Segment-Augmented Long Video Assistant for Targeted Retrieval and Routing in Long-Form Video Analysis Learning transferable visual models from natural language supervi- sion

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:31:06.138060Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T13:31:04.355067Z digest=sha256:5a144905168b1296950573b5da7496bb50507b1781e2fdd9abbfe62ee2addf5b

Observation f8ece6f2-1ed1-4ca2-bd29-0a08c526bacf · outbound

This paper cites Zero: Memory optimizations toward training trillion parameter models.

SALOVA: Segment-Augmented Long Video Assistant for Targeted Retrieval and Routing in Long-Form Video Analysis Zero: Memory optimizations toward training trillion parameter models

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:31:06.118763Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T13:31:04.364815Z digest=sha256:d55e67a90dc53fcb53d359359fbfb59fb452546d45e7437eb04e3ea654c19d5b

Observation 8fa5a5ea-fb52-465e-8c49-5783e0c1999b · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

SALOVA: Segment-Augmented Long Video Assistant for Targeted Retrieval and Routing in Long-Form Video Analysis Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-12T13:31:04.377417Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:31:04.377417Z digest=sha256:d2f9d0b69a244b0e5b83071429b82d082381a1f7297e2e738850fe0fd0c6d542

Observation 487b258f-0b1d-4602-ab57-21fe59e4f393 · outbound

This paper cites Sharegemini: Scaling up video caption data for mul- timodal large language models, 2024.

SALOVA: Segment-Augmented Long Video Assistant for Targeted Retrieval and Routing in Long-Form Video Analysis Sharegemini: Scaling up video caption data for mul- timodal large language models, 2024

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:31:06.090774Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T13:31:04.387877Z digest=sha256:bf68d2e0d6bf461379f9a98a4aeda0b50bef7d06fad3cc1e86a4d1d7cca418ec

Observation 75f9b14d-e66a-46ab-b856-546296abac5e · outbound

This paper cites Conceptual captions: A cleaned, hypernymed, im- age alt-text dataset for automatic image captioning.

SALOVA: Segment-Augmented Long Video Assistant for Targeted Retrieval and Routing in Long-Form Video Analysis Conceptual captions: A cleaned, hypernymed, im- age alt-text dataset for automatic image captioning

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-12T13:31:04.396335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:31:04.396335Z digest=sha256:ed3e58bfb95a90242a31ee002726b49598c3ec5b733f70de42048ac5480bd6e3

Observation 3340ab37-0b6a-41dd-8b98-5c21220e616d · outbound

This paper cites Moviechat: From dense token to sparse memory for long video understanding.

SALOVA: Segment-Augmented Long Video Assistant for Targeted Retrieval and Routing in Long-Form Video Analysis Moviechat: From dense token to sparse memory for long video understanding

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:31:06.043356Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T13:31:04.405996Z digest=sha256:b0ff1f5bfe71f3a14eb8f0d31fa6a8f8e48d2e9c951c60b65502e80b871d9d5a

Observation 0b280126-5486-4eda-a7f5-bc4ad7429f77 · outbound

This paper cites Augmented SBERT: Data augmentation method for improving bi-encoders for pairwise sentence scoring tasks.

SALOVA: Segment-Augmented Long Video Assistant for Targeted Retrieval and Routing in Long-Form Video Analysis Augmented SBERT: Data augmentation method for improving bi-encoders for pairwise sentence scoring tasks

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:31:06.016300Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T13:31:04.418605Z digest=sha256:20943891446c9b7892b40c2a7733610111d4cdf736fc1da03e227eade91992ae

Observation e2b0f3f6-6c41-4eab-9d38-bcf90bae6fd5 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

SALOVA: Segment-Augmented Long Video Assistant for Targeted Retrieval and Routing in Long-Form Video Analysis LLaMA: Open and Efficient Foundation Language Models

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-12T13:31:04.434969Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:31:04.434969Z digest=sha256:5a9ca3b58add0ecbcb6426a48128887e5040be571f8d96406585237bdd678850

Observation 67f76ad6-0355-494f-a814-30bded0c0e2c · outbound

This paper cites InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation.

SALOVA: Segment-Augmented Long Video Assistant for Targeted Retrieval and Routing in Long-Form Video Analysis InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-12T13:31:04.453981Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:31:04.453981Z digest=sha256:1dd14d79ce38c528ddfefcabc8b423e2c3dde139e7fecc5ac181829011e960ed

Observation 3d5832b1-e6aa-43c5-9750-01e677997025 · outbound

This paper cites LongVideoBench: A Benchmark for Long-context Interleaved Video-Language Understanding.

SALOVA: Segment-Augmented Long Video Assistant for Targeted Retrieval and Routing in Long-Form Video Analysis LongVideoBench: A Benchmark for Long-context Interleaved Video-Language Understanding

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-12T13:31:04.464160Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:31:04.464160Z digest=sha256:2e3692a675c5eec687affd5aab400bb5bcebec39859e083234c5678c5ba719c6

Observation 5995798c-f68f-4787-b17f-0f1239bca1b3 · outbound

This paper cites Next-qa: Next phase of question-answering to explaining temporal actions.

SALOVA: Segment-Augmented Long Video Assistant for Targeted Retrieval and Routing in Long-Form Video Analysis Next-qa: Next phase of question-answering to explaining temporal actions

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-12T13:31:04.471598Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:31:04.471598Z digest=sha256:d0f77a941710a8f757cfc6012b0a3efe502eb1a3850eec46ffb126966869d9ff

Observation 5b6bea90-eff3-4416-bbed-44242b21081a · outbound

This paper cites PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning.

SALOVA: Segment-Augmented Long Video Assistant for Targeted Retrieval and Routing in Long-Form Video Analysis PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-12T13:31:04.481024Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:31:04.481024Z digest=sha256:b4a610a8facbcce19493cb82c07221c9ff72241ba317e6ef4712aea7a4c39f99

Observation c536aed2-89b0-4e9b-aff9-8eade030e14f · outbound

This paper cites SlowFast-LLaVA: A Strong Training-Free Baseline for Video Large Language Models.

SALOVA: Segment-Augmented Long Video Assistant for Targeted Retrieval and Routing in Long-Form Video Analysis SlowFast-LLaVA: A Strong Training-Free Baseline for Video Large Language Models

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-12T13:31:04.488520Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:31:04.488520Z digest=sha256:4583e2c72b36d06675a695a65d398f05a511de34a8cffaca2f5d85909883ae24

Observation b221bf69-f184-4f8b-b50b-59585d1ed87a · outbound

This paper cites LongVILA: Scaling Long-Context Visual Language Models for Long Videos.

SALOVA: Segment-Augmented Long Video Assistant for Targeted Retrieval and Routing in Long-Form Video Analysis LongVILA: Scaling Long-Context Visual Language Models for Long Videos

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-12T13:31:04.502907Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:31:04.502907Z digest=sha256:60e2e12cd88395ca9ca978498c99b6972b8e2bc7c5d296f6de6c7232c9302d06

Observation b924dcfb-5840-45b1-b5e0-9d4e77e949c0 · outbound

This paper cites Ad- vancing high-resolution video-language representation with large-scale video transcriptions.

SALOVA: Segment-Augmented Long Video Assistant for Targeted Retrieval and Routing in Long-Form Video Analysis Ad- vancing high-resolution video-language representation with large-scale video transcriptions

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-12T13:31:04.527268Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:31:04.527268Z digest=sha256:78edf8150be844e83257edf7c6ff9b5d1970852f2977cb8447655f4d65339db0

Observation 08fda050-74fa-4653-b030-5d8c6dd087dd · outbound

This paper cites Qwen2 Technical Report.

SALOVA: Segment-Augmented Long Video Assistant for Targeted Retrieval and Routing in Long-Form Video Analysis Qwen2 Technical Report

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-12T13:31:04.539324Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:31:04.539324Z digest=sha256:1b653e27d2f730df54f9d03f7a520d13f2c7dc6e102c2a616d7c379e91b635cf

Observation 59bda63b-3db2-46da-b9b5-a12b5ca0a35f · outbound

This paper cites mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models.

SALOVA: Segment-Augmented Long Video Assistant for Targeted Retrieval and Routing in Long-Form Video Analysis mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-12T13:31:04.551482Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:31:04.551482Z digest=sha256:48a69f6b87e1c7f1cdbc06145a2d26c157289a449f33e617455e1eb0706304d3

Observation 36a8728f-5a81-4d03-8a19-4fdce3b012e3 · outbound

This paper cites mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality.

SALOVA: Segment-Augmented Long Video Assistant for Targeted Retrieval and Routing in Long-Form Video Analysis mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-12T13:31:04.561169Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:31:04.561169Z digest=sha256:ed18a8cdba7b9b57f45612d155617cc08b12f72065c66b88cef728ad2c6bfb81

Observation 37140d12-d812-432d-b6a3-c34f6c01b849 · outbound

This paper cites Activitynet-qa: A dataset for understanding complex web videos via question answering.

SALOVA: Segment-Augmented Long Video Assistant for Targeted Retrieval and Routing in Long-Form Video Analysis Activitynet-qa: A dataset for understanding complex web videos via question answering

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:31:05.958307Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T13:31:04.585042Z digest=sha256:c131b5443b1e0d047484e56c871ed82f88902bef4bd8587df4f38a8b289ef326

Observation 1d21fcc6-94b7-4433-acbb-c6a258cca652 · outbound

This paper cites Sigmoid loss for language image pre-training.

SALOVA: Segment-Augmented Long Video Assistant for Targeted Retrieval and Routing in Long-Form Video Analysis Sigmoid loss for language image pre-training

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:31:05.937092Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T13:31:04.601541Z digest=sha256:c9dd7113cf6ed08f1af65e23e02814d05d7c04ced1a1bc3fe0b38118df7368f7

Observation 8edc5634-90ce-4f6f-91ef-5363bcf79dbd · outbound

This paper cites Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding.

SALOVA: Segment-Augmented Long Video Assistant for Targeted Retrieval and Routing in Long-Form Video Analysis Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-12T13:31:04.607338Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:31:04.607338Z digest=sha256:253363e500e528c4b3178fd21c467167262354213b1827e3d25dfba3675e605c

Observation 843cf036-d434-4165-8927-fece6d5dee9f · outbound

This paper cites InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output.

SALOVA: Segment-Augmented Long Video Assistant for Targeted Retrieval and Routing in Long-Form Video Analysis InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-12T13:31:04.614880Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:31:04.614880Z digest=sha256:3046aa574e5112b8b4bf9685cb4564a2dbd12d052c8b83bc98e4c929d9b30b1f

Observation 4c4f305f-2f17-4920-8a54-88212773c02c · outbound

This paper cites Long Context Transfer from Language to Vision.

SALOVA: Segment-Augmented Long Video Assistant for Targeted Retrieval and Routing in Long-Form Video Analysis Long Context Transfer from Language to Vision

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-12T13:31:04.620891Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:31:04.620891Z digest=sha256:cb1eb6e2a8c0dd178e7bbd66739d647b608060dde920dea13212a62e6ae9124d

Observation 857e9d88-3a7e-43e0-9160-689b506dcb79 · outbound

This paper cites LLaVA-Video: Video Instruction Tuning With Synthetic Data.

SALOVA: Segment-Augmented Long Video Assistant for Targeted Retrieval and Routing in Long-Form Video Analysis LLaVA-Video: Video Instruction Tuning With Synthetic Data

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-12T13:31:04.629482Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:31:04.629482Z digest=sha256:0e3e6f136576e1003fd3caa0bb623c8180eaee7758fb5bd4f7bf5f35b55b5672

Observation eec66a7f-c702-4e86-bd74-b72d55c96ba4 · outbound

This paper cites Beyond LLaVA-HD: Diving into High-Resolution Large Multimodal Models.

SALOVA: Segment-Augmented Long Video Assistant for Targeted Retrieval and Routing in Long-Form Video Analysis Beyond LLaVA-HD: Diving into High-Resolution Large Multimodal Models

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-12T13:31:04.636677Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:31:04.636677Z digest=sha256:b0eaf7d8c9a79605a0a0a5aaba00360a4edb0138ed78ec96ff52b4c503fcb486

Observation 7f9b90d5-68b8-43b7-8f48-91d4183d0a95 · outbound

This paper cites Provide a detailed description of both the visual content and the storyline depicted in the video.

SALOVA: Segment-Augmented Long Video Assistant for Targeted Retrieval and Routing in Long-Form Video Analysis Provide a detailed description of both the visual content and the storyline depicted in the video

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:31:05.907211Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T13:31:04.642937Z digest=sha256:3adb07380efa6e4ae1746d25d517d6c838df4dfebba5c8581d7b6eb4b90d3588

Pith citing papers

Observation e2fbcb1b-fc7c-481e-ad04-a3026aca8db2 · inbound

Look Every Frame All at Once: Video-Ma$^2$mba for Efficient Long-form Video Understanding with Multi-Axis Gradient Checkpointing cites this paper.

Look Every Frame All at Once: Video-Ma$^2$mba for Efficient Long-form Video Understanding with Multi-Axis Gradient Checkpointing SALOVA: Segment-Augmented Long Video Assistant for Targeted Retrieval and Routing in Long-Form Video Analysis

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-12T10:18:48.890691Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:18:48.890691Z digest=sha256:d1373e347b7d964be958dc35ef9626111a5a7d80f6c4c92f0b894c0396ccd294

Observation 3a68fc8c-0d7f-4210-b4c3-d18e7ecdc0d4 · inbound

ReFoCUS: Reinforcement-guided Frame Optimization for Contextual Understanding cites this paper.

ReFoCUS: Reinforcement-guided Frame Optimization for Contextual Understanding SALOVA: Segment-Augmented Long Video Assistant for Targeted Retrieval and Routing in Long-Form Video Analysis

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T11:51:26.782912Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:51:26.782912Z digest=sha256:3ed93da07a55e3e1bfbc6f37a4cf33e1655b5bf45872e043a6c80aa0f0995372

Observation 305c3027-cfc9-4e7a-8d23-f314ba2448b6 · inbound

GRASP: Learning to Ground Social Reasoning in Multi-Person Non-Verbal Interactions cites this paper.

GRASP: Learning to Ground Social Reasoning in Multi-Person Non-Verbal Interactions SALOVA: Segment-Augmented Long Video Assistant for Targeted Retrieval and Routing in Long-Form Video Analysis

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-20T18:53:38.997930Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-20T18:49:18.815456Z digest=sha256:28dcb27846df4580a21765a9aefe9c653b4081e027f5eebe9714cd12196223f9