Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-12T13:31:04.642937Z
Paper Citation Record · LEDGER
As of 20 August 2026, this Paper Citation Record lists 73 of 73 outbound references and 3 inbound Pith citation observations for arXiv:2411.16173.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-12T13:31:04.642937Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-12T10:18:48.890691Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-20T18:53:38.996149Z
73 of 73 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 57e231a7-ea8f-496b-a33d-a6ae56e950b3 · outbound
SALOVA: Segment-Augmented Long Video Assistant for Targeted Retrieval and Routing in Long-Form Video Analysis Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f30c46be-108a-4b7c-a23a-05c66c022612 · outbound
SALOVA: Segment-Augmented Long Video Assistant for Targeted Retrieval and Routing in Long-Form Video Analysis Flamingo: a visual language model for few-shot learning
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 35b0cc40-492e-4f56-b313-504ba88fe11d · outbound
SALOVA: Segment-Augmented Long Video Assistant for Targeted Retrieval and Routing in Long-Form Video Analysis Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 58727375-4b88-4fce-87d1-500447a09a19 · outbound
SALOVA: Segment-Augmented Long Video Assistant for Targeted Retrieval and Routing in Long-Form Video Analysis Frozen in time: A joint video and image encoder for end-to-end retrieval
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 70d5b1df-bbd2-4b14-ad30-864834dd6e21 · outbound
SALOVA: Segment-Augmented Long Video Assistant for Targeted Retrieval and Routing in Long-Form Video Analysis Lan- guage models are few-shot learners
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aeb4c7d7-ef0a-43a8-9e10-a9bd27aab885 · outbound
SALOVA: Segment-Augmented Long Video Assistant for Targeted Retrieval and Routing in Long-Form Video Analysis Activitynet: A large-scale video benchmark for human activity understanding
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 9b4856bd-ad5b-49d8-9f3b-c6ebb171409e · outbound
SALOVA: Segment-Augmented Long Video Assistant for Targeted Retrieval and Routing in Long-Form Video Analysis Honeybee: Locality-enhanced Projector for Multimodal LLM
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d115cd80-01c9-43a4-be4d-04a3b491d210 · outbound
SALOVA: Segment-Augmented Long Video Assistant for Targeted Retrieval and Routing in Long-Form Video Analysis ShareGPT4Video: Improving Video Understanding and Generation with Better Captions
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a7d7b95f-b852-4a5c-9dd7-ac6c129150fa · outbound
SALOVA: Segment-Augmented Long Video Assistant for Targeted Retrieval and Routing in Long-Form Video Analysis Training Deep Nets with Sublinear Memory Cost
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f5ebbed0-bea7-42cd-9ccf-14966dd6a3f3 · outbound
SALOVA: Segment-Augmented Long Video Assistant for Targeted Retrieval and Routing in Long-Form Video Analysis Panda-70m: Captioning 70m videos with multiple cross-modality teachers
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c2f1a3ba-b6cd-4faa-9ae6-b73478f4aff1 · outbound
SALOVA: Segment-Augmented Long Video Assistant for Targeted Retrieval and Routing in Long-Form Video Analysis InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e3a92d58-ee71-4499-ac88-fd22ed3be7b1 · outbound
SALOVA: Segment-Augmented Long Video Assistant for Targeted Retrieval and Routing in Long-Form Video Analysis VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7a423278-109c-4505-a66e-7b863c51a2cd · outbound
SALOVA: Segment-Augmented Long Video Assistant for Targeted Retrieval and Routing in Long-Form Video Analysis Gonzalez, Ion Stoica, and Eric P
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1531c6b0-7512-4441-b98c-a26d17260198 · outbound
SALOVA: Segment-Augmented Long Video Assistant for Targeted Retrieval and Routing in Long-Form Video Analysis InstructBLIP: Towards general-purpose vision- language models with instruction tuning
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation e98ad8a7-fb69-4b87-8731-3af104786d0a · outbound
SALOVA: Segment-Augmented Long Video Assistant for Targeted Retrieval and Routing in Long-Form Video Analysis FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d2c8bc61-368b-4447-b94a-6ff9e3e29571 · outbound
SALOVA: Segment-Augmented Long Video Assistant for Targeted Retrieval and Routing in Long-Form Video Analysis Patch n’pack: Navit, a vision transformer for any aspect ratio and resolution
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 9855ebe5-7360-4b12-bc9a-c8eefe3d9b8d · outbound
SALOVA: Segment-Augmented Long Video Assistant for Targeted Retrieval and Routing in Long-Form Video Analysis InternLM-XComposer2: Mastering Free-form Text-Image Composition and Comprehension in Vision-Language Large Model
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a8d85cd7-4d49-49e3-859d-04e545fedc8c · outbound
SALOVA: Segment-Augmented Long Video Assistant for Targeted Retrieval and Routing in Long-Form Video Analysis The Llama 3 Herd of Models
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a9b923ba-1570-4c45-a686-8919fb0580a9 · outbound
SALOVA: Segment-Augmented Long Video Assistant for Targeted Retrieval and Routing in Long-Form Video Analysis Slowfast networks for video recognition
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation a483f1a2-da3e-4714-808b-17345e50c169 · outbound
SALOVA: Segment-Augmented Long Video Assistant for Targeted Retrieval and Routing in Long-Form Video Analysis Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 31425c59-ec50-4615-a346-5b13beed029b · outbound
SALOVA: Segment-Augmented Long Video Assistant for Targeted Retrieval and Routing in Long-Form Video Analysis Long Story Short: Story-level Video Understanding from 20K Short Films
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 593182dd-316b-4b18-a780-637c6680b74a · outbound
SALOVA: Segment-Augmented Long Video Assistant for Targeted Retrieval and Routing in Long-Form Video Analysis Gemini, 2023
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 67b4aa85-c730-4367-b10c-e3009ae6c505 · outbound
SALOVA: Segment-Augmented Long Video Assistant for Targeted Retrieval and Routing in Long-Form Video Analysis Ma-lmm: Memory-augmented large multimodal model for long-term video understanding
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 955a6b36-8383-4a74-8f11-c3f770b2fc6a · outbound
SALOVA: Segment-Augmented Long Video Assistant for Targeted Retrieval and Routing in Long-Form Video Analysis LoRA: Low-Rank Adaptation of Large Language Models
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b61891ba-2acd-45cf-ba8d-e0846530d678 · outbound
SALOVA: Segment-Augmented Long Video Assistant for Targeted Retrieval and Routing in Long-Form Video Analysis Language is not all you need: Aligning perception with language mod- els
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3b625dcc-6d42-4325-9f5b-550aa38baf81 · outbound
SALOVA: Segment-Augmented Long Video Assistant for Targeted Retrieval and Routing in Long-Form Video Analysis Chat-univi: Unified visual representation em- powers large language models with image and video un- derstanding
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation ecb673a5-47a4-4bd5-815c-1231256f4f9b · outbound
SALOVA: Segment-Augmented Long Video Assistant for Targeted Retrieval and Routing in Long-Form Video Analysis MiraData: A Large-Scale Video Dataset with Long Durations and Structured Captions
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation af23ba6b-26fe-4df3-b4be-2cc75786ba3f · outbound
SALOVA: Segment-Augmented Long Video Assistant for Targeted Retrieval and Routing in Long-Form Video Analysis Retrieval-augmented generation for knowledge-intensive nlp tasks
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4200f90e-e588-4b41-9218-d25a475f9c82 · outbound
SALOVA: Segment-Augmented Long Video Assistant for Targeted Retrieval and Routing in Long-Form Video Analysis LLaVA-OneVision: Easy Visual Task Transfer
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation afd45411-2347-471e-979c-258f4d30e4cb · outbound
SALOVA: Segment-Augmented Long Video Assistant for Targeted Retrieval and Routing in Long-Form Video Analysis LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ce9f6c74-54a6-4763-9dc7-d899d4e50287 · outbound
SALOVA: Segment-Augmented Long Video Assistant for Targeted Retrieval and Routing in Long-Form Video Analysis Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation b6cf23d7-e0a9-4138-8760-5c6d1b56e995 · outbound
SALOVA: Segment-Augmented Long Video Assistant for Targeted Retrieval and Routing in Long-Form Video Analysis Mvbench: A comprehensive multi-modal video understand- ing benchmark
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 7617c986-c22f-4574-8bad-e999925e9145 · outbound
SALOVA: Segment-Augmented Long Video Assistant for Targeted Retrieval and Routing in Long-Form Video Analysis Mo- mentdiff: Generative video moment retrieval from random to real
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation d3a3f9f2-fd55-4d69-886b-072e49de2732 · outbound
SALOVA: Segment-Augmented Long Video Assistant for Targeted Retrieval and Routing in Long-Form Video Analysis Llama-vid: An image is worth 2 tokens in large language models
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 98bd8e06-81c9-436b-8443-e79d3005ca83 · outbound
SALOVA: Segment-Augmented Long Video Assistant for Targeted Retrieval and Routing in Long-Form Video Analysis Video-LLaVA: Learning United Visual Representation by Alignment Before Projection
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 92c44ada-7a89-4e0e-a43d-58809150ce82 · outbound
SALOVA: Segment-Augmented Long Video Assistant for Targeted Retrieval and Routing in Long-Form Video Analysis Vila: On pre-training for vi- sual language models
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation c9180e20-dc9b-4d94-a9c9-7b2a3274d05c · outbound
SALOVA: Segment-Augmented Long Video Assistant for Targeted Retrieval and Routing in Long-Form Video Analysis Improved Baselines with Visual Instruction Tuning
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 187f7b3f-9c74-4ddc-ac85-d5c472fc4290 · outbound
SALOVA: Segment-Augmented Long Video Assistant for Targeted Retrieval and Routing in Long-Form Video Analysis Visual instruction tuning
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 76d12409-f3aa-4272-bbf8-01b22871879c · outbound
SALOVA: Segment-Augmented Long Video Assistant for Targeted Retrieval and Routing in Long-Form Video Analysis St-llm: Large language models are effective tem- poral learners
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 72545038-e80b-4558-ac72-88e87644066f · outbound
SALOVA: Segment-Augmented Long Video Assistant for Targeted Retrieval and Routing in Long-Form Video Analysis Sora: A Review on Background, Technology, Limitations, and Opportunities of Large Vision Models
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5f7cf55e-fe58-4508-8046-b1e46a8e1da8 · outbound
SALOVA: Segment-Augmented Long Video Assistant for Targeted Retrieval and Routing in Long-Form Video Analysis Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6c6c3ab3-4ad6-405e-8bfb-4be734a5e7a6 · outbound
SALOVA: Segment-Augmented Long Video Assistant for Targeted Retrieval and Routing in Long-Form Video Analysis MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c11bf21f-bd96-4bb1-8a9b-1f1e41bdafbb · outbound
SALOVA: Segment-Augmented Long Video Assistant for Targeted Retrieval and Routing in Long-Form Video Analysis Unresolved cited work
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e1bff855-85a9-404b-9f39-3caf5e0eba79 · outbound
SALOVA: Segment-Augmented Long Video Assistant for Targeted Retrieval and Routing in Long-Form Video Analysis Gpt-4 technical report, 2023
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 20d9bed8-ea03-41be-8302-c3beb6c69331 · outbound
SALOVA: Segment-Augmented Long Video Assistant for Targeted Retrieval and Routing in Long-Form Video Analysis GPT-4V(ision) System Card, 2023
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 78b39bb4-a728-49bb-966f-22c22a91afcd · outbound
SALOVA: Segment-Augmented Long Video Assistant for Targeted Retrieval and Routing in Long-Form Video Analysis Hello gpt-4o, 2024
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation dedd7203-000c-475d-b2ea-113194113098 · outbound
SALOVA: Segment-Augmented Long Video Assistant for Targeted Retrieval and Routing in Long-Form Video Analysis Per- ception test: A diagnostic benchmark for multimodal video models
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 8f7b9ade-22f0-46ad-af56-680ccbc8cbae · outbound
SALOVA: Segment-Augmented Long Video Assistant for Targeted Retrieval and Routing in Long-Form Video Analysis Learning transferable visual models from natural language supervi- sion
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation f8ece6f2-1ed1-4ca2-bd29-0a08c526bacf · outbound
SALOVA: Segment-Augmented Long Video Assistant for Targeted Retrieval and Routing in Long-Form Video Analysis Zero: Memory optimizations toward training trillion parameter models
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 8fa5a5ea-fb52-465e-8c49-5783e0c1999b · outbound
SALOVA: Segment-Augmented Long Video Assistant for Targeted Retrieval and Routing in Long-Form Video Analysis Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 487b258f-0b1d-4602-ab57-21fe59e4f393 · outbound
SALOVA: Segment-Augmented Long Video Assistant for Targeted Retrieval and Routing in Long-Form Video Analysis Sharegemini: Scaling up video caption data for mul- timodal large language models, 2024
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 75f9b14d-e66a-46ab-b856-546296abac5e · outbound
SALOVA: Segment-Augmented Long Video Assistant for Targeted Retrieval and Routing in Long-Form Video Analysis Conceptual captions: A cleaned, hypernymed, im- age alt-text dataset for automatic image captioning
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3340ab37-0b6a-41dd-8b98-5c21220e616d · outbound
SALOVA: Segment-Augmented Long Video Assistant for Targeted Retrieval and Routing in Long-Form Video Analysis Moviechat: From dense token to sparse memory for long video understanding
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 0b280126-5486-4eda-a7f5-bc4ad7429f77 · outbound
SALOVA: Segment-Augmented Long Video Assistant for Targeted Retrieval and Routing in Long-Form Video Analysis Augmented SBERT: Data augmentation method for improving bi-encoders for pairwise sentence scoring tasks
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation e2b0f3f6-6c41-4eab-9d38-bcf90bae6fd5 · outbound
SALOVA: Segment-Augmented Long Video Assistant for Targeted Retrieval and Routing in Long-Form Video Analysis LLaMA: Open and Efficient Foundation Language Models
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 67f76ad6-0355-494f-a814-30bded0c0e2c · outbound
SALOVA: Segment-Augmented Long Video Assistant for Targeted Retrieval and Routing in Long-Form Video Analysis InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3d5832b1-e6aa-43c5-9750-01e677997025 · outbound
SALOVA: Segment-Augmented Long Video Assistant for Targeted Retrieval and Routing in Long-Form Video Analysis LongVideoBench: A Benchmark for Long-context Interleaved Video-Language Understanding
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5995798c-f68f-4787-b17f-0f1239bca1b3 · outbound
SALOVA: Segment-Augmented Long Video Assistant for Targeted Retrieval and Routing in Long-Form Video Analysis Next-qa: Next phase of question-answering to explaining temporal actions
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5b6bea90-eff3-4416-bbed-44242b21081a · outbound
SALOVA: Segment-Augmented Long Video Assistant for Targeted Retrieval and Routing in Long-Form Video Analysis PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c536aed2-89b0-4e9b-aff9-8eade030e14f · outbound
SALOVA: Segment-Augmented Long Video Assistant for Targeted Retrieval and Routing in Long-Form Video Analysis SlowFast-LLaVA: A Strong Training-Free Baseline for Video Large Language Models
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b221bf69-f184-4f8b-b50b-59585d1ed87a · outbound
SALOVA: Segment-Augmented Long Video Assistant for Targeted Retrieval and Routing in Long-Form Video Analysis LongVILA: Scaling Long-Context Visual Language Models for Long Videos
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b924dcfb-5840-45b1-b5e0-9d4e77e949c0 · outbound
SALOVA: Segment-Augmented Long Video Assistant for Targeted Retrieval and Routing in Long-Form Video Analysis Ad- vancing high-resolution video-language representation with large-scale video transcriptions
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 08fda050-74fa-4653-b030-5d8c6dd087dd · outbound
SALOVA: Segment-Augmented Long Video Assistant for Targeted Retrieval and Routing in Long-Form Video Analysis Qwen2 Technical Report
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 59bda63b-3db2-46da-b9b5-a12b5ca0a35f · outbound
SALOVA: Segment-Augmented Long Video Assistant for Targeted Retrieval and Routing in Long-Form Video Analysis mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 36a8728f-5a81-4d03-8a19-4fdce3b012e3 · outbound
SALOVA: Segment-Augmented Long Video Assistant for Targeted Retrieval and Routing in Long-Form Video Analysis mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 37140d12-d812-432d-b6a3-c34f6c01b849 · outbound
SALOVA: Segment-Augmented Long Video Assistant for Targeted Retrieval and Routing in Long-Form Video Analysis Activitynet-qa: A dataset for understanding complex web videos via question answering
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 1d21fcc6-94b7-4433-acbb-c6a258cca652 · outbound
SALOVA: Segment-Augmented Long Video Assistant for Targeted Retrieval and Routing in Long-Form Video Analysis Sigmoid loss for language image pre-training
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 8edc5634-90ce-4f6f-91ef-5363bcf79dbd · outbound
SALOVA: Segment-Augmented Long Video Assistant for Targeted Retrieval and Routing in Long-Form Video Analysis Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding
Reference 68
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 843cf036-d434-4165-8927-fece6d5dee9f · outbound
SALOVA: Segment-Augmented Long Video Assistant for Targeted Retrieval and Routing in Long-Form Video Analysis InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output
Reference 69
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4c4f305f-2f17-4920-8a54-88212773c02c · outbound
SALOVA: Segment-Augmented Long Video Assistant for Targeted Retrieval and Routing in Long-Form Video Analysis Long Context Transfer from Language to Vision
Reference 70
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 857e9d88-3a7e-43e0-9160-689b506dcb79 · outbound
SALOVA: Segment-Augmented Long Video Assistant for Targeted Retrieval and Routing in Long-Form Video Analysis LLaVA-Video: Video Instruction Tuning With Synthetic Data
Reference 71
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eec66a7f-c702-4e86-bd74-b72d55c96ba4 · outbound
SALOVA: Segment-Augmented Long Video Assistant for Targeted Retrieval and Routing in Long-Form Video Analysis Beyond LLaVA-HD: Diving into High-Resolution Large Multimodal Models
Reference 72
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7f9b90d5-68b8-43b7-8f48-91d4183d0a95 · outbound
SALOVA: Segment-Augmented Long Video Assistant for Targeted Retrieval and Routing in Long-Form Video Analysis Provide a detailed description of both the visual content and the storyline depicted in the video
Reference 73
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation e2fbcb1b-fc7c-481e-ad04-a3026aca8db2 · inbound
Look Every Frame All at Once: Video-Ma$^2$mba for Efficient Long-form Video Understanding with Multi-Axis Gradient Checkpointing SALOVA: Segment-Augmented Long Video Assistant for Targeted Retrieval and Routing in Long-Form Video Analysis
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3a68fc8c-0d7f-4210-b4c3-d18e7ecdc0d4 · inbound
ReFoCUS: Reinforcement-guided Frame Optimization for Contextual Understanding SALOVA: Segment-Augmented Long Video Assistant for Targeted Retrieval and Routing in Long-Form Video Analysis
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 305c3027-cfc9-4e7a-8d23-f314ba2448b6 · inbound
GRASP: Learning to Ground Social Reasoning in Multi-Person Non-Verbal Interactions SALOVA: Segment-Augmented Long Video Assistant for Targeted Retrieval and Routing in Long-Form Video Analysis
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.