Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-12T19:03:32.584154Z
Paper Citation Record · LEDGER
As of 14 August 2026, this Paper Citation Record lists 56 of 56 outbound references and 7 inbound Pith citation observations for arXiv:2411.11066.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-12T19:03:32.584154Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-06T22:37:38.250264Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z
56 of 56 outbound references displayed
External citation measurements
25
arxiv_reference, observed 2026-08-05T02:28:24.338817Z
Observation 49d949da-4101-48f5-88b1-b60044caf604 · outbound
TS-LLaVA: Constructing Visual Tokens through Thumbnail-and-Sampling for Training-Free Video Large Language Models Qwen Technical Report
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8cf08a90-6df7-4ff9-8eb8-214db30a03f9 · outbound
TS-LLaVA: Constructing Visual Tokens through Thumbnail-and-Sampling for Training-Free Video Large Language Models Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 25aafd0b-38a8-4169-9d1f-998625a32d2d · outbound
TS-LLaVA: Constructing Visual Tokens through Thumbnail-and-Sampling for Training-Free Video Large Language Models Matryoshka Multimodal Models
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 063c4c28-daa6-45c7-b103-b0989029fd18 · outbound
TS-LLaVA: Constructing Visual Tokens through Thumbnail-and-Sampling for Training-Free Video Large Language Models Collecting highly parallel data for paraphrase evaluation
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation b467affe-0ee3-4eef-963e-26673668d44d · outbound
TS-LLaVA: Constructing Visual Tokens through Thumbnail-and-Sampling for Training-Free Video Large Language Models VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a3fc6f6b-dba2-4b7c-a953-ed9cda08e7cd · outbound
TS-LLaVA: Constructing Visual Tokens through Thumbnail-and-Sampling for Training-Free Video Large Language Models Gonzalez, Ion Stoica, and Eric P
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 86dfe8fb-1ca0-44fa-87f5-658e4ade1cf3 · outbound
TS-LLaVA: Constructing Visual Tokens through Thumbnail-and-Sampling for Training-Free Video Large Language Models InstructBLIP: Towards general-purpose vision-language models with instruction tuning
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d835c0fa-309c-44fd-b5cd-7cf8ca2383b5 · outbound
TS-LLaVA: Constructing Visual Tokens through Thumbnail-and-Sampling for Training-Free Video Large Language Models An image is worth 16x16 words: Transformers for image recognition at scale
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a6da4ce0-5e6d-4df6-b0ae-959caac389f5 · outbound
TS-LLaVA: Constructing Visual Tokens through Thumbnail-and-Sampling for Training-Free Video Large Language Models Slowfast networks for video recognition
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 8c4402a2-4819-4d01-9714-55fafd592f34 · outbound
TS-LLaVA: Constructing Visual Tokens through Thumbnail-and-Sampling for Training-Free Video Large Language Models Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 026bce24-78e3-4d60-abe2-19d7abe416a6 · outbound
TS-LLaVA: Constructing Visual Tokens through Thumbnail-and-Sampling for Training-Free Video Large Language Models LITA: Language Instructed Temporal-Localization Assistant
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bf677fb0-5780-4252-8d81-fc61ba1acf29 · outbound
TS-LLaVA: Constructing Visual Tokens through Thumbnail-and-Sampling for Training-Free Video Large Language Models Mixtral of Experts
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 629b42d8-c99d-4fbf-9987-35bcecbf71c9 · outbound
TS-LLaVA: Constructing Visual Tokens through Thumbnail-and-Sampling for Training-Free Video Large Language Models An Image Grid Can Be Worth a Video: Zero-shot Video Question Answering Using a VLM
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6fe6cce8-ec58-4c31-b4d5-83e2d9e174d4 · outbound
TS-LLaVA: Constructing Visual Tokens through Thumbnail-and-Sampling for Training-Free Video Large Language Models BLIP-2: Bootstrapping language-image pre-training with frozen image encoders and large language models
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 99278d1d-5ca3-4f63-9b70-514c2e126e18 · outbound
TS-LLaVA: Constructing Visual Tokens through Thumbnail-and-Sampling for Training-Free Video Large Language Models Inten- tqa: Context-aware video intent reasoning
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 8a26ddf1-7c2f-4a28-b60b-0d06a77b490e · outbound
TS-LLaVA: Constructing Visual Tokens through Thumbnail-and-Sampling for Training-Free Video Large Language Models VideoChat: Chat-Centric Video Understanding
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0eba0374-34d4-4b17-ab94-035a947aa78a · outbound
TS-LLaVA: Constructing Visual Tokens through Thumbnail-and-Sampling for Training-Free Video Large Language Models Mvbench: A comprehensive multi- modal video understanding benchmark
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 3e925e5b-2366-479e-93be-f68f161e4f0a · outbound
TS-LLaVA: Constructing Visual Tokens through Thumbnail-and-Sampling for Training-Free Video Large Language Models Tgif: A new dataset and benchmark on animated gif description
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 48d78477-2644-40f0-8843-779fb6822449 · outbound
TS-LLaVA: Constructing Visual Tokens through Thumbnail-and-Sampling for Training-Free Video Large Language Models Llama-vid: An image is worth 2 tokens in large language models
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 966c248f-6e73-4062-bcb7-fd9abb640bb9 · outbound
TS-LLaVA: Constructing Visual Tokens through Thumbnail-and-Sampling for Training-Free Video Large Language Models Video-LLaVA: Learning United Visual Representation by Alignment Before Projection
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 065f7bc0-e07a-4975-ab56-355775f5a668 · outbound
TS-LLaVA: Constructing Visual Tokens through Thumbnail-and-Sampling for Training-Free Video Large Language Models Visual instruction tuning
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation ea3865ee-5d69-4452-8071-078583c44807 · outbound
TS-LLaVA: Constructing Visual Tokens through Thumbnail-and-Sampling for Training-Free Video Large Language Models Improved baselines with visual instruction tuning
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 6a8855a7-698c-492d-b184-4795c63dc146 · outbound
TS-LLaVA: Constructing Visual Tokens through Thumbnail-and-Sampling for Training-Free Video Large Language Models Llava-next: Im- proved reasoning, ocr, and world knowledge, 2024
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 6fdb8bc3-08c8-4df8-8837-be660e658518 · outbound
TS-LLaVA: Constructing Visual Tokens through Thumbnail-and-Sampling for Training-Free Video Large Language Models St-llm: Large language models are effective tem- poral learners
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 8b501599-1bf7-4069-90dc-9a65dee36862 · outbound
TS-LLaVA: Constructing Visual Tokens through Thumbnail-and-Sampling for Training-Free Video Large Language Models Vista-llama: Reducing hallucination in video language models via equal distance to visual tokens
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 80b922fb-32be-4c83-9ae1-26c6f76d7b00 · outbound
TS-LLaVA: Constructing Visual Tokens through Thumbnail-and-Sampling for Training-Free Video Large Language Models Video-ChatGPT: Towards detailed video un- derstanding via large vision and language models
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation dfa5c33a-bcac-4d8b-ac98-b59275345f64 · outbound
TS-LLaVA: Constructing Visual Tokens through Thumbnail-and-Sampling for Training-Free Video Large Language Models Egoschema: A diagnostic benchmark for very long- form video language understanding
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 4d65746e-6007-420e-885f-691aa1ddb11d · outbound
TS-LLaVA: Constructing Visual Tokens through Thumbnail-and-Sampling for Training-Free Video Large Language Models DeepStack: Deeply Stacking Visual Tokens is Surprisingly Simple and Effective for LMMs
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 64b4855d-4bd5-437b-a49b-481ba80a0497 · outbound
TS-LLaVA: Constructing Visual Tokens through Thumbnail-and-Sampling for Training-Free Video Large Language Models Nous-hermes-2-yi-34b model card, 2023
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation f44c9f46-5076-4aaa-b059-05d27b8992af · outbound
TS-LLaVA: Constructing Visual Tokens through Thumbnail-and-Sampling for Training-Free Video Large Language Models Gpt-4v(ision) system card, 2023
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 52ae98c6-b9ae-429d-85f0-971bf09715f3 · outbound
TS-LLaVA: Constructing Visual Tokens through Thumbnail-and-Sampling for Training-Free Video Large Language Models GPT-4 Technical Report
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 062d0e32-bcfd-44f0-b41a-f06eeae5b588 · outbound
TS-LLaVA: Constructing Visual Tokens through Thumbnail-and-Sampling for Training-Free Video Large Language Models Introducing routing functions to vision-language parameter- efficient fine-tuning with low-rank bottlenecks
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 86c645ed-e928-4f3b-b381-1124b575ed83 · outbound
TS-LLaVA: Constructing Visual Tokens through Thumbnail-and-Sampling for Training-Free Video Large Language Models Learning transferable visual models from natural language supervision
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8dadc60d-6921-45fb-945e-4233ee06fd1a · outbound
TS-LLaVA: Constructing Visual Tokens through Thumbnail-and-Sampling for Training-Free Video Large Language Models Direct preference optimization: Your language model is secretly a reward model
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 07b575ed-eb43-4530-9b83-7eb9ad6cbc9f · outbound
TS-LLaVA: Constructing Visual Tokens through Thumbnail-and-Sampling for Training-Free Video Large Language Models Moviechat: From dense token to sparse memory for long video understanding
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation c59ad5fa-7152-4191-9dd5-9fa4c2acdc0d · outbound
TS-LLaVA: Constructing Visual Tokens through Thumbnail-and-Sampling for Training-Free Video Large Language Models MovieChat+: Question-aware Sparse Memory for Long Video Question Answering
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 357115af-0b2a-40e2-b44d-f4237438b1d1 · outbound
TS-LLaVA: Constructing Visual Tokens through Thumbnail-and-Sampling for Training-Free Video Large Language Models LLaMA: Open and Efficient Foundation Language Models
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 30c069cb-d8a1-4ee6-b41d-681485490811 · outbound
TS-LLaVA: Constructing Visual Tokens through Thumbnail-and-Sampling for Training-Free Video Large Language Models Videoagent: Long-form video understanding with large language model as agent
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation bd605700-9e6c-4c29-8b8d-1dabd561fca6 · outbound
TS-LLaVA: Constructing Visual Tokens through Thumbnail-and-Sampling for Training-Free Video Large Language Models Videotree: Adaptive tree-based video representation for llm reasoning on long videos
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation df1c2007-d344-4dd7-9213-d609fb4ed462 · outbound
TS-LLaVA: Constructing Visual Tokens through Thumbnail-and-Sampling for Training-Free Video Large Language Models FreeVA: Offline MLLM as Training-Free Video Assistant
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d8aa5db7-4889-491b-9f80-ac9b2b609948 · outbound
TS-LLaVA: Constructing Visual Tokens through Thumbnail-and-Sampling for Training-Free Video Large Language Models Audiovisual SlowFast Networks for Video Recognition
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9e9814f2-3e7a-495e-96e7-e2c7504d0be3 · outbound
TS-LLaVA: Constructing Visual Tokens through Thumbnail-and-Sampling for Training-Free Video Large Language Models Next-qa: Next phase of question-answering to explaining temporal actions
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 454efdcd-bade-450b-947b-92cdef74de9b · outbound
TS-LLaVA: Constructing Visual Tokens through Thumbnail-and-Sampling for Training-Free Video Large Language Models Msr-vtt: A large video description dataset for bridging video and language
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 960ececd-5f18-4a8c-8b78-dfb1f4466133 · outbound
TS-LLaVA: Constructing Visual Tokens through Thumbnail-and-Sampling for Training-Free Video Large Language Models PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 90f4958e-b8eb-4ca9-80d1-ca51d38a00e1 · outbound
TS-LLaVA: Constructing Visual Tokens through Thumbnail-and-Sampling for Training-Free Video Large Language Models SlowFast-LLaVA: A Strong Training-Free Baseline for Video Large Language Models
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6bf81f57-f298-49cd-a8ba-365e17bce8b4 · outbound
TS-LLaVA: Constructing Visual Tokens through Thumbnail-and-Sampling for Training-Free Video Large Language Models mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f7504c50-82e8-4672-a0e5-d1691be65961 · outbound
TS-LLaVA: Constructing Visual Tokens through Thumbnail-and-Sampling for Training-Free Video Large Language Models Activitynet-qa: A dataset for understanding complex web videos via question answering
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 0c2cd52a-0aeb-411f-8a50-df560546a403 · outbound
TS-LLaVA: Constructing Visual Tokens through Thumbnail-and-Sampling for Training-Free Video Large Language Models A Simple LLM Framework for Long-Range Video Question-Answering
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5251b9e7-9c04-4f8d-80a2-baca23deffc2 · outbound
TS-LLaVA: Constructing Visual Tokens through Thumbnail-and-Sampling for Training-Free Video Large Language Models Video-LLaMA: An instruction-tuned audio-visual language model for video un- derstanding
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation d5f31825-3ca4-46ec-9b5c-2aa88069d694 · outbound
TS-LLaVA: Constructing Visual Tokens through Thumbnail-and-Sampling for Training-Free Video Large Language Models Long Context Transfer from Language to Vision
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ff5e4446-e2bd-4034-9277-7defca510274 · outbound
TS-LLaVA: Constructing Visual Tokens through Thumbnail-and-Sampling for Training-Free Video Large Language Models LLaMA-adapter: Efficient fine-tuning of large language models with zero- initialized attention
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e7f5868a-7c34-4aeb-9f05-91b1058f5499 · outbound
TS-LLaVA: Constructing Visual Tokens through Thumbnail-and-Sampling for Training-Free Video Large Language Models Llava- next: A strong zero-shot video understanding model, 2024
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 9290f609-327a-4826-870a-9b568ad1d897 · outbound
TS-LLaVA: Constructing Visual Tokens through Thumbnail-and-Sampling for Training-Free Video Large Language Models MLVU: Benchmarking Multi-task Long Video Understanding
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9bd6bf4a-077d-435c-897b-011540ce64f1 · outbound
TS-LLaVA: Constructing Visual Tokens through Thumbnail-and-Sampling for Training-Free Video Large Language Models MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5fb455bb-ce26-4a11-aed0-35a49164e72f · outbound
TS-LLaVA: Constructing Visual Tokens through Thumbnail-and-Sampling for Training-Free Video Large Language Models Each frame is resized accordingly to fit within the resulting 336 ×336 thumbnail image
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 6f72a5fd-f31a-4f6d-8dfb-366a28852cda · outbound
TS-LLaVA: Constructing Visual Tokens through Thumbnail-and-Sampling for Training-Free Video Large Language Models We start with additional experiments conducted for the study on compression strategies
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation f1827bc4-9c67-4cb6-ace2-6ab52c299c2a · inbound
IPFormer-VideoLLM: Enhancing Multi-modal Video Understanding for Multi-shot Scenes TS-LLaVA: Constructing Visual Tokens through Thumbnail-and-Sampling for Training-Free Video Large Language Models
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3bb5536f-a601-4a5b-a460-1e52aa4c1314 · inbound
Direct RNA sequence design under codon constraints using expressive tensor-based secondary structure models TS-LLaVA: Constructing Visual Tokens through Thumbnail-and-Sampling for Training-Free Video Large Language Models
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation bbac7741-6968-45af-9c5e-35fc5092ab46 · inbound
WindowQuant: Mixed-Precision KV Cache Quantization based on Window-Level Similarity for VLMs Inference Optimization TS-LLaVA: Constructing Visual Tokens through Thumbnail-and-Sampling for Training-Free Video Large Language Models
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation c5335c59-93c7-493c-a6f2-bcc8ab59db26 · inbound
Enhancing Visual Token Representations for Video Large Language Models via Training-Free Spatial-Temporal Pooling and Gridding TS-LLaVA: Constructing Visual Tokens through Thumbnail-and-Sampling for Training-Free Video Large Language Models
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 5348825f-e32a-4020-9bf9-db996aa84538 · inbound
Zero-Shot 3D Question Answering via Hierarchical View-to-Token Transportation TS-LLaVA: Constructing Visual Tokens through Thumbnail-and-Sampling for Training-Free Video Large Language Models
Reference 79
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 9e73c270-c79c-40ff-b7cc-3c2a95686929 · inbound
MemoryCard: Topic-Aware Multi-Modal Clue Compression for Long-Video Question Answering TS-LLaVA: Constructing Visual Tokens through Thumbnail-and-Sampling for Training-Free Video Large Language Models
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 6464f441-c331-49fd-ad4d-a67d7c43f29b · inbound
TreeSoc: Tree-Structured Dynamic Reasoning and Tool Synergy for Soccer Video Understanding TS-LLaVA: Constructing Visual Tokens through Thumbnail-and-Sampling for Training-Free Video Large Language Models
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.