Pith. sign in

Paper Citation Record · LEDGER

TS-LLaVA: Constructing Visual Tokens through Thumbnail-and-Sampling for Training-Free Video Large Language Models

As of 14 August 2026, this Paper Citation Record lists 56 of 56 outbound references and 7 inbound Pith citation observations for arXiv:2411.11066.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.11066 v1

Coverage vector

measured 56 of 56 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T19:03:32.584154Z

measured 63 of 63 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 7 of 7 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T22:37:38.250264Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

56 of 56 outbound references displayed

  • verified exact0
  • verified fuzzy26
  • unresolved30
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

25
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 49d949da-4101-48f5-88b1-b60044caf604 · outbound

This paper cites Qwen Technical Report.

TS-LLaVA: Constructing Visual Tokens through Thumbnail-and-Sampling for Training-Free Video Large Language Models Qwen Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-12T19:03:31.904212Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:03:31.904212Z digest=sha256:ea4e8bbee5a251270c403b4f99c422c7e7c739a91a1a7bda8dfa64ff33484a19

Observation 8cf08a90-6df7-4ff9-8eb8-214db30a03f9 · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

TS-LLaVA: Constructing Visual Tokens through Thumbnail-and-Sampling for Training-Free Video Large Language Models Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-12T19:03:31.930824Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:03:31.930824Z digest=sha256:d7cbb1834e7a63cab387e4882378d33c5cb8804af0e8a5489831f42b11cec915

Observation 25aafd0b-38a8-4169-9d1f-998625a32d2d · outbound

This paper cites Matryoshka Multimodal Models.

TS-LLaVA: Constructing Visual Tokens through Thumbnail-and-Sampling for Training-Free Video Large Language Models Matryoshka Multimodal Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-12T19:03:31.940342Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:03:31.940342Z digest=sha256:00e1fdd983d2013301d74925ae9f7282b7fb45e968a79934516bf0c57927e925

Observation 063c4c28-daa6-45c7-b103-b0989029fd18 · outbound

This paper cites Collecting highly parallel data for paraphrase evaluation.

TS-LLaVA: Constructing Visual Tokens through Thumbnail-and-Sampling for Training-Free Video Large Language Models Collecting highly parallel data for paraphrase evaluation

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:03:33.620276Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T19:03:31.943621Z digest=sha256:9b5199b45308281ac1afa7fbca96d1e50619a58cc70fddac6a7b4aab1c94a207

Observation b467affe-0ee3-4eef-963e-26673668d44d · outbound

This paper cites VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs.

TS-LLaVA: Constructing Visual Tokens through Thumbnail-and-Sampling for Training-Free Video Large Language Models VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-12T19:03:31.946661Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:03:31.946661Z digest=sha256:2253146f98bc3cfb8089a577efdbca8d1d98dcb2893f992c3b5265d98a0fd375

Observation a3fc6f6b-dba2-4b7c-a953-ed9cda08e7cd · outbound

This paper cites Gonzalez, Ion Stoica, and Eric P.

TS-LLaVA: Constructing Visual Tokens through Thumbnail-and-Sampling for Training-Free Video Large Language Models Gonzalez, Ion Stoica, and Eric P

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:03:33.550546Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T19:03:31.950260Z digest=sha256:f03840a9d07bceee41b4d58dc29d034ee43d06c16f908dcfa3582fd1c6fab52d

Observation 86dfe8fb-1ca0-44fa-87f5-658e4ade1cf3 · outbound

This paper cites InstructBLIP: Towards general-purpose vision-language models with instruction tuning.

TS-LLaVA: Constructing Visual Tokens through Thumbnail-and-Sampling for Training-Free Video Large Language Models InstructBLIP: Towards general-purpose vision-language models with instruction tuning

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-12T19:03:31.954093Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:03:31.954093Z digest=sha256:61c49c3c6d121bc8185130d8197def03b1768376984ac98bc6a5366f02b37c36

Observation d835c0fa-309c-44fd-b5cd-7cf8ca2383b5 · outbound

This paper cites An image is worth 16x16 words: Transformers for image recognition at scale.

TS-LLaVA: Constructing Visual Tokens through Thumbnail-and-Sampling for Training-Free Video Large Language Models An image is worth 16x16 words: Transformers for image recognition at scale

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-12T19:03:31.957310Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:03:31.957310Z digest=sha256:0f4c7c0760c391fe4f8eceb661e9df9b5e6a9f9738a0d7e8e6fd09314cc32c6b

Observation a6da4ce0-5e6d-4df6-b0ae-959caac389f5 · outbound

This paper cites Slowfast networks for video recognition.

TS-LLaVA: Constructing Visual Tokens through Thumbnail-and-Sampling for Training-Free Video Large Language Models Slowfast networks for video recognition

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:03:33.531487Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T19:03:31.960775Z digest=sha256:9a30a66623d9e36819bbe06914937a8615a54757560125867cb01283b36cec0b

Observation 8c4402a2-4819-4d01-9714-55fafd592f34 · outbound

This paper cites Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis.

TS-LLaVA: Constructing Visual Tokens through Thumbnail-and-Sampling for Training-Free Video Large Language Models Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-12T19:03:31.964048Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:03:31.964048Z digest=sha256:cf9c906864376eb45792e5e516f68e749970e0df59292e4f594855d5eaf44c93

Observation 026bce24-78e3-4d60-abe2-19d7abe416a6 · outbound

This paper cites LITA: Language Instructed Temporal-Localization Assistant.

TS-LLaVA: Constructing Visual Tokens through Thumbnail-and-Sampling for Training-Free Video Large Language Models LITA: Language Instructed Temporal-Localization Assistant

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-12T19:03:31.982517Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:03:31.982517Z digest=sha256:d0a74997c4ec0b6ae70b86b1392b252dc2d5e66a648e5468c2302eb2012a4e4d

Observation bf677fb0-5780-4252-8d81-fc61ba1acf29 · outbound

This paper cites Mixtral of Experts.

TS-LLaVA: Constructing Visual Tokens through Thumbnail-and-Sampling for Training-Free Video Large Language Models Mixtral of Experts

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-12T19:03:32.019710Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:03:32.019710Z digest=sha256:b677217c29d0b7094d41aa4335b6f63c9f80a6c2dc2500d76763a7c6ae6a3727

Observation 629b42d8-c99d-4fbf-9987-35bcecbf71c9 · outbound

This paper cites An Image Grid Can Be Worth a Video: Zero-shot Video Question Answering Using a VLM.

TS-LLaVA: Constructing Visual Tokens through Thumbnail-and-Sampling for Training-Free Video Large Language Models An Image Grid Can Be Worth a Video: Zero-shot Video Question Answering Using a VLM

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-12T19:03:32.050286Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:03:32.050286Z digest=sha256:42b7f39e3fda9c51665e9fd427ec40eefccbed5b9796d0000c47c74b2d523eb9

Observation 6fe6cce8-ec58-4c31-b4d5-83e2d9e174d4 · outbound

This paper cites BLIP-2: Bootstrapping language-image pre-training with frozen image encoders and large language models.

TS-LLaVA: Constructing Visual Tokens through Thumbnail-and-Sampling for Training-Free Video Large Language Models BLIP-2: Bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:03:33.523341Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T19:03:32.089359Z digest=sha256:fbacd04ffc13ac3deda8f9102d1b0311589067a7e3045135a900dff771edbce6

Observation 99278d1d-5ca3-4f63-9b70-514c2e126e18 · outbound

This paper cites Inten- tqa: Context-aware video intent reasoning.

TS-LLaVA: Constructing Visual Tokens through Thumbnail-and-Sampling for Training-Free Video Large Language Models Inten- tqa: Context-aware video intent reasoning

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:03:33.515890Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T19:03:32.110166Z digest=sha256:31fe4eb49181640869b666f6ba2112b6913256b441f8c91bcf68e3e2ee6b3156

Observation 8a26ddf1-7c2f-4a28-b60b-0d06a77b490e · outbound

This paper cites VideoChat: Chat-Centric Video Understanding.

TS-LLaVA: Constructing Visual Tokens through Thumbnail-and-Sampling for Training-Free Video Large Language Models VideoChat: Chat-Centric Video Understanding

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-12T19:03:32.113190Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:03:32.113190Z digest=sha256:ec62d7c5a0d1454ea0eb845ad8c613f3432f87fdee0a2d033ddf2a313eb80f29

Observation 0eba0374-34d4-4b17-ab94-035a947aa78a · outbound

This paper cites Mvbench: A comprehensive multi- modal video understanding benchmark.

TS-LLaVA: Constructing Visual Tokens through Thumbnail-and-Sampling for Training-Free Video Large Language Models Mvbench: A comprehensive multi- modal video understanding benchmark

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:03:33.508446Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T19:03:32.126048Z digest=sha256:684e7beff7ebc2fa22fe65e5df1f01d0369ae3a57d01eb38a05b4a184e9b4e5f

Observation 3e925e5b-2366-479e-93be-f68f161e4f0a · outbound

This paper cites Tgif: A new dataset and benchmark on animated gif description.

TS-LLaVA: Constructing Visual Tokens through Thumbnail-and-Sampling for Training-Free Video Large Language Models Tgif: A new dataset and benchmark on animated gif description

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:03:33.487482Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T19:03:32.156621Z digest=sha256:df2ab9a7407ef40741c37ce7f36b98d25f62b2cc8a34cb8a2aa33b1f8d5a7d6b

Observation 48d78477-2644-40f0-8843-779fb6822449 · outbound

This paper cites Llama-vid: An image is worth 2 tokens in large language models.

TS-LLaVA: Constructing Visual Tokens through Thumbnail-and-Sampling for Training-Free Video Large Language Models Llama-vid: An image is worth 2 tokens in large language models

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:03:33.455390Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T19:03:32.162477Z digest=sha256:128a9124cc4234e10fdc67ac9903823855e15300e336b77a59ff70382b898132

Observation 966c248f-6e73-4062-bcb7-fd9abb640bb9 · outbound

This paper cites Video-LLaVA: Learning United Visual Representation by Alignment Before Projection.

TS-LLaVA: Constructing Visual Tokens through Thumbnail-and-Sampling for Training-Free Video Large Language Models Video-LLaVA: Learning United Visual Representation by Alignment Before Projection

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-12T19:03:32.165398Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:03:32.165398Z digest=sha256:dbd8e5576ed5ca1a05d3c99548dbbf2c835346fbb48ae7100cc7cba93edca67a

Observation 065f7bc0-e07a-4975-ab56-355775f5a668 · outbound

This paper cites Visual instruction tuning.

TS-LLaVA: Constructing Visual Tokens through Thumbnail-and-Sampling for Training-Free Video Large Language Models Visual instruction tuning

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:03:33.376748Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T19:03:32.168388Z digest=sha256:eb639b495a69de2a9018263bb2545c064c4256e986c8f638e7e175970351255f

Observation ea3865ee-5d69-4452-8071-078583c44807 · outbound

This paper cites Improved baselines with visual instruction tuning.

TS-LLaVA: Constructing Visual Tokens through Thumbnail-and-Sampling for Training-Free Video Large Language Models Improved baselines with visual instruction tuning

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:03:33.298728Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T19:03:32.170871Z digest=sha256:d52cb84b363a7798e4447cee0e7db6765c1d609e8dbcb1cf329f5a474544885a

Observation 6a8855a7-698c-492d-b184-4795c63dc146 · outbound

This paper cites Llava-next: Im- proved reasoning, ocr, and world knowledge, 2024.

TS-LLaVA: Constructing Visual Tokens through Thumbnail-and-Sampling for Training-Free Video Large Language Models Llava-next: Im- proved reasoning, ocr, and world knowledge, 2024

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:03:33.225453Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T19:03:32.173551Z digest=sha256:768683601188dac74823d42ccc4714f8c2635d355133eb5852a928686ad25855

Observation 6fdb8bc3-08c8-4df8-8837-be660e658518 · outbound

This paper cites St-llm: Large language models are effective tem- poral learners.

TS-LLaVA: Constructing Visual Tokens through Thumbnail-and-Sampling for Training-Free Video Large Language Models St-llm: Large language models are effective tem- poral learners

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:03:33.171618Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T19:03:32.176107Z digest=sha256:0e77c6a236f3c5d82fc8318d540bf11ace8df1ae3d18fb2b4f885d8f41260593

Observation 8b501599-1bf7-4069-90dc-9a65dee36862 · outbound

This paper cites Vista-llama: Reducing hallucination in video language models via equal distance to visual tokens.

TS-LLaVA: Constructing Visual Tokens through Thumbnail-and-Sampling for Training-Free Video Large Language Models Vista-llama: Reducing hallucination in video language models via equal distance to visual tokens

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:03:33.163744Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T19:03:32.178588Z digest=sha256:c340fb125ff8618b11eae5e2574d1743d3d03d32b725b09539e1150099df4c56

Observation 80b922fb-32be-4c83-9ae1-26c6f76d7b00 · outbound

This paper cites Video-ChatGPT: Towards detailed video un- derstanding via large vision and language models.

TS-LLaVA: Constructing Visual Tokens through Thumbnail-and-Sampling for Training-Free Video Large Language Models Video-ChatGPT: Towards detailed video un- derstanding via large vision and language models

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:03:33.155251Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T19:03:32.181013Z digest=sha256:4a99f2576a104dba14cd14adeac8ec312124257186939afb1d5d2b128b2c7b24

Observation dfa5c33a-bcac-4d8b-ac98-b59275345f64 · outbound

This paper cites Egoschema: A diagnostic benchmark for very long- form video language understanding.

TS-LLaVA: Constructing Visual Tokens through Thumbnail-and-Sampling for Training-Free Video Large Language Models Egoschema: A diagnostic benchmark for very long- form video language understanding

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:03:33.146272Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T19:03:32.184057Z digest=sha256:371894b15d7a5fb67efe0327918b0674238f95eff620f8a0f429681a44c88ad4

Observation 4d65746e-6007-420e-885f-691aa1ddb11d · outbound

This paper cites DeepStack: Deeply Stacking Visual Tokens is Surprisingly Simple and Effective for LMMs.

TS-LLaVA: Constructing Visual Tokens through Thumbnail-and-Sampling for Training-Free Video Large Language Models DeepStack: Deeply Stacking Visual Tokens is Surprisingly Simple and Effective for LMMs

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-12T19:03:32.186910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:03:32.186910Z digest=sha256:252dc720f63381e7bfedff916cc0b6506c5fe30b436ffedc934c6e23df0187c5

Observation 64b4855d-4bd5-437b-a49b-481ba80a0497 · outbound

This paper cites Nous-hermes-2-yi-34b model card, 2023.

TS-LLaVA: Constructing Visual Tokens through Thumbnail-and-Sampling for Training-Free Video Large Language Models Nous-hermes-2-yi-34b model card, 2023

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:03:33.136959Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T19:03:32.190184Z digest=sha256:e467ed26e9595184c2a150a89a6544fb3454e4125a8041ab19eb1fe0e6c4c2b9

Observation f44c9f46-5076-4aaa-b059-05d27b8992af · outbound

This paper cites Gpt-4v(ision) system card, 2023.

TS-LLaVA: Constructing Visual Tokens through Thumbnail-and-Sampling for Training-Free Video Large Language Models Gpt-4v(ision) system card, 2023

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-12T19:03:32.193233Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:03:32.193233Z digest=sha256:8d12fe1d8aeea9482c0d9c3ae65576620e883f6e4b414a5612b8ab96c94d9341

Observation 52ae98c6-b9ae-429d-85f0-971bf09715f3 · outbound

This paper cites GPT-4 Technical Report.

TS-LLaVA: Constructing Visual Tokens through Thumbnail-and-Sampling for Training-Free Video Large Language Models GPT-4 Technical Report

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-12T19:03:32.195669Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:03:32.195669Z digest=sha256:1679848ce531da17cc43a48e392d90a76271968a2f41178c42038c7d60d8dc1e

Observation 062d0e32-bcfd-44f0-b41a-f06eeae5b588 · outbound

This paper cites Introducing routing functions to vision-language parameter- efficient fine-tuning with low-rank bottlenecks.

TS-LLaVA: Constructing Visual Tokens through Thumbnail-and-Sampling for Training-Free Video Large Language Models Introducing routing functions to vision-language parameter- efficient fine-tuning with low-rank bottlenecks

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:03:33.089594Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T19:03:32.198582Z digest=sha256:368be648ea5d1d54e5347336c1a8ed63e50681e51da23de97da17b9799582f8b

Observation 86c645ed-e928-4f3b-b381-1124b575ed83 · outbound

This paper cites Learning transferable visual models from natural language supervision.

TS-LLaVA: Constructing Visual Tokens through Thumbnail-and-Sampling for Training-Free Video Large Language Models Learning transferable visual models from natural language supervision

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-12T19:03:32.201141Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:03:32.201141Z digest=sha256:b0f07fb795bbf64e1fb80e521b9def2153abbc4b34889b6ccf98e1c1c3990fae

Observation 8dadc60d-6921-45fb-945e-4233ee06fd1a · outbound

This paper cites Direct preference optimization: Your language model is secretly a reward model.

TS-LLaVA: Constructing Visual Tokens through Thumbnail-and-Sampling for Training-Free Video Large Language Models Direct preference optimization: Your language model is secretly a reward model

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-12T19:03:32.204134Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:03:32.204134Z digest=sha256:9f6f0e430b5aefebd7fefb0662cec84177de79dc5d7e75e78ec71e6fb911de0d

Observation 07b575ed-eb43-4530-9b83-7eb9ad6cbc9f · outbound

This paper cites Moviechat: From dense token to sparse memory for long video understanding.

TS-LLaVA: Constructing Visual Tokens through Thumbnail-and-Sampling for Training-Free Video Large Language Models Moviechat: From dense token to sparse memory for long video understanding

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:03:32.997350Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T19:03:32.207504Z digest=sha256:be7b0ee80dd735cac4f88d5305df4deb721389f3ca3168462ba2ac202392a918

Observation c59ad5fa-7152-4191-9dd5-9fa4c2acdc0d · outbound

This paper cites MovieChat+: Question-aware Sparse Memory for Long Video Question Answering.

TS-LLaVA: Constructing Visual Tokens through Thumbnail-and-Sampling for Training-Free Video Large Language Models MovieChat+: Question-aware Sparse Memory for Long Video Question Answering

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-12T19:03:32.209929Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:03:32.209929Z digest=sha256:913a72cac7cda1dbb5053cb8d079b44e2a0ef6332cd8ce7692e89b97fc96acc8

Observation 357115af-0b2a-40e2-b44d-f4237438b1d1 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

TS-LLaVA: Constructing Visual Tokens through Thumbnail-and-Sampling for Training-Free Video Large Language Models LLaMA: Open and Efficient Foundation Language Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-12T19:03:32.245623Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:03:32.245623Z digest=sha256:3d34264b15550cb554f1210f53eb5219a2eaefd189356ad4d9d94a7d50679bfe

Observation 30c069cb-d8a1-4ee6-b41d-681485490811 · outbound

This paper cites Videoagent: Long-form video understanding with large language model as agent.

TS-LLaVA: Constructing Visual Tokens through Thumbnail-and-Sampling for Training-Free Video Large Language Models Videoagent: Long-form video understanding with large language model as agent

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:03:32.966106Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T19:03:32.290029Z digest=sha256:2366a327d9c46bab725bb2913a6ea8cb79459a16620b6374442ecfd62e027375

Observation bd605700-9e6c-4c29-8b8d-1dabd561fca6 · outbound

This paper cites Videotree: Adaptive tree-based video representation for llm reasoning on long videos.

TS-LLaVA: Constructing Visual Tokens through Thumbnail-and-Sampling for Training-Free Video Large Language Models Videotree: Adaptive tree-based video representation for llm reasoning on long videos

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:03:32.940983Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T19:03:32.393563Z digest=sha256:dc3f20c32c3301b94b82d284b5d632819702364a701a660a5b44f619717dab26

Observation df1c2007-d344-4dd7-9213-d609fb4ed462 · outbound

This paper cites FreeVA: Offline MLLM as Training-Free Video Assistant.

TS-LLaVA: Constructing Visual Tokens through Thumbnail-and-Sampling for Training-Free Video Large Language Models FreeVA: Offline MLLM as Training-Free Video Assistant

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-12T19:03:32.413331Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:03:32.413331Z digest=sha256:0135f98be865104f609524cbdadb9f423a5e626eea6de29ab4649c8dbcdbdf1f

Observation d8aa5db7-4889-491b-9f80-ac9b2b609948 · outbound

This paper cites Audiovisual SlowFast Networks for Video Recognition.

TS-LLaVA: Constructing Visual Tokens through Thumbnail-and-Sampling for Training-Free Video Large Language Models Audiovisual SlowFast Networks for Video Recognition

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-12T19:03:32.416966Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:03:32.416966Z digest=sha256:76360a45feffcde4a17137497e8159b16aadb877da1555207b03e152fcc82bbc

Observation 9e9814f2-3e7a-495e-96e7-e2c7504d0be3 · outbound

This paper cites Next-qa: Next phase of question-answering to explaining temporal actions.

TS-LLaVA: Constructing Visual Tokens through Thumbnail-and-Sampling for Training-Free Video Large Language Models Next-qa: Next phase of question-answering to explaining temporal actions

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-12T19:03:32.420208Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:03:32.420208Z digest=sha256:c56b9ec8838c6c46b672b16882dc5a34a63c1b96eace2d67baecfe9272bad8fc

Observation 454efdcd-bade-450b-947b-92cdef74de9b · outbound

This paper cites Msr-vtt: A large video description dataset for bridging video and language.

TS-LLaVA: Constructing Visual Tokens through Thumbnail-and-Sampling for Training-Free Video Large Language Models Msr-vtt: A large video description dataset for bridging video and language

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:03:32.897591Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T19:03:32.424070Z digest=sha256:0e3fb041021e421bc8a63ce0fd8df8a1c10b72993435d740fc23477424e3dcaa

Observation 960ececd-5f18-4a8c-8b78-dfb1f4466133 · outbound

This paper cites PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning.

TS-LLaVA: Constructing Visual Tokens through Thumbnail-and-Sampling for Training-Free Video Large Language Models PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-12T19:03:32.426845Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:03:32.426845Z digest=sha256:58791d1da9544264295f424fd8ba0bb5728f88e4edd570ba00e02289291855ad

Observation 90f4958e-b8eb-4ca9-80d1-ca51d38a00e1 · outbound

This paper cites SlowFast-LLaVA: A Strong Training-Free Baseline for Video Large Language Models.

TS-LLaVA: Constructing Visual Tokens through Thumbnail-and-Sampling for Training-Free Video Large Language Models SlowFast-LLaVA: A Strong Training-Free Baseline for Video Large Language Models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-12T19:03:32.430071Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:03:32.430071Z digest=sha256:8b0c74c411b30ae9a1c4728276d7b3a0a3f77c7436d9440b03f7bbd5d0c3c7c5

Observation 6bf81f57-f298-49cd-a8ba-365e17bce8b4 · outbound

This paper cites mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality.

TS-LLaVA: Constructing Visual Tokens through Thumbnail-and-Sampling for Training-Free Video Large Language Models mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-12T19:03:32.433441Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:03:32.433441Z digest=sha256:3f9ab9ddc787b3f7e4c98f7dfeed21702ba27d6d8b580e50ca4754e510ff860b

Observation f7504c50-82e8-4672-a0e5-d1691be65961 · outbound

This paper cites Activitynet-qa: A dataset for understanding complex web videos via question answering.

TS-LLaVA: Constructing Visual Tokens through Thumbnail-and-Sampling for Training-Free Video Large Language Models Activitynet-qa: A dataset for understanding complex web videos via question answering

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:03:32.888047Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T19:03:32.436405Z digest=sha256:0874f8ac59f135dc9d2ef30088f0bc7219a245d7ed1611d79d110690e4eab2c3

Observation 0c2cd52a-0aeb-411f-8a50-df560546a403 · outbound

This paper cites A Simple LLM Framework for Long-Range Video Question-Answering.

TS-LLaVA: Constructing Visual Tokens through Thumbnail-and-Sampling for Training-Free Video Large Language Models A Simple LLM Framework for Long-Range Video Question-Answering

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-12T19:03:32.439057Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:03:32.439057Z digest=sha256:cd9b3b36080be8d62f2742456938b4f1850f9797f95cd4401fa2925cacc9ce54

Observation 5251b9e7-9c04-4f8d-80a2-baca23deffc2 · outbound

This paper cites Video-LLaMA: An instruction-tuned audio-visual language model for video un- derstanding.

TS-LLaVA: Constructing Visual Tokens through Thumbnail-and-Sampling for Training-Free Video Large Language Models Video-LLaMA: An instruction-tuned audio-visual language model for video un- derstanding

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:03:32.878571Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T19:03:32.442459Z digest=sha256:6f2f9faf69d53c0842a38373501873961a38c29bc7a428eb05eeb025fa2c6bda

Observation d5f31825-3ca4-46ec-9b5c-2aa88069d694 · outbound

This paper cites Long Context Transfer from Language to Vision.

TS-LLaVA: Constructing Visual Tokens through Thumbnail-and-Sampling for Training-Free Video Large Language Models Long Context Transfer from Language to Vision

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-12T19:03:32.445716Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:03:32.445716Z digest=sha256:4681e9a8d6ac2c2d4e9157c93214037f9a79a393acddeb9aad839f52ab28351e

Observation ff5e4446-e2bd-4034-9277-7defca510274 · outbound

This paper cites LLaMA-adapter: Efficient fine-tuning of large language models with zero- initialized attention.

TS-LLaVA: Constructing Visual Tokens through Thumbnail-and-Sampling for Training-Free Video Large Language Models LLaMA-adapter: Efficient fine-tuning of large language models with zero- initialized attention

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-12T19:03:32.449162Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:03:32.449162Z digest=sha256:51281e64019f482b4af3355310b592c8807711afaeb1ec2a6261f3270a1840e0

Observation e7f5868a-7c34-4aeb-9f05-91b1058f5499 · outbound

This paper cites Llava- next: A strong zero-shot video understanding model, 2024.

TS-LLaVA: Constructing Visual Tokens through Thumbnail-and-Sampling for Training-Free Video Large Language Models Llava- next: A strong zero-shot video understanding model, 2024

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:03:32.861900Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T19:03:32.452314Z digest=sha256:9218aaee1bc814b8573b021c4a9a90d6fe48bc40c0b3702f636fc5a74c735264

Observation 9290f609-327a-4826-870a-9b568ad1d897 · outbound

This paper cites MLVU: Benchmarking Multi-task Long Video Understanding.

TS-LLaVA: Constructing Visual Tokens through Thumbnail-and-Sampling for Training-Free Video Large Language Models MLVU: Benchmarking Multi-task Long Video Understanding

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-12T19:03:32.455173Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:03:32.455173Z digest=sha256:2f610c0f62bd72753b966bb80676648e87d8c31299a4f610e7f4247880cdb022

Observation 9bd6bf4a-077d-435c-897b-011540ce64f1 · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

TS-LLaVA: Constructing Visual Tokens through Thumbnail-and-Sampling for Training-Free Video Large Language Models MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-12T19:03:32.482619Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:03:32.482619Z digest=sha256:46706121420af9b35e77fd702f653a159317bfcee5a58f1642ecea5e93258bad

Observation 5fb455bb-ce26-4a11-aed0-35a49164e72f · outbound

This paper cites Each frame is resized accordingly to fit within the resulting 336 ×336 thumbnail image.

TS-LLaVA: Constructing Visual Tokens through Thumbnail-and-Sampling for Training-Free Video Large Language Models Each frame is resized accordingly to fit within the resulting 336 ×336 thumbnail image

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:03:32.850844Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T19:03:32.513984Z digest=sha256:de68bb2ccda8cecd9db2a00b16d83780218f6f9413ced37a6634073b57a9a581

Observation 6f72a5fd-f31a-4f6d-8dfb-366a28852cda · outbound

This paper cites We start with additional experiments conducted for the study on compression strategies.

TS-LLaVA: Constructing Visual Tokens through Thumbnail-and-Sampling for Training-Free Video Large Language Models We start with additional experiments conducted for the study on compression strategies

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:03:32.839951Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T19:03:32.584154Z digest=sha256:71b17e4ebb0b2301dd0b9690e4091a221adc27d475d959f2eb55985939853feb

Pith citing papers

Observation f1827bc4-9c67-4cb6-ace2-6ab52c299c2a · inbound

IPFormer-VideoLLM: Enhancing Multi-modal Video Understanding for Multi-shot Scenes cites this paper.

IPFormer-VideoLLM: Enhancing Multi-modal Video Understanding for Multi-shot Scenes TS-LLaVA: Constructing Visual Tokens through Thumbnail-and-Sampling for Training-Free Video Large Language Models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T22:37:38.250264Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:37:38.250264Z digest=sha256:c1e3d6e6297f0116743a906243b87d2d2edf501b08934e810abfbe9ca109a812

Observation 3bb5536f-a601-4a5b-a460-1e52aa4c1314 · inbound

Direct RNA sequence design under codon constraints using expressive tensor-based secondary structure models cites this paper.

Direct RNA sequence design under codon constraints using expressive tensor-based secondary structure models TS-LLaVA: Constructing Visual Tokens through Thumbnail-and-Sampling for Training-Free Video Large Language Models

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-10T00:29:46.584977Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-10T00:28:35.611378Z digest=sha256:1b05cb6f880c30948acf1cacbdc0402ff610ce3ff247955aad242089a29d5a43

Observation bbac7741-6968-45af-9c5e-35fc5092ab46 · inbound

WindowQuant: Mixed-Precision KV Cache Quantization based on Window-Level Similarity for VLMs Inference Optimization cites this paper.

WindowQuant: Mixed-Precision KV Cache Quantization based on Window-Level Similarity for VLMs Inference Optimization TS-LLaVA: Constructing Visual Tokens through Thumbnail-and-Sampling for Training-Free Video Large Language Models

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-11T16:36:07.914366Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-09T16:06:26.450483Z digest=sha256:440624b4f13004682d11a7f1b47d06263724fa16ef7fbbf175c477cce7f9e1d6

Observation c5335c59-93c7-493c-a6f2-bcc8ab59db26 · inbound

Enhancing Visual Token Representations for Video Large Language Models via Training-Free Spatial-Temporal Pooling and Gridding cites this paper.

Enhancing Visual Token Representations for Video Large Language Models via Training-Free Spatial-Temporal Pooling and Gridding TS-LLaVA: Constructing Visual Tokens through Thumbnail-and-Sampling for Training-Free Video Large Language Models

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-22T06:21:10.852517Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-22T06:16:32.972099Z digest=sha256:b057e6c82a4c539c2bf28a3713ba256a49b92831ad7fd36d5c675deaca175c41

Observation 5348825f-e32a-4020-9bf9-db996aa84538 · inbound

Zero-Shot 3D Question Answering via Hierarchical View-to-Token Transportation cites this paper.

Zero-Shot 3D Question Answering via Hierarchical View-to-Token Transportation TS-LLaVA: Constructing Visual Tokens through Thumbnail-and-Sampling for Training-Free Video Large Language Models

Reference 79

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T02:26:27.047994Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-06-28T10:54:02.188634Z digest=sha256:a49409987a9f6dda3537619a5d3908a8e339a7be51c13cb1421196d832f39dbc

Observation 9e73c270-c79c-40ff-b7cc-3c2a95686929 · inbound

MemoryCard: Topic-Aware Multi-Modal Clue Compression for Long-Video Question Answering cites this paper.

MemoryCard: Topic-Aware Multi-Modal Clue Compression for Long-Video Question Answering TS-LLaVA: Constructing Visual Tokens through Thumbnail-and-Sampling for Training-Free Video Large Language Models

Reference 33

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T12:46:56.879237Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-06-28T01:52:33.768494Z digest=sha256:deba064d6e9f8c26bf3540203684ff6812d758248df7f5b2050840d26201fea6

Observation 6464f441-c331-49fd-ad4d-a67d7c43f29b · inbound

TreeSoc: Tree-Structured Dynamic Reasoning and Tool Synergy for Soccer Video Understanding cites this paper.

TreeSoc: Tree-Structured Dynamic Reasoning and Tool Synergy for Soccer Video Understanding TS-LLaVA: Constructing Visual Tokens through Thumbnail-and-Sampling for Training-Free Video Large Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-07-14T07:49:34.639380Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T07:49:34.639380Z digest=sha256:17da580fb3f694c0236edec1a57c0324857e93f8066048864e68ba2134b7443d