Pith. sign in

Paper Citation Record · LEDGER

SlowFast-LLaVA: A Strong Training-Free Baseline for Video Large Language Models

As of 12 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 42 inbound Pith citation observations for arXiv:2407.15841.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2407.15841 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 42 of 42 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 42 of 42 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-12T05:44:05.836761Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T03:19:29.898625Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 05397695-4940-4979-8b87-c9bb464a2e02 · inbound

LLaVA-Video: Video Instruction Tuning With Synthetic Data cites this paper.

LLaVA-Video: Video Instruction Tuning With Synthetic Data SlowFast-LLaVA: A Strong Training-Free Baseline for Video Large Language Models

Reference 213

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T23:20:33.156675Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-10T23:20:32.330351Z digest=sha256:44f8440e789390e3f01dfa0f6eb1cd777e05b39f510b3f4c90e6d0cd226aa837

Observation 38b01930-0c6e-4c15-8049-3efbe895413d · inbound

LongVU: Spatiotemporal Adaptive Compression for Long Video-Language Understanding cites this paper.

LongVU: Spatiotemporal Adaptive Compression for Long Video-Language Understanding SlowFast-LLaVA: A Strong Training-Free Baseline for Video Large Language Models

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-16T13:53:33.711388Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-16T13:53:33.585035Z digest=sha256:2bc3a95a4b4649d1850de98dbd9a093528497bbb21792903ef95a2b2b820c074

Observation d96dadec-8a78-48f9-950f-320a2629dd55 · inbound

Sparrow: Data-Efficient Video-LLM with Text-to-Image Augmentation cites this paper.

Sparrow: Data-Efficient Video-LLM with Text-to-Image Augmentation SlowFast-LLaVA: A Strong Training-Free Baseline for Video Large Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-12T05:44:05.836761Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:44:05.836761Z digest=sha256:43af9bb85e3eb895d67dc1998b9eedba3821655eb99487a52fb21d9c2c8d829f

Observation 0a008773-a7cf-4d40-8821-176c5e918590 · inbound

A Stitch in Time Saves Nine: Small VLM is a Precise Guidance for Accelerating Large VLMs cites this paper.

A Stitch in Time Saves Nine: Small VLM is a Precise Guidance for Accelerating Large VLMs SlowFast-LLaVA: A Strong Training-Free Baseline for Video Large Language Models

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-11T22:35:24.362275Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:35:24.362275Z digest=sha256:caf7d4117f89833e9cdb1c8de6086e715238125694a7cb9ce5e61a5b6d04c657

Observation 1574aa84-21a1-442e-9207-b8afb5df4826 · inbound

NVILA: Efficient Frontier Visual Language Models cites this paper.

NVILA: Efficient Frontier Visual Language Models SlowFast-LLaVA: A Strong Training-Free Baseline for Video Large Language Models

Reference 82

Resolution
verified exact
arxiv_id, observed 2026-05-23T07:42:43.161469Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-23T07:42:22.478647Z digest=sha256:8795bb7c878dcc8d6799a4c20fd465ebeff249a82069934869a74e5b7865bd40

Observation 47d7345f-2cd6-49c2-ae8a-22677989dc35 · inbound

LinVT: Empower Your Image-level Large Language Model to Understand Videos cites this paper.

LinVT: Empower Your Image-level Large Language Model to Understand Videos SlowFast-LLaVA: A Strong Training-Free Baseline for Video Large Language Models

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-11T20:54:16.543027Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:54:16.543027Z digest=sha256:78b8e93a22e6ffda9426c56bb083928234f680340eabe9c8328da0184eb5f0e1

Observation 87d0de7b-e66a-447e-bf0d-8fa8f8cb0c7b · inbound

Dynamic-VLM: Simple Dynamic Visual Token Compression for VideoLLM cites this paper.

Dynamic-VLM: Simple Dynamic Visual Token Compression for VideoLLM SlowFast-LLaVA: A Strong Training-Free Baseline for Video Large Language Models

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-11T17:02:17.010894Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:02:17.010894Z digest=sha256:58d7cf45a13f02cadaeff86aa89b5e6aeb9a4ce1f55c8c99a275ce6b68510c72

Observation ebe804c1-4aeb-4ef2-81e3-a8dc5116d457 · inbound

Apollo: An Exploration of Video Understanding in Large Multimodal Models cites this paper.

Apollo: An Exploration of Video Understanding in Large Multimodal Models SlowFast-LLaVA: A Strong Training-Free Baseline for Video Large Language Models

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-11T16:11:10.710115Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:11:10.710115Z digest=sha256:7ad6daba35f1c5b7f585629f671cf2da32f00a7ea6bddabecb5d6595ee504f29

Observation 2645119b-22bc-472e-8c14-06d5518652da · inbound

FocusChat: Text-guided Long Video Understanding via Spatiotemporal Information Filtering cites this paper.

FocusChat: Text-guided Long Video Understanding via Spatiotemporal Information Filtering SlowFast-LLaVA: A Strong Training-Free Baseline for Video Large Language Models

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-11T13:44:50.577987Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:44:50.577987Z digest=sha256:3cb255b1907e5aea7908395a256e4a58355db3336d5fdce1b02f0efa45c7a796

Observation 5252730d-c4ac-44a8-9b52-b2d50ffadf6b · inbound

HLV-1K: A Large-scale Hour-Long Video Benchmark for Time-Specific Long Video Understanding cites this paper.

HLV-1K: A Large-scale Hour-Long Video Benchmark for Time-Specific Long Video Understanding SlowFast-LLaVA: A Strong Training-Free Baseline for Video Large Language Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-10T22:27:56.445639Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:27:56.445639Z digest=sha256:dc4d12be2cce3b8fa4a9f7ec22a33dd63b56fff26f52b15ed0fb56fd2f22a929

Observation 8cdbed8d-6423-4f4c-9c10-2ec71ced3420 · inbound

FrameFusion: Combining Similarity and Importance for Video Token Reduction on Large Vision Language Models cites this paper.

FrameFusion: Combining Similarity and Importance for Video Token Reduction on Large Vision Language Models SlowFast-LLaVA: A Strong Training-Free Baseline for Video Large Language Models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-10T23:09:25.126995Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:09:25.126995Z digest=sha256:a9bf20ab00c90dde421d27ecb15bac75f17a70628f990867faad6246b1fff116

Observation d84ca329-0445-4314-8988-8bf20ebd8055 · inbound

VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding cites this paper.

VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding SlowFast-LLaVA: A Strong Training-Free Baseline for Video Large Language Models

Reference 147

Resolution
verified exact
arxiv_id, observed 2026-05-11T01:20:00.062242Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-11T01:19:59.603343Z digest=sha256:e829810c7bc01b56e8f98371488c7f7b897534075e5c5c008712dd1c5e83518a

Observation 8e6a65e6-3602-4ea7-b3d4-6eab3cc07ab7 · inbound

WorldSense: Evaluating Real-world Omnimodal Understanding for Multimodal LLMs cites this paper.

WorldSense: Evaluating Real-world Omnimodal Understanding for Multimodal LLMs SlowFast-LLaVA: A Strong Training-Free Baseline for Video Large Language Models

Reference 76

Resolution
verified exact
arxiv_id, observed 2026-05-17T05:53:26.197276Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-17T05:53:26.066674Z digest=sha256:6b777ff2095391a5d7692d5ed926e1e08494ff33a632efb930b77bc56dbed1e3

Observation 7a46c902-f8b5-4033-88b7-c49360ce7c9a · inbound

Breaking Down Video LLM Benchmarks: Knowledge, Spatial Perception, or True Temporal Understanding? cites this paper.

Breaking Down Video LLM Benchmarks: Knowledge, Spatial Perception, or True Temporal Understanding? SlowFast-LLaVA: A Strong Training-Free Baseline for Video Large Language Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T15:41:09.214825Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:41:09.214825Z digest=sha256:84b4a6c56cd5001a8d5ef2d3d0161a0b44d8e1b2b407b44288595fa6e9a697bd

Observation 0121dee9-6d54-4503-89a9-2066e849809e · inbound

UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation cites this paper.

UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation SlowFast-LLaVA: A Strong Training-Free Baseline for Video Large Language Models

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:55.852874Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:55.852874Z digest=sha256:882fa5551ad4b85e02c7e0fa1780ce624145de4667f56b5b1063942dead4f8db

Observation 91fb4939-f0b4-4ccc-a4c0-9fc639f510d9 · inbound

LiveVLM: Efficient Online Video Understanding via Streaming-Oriented KV Cache and Retrieval cites this paper.

LiveVLM: Efficient Online Video Understanding via Streaming-Oriented KV Cache and Retrieval SlowFast-LLaVA: A Strong Training-Free Baseline for Video Large Language Models

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-22T14:31:40.800513Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-22T14:26:59.015559Z digest=sha256:3945762ac8685cad885a76bc0ae7bd29950783e1519973945e456a1456a604a2

Observation 0a170583-a703-4bcf-9e43-4138824d1160 · inbound

Clapper: Compact Learning and Video Representation in VLMs cites this paper.

Clapper: Compact Learning and Video Representation in VLMs SlowFast-LLaVA: A Strong Training-Free Baseline for Video Large Language Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T15:20:47.549466Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:20:47.549466Z digest=sha256:ca8d7cc6beedd1cda120fe00d3dee0c24cc5c74f910b4e8388a5f1a1b066f0a6

Observation cba55f33-db7f-408d-99fa-78317f5baaca · inbound

Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders cites this paper.

Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders SlowFast-LLaVA: A Strong Training-Free Baseline for Video Large Language Models

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-07T12:37:23.032653Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:37:23.032653Z digest=sha256:b692821d9fc896f7a4893c9aacc90bb1419dff344b9f6d43f12f0b1510935986

Observation 2ca677bd-85df-4f23-bf31-c6e4b77df4bb · inbound

MOOSE: Pay Attention to Temporal Dynamics for Video Understanding via Optical Flows cites this paper.

MOOSE: Pay Attention to Temporal Dynamics for Video Understanding via Optical Flows SlowFast-LLaVA: A Strong Training-Free Baseline for Video Large Language Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T11:55:28.590775Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:55:28.590775Z digest=sha256:ea8f37a0da9a8b47f6c8fe1429e941cc3378a9b1cc66ec39e08d099569eb0f14

Observation 489e65a4-c689-416e-a146-ba4a2a8813dd · inbound

IPFormer-VideoLLM: Enhancing Multi-modal Video Understanding for Multi-shot Scenes cites this paper.

IPFormer-VideoLLM: Enhancing Multi-modal Video Understanding for Multi-shot Scenes SlowFast-LLaVA: A Strong Training-Free Baseline for Video Large Language Models

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-06T22:37:39.555689Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:37:39.555689Z digest=sha256:6dd62eae535012e259bc0cd48778b437bcfb198c802ea2f5f9c419fa6bf441c2

Observation 9ec28918-801c-48e3-a1b7-284e41643535 · inbound

Region-Aware Multimodal Large Language Model via SlowFast Tokenization and Pseudo-Mask Guidance for 3D CT Report Generation cites this paper.

Region-Aware Multimodal Large Language Model via SlowFast Tokenization and Pseudo-Mask Guidance for 3D CT Report Generation SlowFast-LLaVA: A Strong Training-Free Baseline for Video Large Language Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T21:55:47.885592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:55:47.885592Z digest=sha256:df69ff06d77fba6f7877243412763a8cc47b8999c6ca4b1dbfd9055efb4f58a9

Observation 52342c69-8d84-47d4-908e-36077b7c5952 · inbound

MANTA: Cross-Modal Semantic Alignment and Information-Theoretic Optimization for Long-form Multimodal Understanding cites this paper.

MANTA: Cross-Modal Semantic Alignment and Information-Theoretic Optimization for Long-form Multimodal Understanding SlowFast-LLaVA: A Strong Training-Free Baseline for Video Large Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T22:00:07.796070Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:00:07.796070Z digest=sha256:c9800ce550f6b6815cd419ba8698c1b20d2902d187386687ddd7cbd34c3e0036

Observation 53ec807f-a90b-4c6d-a069-f22338da95a8 · inbound

IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning cites this paper.

IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning SlowFast-LLaVA: A Strong Training-Free Baseline for Video Large Language Models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T14:36:29.673490Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:36:29.673490Z digest=sha256:3a1d5a1a3f76c78610516007456ffbb7ece9a77b4e10f1f3dd30a8c195dba748

Observation e49916f7-c529-4d14-8a85-dbf10b8ef374 · inbound

Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data cites this paper.

Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data SlowFast-LLaVA: A Strong Training-Free Baseline for Video Large Language Models

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-05T10:58:18.397448Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:58:18.397448Z digest=sha256:fa40f296386126f54c84539bee7328f93d19f787b11dcd9535c72586eba20d1f

Observation 13d067ec-0901-4417-ae9a-7df704bc5cd6 · inbound

Video Parallel Scaling: Aggregating Diverse Frame Subsets for VideoLLMs cites this paper.

Video Parallel Scaling: Aggregating Diverse Frame Subsets for VideoLLMs SlowFast-LLaVA: A Strong Training-Free Baseline for Video Large Language Models

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-18T18:36:44.160046Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-18T18:35:01.328250Z digest=sha256:c63f7c024c1a0cc9d7e3c6ba9f5240542c76a09517e8f520da0a276372958871

Observation b41929b7-75f7-4fcb-b25d-25142ef52b0c · inbound

MASS: Motion-Aware Spatial-Temporal Grounding for Physics Reasoning and Comprehension in Vision-Language Models cites this paper.

MASS: Motion-Aware Spatial-Temporal Grounding for Physics Reasoning and Comprehension in Vision-Language Models SlowFast-LLaVA: A Strong Training-Free Baseline for Video Large Language Models

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-05-17T05:59:08.712440Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-17T05:55:11.495430Z digest=sha256:a54f3391a10945c65bb6624b758eba15362fc7f18f42ec10ff2073aa580c15ed

Observation e54c71a7-b589-40e4-8698-aa6b8ab7bf57 · inbound

Molmo2: Open Weights and Data for Vision-Language Models with Video Understanding and Grounding cites this paper.

Molmo2: Open Weights and Data for Vision-Language Models with Video Understanding and Grounding SlowFast-LLaVA: A Strong Training-Free Baseline for Video Large Language Models

Reference 164

Resolution
verified exact
arxiv_id, observed 2026-05-16T04:21:29.783947Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-16T04:21:29.526008Z digest=sha256:745d2878dee388d31489ffa72be1bfa738b7f45373122c850dfea9be6526087e

Observation e40e8b56-6ef9-4885-8737-9f6dc53236f1 · inbound

APB-V: Accelerating Long-Video Understanding via Sequence-Parallelism-aware Approximate Attention cites this paper.

APB-V: Accelerating Long-Video Understanding via Sequence-Parallelism-aware Approximate Attention SlowFast-LLaVA: A Strong Training-Free Baseline for Video Large Language Models

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-03T07:02:08.774478Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T07:02:08.774478Z digest=sha256:dfa791d89f86ff98f2bc55bcf0cb09185ea3ada813b9149f80b9e4abb7b2b9cb

Observation 4016aeb4-078b-482d-bb10-a143f7678108 · inbound

TrajTok: Learning Trajectory Tokens enables better Video Understanding cites this paper.

TrajTok: Learning Trajectory Tokens enables better Video Understanding SlowFast-LLaVA: A Strong Training-Free Baseline for Video Large Language Models

Reference 76

Resolution
verified exact
arxiv_id, observed 2026-05-15T19:16:31.733752Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-15T19:11:52.694778Z digest=sha256:8d24e58e7e57933fe7f3eb6f3f9ab6c344bde5776082aba80665382fcbed5cc1

Observation 835a5c98-cc5b-483f-a37e-9daf7baccd97 · inbound

TrajTok: Learning Trajectory Tokens enables better Video Understanding cites this paper.

TrajTok: Learning Trajectory Tokens enables better Video Understanding SlowFast-LLaVA: A Strong Training-Free Baseline for Video Large Language Models

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-02T20:38:44.615002Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T20:38:44.615002Z digest=sha256:3bf2ce80da5a6286b14722d6b4016ec024a3a81cab8e74290988b3bd1d13c7df

Observation 5962e419-38d9-4350-b3ba-319120e1c9f6 · inbound

One Token per Highly Selective Frame: Towards Extreme Compression for Long Video Understanding cites this paper.

One Token per Highly Selective Frame: Towards Extreme Compression for Long Video Understanding SlowFast-LLaVA: A Strong Training-Free Baseline for Video Large Language Models

Reference 70

Resolution
verified exact
arxiv_id, observed 2026-05-10T13:30:26.516769Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T13:28:58.920442Z digest=sha256:28c47072cacd7b76f7a437ee5bb65e832f686ac59e3548739d7127c6e26a2135

Observation 746f123b-cbf3-4c36-b8e7-f66d6fec7f4b · inbound

EgoSelf: From Memory to Personalized Egocentric Assistant cites this paper.

EgoSelf: From Memory to Personalized Egocentric Assistant SlowFast-LLaVA: A Strong Training-Free Baseline for Video Large Language Models

Reference 60

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T13:06:03.801398Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T02:23:21.119521Z digest=sha256:59791efae6c78e2df370bb3b9db354833d330e2031cb10942a9dca7e647d7866

Observation 8ed0375a-ec45-4b75-ba15-d609e1622694 · inbound

WindowQuant: Mixed-Precision KV Cache Quantization based on Window-Level Similarity for VLMs Inference Optimization cites this paper.

WindowQuant: Mixed-Precision KV Cache Quantization based on Window-Level Similarity for VLMs Inference Optimization SlowFast-LLaVA: A Strong Training-Free Baseline for Video Large Language Models

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-05-11T16:36:07.887239Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-09T16:06:26.450483Z digest=sha256:c3b09eb760e20c90cfe08b8d533aebb99402259c56a33b315ac03bd796ee1867

Observation 1e76134c-26d1-44d4-bacd-125336067b52 · inbound

VidPrism: Heterogeneous Mixture of Experts for Image-to-Video Transfer cites this paper.

VidPrism: Heterogeneous Mixture of Experts for Image-to-Video Transfer SlowFast-LLaVA: A Strong Training-Free Baseline for Video Large Language Models

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-06-29T13:03:25.914694Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-29T13:01:42.880738Z digest=sha256:ae0f9b3484124e8cb6cc2a575600510c872b648a37e8e4f8b4298c81927c2b04

Observation 265f85fd-529f-44cc-b304-b4fccd0d919c · inbound

Linear Scaling Video VLMs for Long Video Understanding cites this paper.

Linear Scaling Video VLMs for Long Video Understanding SlowFast-LLaVA: A Strong Training-Free Baseline for Video Large Language Models

Reference 72

Resolution
metadata mismatch
arxiv_id, observed 2026-06-28T23:02:46.236920Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-28T23:00:11.246232Z digest=sha256:f5ef2168690805ebb0538faf06bb2201ccdf2dbc3709765495bbec341ce74360

Observation e6c52ccc-1d76-4dfb-a9aa-00019a82e63c · inbound

V-LynX: Token Interface Alignment for Video+X LLMs cites this paper.

V-LynX: Token Interface Alignment for Video+X LLMs SlowFast-LLaVA: A Strong Training-Free Baseline for Video Large Language Models

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-06-28T19:02:34.161789Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-28T19:00:18.702121Z digest=sha256:0de993d5ee64de9c4e0670091aea7286d515acaca88b79427eb4f6991900d2f5

Observation d43f3e47-6db7-4ff9-a3f3-42971a4175bf · inbound

MemoryCard: Topic-Aware Multi-Modal Clue Compression for Long-Video Question Answering cites this paper.

MemoryCard: Topic-Aware Multi-Modal Clue Compression for Long-Video Question Answering SlowFast-LLaVA: A Strong Training-Free Baseline for Video Large Language Models

Reference 28

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T12:46:56.883859Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-06-28T01:52:33.768494Z digest=sha256:ec39634cb69736925353fac5870f792e06f5170651a9c944b5c48b5ec234fa95

Observation e9a935e4-350f-4f6a-90ff-8778fee50540 · inbound

Q-Fold: Query-Aware Focus-Context Spatio-Temporal Folding for Long Video Understanding cites this paper.

Q-Fold: Query-Aware Focus-Context Spatio-Temporal Folding for Long Video Understanding SlowFast-LLaVA: A Strong Training-Free Baseline for Video Large Language Models

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-07-03T10:27:56.050250Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-27T10:04:29.739632Z digest=sha256:41e6bcabdfcf3c5e51064903ed636ae9d9c47f787a4e943eaeb587562f46058b

Observation cf97c963-6366-4263-830e-b9b627f77316 · inbound

ViCoStream: Streaming VideoLLMs Can Run Beyond 100 FPS with Stage-Wise Coordinated Inference cites this paper.

ViCoStream: Streaming VideoLLMs Can Run Beyond 100 FPS with Stage-Wise Coordinated Inference SlowFast-LLaVA: A Strong Training-Free Baseline for Video Large Language Models

Reference 17

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T03:19:29.901021Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-06-26T18:17:53.013043Z digest=sha256:8738db026ee722f249f929641624e388f39234d11b4e462e341807c2e994ae2a

Observation 16129752-a127-47ff-a819-6c64e0f6fe8d · inbound

Efficient Frame Selection for Long Videos at Test Time with Attention-Based MLLM Selectors cites this paper.

Efficient Frame Selection for Long Videos at Test Time with Attention-Based MLLM Selectors SlowFast-LLaVA: A Strong Training-Free Baseline for Video Large Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-01T22:38:42.003198Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T22:38:42.003198Z digest=sha256:936def1e32ee01d854543eb5c8e7632039b224cd89675702e28b6ee41ac279c5

Observation 9ce06e3c-e938-4b4b-ab75-acf4c575595c · inbound

SkyVLaM: Multimodal Large Language Model for UAV Video Understanding in Remote Sensing cites this paper.

SkyVLaM: Multimodal Large Language Model for UAV Video Understanding in Remote Sensing SlowFast-LLaVA: A Strong Training-Free Baseline for Video Large Language Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-01T18:11:03.486399Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:11:03.486399Z digest=sha256:9f6b04bc619789f59957973fef93f945c14a07d07b435a53b15a2f3884ee271d

Observation e7103d0d-f290-430d-9da5-604b896cb835 · inbound

Mage-VL: An Efficient Codec-Native Streaming Multimodal Foundation Model cites this paper.

Mage-VL: An Efficient Codec-Native Streaming Multimodal Foundation Model SlowFast-LLaVA: A Strong Training-Free Baseline for Video Large Language Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-07-31T06:20:13.737936Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T06:20:13.737936Z digest=sha256:4989d58fac98658b6155c92a735106098dd8310a85d7c30a20854bda94a56fe1