Pith. sign in

Paper Citation Record · LEDGER

SlowFast-LLaVA-1.5: A Family of Token-Efficient Video Large Language Models for Long-Form Video Understanding

As of 19 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 10 inbound Pith citation observations for arXiv:2503.18943.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2503.18943 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 10 of 10 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 10 of 10 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:34:55.856714Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-29T23:14:02.175777Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 765079fd-8673-4ee6-80de-35b19f652f4c · inbound

UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation cites this paper.

UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation SlowFast-LLaVA-1.5: A Family of Token-Efficient Video Large Language Models for Long-Form Video Understanding

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:55.856714Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:55.856714Z digest=sha256:c6081053b6a6bcb2e35cadb3d917f7994b4b28cd0673b72b00392bc45989c7ae

Observation fdd0506c-935c-43ec-b472-bfb3ddaf2094 · inbound

FlexSelect: Flexible Token Selection for Efficient Long Video Understanding cites this paper.

FlexSelect: Flexible Token Selection for Efficient Long Video Understanding SlowFast-LLaVA-1.5: A Family of Token-Efficient Video Large Language Models for Long-Form Video Understanding

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T11:59:09.935958Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:59:09.935958Z digest=sha256:03bba0d001524802b1315cccb87e46bb758e55950da28047954cf48679e12eee

Observation cda3bfca-1bd2-4a86-b1ff-576b3bf1202b · inbound

Context-Aware Multi-Turn Visual-Textual Reasoning in LVLMs via Dynamic Memory and Adaptive Visual Guidance cites this paper.

Context-Aware Multi-Turn Visual-Textual Reasoning in LVLMs via Dynamic Memory and Adaptive Visual Guidance SlowFast-LLaVA-1.5: A Family of Token-Efficient Video Large Language Models for Long-Form Video Understanding

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-05T05:16:57.525652Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:16:57.525652Z digest=sha256:d8fd8f03aba72e5d60a64e442e8af0301e167810d537641ffe98a28ee7edc936

Observation 8cab5815-04f5-4877-b350-280990ddda3d · inbound

Stateful Token Reduction for Long-Video Hybrid VLMs cites this paper.

Stateful Token Reduction for Long-Video Hybrid VLMs SlowFast-LLaVA-1.5: A Family of Token-Efficient Video Large Language Models for Long-Form Video Understanding

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-02T20:14:05.162651Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T20:14:05.162651Z digest=sha256:35704b3989858f5cae574ef954006248ea98eac30d4a795cfa7126aae6535a75

Observation c26e1f9c-aca7-40f5-96cd-ea1a1892c99b · inbound

Towards Temporal Compositional Reasoning in Long-Form Sports Videos cites this paper.

Towards Temporal Compositional Reasoning in Long-Form Sports Videos SlowFast-LLaVA-1.5: A Family of Token-Efficient Video Large Language Models for Long-Form Video Understanding

Reference 43

Resolution
unresolved
no resolver link, observed 2026-07-14T19:27:29.843866Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T19:27:29.843866Z digest=sha256:b5fac30a2a0f893ebe4e2f5fec5dabd8da1f27ffb29b0bb2826a935981f11b88

Observation 9b7ca5d7-6d06-4dcf-b72c-144ae5a59245 · inbound

LookWhen? Fast Video Recognition by Learning When, Where, and What to Compute cites this paper.

LookWhen? Fast Video Recognition by Learning When, Where, and What to Compute SlowFast-LLaVA-1.5: A Family of Token-Efficient Video Large Language Models for Long-Form Video Understanding

Reference 54

Resolution
verified exact
arxiv_id, observed 2026-05-11T01:40:52.137771Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-11T01:36:50.406353Z digest=sha256:a4118ed7a96fe871c897ee800df4dcd2092c1d94f729de8e7bf3e0b078af5e11

Observation d36f150d-e770-4291-9c13-db06a1cf51b9 · inbound

Swift Sampling: Selecting Temporal Surprises via Taylor Series cites this paper.

Swift Sampling: Selecting Temporal Surprises via Taylor Series SlowFast-LLaVA-1.5: A Family of Token-Efficient Video Large Language Models for Long-Form Video Understanding

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-22T05:56:07.903630Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-22T05:55:23.479344Z digest=sha256:00f486a80d53be5c70a739f605ab9944e6cc4199624fc6e1e22d8d96656071a4

Observation 9c8bb1ec-0283-48c4-9bae-2a968b2deed9 · inbound

STORM: Internalized Modeling for Spatial-Temporal Reasoning in Video-Language Models cites this paper.

STORM: Internalized Modeling for Spatial-Temporal Reasoning in Video-Language Models SlowFast-LLaVA-1.5: A Family of Token-Efficient Video Large Language Models for Long-Form Video Understanding

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-06-29T23:14:02.177596Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-29T23:04:21.463842Z digest=sha256:f488dcc542cba4c4095b675fa826fcd9c17aeec84487cad7439db0983fec4252

Observation 9589498d-9658-46b0-9588-7730b46fa04e · inbound

IPIBench: Evaluating Interactive Proactive Intelligence of MLLMs under Continuous Streams cites this paper.

IPIBench: Evaluating Interactive Proactive Intelligence of MLLMs under Continuous Streams SlowFast-LLaVA-1.5: A Family of Token-Efficient Video Large Language Models for Long-Form Video Understanding

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-06-29T18:33:51.049541Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-29T18:24:57.881644Z digest=sha256:14431201225417a12a27c2f87d80571e4bf6cc434b7729881539de7bcab50c03

Observation b07bc466-d2f3-4511-bdac-08875e0c8160 · inbound

VidPrism: Heterogeneous Mixture of Experts for Image-to-Video Transfer cites this paper.

VidPrism: Heterogeneous Mixture of Experts for Image-to-Video Transfer SlowFast-LLaVA-1.5: A Family of Token-Efficient Video Large Language Models for Long-Form Video Understanding

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-06-29T13:03:25.906481Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-29T13:01:42.880738Z digest=sha256:dbc480606be83b3c9b9d59d48fae1f0d036735d20ae9c6bac7199ece9d56f48e