Pith. sign in

Paper Citation Record · LEDGER

LongViTU: Instruction Tuning for Long-Form Video Understanding

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 5 inbound Pith citation observations for arXiv:2501.05037.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.05037 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 5 of 5 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 5 of 5 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:20:59.162625Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T07:16:45.171051Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 1445af28-cabc-47dc-8c25-d4d3be3d296f · inbound

ViaRL: Adaptive Temporal Grounding via Visual Iterated Amplification Reinforcement Learning cites this paper.

ViaRL: Adaptive Temporal Grounding via Visual Iterated Amplification Reinforcement Learning LongViTU: Instruction Tuning for Long-Form Video Understanding

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T15:20:59.162625Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:20:59.162625Z digest=sha256:9a98e16b4ce2217d6f93a2738bde497530c9ade8bc48cbfcd40101dfd671681c

Observation aea3764e-83dd-40b8-9999-1eb2fca7070d · inbound

Infinite Video Understanding cites this paper.

Infinite Video Understanding LongViTU: Instruction Tuning for Long-Form Video Understanding

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-06T18:10:14.416661Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:10:14.416661Z digest=sha256:3efd9ae0969ba0b6b184acd0ed061c699149c3d843eed23848814a422b1ba717

Observation 88627f00-7b73-4104-9607-25201e2aaa50 · inbound

Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data cites this paper.

Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data LongViTU: Instruction Tuning for Long-Form Video Understanding

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-05T10:58:18.003545Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:58:18.003545Z digest=sha256:260deef33936c45fcfa4893ff111f4b37434f213921a51942cddddf77ef28da7

Observation e39ccf5c-5ad5-4945-9d88-75828525e7ab · inbound

Less Data, Faster Convergence: Goal-Driven Data Optimization for Multimodal Instruction Tuning cites this paper.

Less Data, Faster Convergence: Goal-Driven Data Optimization for Multimodal Instruction Tuning LongViTU: Instruction Tuning for Long-Form Video Understanding

Reference 58

Resolution
unresolved
no resolver link, observed 2026-07-14T22:18:56.115559Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T22:18:56.115559Z digest=sha256:ecc5905fee8582d2eec81a084212d8bb8aa026a149da25b65246f212b115b836

Observation 2c57ab1e-4cf6-4588-a651-db91dc8b6080 · inbound

VideoKR: Towards Knowledge- and Reasoning-Intensive Video Understanding cites this paper.

VideoKR: Towards Knowledge- and Reasoning-Intensive Video Understanding LongViTU: Instruction Tuning for Long-Form Video Understanding

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-07-02T07:16:45.173173Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-28T07:00:21.192082Z digest=sha256:1f8b8e8090d61dfccb2297542549f368d16471ee302726cdc417bd6f19af8cb6