Pith. sign in

Paper Citation Record · LEDGER

Beyond Frame Selection: Rethinking Long-Video Understanding with MLLMs

As of 10 August 2026, this Paper Citation Record lists 13 of 13 outbound references and 0 inbound Pith citation observations for arXiv:2608.05592.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.05592 v1

Coverage vector

measured 13 of 13 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-08T05:58:12.511658Z

measured 13 of 13 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

13 of 13 outbound references displayed

  • verified exact0
  • verified fuzzy1
  • unresolved12
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 6def475b-324d-4410-88f9-f47b87a9bc99 · outbound

This paper cites PaLM 2 Technical Report.

Beyond Frame Selection: Rethinking Long-Video Understanding with MLLMs PaLM 2 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-08T05:58:12.472240Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:58:12.472240Z digest=sha256:7d9f9570a200d8278e6974a5f329dcce612745e2640438ed044902ee546c461e

Observation ecaf5288-9df8-4e82-a509-3ba87408f1fc · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

Beyond Frame Selection: Rethinking Long-Video Understanding with MLLMs Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-08T05:58:12.482845Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:58:12.482845Z digest=sha256:205d2a3afd3a269703ef59325b408ffd52f7ec1a515578e2a295ab2170a933d0

Observation 62a6bca5-905c-4c51-a7e7-0ad0095cd7bc · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

Beyond Frame Selection: Rethinking Long-Video Understanding with MLLMs LLaVA-OneVision: Easy Visual Task Transfer

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-08T05:58:12.489448Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:58:12.489448Z digest=sha256:b02a53ca4d34b4bd34c0136a25663d5197132431c2fae5aed2efbe03af1baafd

Observation 0c75d6b4-4d1b-4ecc-80f8-c31288db3ffa · outbound

This paper cites Commonsense video question answering through video-grounded entailment tree reasoning.

Beyond Frame Selection: Rethinking Long-Video Understanding with MLLMs Commonsense video question answering through video-grounded entailment tree reasoning

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-08T05:58:12.492986Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:58:12.492986Z digest=sha256:5042bed5328abb4259d3ad03e32420fe312d5a27fffb5bdde92023ee7dbf7ff7

Observation 2963d8cb-b5c7-4071-a8bd-c57a84f4359e · outbound

This paper cites Video-MTR: Reinforced Multi-Turn Reasoning for Long Video Understanding.

Beyond Frame Selection: Rethinking Long-Video Understanding with MLLMs Video-MTR: Reinforced Multi-Turn Reasoning for Long Video Understanding

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-08T05:58:12.496025Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:58:12.496025Z digest=sha256:4b613560d2fe3f9f1efa1e767cb00a663566d78a3cdbb121f856e2a00a215d7d

Observation ee3767d9-c33d-43d2-bb66-3cda6ad7a77e · outbound

This paper cites Qwen2 Technical Report.

Beyond Frame Selection: Rethinking Long-Video Understanding with MLLMs Qwen2 Technical Report

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-08T05:58:12.499254Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:58:12.499254Z digest=sha256:6a62f290396a5cf980c3001e77586ea1aae2d8a03d28f93a5da2a70e7159b724

Observation f23307aa-93c6-45e4-91f0-af19404e5a18 · outbound

This paper cites mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality.

Beyond Frame Selection: Rethinking Long-Video Understanding with MLLMs mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-08T05:58:12.502565Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:58:12.502565Z digest=sha256:226cd7287fbf4d2730a911a3e5903f50540b3d36bb5dfbb67a871a2d7e8435c2

Observation 22419d38-e134-4199-bea9-b3855609dab1 · outbound

This paper cites Thinking With Videos: Multimodal Tool-Augmented Reinforcement Learning for Long Video Reasoning.

Beyond Frame Selection: Rethinking Long-Video Understanding with MLLMs Thinking With Videos: Multimodal Tool-Augmented Reinforcement Learning for Long Video Reasoning

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-08T05:58:12.505711Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:58:12.505711Z digest=sha256:fe9b5eb388c19eeb55d03f31e1d571aaa40448db3dbd5e07e55f5a066573ec1e

Observation 53c686ca-3b8d-43fd-be20-a2701b80165b · outbound

This paper cites Videolucy: Deep memory backtracking for long video understanding.arXiv:2510.12422,.

Beyond Frame Selection: Rethinking Long-Video Understanding with MLLMs Videolucy: Deep memory backtracking for long video understanding.arXiv:2510.12422,

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-08T05:58:12.508728Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:58:12.508728Z digest=sha256:5f89de9aebfcad410d7a853a721bbfcf0f6d5fa9bcc8519c8d7d20429c2949c0

Observation dbf2bbd4-9c1f-40b3-8550-3b437f0bf446 · outbound

This paper cites 14 C.2 Ablation on the Global Representation Depth.

Beyond Frame Selection: Rethinking Long-Video Understanding with MLLMs 14 C.2 Ablation on the Global Representation Depth

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T05:58:12.973330Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T05:58:12.511658Z digest=sha256:453e8a926b195c766d8487fd550dcc453bf51a5a6ef2955cc8f3e4a3b0131e79

Observation b0d01f99-df2b-4f19-80e9-377357e4ccb3 · outbound

This paper cites Temporal Chain of Thought: Long-Video Understanding by Thinking in Frames.

Beyond Frame Selection: Rethinking Long-Video Understanding with MLLMs Temporal Chain of Thought: Long-Video Understanding by Thinking in Frames

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-08T05:58:12.475954Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:58:12.475954Z digest=sha256:30718285f487d01a5c6bff1522c5e710e2f723e7376428f369abea7e80aed30b

Observation fffa4d07-097b-4e3e-aa76-2dfb45a3bd33 · outbound

This paper cites GPT-4o System Card.

Beyond Frame Selection: Rethinking Long-Video Understanding with MLLMs GPT-4o System Card

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-08T05:58:12.486291Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:58:12.486291Z digest=sha256:1f40a2b7ca1b239c1d59590f1dd56ca22d44af17532b62e1a88129e578b0dbcd

Observation 884443e7-ef25-4171-a095-1834aa17c2ea · outbound

This paper cites MiniGPT4-Video: Advancing Multimodal LLMs for Video Understanding with Interleaved Visual-Textual Tokens.

Beyond Frame Selection: Rethinking Long-Video Understanding with MLLMs MiniGPT4-Video: Advancing Multimodal LLMs for Video Understanding with Interleaved Visual-Textual Tokens

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-08T05:58:12.479352Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:58:12.479352Z digest=sha256:c41498b57f7c4999af354a9ce8041b24f9bdca7da3116da1ce2172ba4c5b3308

Pith citing papers

No inbound Pith citation observations are available.