Pith. sign in

Paper Citation Record · LEDGER

KeyVideoLLM: Towards Large-scale Video Keyframe Selection

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 18 inbound Pith citation observations for arXiv:2407.03104.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2407.03104 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 18 of 18 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 18 of 18 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T23:42:35.271317Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

1
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation da9d7c4a-a09d-46f1-aee2-ea29148cb717 · inbound

ViaRL: Adaptive Temporal Grounding via Visual Iterated Amplification Reinforcement Learning cites this paper.

ViaRL: Adaptive Temporal Grounding via Visual Iterated Amplification Reinforcement Learning KeyVideoLLM: Towards Large-scale Video Keyframe Selection

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T15:20:57.485440Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:20:57.485440Z digest=sha256:3b88581acd417c88591baadd2668e67bf70ff928aef974de4769a71512ace69e

Observation 34c1edd3-b612-4a39-9548-dd7104c7a1fd · inbound

TriPSS: A Tri-Modal Keyframe Extraction Framework Using Perceptual, Structural, and Semantic Representations cites this paper.

TriPSS: A Tri-Modal Keyframe Extraction Framework Using Perceptual, Structural, and Semantic Representations KeyVideoLLM: Towards Large-scale Video Keyframe Selection

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T11:11:41.088190Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:11:41.088190Z digest=sha256:61749ce2bbb377e6275223c9f78ff6c17f565b59f5b88857b22fba49173e9fc0

Observation b2aa391c-367c-48d0-ac88-a8272f731738 · inbound

Flattery in Motion: Benchmarking and Analyzing Sycophancy in Video-LLMs cites this paper.

Flattery in Motion: Benchmarking and Analyzing Sycophancy in Video-LLMs KeyVideoLLM: Towards Large-scale Video Keyframe Selection

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-19T11:07:15.442129Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-19T11:03:59.222849Z digest=sha256:aada4baa9ba91baa80a19b5ac4c7d3c15c69a9ec6e1e9cc20d9cc2f60836b5ec

Observation 76dc95d8-f7e4-4b88-9a02-94e10a7c9b77 · inbound

Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs cites this paper.

Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs KeyVideoLLM: Towards Large-scale Video Keyframe Selection

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T22:15:02.922173Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:15:02.922173Z digest=sha256:ae3d0c6a2d69b0f3e9eb2750587cfd3ad1e72c53af1883b5408b6ecb59099581

Observation 3fd6c35e-4a3e-4025-9b32-9ad7adee1bcf · inbound

AnchorSync: Global Consistency Optimization for Long Video Editing cites this paper.

AnchorSync: Global Consistency Optimization for Long Video Editing KeyVideoLLM: Towards Large-scale Video Keyframe Selection

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-05T18:29:01.422233Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T18:29:01.422233Z digest=sha256:6c25254eb098cb9496f7ae2b7c902d6f577e89e4a0c1d54b4248b3aced7288d4

Observation 7e66f06b-3700-47bf-8853-b5f39ea919ff · inbound

OpenWorldLib: A Unified Codebase and Definition of Advanced World Models cites this paper.

OpenWorldLib: A Unified Codebase and Definition of Advanced World Models KeyVideoLLM: Towards Large-scale Video Keyframe Selection

Reference 73

Resolution
verified exact
arxiv_id, observed 2026-05-10T22:45:48.672481Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T19:36:42.100191Z digest=sha256:79b5e3fed48f8566ad5b7e880a2f0583b27078b19bed84762095088426525197

Observation 9707cea1-f099-4dc8-96b8-b678058fa7b8 · inbound

OpenWorldLib: A Unified Codebase and Definition of Advanced World Models cites this paper.

OpenWorldLib: A Unified Codebase and Definition of Advanced World Models KeyVideoLLM: Towards Large-scale Video Keyframe Selection

Reference 73

Resolution
unresolved
no resolver link, observed 2026-07-13T09:42:23.808691Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T09:42:23.808691Z digest=sha256:8ee336f42049e06dda128007e98513f7b6ba014b161d479e2090ce4331cb130e

Observation c99e4a4a-b0d5-494b-bd4a-ba7a6dc0c52d · inbound

Efficient Inference for Large Vision-Language Models: Bottlenecks, Techniques, and Prospects cites this paper.

Efficient Inference for Large Vision-Language Models: Bottlenecks, Techniques, and Prospects KeyVideoLLM: Towards Large-scale Video Keyframe Selection

Reference 5

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T23:45:50.934752Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T18:54:04.104227Z digest=sha256:d00baac975b3e25da375a9f634bd4c42e3a85efff561220ccdf2279da77619b0

Observation 284dbbde-cc6e-41b3-aa94-546ccfcbc54c · inbound

LDDR: Linear-DPP-Based Dynamic-Resolution Frame Sampling for Video MLLMs cites this paper.

LDDR: Linear-DPP-Based Dynamic-Resolution Frame Sampling for Video MLLMs KeyVideoLLM: Towards Large-scale Video Keyframe Selection

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-13T03:22:10.129873Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-13T03:20:42.684286Z digest=sha256:7bce23eb4ebd2b46ac60e536ca71e86273a549c717e626e7f7c2251d9080e5f4

Observation 02682ce4-8788-49cd-bd2e-0e84da35cd75 · inbound

Multi-Scale Separable Fourier Neural Networks for Solving High-Frequency PDEs cites this paper.

Multi-Scale Separable Fourier Neural Networks for Solving High-Frequency PDEs KeyVideoLLM: Towards Large-scale Video Keyframe Selection

Reference 14

Resolution
unresolved
no resolver link, observed 2026-07-12T15:34:37.002001Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T15:34:37.002001Z digest=sha256:fbbb426244b1aee8738660ef3fe08497bcac906a9a01f52caeb530915e5d14d5

Observation 457b445e-5618-4d49-83f0-6d365ce1734a · inbound

Multi-Scale Separable Fourier Neural Networks for Solving High-Frequency PDEs cites this paper.

Multi-Scale Separable Fourier Neural Networks for Solving High-Frequency PDEs KeyVideoLLM: Towards Large-scale Video Keyframe Selection

Reference 14

Resolution
unresolved
no resolver link, observed 2026-07-14T18:39:24.915547Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T18:39:24.915547Z digest=sha256:336d6f92747fcdd4781ee703139ad295ff7e3d2000eec52c013990fceff93255

Observation 77a491e7-d551-4df2-bbed-52b973c054ee · inbound

PEEK: Picking Essential frames via Efficient Knowledge distillation cites this paper.

PEEK: Picking Essential frames via Efficient Knowledge distillation KeyVideoLLM: Towards Large-scale Video Keyframe Selection

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-07-01T19:16:01.296697Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-28T22:46:53.704892Z digest=sha256:8b2a93df1e3577097523615c3075a7facfbe9972676eca136b9a90e9ec9ae343

Observation 73a56052-63e4-4ab0-96e0-4439b22c03bb · inbound

AdaCodec: A Predictive Visual Code for Video MLLMs cites this paper.

AdaCodec: A Predictive Visual Code for Video MLLMs KeyVideoLLM: Towards Large-scale Video Keyframe Selection

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-06-28T15:22:19.548791Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-28T15:20:48.248576Z digest=sha256:afc409d74daa1378f3cf232698260d01cedf1ac794cdbb95bc6cccb15f736f2a

Observation 3c27c65e-00ef-4d6b-a49f-1eadd7d241e0 · inbound

Q-Fold: Query-Aware Focus-Context Spatio-Temporal Folding for Long Video Understanding cites this paper.

Q-Fold: Query-Aware Focus-Context Spatio-Temporal Folding for Long Video Understanding KeyVideoLLM: Towards Large-scale Video Keyframe Selection

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-07-03T10:27:56.042192Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-27T10:04:29.739632Z digest=sha256:b3a1bd74bd7dd25f0591c867ea3e6a9bb3527fcdad02509077173c526f72a501

Observation e7ad3a05-4fa2-49c3-ab26-2558ecaec89c · inbound

QCA: Query- and Content-Aware Keyframe Selection for Long Video Understanding cites this paper.

QCA: Query- and Content-Aware Keyframe Selection for Long Video Understanding KeyVideoLLM: Towards Large-scale Video Keyframe Selection

Reference 20

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T14:17:02.660613Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-02T14:09:54.549499Z digest=sha256:f8a39ebb6240f57b257df30132abebf4075303813e286bdecd3cf23449a07726

Observation 278fde52-fa8a-4fbb-9267-5625a466fbdc · inbound

Reasoning with Memory: A Temporal Granularity-Adaptive Framework for Training-Free Long Video Understanding cites this paper.

Reasoning with Memory: A Temporal Granularity-Adaptive Framework for Training-Free Long Video Understanding KeyVideoLLM: Towards Large-scale Video Keyframe Selection

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-02T09:17:00.597095Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T09:17:00.597095Z digest=sha256:5b08e91a5f29eb4ca675aa6f7c8bae25c96e068096e7518bf340668518d34394

Observation e8524394-313c-4732-8c3b-06a91ccf5ab6 · inbound

VisualRouter: Query-Grounded Visual Sampling for Long Video Understanding cites this paper.

VisualRouter: Query-Grounded Visual Sampling for Long Video Understanding KeyVideoLLM: Towards Large-scale Video Keyframe Selection

Reference 15

Resolution
unresolved
no resolver link, observed 2026-07-31T06:37:32.960181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T06:37:32.960181Z digest=sha256:8950688249e4b832b829e040a1244d50ca870d4338c5f7168b6fa1caecb1e7d1

Observation d5d5aada-913e-4dd2-af7c-e63665cd10d5 · inbound

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding cites this paper.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding KeyVideoLLM: Towards Large-scale Video Keyframe Selection

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-07T23:42:35.271317Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:42:35.271317Z digest=sha256:58003978e1c33bb65da3ed6c04144efdcffb138ff9783747f3ee8b186159e012