Pith. sign in

Paper Citation Record · LEDGER

Infinite Video Understanding

As of 14 August 2026, this Paper Citation Record lists 64 of 64 outbound references and 0 inbound Pith citation observations for arXiv:2507.09068.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.09068 v2

Coverage vector

measured 64 of 64 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T18:10:14.919272Z

measured 64 of 64 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

64 of 64 outbound references displayed

  • verified exact8
  • verified fuzzy1
  • unresolved54
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation ff0a1ede-3856-4a74-ad1e-e3544cd0a8b4 · outbound

This paper cites CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders.

Infinite Video Understanding CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-08-06T18:10:16.731929Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T18:09:18.061419Z digest=sha256:121a5193885cb07dc319ece1e70ba855c06b9ef1e2514d5e9a292c80d58a5405

Observation c25e915d-4ba9-493f-847b-51365e8f7fda · outbound

This paper cites an unresolved cited work.

Infinite Video Understanding Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:10:18.299149Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T18:09:18.114356Z digest=sha256:feb3a9e8949fe74118d6580a22cd043f5ce650aef90d3c414a35afe0169908df

Observation 01bcc5ab-e0c7-4408-a60e-48774759a8fe · outbound

This paper cites HourVideo: 1-Hour Video-Language Understanding.

Infinite Video Understanding HourVideo: 1-Hour Video-Language Understanding

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T18:09:18.211447Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:09:18.211447Z digest=sha256:e400bee3d12396fb12532edd757a9ae266241706d0b3daaf574ef251014f728d

Observation 339e7bb3-dc67-4a6f-8af4-aff75a9dfa3c · outbound

This paper cites LongVILA: Scaling Long-Context Visual Language Models for Long Videos.

Infinite Video Understanding LongVILA: Scaling Long-Context Visual Language Models for Long Videos

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T18:09:18.313230Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:09:18.313230Z digest=sha256:216787689d7533d3a1a724374ee80cca2c8c144ef21fba092dbbbc0717581c53

Observation 809b2dcf-8bb2-479b-964a-3488ccdb6144 · outbound

This paper cites Scaling Video-Language Models to 10K Frames via Hierarchical Differential Distillation.

Infinite Video Understanding Scaling Video-Language Models to 10K Frames via Hierarchical Differential Distillation

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T18:09:18.568750Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:09:18.568750Z digest=sha256:96127714a56434cdcdec8f2f1b4917de5af4c7b5c7f30c1ca90f3895c38096a6

Observation 8d053b18-1eb2-430a-9e88-2a6c782636e8 · outbound

This paper cites Video-Holmes: Can MLLM Think Like Holmes for Complex Video Reasoning?.

Infinite Video Understanding Video-Holmes: Can MLLM Think Like Holmes for Complex Video Reasoning?

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T18:09:18.954541Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:09:18.954541Z digest=sha256:482e0c4f5749a016dd76ede08ece03e130b9add526e33439599f9a922ed36067

Observation dbef02ca-88ff-488e-a914-02db772cb241 · outbound

This paper cites Zero-Shot Video Question Answering with Procedural Programs.

Infinite Video Understanding Zero-Shot Video Question Answering with Procedural Programs

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T18:09:19.409539Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:09:19.409539Z digest=sha256:43fb152763549f8d55a4fa6513ae4c6fc58a2953423d3ed5972520d69abfd615

Observation 68494767-b952-4b9b-bba1-bbddf06ed232 · outbound

This paper cites Project Aria: A New Tool for Egocentric Multi-Modal AI Research.

Infinite Video Understanding Project Aria: A New Tool for Egocentric Multi-Modal AI Research

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T18:10:11.725676Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:10:11.725676Z digest=sha256:a4119e1d8a3467553a968a16b37d9f8c2aae513744f7c1dd0675067f222008e4

Observation 545305b2-f2bb-4383-9be8-639125a12ebf · outbound

This paper cites Efficient Image Pre-Training with Siamese Cropped Masked Autoencoders.

Infinite Video Understanding Efficient Image Pre-Training with Siamese Cropped Masked Autoencoders

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-08-06T18:10:17.212457Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T18:10:11.854309Z digest=sha256:487427abbaf2b3f55ef0ba2df91b806fb3d519bf815d8e862875bfed10d69aff

Observation 316ad4f0-026d-4451-84aa-0b1e8874df9e · outbound

This paper cites an unresolved cited work.

Infinite Video Understanding Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:10:18.130805Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T18:10:11.899351Z digest=sha256:17134dc3770c97e4d3abe3b1065b1c2dcd8b4a17a42f150223603c402599b0c7

Observation 276abd28-1608-4fa1-8d37-0830654d553a · outbound

This paper cites Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis.

Infinite Video Understanding Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T18:10:12.056934Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:10:12.056934Z digest=sha256:fa3c75c9b0942094b1462d4612711e6202b197b0a34f00b370f3e3f4251bd528

Observation 4cbccd92-ffbe-4212-98e0-67fa7f83e1a5 · outbound

This paper cites The Ethics of Advanced AI Assistants.

Infinite Video Understanding The Ethics of Advanced AI Assistants

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T18:10:12.125589Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:10:12.125589Z digest=sha256:12b8174b32c60968edf75e3a686b0a49d72f175d1a7073182ef83975b02fbf28

Observation 963c17ec-82c8-4343-91e9-d85d06aa17ca · outbound

This paper cites Siamese Masked Autoencoders.

Infinite Video Understanding Siamese Masked Autoencoders

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T18:10:12.195376Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:10:12.195376Z digest=sha256:c2505a24cf456a2c0363ccfda5b025d7ba2973503e50a18715fc508c1e5bbc79

Observation 12718065-f4c4-4ac8-9980-1d22368d0953 · outbound

This paper cites RULER: What's the Real Context Size of Your Long-Context Language Models?.

Infinite Video Understanding RULER: What's the Real Context Size of Your Long-Context Language Models?

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T18:10:12.307299Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:10:12.307299Z digest=sha256:9f1fea70da6b0facf462d46b0838af3056bbc870fe859129a39fae76480547b9

Observation 9110064f-0fc6-4002-a20d-c5fe49e3a397 · outbound

This paper cites Visual Representation Learning with Stochastic Frame Prediction.

Infinite Video Understanding Visual Representation Learning with Stochastic Frame Prediction

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-08-06T18:10:16.421305Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T18:10:12.401921Z digest=sha256:e8899c205fd5dfd4de4203310e0c96de8def550c8dd329f7808a0cdae10a6833

Observation 2b5d77ea-50d5-4d97-8f44-c7bd15d9693d · outbound

This paper cites Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning.

Infinite Video Understanding Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T18:10:12.472087Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:10:12.472087Z digest=sha256:991f99ad8181fb94030fcc510f115cd127eea5cfdfb9cc825f59d1e7081950f0

Observation 0b728617-5100-420c-887d-8c52582f04f6 · outbound

This paper cites SNeRV: Spectra-preserving Neural Representation for Video.

Infinite Video Understanding SNeRV: Spectra-preserving Neural Representation for Video

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-08-06T18:10:16.964469Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T18:10:12.519667Z digest=sha256:2bf17fa12de03ff8b516f418335fc73311140383f5aa08dc392cc13e4a0163f6

Observation fb2b3eb7-0cd3-4be0-bd16-9ca2d015fd81 · outbound

This paper cites an unresolved cited work.

Infinite Video Understanding Unresolved cited work

Reference 18

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:10:17.987768Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T18:10:12.559183Z digest=sha256:635ad7b8991ba256bc7143492074de6095fb642b60f285bbda90df9d52a570f8

Observation 7b403170-dbfe-452a-bdbd-19522ed26d68 · outbound

This paper cites an unresolved cited work.

Infinite Video Understanding Unresolved cited work

Reference 19

Resolution
verified exact
doi, observed 2026-08-06T18:10:16.232019Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T18:10:12.643858Z digest=sha256:bcb1b6606631a8a6442d2dc12e7f5a90c01631ed3fd7845c21d0008aaf009834

Observation 43e3fa0f-361d-4cba-9e31-bc1fd9e80412 · outbound

This paper cites START: Self-taught Reasoner with Tools.

Infinite Video Understanding START: Self-taught Reasoner with Tools

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T18:10:12.598324Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:10:12.598324Z digest=sha256:541355c908dcf496f32aa5e5329ea51480cf62cc741eec993cfe16d1b4f88267

Observation 97ac225c-a44b-48e7-b367-d1e266ecfdf5 · outbound

This paper cites VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling.

Infinite Video Understanding VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T18:10:12.714751Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:10:12.714751Z digest=sha256:c0a50a1f66d72071326a2f0e1da0bf28ae3b15daad0b9242ce659cf886a35b67

Observation cc6d9d5e-c795-46dd-b38d-5aa40d4a9983 · outbound

This paper cites MVBench: A Comprehensive Multi-modal Video Understanding Benchmark.

Infinite Video Understanding MVBench: A Comprehensive Multi-modal Video Understanding Benchmark

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T18:10:12.662954Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:10:12.662954Z digest=sha256:485ff710427f806d4b2c68e2e5a9cfb4ebc779f4f7cff49ed8f4ad7c9625dafd

Observation 3633c010-534b-47ba-b99c-22c9cd1f6f03 · outbound

This paper cites Video-XL-Pro: Reconstructive Token Compression for Extremely Long Video Understanding.

Infinite Video Understanding Video-XL-Pro: Reconstructive Token Compression for Extremely Long Video Understanding

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T18:10:12.773039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:10:12.773039Z digest=sha256:094eb2499552c12447af9d20487922cf4101993b0618594ce65d890400058ea2

Observation 448a4907-dd0e-4419-9ed5-fba4759143ce · outbound

This paper cites World Model on Million-Length Video And Language With Blockwise RingAttention.

Infinite Video Understanding World Model on Million-Length Video And Language With Blockwise RingAttention

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T18:10:12.748486Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:10:12.748486Z digest=sha256:c0d1857c4332deae72f36120f4d50f4bf82b874f1645edd4b4651fcbeb545271

Observation e0001093-1c8d-428c-8368-db9820138888 · outbound

This paper cites EgoSchema: A Diagnostic Benchmark for Very Long-form Video Language Understanding.

Infinite Video Understanding EgoSchema: A Diagnostic Benchmark for Very Long-form Video Language Understanding

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T18:10:12.839920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:10:12.839920Z digest=sha256:ed1d40c9e94f9c9b2a097a6aa5bec955663af0d01076bddc69003a69afabbf46

Observation 6dcac48e-b7e6-4df2-9834-febd2b9658ff · outbound

This paper cites an unresolved cited work.

Infinite Video Understanding Unresolved cited work

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T18:10:12.786822Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:10:12.786822Z digest=sha256:0ff4ea624c328b572f5f9082aec77e20bb6671ebe3f1e4bb30f0c7dca7f767df

Observation 01d0693d-1a66-4237-85d4-664f63fa52b0 · outbound

This paper cites Cosmos World Foundation Model Platform for Physical AI.

Infinite Video Understanding Cosmos World Foundation Model Platform for Physical AI

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T18:10:12.942821Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:10:12.942821Z digest=sha256:1d636a01736cde7bfdc21193e01734978bb921fc640af924b91bcc93bf2684ef

Observation 20a4e7fb-271f-4a54-b77d-5c48378e0b2e · outbound

This paper cites an unresolved cited work.

Infinite Video Understanding Unresolved cited work

Reference 28

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:10:17.807895Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T18:10:12.877764Z digest=sha256:1ab564cc8d49f14e39bc798b2d71710c0818ebe2008e05cd81fdde0cf1fcdb22

Observation af469056-2597-4916-802c-3e9a4e6eb44a · outbound

This paper cites Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification.

Infinite Video Understanding Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T18:10:13.072860Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:10:13.072860Z digest=sha256:55771a48e27f8970c720c8f21a311f85daf8ba52e55143bb62e7cfac10f85e5d

Observation e76d6843-4626-47ad-ba97-c6f7c48102fe · outbound

This paper cites Streaming Long Video Understanding with Large Language Models.

Infinite Video Understanding Streaming Long Video Understanding with Large Language Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T18:10:13.022336Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:10:13.022336Z digest=sha256:efc16d22a10532090974777b6dcb4352303f9c2663768f3257c34d6b2e67dc5a

Observation e32d3cd9-dbef-4bfe-8da5-bd1a7f618ab4 · outbound

This paper cites Learning Transferable Visual Models From Natural Language Supervision.

Infinite Video Understanding Learning Transferable Visual Models From Natural Language Supervision

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T18:10:13.140709Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:10:13.140709Z digest=sha256:bef4ea25ee4643412b7532c5fd0722e27e55373dadcb353489bd5cbee649c4d2

Observation 2aaa3216-7106-47ce-b2e1-40f190617bbc · outbound

This paper cites Recent Advances of Continual Learning in Computer Vision: An Overview.

Infinite Video Understanding Recent Advances of Continual Learning in Computer Vision: An Overview

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T18:10:13.104601Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:10:13.104601Z digest=sha256:b99d877448894d9616a97569aae9bbf383e42da67ad951330b03eaa41d72f3a2

Observation ea010615-2ce1-47d8-9be6-4ea8823abe73 · outbound

This paper cites LongVU: Spatiotemporal Adaptive Compression for Long Video-Language Understanding.

Infinite Video Understanding LongVU: Spatiotemporal Adaptive Compression for Long Video-Language Understanding

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T18:10:13.260398Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:10:13.260398Z digest=sha256:fa35a8b43e1dc2841d923e3208615c03d92d4f2ebcf5db00c48290212a97e8ba

Observation edc308c6-cf1f-431d-b0d1-b121878aafd8 · outbound

This paper cites $\infty$-Video: A Training-Free Approach to Long Video Understanding via Continuous-Time Memory Consolidation.

Infinite Video Understanding $\infty$-Video: A Training-Free Approach to Long Video Understanding via Continuous-Time Memory Consolidation

Reference 34

Resolution
metadata mismatch
local_arxiv, observed 2026-08-06T18:10:15.807015Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T18:10:13.200174Z digest=sha256:b289a3d3f133d0f6c4249c66870a369d712a230e7d5b052ae0af94677dce8ae6

Observation e76d195d-b50d-4df2-969f-288c3370d330 · outbound

This paper cites an unresolved cited work.

Infinite Video Understanding Unresolved cited work

Reference 35

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:10:17.698230Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T18:10:13.328713Z digest=sha256:e9e81d5f7b5774c0095a8188fd4c767e47647396faab9d3c39a46595bedeec43

Observation c2df738d-2eb6-4bef-b7fd-a788422984d2 · outbound

This paper cites Video-XL: Extra-Long Vision Language Model for Hour-Scale Video Understanding.

Infinite Video Understanding Video-XL: Extra-Long Vision Language Model for Hour-Scale Video Understanding

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T18:10:13.296306Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:10:13.296306Z digest=sha256:b909a5a5a8caabaa9f5ae1a5e6d8eaebc48f12f5699b10b1e5781e30eed65792

Observation 6daaacba-cd5e-4c86-b121-09b740e1fcb6 · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

Infinite Video Understanding Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T18:10:13.446720Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:10:13.446720Z digest=sha256:6355ace9ffe6c55da2fc5719659fcfbd9855bb5cff8ad9ab83121d299b52b4f7

Observation 102ccd94-8ead-4ada-9021-3f44632e9396 · outbound

This paper cites MovieChat: From Dense Token to Sparse Memory for Long Video Understanding.

Infinite Video Understanding MovieChat: From Dense Token to Sparse Memory for Long Video Understanding

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T18:10:13.357356Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:10:13.357356Z digest=sha256:b49e996a0f4c2b54586bbb227f80d7f88b9b58d596ac024411e36344c36019b3

Observation 1d59b233-d2bb-4843-81a8-9b4565f0816f · outbound

This paper cites SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features.

Infinite Video Understanding SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T18:10:13.651494Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:10:13.651494Z digest=sha256:8987e4bf6ce2361f9b15e8e4e792e3a04745bee28e6939bb5dc1da6feabb510d

Observation 8627080a-8146-4fd5-893f-d0b40c9a5e1c · outbound

This paper cites VideoMAE: Masked Autoencoders are Data-Efficient Learners for Self-Supervised Video Pre-Training.

Infinite Video Understanding VideoMAE: Masked Autoencoders are Data-Efficient Learners for Self-Supervised Video Pre-Training

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T18:10:13.574181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:10:13.574181Z digest=sha256:cb235e95a501e18218c20eaa861dce5d034372c1ce019374ee1faa720106e9df

Observation 5c8a1017-73f4-4f2c-9e4d-a780ab36782d · outbound

This paper cites A Comprehensive Survey of Continual Learning: Theory, Method and Application.

Infinite Video Understanding A Comprehensive Survey of Continual Learning: Theory, Method and Application

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T18:10:13.775787Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:10:13.775787Z digest=sha256:9a8ff24328a2886cc627ec5e7282122ab6cdb26a3560f46f0bbf470ae62f4840

Observation 06bfa572-b497-4805-9795-6453d48c3512 · outbound

This paper cites VideoMAE V2: Scaling Video Masked Autoencoders with Dual Masking.

Infinite Video Understanding VideoMAE V2: Scaling Video Masked Autoencoders with Dual Masking

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T18:10:13.711528Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:10:13.711528Z digest=sha256:d295c11383d33566016f93f45c69ef80d818040d3d4a4d15a3031089ce537f26

Observation ef828168-849f-4259-8d45-f7f2490571e5 · outbound

This paper cites VidTwin: Video VAE with Decoupled Structure and Dynamics.

Infinite Video Understanding VidTwin: Video VAE with Decoupled Structure and Dynamics

Reference 43

Resolution
verified exact
local_arxiv, observed 2026-08-06T18:10:15.572551Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T18:10:14.101777Z digest=sha256:53342f78024ac7785973ec7d8e480e62eb70afc4055177595f74cd97821dc3aa

Observation a69f610b-c659-49d6-9892-4f1bf7f702e1 · outbound

This paper cites an unresolved cited work.

Infinite Video Understanding Unresolved cited work

Reference 44

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:10:17.529367Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T18:10:13.937532Z digest=sha256:c7393d1435a7d511d63684ddddce67c9a97b4cfa81e803a5cf60e0859b669387

Observation f519b7f8-72c3-4347-acda-48ca2a008599 · outbound

This paper cites VideoRoPE: What Makes for Good Video Rotary Position Embedding?.

Infinite Video Understanding VideoRoPE: What Makes for Good Video Rotary Position Embedding?

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T18:10:14.257348Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:10:14.257348Z digest=sha256:5b61a89c79901e18c9b07bd0cbe4d12346999d596a3f0938d112dc1cefca7fa2

Observation a30160aa-9053-40af-9c42-bef00968d3a7 · outbound

This paper cites LongVLM: Efficient Long Video Understanding via Large Language Models.

Infinite Video Understanding LongVLM: Efficient Long Video Understanding via Large Language Models

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T18:10:14.298291Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:10:14.298291Z digest=sha256:04323c264f016071da47e2cfeb9099964f245e1b20524001b360907cf3e9700a

Observation 5d65b882-b812-4eb2-9817-d7922d9a8c0b · outbound

This paper cites VideoTree: Adaptive Tree-based Video Representation for LLM Reasoning on Long Videos.

Infinite Video Understanding VideoTree: Adaptive Tree-based Video Representation for LLM Reasoning on Long Videos

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-06T18:10:14.223529Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:10:14.223529Z digest=sha256:1abc98766ee8ba0892632fd1dbafcd95246136ecf4d97224849f2a3eb97bcf77

Observation 88a9e838-e338-44d7-b0b5-2cabd1433fbc · outbound

This paper cites Agentic Reasoning: A Streamlined Framework for Enhancing LLM Reasoning with Agentic Tools.

Infinite Video Understanding Agentic Reasoning: A Streamlined Framework for Enhancing LLM Reasoning with Agentic Tools

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-06T18:10:14.365798Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:10:14.365798Z digest=sha256:874191ef6b381f1a1dd499328dad5ee61b9e347ad847a9fc870e6b700880f6fa

Observation aea3764e-83dd-40b8-9999-1eb2fca7070d · outbound

This paper cites LongViTU: Instruction Tuning for Long-Form Video Understanding.

Infinite Video Understanding LongViTU: Instruction Tuning for Long-Form Video Understanding

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-06T18:10:14.416661Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:10:14.416661Z digest=sha256:946d0943aca3feda46988aa25971d0a325658cf43d4f5644f14f21148fd0abe5

Observation b6143b35-bfaa-484d-9084-4430ec886d4b · outbound

This paper cites LongVideoBench: A Benchmark for Long-context Interleaved Video-Language Understanding.

Infinite Video Understanding LongVideoBench: A Benchmark for Long-context Interleaved Video-Language Understanding

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-06T18:10:14.340618Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:10:14.340618Z digest=sha256:a254a1e2f1821a70f68c2ff89c6be1e32d77061adf322f939ddf3d1de4286e1f

Observation 9dda5626-d89f-474a-92d8-7620a5c62f59 · outbound

This paper cites MAGVIT: Masked Generative Video Transformer.

Infinite Video Understanding MAGVIT: Masked Generative Video Transformer

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-06T18:10:14.507734Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:10:14.507734Z digest=sha256:8d2a8418006c09e64c823b0bb3917a21e5662caf3e34e164a247ef53fbc2ed47

Observation a4883221-886c-418b-9b68-445e2b4784e5 · outbound

This paper cites Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation.

Infinite Video Understanding Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-06T18:10:14.568887Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:10:14.568887Z digest=sha256:18abeda349aea1eb40105b7b98826ac8ae744a1a73dc67c01b5ac84c4872596c

Observation 5ab42a16-8fdb-4d71-84eb-02b6ab549f18 · outbound

This paper cites an unresolved cited work.

Infinite Video Understanding Unresolved cited work

Reference 53

Resolution
verified exact
doi, observed 2026-08-06T18:10:15.381307Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T18:10:14.459822Z digest=sha256:81aa0d7369637400e7c1e111b9c6f74e7c3f0eb6d52170c1da758591c732181a

Observation ef731b3e-8c43-42d9-aa5e-109c6274611b · outbound

This paper cites an unresolved cited work.

Infinite Video Understanding Unresolved cited work

Reference 54

Resolution
verified exact
doi, observed 2026-08-06T18:10:15.203453Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T18:10:14.678140Z digest=sha256:64d98af431bcfea72e02a62294f9c893c244e208c5d75428ab957c42edde8b7e

Observation 0c444863-559e-43e1-8eb9-3f9de9561d65 · outbound

This paper cites Gundavarapu, Liangzhe Yuan, Hao Zhou, Shen Yan, Jennifer J.

Infinite Video Understanding Gundavarapu, Liangzhe Yuan, Hao Zhou, Shen Yan, Jennifer J

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:10:17.373787Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T18:10:14.711077Z digest=sha256:a587cbd9728c09142eaf813340dbf0e22ffdcbd5c1a3518715e503d8bfeffe82

Observation 631dce2d-c1ed-46e3-a663-7fb3e379b154 · outbound

This paper cites Sigmoid Loss for Language Image Pre-Training.

Infinite Video Understanding Sigmoid Loss for Language Image Pre-Training

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-06T18:10:14.635404Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:10:14.635404Z digest=sha256:06bf0ae22f42b1168a1b5835f5562d7cb45e395fa8b248357fa23e44793f108b

Observation bc707462-8bba-42d4-b63e-912ab0c3dfc9 · outbound

This paper cites Towards Lifelong Learning of Large Language Models: A Survey.

Infinite Video Understanding Towards Lifelong Learning of Large Language Models: A Survey

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-06T18:10:14.841489Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:10:14.841489Z digest=sha256:556ce7b8e069a0895e2728158061f5aa298566e8c7f11657fa5d7805936de01c

Observation 12c78463-a446-463e-8f21-9ae51e99a337 · outbound

This paper cites MLVU: Benchmarking Multi-task Long Video Understanding.

Infinite Video Understanding MLVU: Benchmarking Multi-task Long Video Understanding

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-06T18:10:14.864857Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:10:14.864857Z digest=sha256:dd0e9822421099e0632e151d0dfe5bb595d12fddab12a60011aa0a23b79d6569

Observation 398801a5-873c-4169-abf0-73e544e5bb62 · outbound

This paper cites VideoPrism: A Foundational Visual Encoder for Video Understanding.

Infinite Video Understanding VideoPrism: A Foundational Visual Encoder for Video Understanding

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-06T18:10:14.749813Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:10:14.749813Z digest=sha256:36dc5cffbb6544aca4ef2221b9a6193ed70684cd7bb7488c194e57314279fb37

Observation 5da26043-a850-4fdd-8944-8fc06fdf7b6a · outbound

This paper cites Needle In A Video Haystack: A Scalable Synthetic Evaluator for Video MLLMs.

Infinite Video Understanding Needle In A Video Haystack: A Scalable Synthetic Evaluator for Video MLLMs

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-06T18:10:14.777464Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:10:14.777464Z digest=sha256:c574d046ad98ba1873b1eb7785ffb13626f9b4c7d944cdee4b66d500454b53f2

Observation 168f0a8e-2993-48df-b4f0-266fb42fca75 · outbound

This paper cites an unresolved cited work.

Infinite Video Understanding Unresolved cited work

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-06T18:10:14.887539Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:10:14.887539Z digest=sha256:e70fef2614a87de62bc63b1eee563af4b9dffd87be518af996e1fbc4d32a6266

Observation ff6bf875-6dfb-403e-aa65-9bd3504869d2 · outbound

This paper cites From Seconds to Hours: Reviewing MultiModal Large Language Models on Comprehensive Long Video Understanding.

Infinite Video Understanding From Seconds to Hours: Reviewing MultiModal Large Language Models on Comprehensive Long Video Understanding

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-06T18:10:14.919272Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:10:14.919272Z digest=sha256:697623c35d3f1393241e48ca2d9d2bc965f3575e01871859923a35fd6aa4480d

Observation 004a05d9-6451-40e2-afb9-1e60ec5aa9fb · outbound

This paper cites LVBench: An Extreme Long Video Understanding Benchmark.

Infinite Video Understanding LVBench: An Extreme Long Video Understanding Benchmark

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-06T18:10:13.986726Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:10:13.986726Z digest=sha256:8a42e43c9b1e4e04255680830040824a21250e0077ac69815b053122627b37df

Observation 5dfc5ac4-c02b-4cfa-9645-4613c33f3a71 · outbound

This paper cites Video-R1: Reinforcing Video Reasoning in MLLMs.

Infinite Video Understanding Video-R1: Reinforcing Video Reasoning in MLLMs

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-06T18:10:11.970598Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:10:11.970598Z digest=sha256:93330f532e5f825e0a2b6bb4e0f3d7f96240c62810f6fe9073a4a3471747702e

Pith citing papers

No inbound Pith citation observations are available.