Pith. sign in

Paper Citation Record · LEDGER

Infinite Video Understanding

As of 15 August 2026, this Paper Citation Record lists 64 of 64 outbound references and 0 inbound Pith citation observations for arXiv:2507.09068.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.09068 v2

Coverage vector

measured 64 of 64 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T18:10:14.919272Z

measured 64 of 64 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

64 of 64 outbound references displayed

  • verified exact8
  • verified fuzzy1
  • unresolved54
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation ff0a1ede-3856-4a74-ad1e-e3544cd0a8b4 · outbound

This paper cites CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders.

Infinite Video Understanding CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-08-06T18:10:16.731929Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T18:09:18.061419Z digest=sha256:19b017d219f3eb047d328306cc176cc19d6b61d37588a78e9698335f2eeee3b7

Observation c25e915d-4ba9-493f-847b-51365e8f7fda · outbound

This paper cites an unresolved cited work.

Infinite Video Understanding Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:10:18.299149Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T18:09:18.114356Z digest=sha256:49953bc4118af2b0372cd8ba828e7970abeb64594ee8840c086c80488fafa0ad

Observation 01bcc5ab-e0c7-4408-a60e-48774759a8fe · outbound

This paper cites HourVideo: 1-Hour Video-Language Understanding.

Infinite Video Understanding HourVideo: 1-Hour Video-Language Understanding

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T18:09:18.211447Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:09:18.211447Z digest=sha256:33fdfeb5e53d612c4f870baa681baf26e3a4c2cddc1e1111606b3e62d2abb317

Observation 339e7bb3-dc67-4a6f-8af4-aff75a9dfa3c · outbound

This paper cites LongVILA: Scaling Long-Context Visual Language Models for Long Videos.

Infinite Video Understanding LongVILA: Scaling Long-Context Visual Language Models for Long Videos

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T18:09:18.313230Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:09:18.313230Z digest=sha256:368a8a3966e1cffafd9d8b775d5bb4334e8d44cc1118f0900fcf4605b76286ec

Observation 809b2dcf-8bb2-479b-964a-3488ccdb6144 · outbound

This paper cites Scaling Video-Language Models to 10K Frames via Hierarchical Differential Distillation.

Infinite Video Understanding Scaling Video-Language Models to 10K Frames via Hierarchical Differential Distillation

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T18:09:18.568750Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:09:18.568750Z digest=sha256:a0e8dfd4d5f81224f42b69c41f7c75618e8f19d68c18d305267a4c1e755adecd

Observation 8d053b18-1eb2-430a-9e88-2a6c782636e8 · outbound

This paper cites Video-Holmes: Can MLLM Think Like Holmes for Complex Video Reasoning?.

Infinite Video Understanding Video-Holmes: Can MLLM Think Like Holmes for Complex Video Reasoning?

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T18:09:18.954541Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:09:18.954541Z digest=sha256:0e57651339d9e839cc3630ad9aea2ddd22926ed79a085ea9e56ab49f3bc987fd

Observation dbef02ca-88ff-488e-a914-02db772cb241 · outbound

This paper cites Zero-Shot Video Question Answering with Procedural Programs.

Infinite Video Understanding Zero-Shot Video Question Answering with Procedural Programs

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T18:09:19.409539Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:09:19.409539Z digest=sha256:acfbec5892cfd76e392f901fd2174a1a42332cf9ded58824cfe4905de0763936

Observation 68494767-b952-4b9b-bba1-bbddf06ed232 · outbound

This paper cites Project Aria: A New Tool for Egocentric Multi-Modal AI Research.

Infinite Video Understanding Project Aria: A New Tool for Egocentric Multi-Modal AI Research

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T18:10:11.725676Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:10:11.725676Z digest=sha256:ef6fe6a9c4e41871e9dc8acbf5a289dea5ed01d394e059a130e80ad4c6e989d6

Observation 545305b2-f2bb-4383-9be8-639125a12ebf · outbound

This paper cites Efficient Image Pre-Training with Siamese Cropped Masked Autoencoders.

Infinite Video Understanding Efficient Image Pre-Training with Siamese Cropped Masked Autoencoders

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-08-06T18:10:17.212457Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T18:10:11.854309Z digest=sha256:30c4fad59520e1d33a0308149f476dccf573dcea2de4e74c4d64c80a90e8ee95

Observation 316ad4f0-026d-4451-84aa-0b1e8874df9e · outbound

This paper cites an unresolved cited work.

Infinite Video Understanding Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:10:18.130805Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T18:10:11.899351Z digest=sha256:c3e24bf8389622d394d7d626024dd3fe0a5dcaa840f84dbcc79d9fa4e855fd70

Observation 276abd28-1608-4fa1-8d37-0830654d553a · outbound

This paper cites Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis.

Infinite Video Understanding Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T18:10:12.056934Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:10:12.056934Z digest=sha256:c24ff0d8272e2d7edc3205a52adf49ff6b93325e693018f8e7def0601eeb960f

Observation 4cbccd92-ffbe-4212-98e0-67fa7f83e1a5 · outbound

This paper cites The Ethics of Advanced AI Assistants.

Infinite Video Understanding The Ethics of Advanced AI Assistants

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T18:10:12.125589Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:10:12.125589Z digest=sha256:dda6615b15031ba262789fa0e92139cd68189248842d1b08f4a69b3141161995

Observation 963c17ec-82c8-4343-91e9-d85d06aa17ca · outbound

This paper cites Siamese Masked Autoencoders.

Infinite Video Understanding Siamese Masked Autoencoders

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T18:10:12.195376Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:10:12.195376Z digest=sha256:cb8ae9d97e16b5b6c2ca7b1171d7a3ef8d8648eeb246c94a8035aba4d1306a8c

Observation 12718065-f4c4-4ac8-9980-1d22368d0953 · outbound

This paper cites RULER: What's the Real Context Size of Your Long-Context Language Models?.

Infinite Video Understanding RULER: What's the Real Context Size of Your Long-Context Language Models?

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T18:10:12.307299Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:10:12.307299Z digest=sha256:4a7bba453affd15e058c85c684c7e56e08b794be466a56afe2be6842671bf4f9

Observation 9110064f-0fc6-4002-a20d-c5fe49e3a397 · outbound

This paper cites Visual Representation Learning with Stochastic Frame Prediction.

Infinite Video Understanding Visual Representation Learning with Stochastic Frame Prediction

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-08-06T18:10:16.421305Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T18:10:12.401921Z digest=sha256:9c10f8a7617d4ef1d6e6301cc0159d1e9cf0ac29a609d92dc7e2b4a9dd5d2556

Observation 2b5d77ea-50d5-4d97-8f44-c7bd15d9693d · outbound

This paper cites Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning.

Infinite Video Understanding Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T18:10:12.472087Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:10:12.472087Z digest=sha256:e2c401ee34fedb010f11205e231f95b68930d5cd2213840ba1c5b1f6b9e3a18b

Observation 0b728617-5100-420c-887d-8c52582f04f6 · outbound

This paper cites SNeRV: Spectra-preserving Neural Representation for Video.

Infinite Video Understanding SNeRV: Spectra-preserving Neural Representation for Video

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-08-06T18:10:16.964469Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T18:10:12.519667Z digest=sha256:8705b7eebb576937d86bc6d0aa17fa9f3f73461ca770b1778c50fb1954c2e044

Observation fb2b3eb7-0cd3-4be0-bd16-9ca2d015fd81 · outbound

This paper cites an unresolved cited work.

Infinite Video Understanding Unresolved cited work

Reference 18

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:10:17.987768Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T18:10:12.559183Z digest=sha256:3ed9c6b0321d8d33fd5aa1cde9fa8f8ca6e1107b3397f73f8df355ddce7490e1

Observation 7b403170-dbfe-452a-bdbd-19522ed26d68 · outbound

This paper cites an unresolved cited work.

Infinite Video Understanding Unresolved cited work

Reference 19

Resolution
verified exact
doi, observed 2026-08-06T18:10:16.232019Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T18:10:12.643858Z digest=sha256:b60e1c77ed492502eba2a3771ca58552f08ff16004ccda51cb4e266082ebcabb

Observation 43e3fa0f-361d-4cba-9e31-bc1fd9e80412 · outbound

This paper cites START: Self-taught Reasoner with Tools.

Infinite Video Understanding START: Self-taught Reasoner with Tools

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T18:10:12.598324Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:10:12.598324Z digest=sha256:a0b0736198bc4ca4bd9523d93e7e308c81e9dda83087e164ed4cb648413bed5d

Observation 97ac225c-a44b-48e7-b367-d1e266ecfdf5 · outbound

This paper cites VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling.

Infinite Video Understanding VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T18:10:12.714751Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:10:12.714751Z digest=sha256:ba11b8a26ecec3415a8e4bdf28362b028d6f6b350d3d1eb8d58c0cb2d9974099

Observation cc6d9d5e-c795-46dd-b38d-5aa40d4a9983 · outbound

This paper cites MVBench: A Comprehensive Multi-modal Video Understanding Benchmark.

Infinite Video Understanding MVBench: A Comprehensive Multi-modal Video Understanding Benchmark

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T18:10:12.662954Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:10:12.662954Z digest=sha256:9bf813f8daa38223eb6e7d8c5f218be225875dc43429b5c30005097746863b56

Observation 3633c010-534b-47ba-b99c-22c9cd1f6f03 · outbound

This paper cites Video-XL-Pro: Reconstructive Token Compression for Extremely Long Video Understanding.

Infinite Video Understanding Video-XL-Pro: Reconstructive Token Compression for Extremely Long Video Understanding

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T18:10:12.773039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:10:12.773039Z digest=sha256:de989c4c0892a134b176efcb896ea00104923c81bd16a83aad64ef50e699b524

Observation 448a4907-dd0e-4419-9ed5-fba4759143ce · outbound

This paper cites World Model on Million-Length Video And Language With Blockwise RingAttention.

Infinite Video Understanding World Model on Million-Length Video And Language With Blockwise RingAttention

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T18:10:12.748486Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:10:12.748486Z digest=sha256:ef23d6f05531dc5d468a774e1c43b0ddef3364564193830163d2a7cbc5ca3aac

Observation e0001093-1c8d-428c-8368-db9820138888 · outbound

This paper cites EgoSchema: A Diagnostic Benchmark for Very Long-form Video Language Understanding.

Infinite Video Understanding EgoSchema: A Diagnostic Benchmark for Very Long-form Video Language Understanding

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T18:10:12.839920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:10:12.839920Z digest=sha256:fd867b2c740d9530de160610e2a41217dded95a95280d87dcf9265880f85a208

Observation 6dcac48e-b7e6-4df2-9834-febd2b9658ff · outbound

This paper cites an unresolved cited work.

Infinite Video Understanding Unresolved cited work

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T18:10:12.786822Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:10:12.786822Z digest=sha256:fda05562c460160809b08a7edf3633fa204351364b8a76855aca3752b0698200

Observation 01d0693d-1a66-4237-85d4-664f63fa52b0 · outbound

This paper cites Cosmos World Foundation Model Platform for Physical AI.

Infinite Video Understanding Cosmos World Foundation Model Platform for Physical AI

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T18:10:12.942821Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:10:12.942821Z digest=sha256:150a26855342b55478c12f68a154f9f7ba378ff45bf3209f37f084ae854dbe27

Observation 20a4e7fb-271f-4a54-b77d-5c48378e0b2e · outbound

This paper cites an unresolved cited work.

Infinite Video Understanding Unresolved cited work

Reference 28

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:10:17.807895Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T18:10:12.877764Z digest=sha256:c920f34eef31ee96c9e63b0e0f83fc29ae016c3f3253c9bd1d432dacf89037b4

Observation af469056-2597-4916-802c-3e9a4e6eb44a · outbound

This paper cites Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification.

Infinite Video Understanding Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T18:10:13.072860Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:10:13.072860Z digest=sha256:df2148e858db39e0f85eb16863e58a322bba72f090742d8c93aef293a665bccf

Observation e76d6843-4626-47ad-ba97-c6f7c48102fe · outbound

This paper cites Streaming Long Video Understanding with Large Language Models.

Infinite Video Understanding Streaming Long Video Understanding with Large Language Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T18:10:13.022336Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:10:13.022336Z digest=sha256:40dc8470194189bd612e8694403d1f7a2ac62e0ce2b11a6699fa43a437dc027d

Observation e32d3cd9-dbef-4bfe-8da5-bd1a7f618ab4 · outbound

This paper cites Learning Transferable Visual Models From Natural Language Supervision.

Infinite Video Understanding Learning Transferable Visual Models From Natural Language Supervision

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T18:10:13.140709Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:10:13.140709Z digest=sha256:f439c06be1490b5bc39b02dae3dcdd1b934dd89ea4cb74d0d380b7f1aca515fc

Observation 2aaa3216-7106-47ce-b2e1-40f190617bbc · outbound

This paper cites Recent Advances of Continual Learning in Computer Vision: An Overview.

Infinite Video Understanding Recent Advances of Continual Learning in Computer Vision: An Overview

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T18:10:13.104601Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:10:13.104601Z digest=sha256:a4b4f53e91fca7e4d3785d9228e5057f10f5da29562cd5580249915c31aed6dc

Observation ea010615-2ce1-47d8-9be6-4ea8823abe73 · outbound

This paper cites LongVU: Spatiotemporal Adaptive Compression for Long Video-Language Understanding.

Infinite Video Understanding LongVU: Spatiotemporal Adaptive Compression for Long Video-Language Understanding

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T18:10:13.260398Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:10:13.260398Z digest=sha256:fa7123b181545fadba46c54acf63d307ebde57436bc9a0a9ae9d004a9aa67b34

Observation edc308c6-cf1f-431d-b0d1-b121878aafd8 · outbound

This paper cites $\infty$-Video: A Training-Free Approach to Long Video Understanding via Continuous-Time Memory Consolidation.

Infinite Video Understanding $\infty$-Video: A Training-Free Approach to Long Video Understanding via Continuous-Time Memory Consolidation

Reference 34

Resolution
metadata mismatch
local_arxiv, observed 2026-08-06T18:10:15.807015Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T18:10:13.200174Z digest=sha256:ce8d29c341975ba46a591938f875d3efd15ecd4d03209e132c5956ce4722a8f0

Observation e76d195d-b50d-4df2-969f-288c3370d330 · outbound

This paper cites an unresolved cited work.

Infinite Video Understanding Unresolved cited work

Reference 35

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:10:17.698230Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T18:10:13.328713Z digest=sha256:275096977e64591ece2110b175a2a1819dec39092d80c01cf5ed67b985e7b096

Observation c2df738d-2eb6-4bef-b7fd-a788422984d2 · outbound

This paper cites Video-XL: Extra-Long Vision Language Model for Hour-Scale Video Understanding.

Infinite Video Understanding Video-XL: Extra-Long Vision Language Model for Hour-Scale Video Understanding

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T18:10:13.296306Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:10:13.296306Z digest=sha256:1105873be821c28d7d5cdfd59aee5adc61bff564eb92e62c8ef4cf07c2389a8b

Observation 6daaacba-cd5e-4c86-b121-09b740e1fcb6 · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

Infinite Video Understanding Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T18:10:13.446720Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:10:13.446720Z digest=sha256:c4e16bd9e84f459c6edb00c0338bf1750fed64339d73237f6ea29d08d202c7d0

Observation 102ccd94-8ead-4ada-9021-3f44632e9396 · outbound

This paper cites MovieChat: From Dense Token to Sparse Memory for Long Video Understanding.

Infinite Video Understanding MovieChat: From Dense Token to Sparse Memory for Long Video Understanding

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T18:10:13.357356Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:10:13.357356Z digest=sha256:3f2731f2b3517260a6819b976b8684d26986acf7b444b8059b56b1ddd968928a

Observation 1d59b233-d2bb-4843-81a8-9b4565f0816f · outbound

This paper cites SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features.

Infinite Video Understanding SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T18:10:13.651494Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:10:13.651494Z digest=sha256:2c0f27a69ceb39897c6878c46d3da844b24d3384ff993154029cfa7b059bb0a1

Observation 8627080a-8146-4fd5-893f-d0b40c9a5e1c · outbound

This paper cites VideoMAE: Masked Autoencoders are Data-Efficient Learners for Self-Supervised Video Pre-Training.

Infinite Video Understanding VideoMAE: Masked Autoencoders are Data-Efficient Learners for Self-Supervised Video Pre-Training

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T18:10:13.574181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:10:13.574181Z digest=sha256:90ffa84aba9c94afdb58bffbc0026fa23a59c145c2c55b7f4e87f0e50632c506

Observation 5c8a1017-73f4-4f2c-9e4d-a780ab36782d · outbound

This paper cites A Comprehensive Survey of Continual Learning: Theory, Method and Application.

Infinite Video Understanding A Comprehensive Survey of Continual Learning: Theory, Method and Application

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T18:10:13.775787Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:10:13.775787Z digest=sha256:7940fec246cfbc25cb98c21be887320126356a7dfa53e10eb7694dc9cb3dcc25

Observation 06bfa572-b497-4805-9795-6453d48c3512 · outbound

This paper cites VideoMAE V2: Scaling Video Masked Autoencoders with Dual Masking.

Infinite Video Understanding VideoMAE V2: Scaling Video Masked Autoencoders with Dual Masking

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T18:10:13.711528Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:10:13.711528Z digest=sha256:1abe8af9e3c25e522d8a3742579ad765d455b5f54692dd911b7ac3733f590b9e

Observation ef828168-849f-4259-8d45-f7f2490571e5 · outbound

This paper cites VidTwin: Video VAE with Decoupled Structure and Dynamics.

Infinite Video Understanding VidTwin: Video VAE with Decoupled Structure and Dynamics

Reference 43

Resolution
verified exact
local_arxiv, observed 2026-08-06T18:10:15.572551Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T18:10:14.101777Z digest=sha256:be117d64881d953f7e9fba9210634d6bcce4926a1f678b9e3fdd10fb8894f379

Observation a69f610b-c659-49d6-9892-4f1bf7f702e1 · outbound

This paper cites an unresolved cited work.

Infinite Video Understanding Unresolved cited work

Reference 44

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:10:17.529367Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T18:10:13.937532Z digest=sha256:a72a0708e8d31f2e33b4b6d7e9d033590a001ea24c976bd075f4e6935975b7cb

Observation f519b7f8-72c3-4347-acda-48ca2a008599 · outbound

This paper cites VideoRoPE: What Makes for Good Video Rotary Position Embedding?.

Infinite Video Understanding VideoRoPE: What Makes for Good Video Rotary Position Embedding?

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T18:10:14.257348Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:10:14.257348Z digest=sha256:250185b2bfef246c029a70a290088c5e15b2fb085aff1c03a26d3746e2ed90bf

Observation a30160aa-9053-40af-9c42-bef00968d3a7 · outbound

This paper cites LongVLM: Efficient Long Video Understanding via Large Language Models.

Infinite Video Understanding LongVLM: Efficient Long Video Understanding via Large Language Models

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T18:10:14.298291Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:10:14.298291Z digest=sha256:0bd1694de7c5c3d0e0dc7c5b23e3c17a996f64e245c4a3b11f91edc64b5941a1

Observation 5d65b882-b812-4eb2-9817-d7922d9a8c0b · outbound

This paper cites VideoTree: Adaptive Tree-based Video Representation for LLM Reasoning on Long Videos.

Infinite Video Understanding VideoTree: Adaptive Tree-based Video Representation for LLM Reasoning on Long Videos

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-06T18:10:14.223529Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:10:14.223529Z digest=sha256:5017f18d4f3f627cb0488d3e1080355f78dca47336be04435217d41d4c74d9ac

Observation 88a9e838-e338-44d7-b0b5-2cabd1433fbc · outbound

This paper cites Agentic Reasoning: A Streamlined Framework for Enhancing LLM Reasoning with Agentic Tools.

Infinite Video Understanding Agentic Reasoning: A Streamlined Framework for Enhancing LLM Reasoning with Agentic Tools

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-06T18:10:14.365798Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:10:14.365798Z digest=sha256:641838b02451c51a3523e7e05af0c2a2dd96114a96085caeae67620b3417fcb4

Observation aea3764e-83dd-40b8-9999-1eb2fca7070d · outbound

This paper cites LongViTU: Instruction Tuning for Long-Form Video Understanding.

Infinite Video Understanding LongViTU: Instruction Tuning for Long-Form Video Understanding

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-06T18:10:14.416661Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:10:14.416661Z digest=sha256:0d3676f32ff659f3da2cc075ebe9cbdcfe55ba8f92c0cd4aaf1fe4f1b0f967bc

Observation b6143b35-bfaa-484d-9084-4430ec886d4b · outbound

This paper cites LongVideoBench: A Benchmark for Long-context Interleaved Video-Language Understanding.

Infinite Video Understanding LongVideoBench: A Benchmark for Long-context Interleaved Video-Language Understanding

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-06T18:10:14.340618Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:10:14.340618Z digest=sha256:07798b82d03ae3b0ed465efc36f572c62a76b9bac6e28fc12f506ae2481ffbc4

Observation 9dda5626-d89f-474a-92d8-7620a5c62f59 · outbound

This paper cites MAGVIT: Masked Generative Video Transformer.

Infinite Video Understanding MAGVIT: Masked Generative Video Transformer

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-06T18:10:14.507734Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:10:14.507734Z digest=sha256:8e1fae4fc8d643212a8f4e5a763e466491e42546ae80a34bcbc472409808ff5f

Observation a4883221-886c-418b-9b68-445e2b4784e5 · outbound

This paper cites Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation.

Infinite Video Understanding Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-06T18:10:14.568887Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:10:14.568887Z digest=sha256:2fa46d314e3d205d15ecd99fa3376c58b5b6d2f302f727d326cf57c774c7eaa5

Observation 5ab42a16-8fdb-4d71-84eb-02b6ab549f18 · outbound

This paper cites an unresolved cited work.

Infinite Video Understanding Unresolved cited work

Reference 53

Resolution
verified exact
doi, observed 2026-08-06T18:10:15.381307Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T18:10:14.459822Z digest=sha256:4be9d7d642d3a13ef56eb4b511b21385e62bacd85416a03c3b607534468b9217

Observation ef731b3e-8c43-42d9-aa5e-109c6274611b · outbound

This paper cites an unresolved cited work.

Infinite Video Understanding Unresolved cited work

Reference 54

Resolution
verified exact
doi, observed 2026-08-06T18:10:15.203453Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T18:10:14.678140Z digest=sha256:973b8d9e960c9957002e9b9f9f76c4447b83cf3a8e16e7c8d1c200e6f4df9f97

Observation 0c444863-559e-43e1-8eb9-3f9de9561d65 · outbound

This paper cites Gundavarapu, Liangzhe Yuan, Hao Zhou, Shen Yan, Jennifer J.

Infinite Video Understanding Gundavarapu, Liangzhe Yuan, Hao Zhou, Shen Yan, Jennifer J

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:10:17.373787Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T18:10:14.711077Z digest=sha256:d1a70828ca493bd97a71a107a1a7793139fd1ce208765a706601dd93c76dc20f

Observation 631dce2d-c1ed-46e3-a663-7fb3e379b154 · outbound

This paper cites Sigmoid Loss for Language Image Pre-Training.

Infinite Video Understanding Sigmoid Loss for Language Image Pre-Training

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-06T18:10:14.635404Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:10:14.635404Z digest=sha256:a4e15577c21d29c0327b3277a6e501f87fbf00d4f26fc66034f522c2a9772adc

Observation bc707462-8bba-42d4-b63e-912ab0c3dfc9 · outbound

This paper cites Towards Lifelong Learning of Large Language Models: A Survey.

Infinite Video Understanding Towards Lifelong Learning of Large Language Models: A Survey

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-06T18:10:14.841489Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:10:14.841489Z digest=sha256:f89fa7a7c7d90823a6bace3b6d782848ef5ac3bb7d86c1afa7a332497125ad6d

Observation 12c78463-a446-463e-8f21-9ae51e99a337 · outbound

This paper cites MLVU: Benchmarking Multi-task Long Video Understanding.

Infinite Video Understanding MLVU: Benchmarking Multi-task Long Video Understanding

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-06T18:10:14.864857Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:10:14.864857Z digest=sha256:65bae9834f5f95909249ff8373ba4b21f0a2179c78b0cf4fdf33e5fa2b998f63

Observation 398801a5-873c-4169-abf0-73e544e5bb62 · outbound

This paper cites VideoPrism: A Foundational Visual Encoder for Video Understanding.

Infinite Video Understanding VideoPrism: A Foundational Visual Encoder for Video Understanding

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-06T18:10:14.749813Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:10:14.749813Z digest=sha256:f3be50832e669ac9ad0036024c0b61f2a7d994dccca1e3fe3e75f6a0237c28da

Observation 5da26043-a850-4fdd-8944-8fc06fdf7b6a · outbound

This paper cites Needle In A Video Haystack: A Scalable Synthetic Evaluator for Video MLLMs.

Infinite Video Understanding Needle In A Video Haystack: A Scalable Synthetic Evaluator for Video MLLMs

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-06T18:10:14.777464Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:10:14.777464Z digest=sha256:ee662732af52fb402c6ba9d7cdd4cf03709a3a7e2c172fd45ef1cc6c6036021e

Observation 168f0a8e-2993-48df-b4f0-266fb42fca75 · outbound

This paper cites an unresolved cited work.

Infinite Video Understanding Unresolved cited work

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-06T18:10:14.887539Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:10:14.887539Z digest=sha256:3bec4fbf6e2894bf1c82d3e990e34e6d869c2044e6901a46c6804bf29af72742

Observation ff6bf875-6dfb-403e-aa65-9bd3504869d2 · outbound

This paper cites From Seconds to Hours: Reviewing MultiModal Large Language Models on Comprehensive Long Video Understanding.

Infinite Video Understanding From Seconds to Hours: Reviewing MultiModal Large Language Models on Comprehensive Long Video Understanding

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-06T18:10:14.919272Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:10:14.919272Z digest=sha256:22af34b9920ee31d6cbff1a72cc08f54c06a50dfefb98d15857f885afd4abf49

Observation 004a05d9-6451-40e2-afb9-1e60ec5aa9fb · outbound

This paper cites LVBench: An Extreme Long Video Understanding Benchmark.

Infinite Video Understanding LVBench: An Extreme Long Video Understanding Benchmark

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-06T18:10:13.986726Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:10:13.986726Z digest=sha256:3575419417e17c2bcc84fd4cce4898eafba10a14bc0b47df75acab65dc73bb70

Observation 5dfc5ac4-c02b-4cfa-9645-4613c33f3a71 · outbound

This paper cites Video-R1: Reinforcing Video Reasoning in MLLMs.

Infinite Video Understanding Video-R1: Reinforcing Video Reasoning in MLLMs

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-06T18:10:11.970598Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:10:11.970598Z digest=sha256:c5129c78aa043ae8f77181f0d075f9a4aba25735a26bd8b1364de205f2be80cd

Pith citing papers

No inbound Pith citation observations are available.