Pith. sign in

Paper Citation Record · LEDGER

EgoSchema: A Diagnostic Benchmark for Very Long-form Video Language Understanding

As of 22 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 27 inbound Pith citation observations for arXiv:2308.09126.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2308.09126 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 27 of 27 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 27 of 27 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T10:50:21.460707Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T16:39:57.306434Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 64afd474-cd84-470f-a8c7-086f8848c8d4 · inbound

MVBench: A Comprehensive Multi-modal Video Understanding Benchmark cites this paper.

MVBench: A Comprehensive Multi-modal Video Understanding Benchmark EgoSchema: A Diagnostic Benchmark for Very Long-form Video Language Understanding

Reference 52

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:22:35.036418Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-17T20:22:34.954228Z digest=sha256:5f493bac74cefba2fef55ebc45e70558fa74878b6456d23c2d45526a228a1b3c

Observation 46e6c3fb-224b-45ad-9dca-67673151045b · inbound

VCA: Video Curious Agent for Long Video Understanding cites this paper.

VCA: Video Curious Agent for Long Video Understanding EgoSchema: A Diagnostic Benchmark for Very Long-form Video Language Understanding

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-11T16:51:15.224560Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:51:15.224560Z digest=sha256:3902faab2b20925fb13637ac557b2e7a450f341e9add284f05fe9d17ade8e971

Observation 8e5744cd-a924-45b2-8ab4-c4eb5b1f5cca · inbound

Online Video Understanding: OVBench and VideoChat-Online cites this paper.

Online Video Understanding: OVBench and VideoChat-Online EgoSchema: A Diagnostic Benchmark for Very Long-form Video Language Understanding

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-10T22:57:40.076720Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:57:40.076720Z digest=sha256:78b3d3c4eccaa86ed6bc31476c7dc8db6d183e21287a9d80f0ef070ea575ef44

Observation 19a7a086-642e-4a1c-8375-70f6a74947f8 · inbound

$\infty$-Video: A Training-Free Approach to Long Video Understanding via Continuous-Time Memory Consolidation cites this paper.

$\infty$-Video: A Training-Free Approach to Long Video Understanding via Continuous-Time Memory Consolidation EgoSchema: A Diagnostic Benchmark for Very Long-form Video Language Understanding

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-09T21:21:45.710426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T21:21:45.710426Z digest=sha256:4946d0c46d6ae42963d313216c1acaff01b0fea166533fd0669059478e688f6a

Observation 68a5db79-232f-41c0-be4a-2b388ff281c4 · inbound

MASR: Self-Reflective Reasoning through Multimodal Hierarchical Attention Focusing for Agent-based Video Understanding cites this paper.

MASR: Self-Reflective Reasoning through Multimodal Hierarchical Attention Focusing for Agent-based Video Understanding EgoSchema: A Diagnostic Benchmark for Very Long-form Video Language Understanding

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-16T10:50:21.460707Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:50:21.460707Z digest=sha256:79a284578e5b3dcafd8991a0998cf00086d3c68c820b0fd814e6157f76e8584b

Observation 2942772e-a2f3-4211-b5d8-6c2a222f5199 · inbound

VideoLLM Benchmarks and Evaluation: A Survey cites this paper.

VideoLLM Benchmarks and Evaluation: A Survey EgoSchema: A Diagnostic Benchmark for Very Long-form Video Language Understanding

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-16T04:10:57.295597Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:10:57.295597Z digest=sha256:4f48706642218bf7c3d18485543de6b8f3ad4d81b3b7af6c3f8c6de76b806649

Observation d3ebb7c6-b8f5-4554-81cc-2595dd0a4c2b · inbound

Movie Facts and Fibs (MF$^2$): A Benchmark for Long Movie Understanding cites this paper.

Movie Facts and Fibs (MF$^2$): A Benchmark for Long Movie Understanding EgoSchema: A Diagnostic Benchmark for Very Long-form Video Language Understanding

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:56.930356Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:56.930356Z digest=sha256:c192450fe953e27535dfc4d56501e7c0e1045480e5c3ef1c82301f0e0d4df14c

Observation 9e01ba51-5333-48b2-be59-6a13003a0f9b · inbound

CausalVQA: A Physically Grounded Causal Reasoning Benchmark for Video Models cites this paper.

CausalVQA: A Physically Grounded Causal Reasoning Benchmark for Video Models EgoSchema: A Diagnostic Benchmark for Very Long-form Video Language Understanding

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T04:45:00.265342Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:45:00.265342Z digest=sha256:a8e5f7d25b64b00e422804edd39f352d6db72b7c85db041535eeaec6efb7b15c

Observation 468d22ad-8ab8-4b9f-b59b-9d312cb0db29 · inbound

LaCo: Efficient Layer-wise Compression of Visual Tokens for Multimodal Large Language Models cites this paper.

LaCo: Efficient Layer-wise Compression of Visual Tokens for Multimodal Large Language Models EgoSchema: A Diagnostic Benchmark for Very Long-form Video Language Understanding

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T20:38:30.134630Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:38:30.134630Z digest=sha256:4652deacab25f05c891e3745a430cd3a22f7d64ba561f519c0d8690614de4c41

Observation e0001093-1c8d-428c-8368-db9820138888 · inbound

Infinite Video Understanding cites this paper.

Infinite Video Understanding EgoSchema: A Diagnostic Benchmark for Very Long-form Video Language Understanding

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T18:10:12.839920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:10:12.839920Z digest=sha256:a764cfb1a2a91b45323cef4ccfbd998e6ff67afab7d14048c23d31d22155e0b2

Observation 86434552-e4a1-40e7-b298-5d29ddd11a74 · inbound

HERMES: KV Cache as Hierarchical Memory for Efficient Streaming Video Understanding cites this paper.

HERMES: KV Cache as Hierarchical Memory for Efficient Streaming Video Understanding EgoSchema: A Diagnostic Benchmark for Very Long-form Video Language Understanding

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-16T12:57:53.948203Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-16T12:55:04.564442Z digest=sha256:7d59036811c89d5601a888b3b7dc47bfeb24bbb0b19172359ea1e8138f90ec9a

Observation 05cf38b2-0bf8-4475-b41a-9d86e78056f9 · inbound

Adaptive Greedy Frame Selection for Long Video Understanding cites this paper.

Adaptive Greedy Frame Selection for Long Video Understanding EgoSchema: A Diagnostic Benchmark for Very Long-form Video Language Understanding

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-15T08:25:18.980287Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-15T08:22:24.872665Z digest=sha256:58d0fc94e25b18fecc572d83ab522b4dc55ed5981ab1e7bb88cbdb38c9db8fbf

Observation e0e18d26-0bf9-486f-b1f2-9c81172d79f5 · inbound

MACS: Modality-Aware Capacity Scaling for Efficient Multimodal MoE Inference cites this paper.

MACS: Modality-Aware Capacity Scaling for Efficient Multimodal MoE Inference EgoSchema: A Diagnostic Benchmark for Very Long-form Video Language Understanding

Reference 35

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T01:45:52.094055Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-11T01:29:26.298131Z digest=sha256:25aae75b34314367369c4d99de7d66449e9508a0b864b49141923619dd573edc

Observation c04ffac2-679e-424d-a546-0c9ed249b1ff · inbound

HumanNet: Scaling Human-centric Video Learning to One Million Hours cites this paper.

HumanNet: Scaling Human-centric Video Learning to One Million Hours EgoSchema: A Diagnostic Benchmark for Very Long-form Video Language Understanding

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:10:54.191355Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-11T00:51:08.414394Z digest=sha256:e45cd8c56a1d408b20a0c7d65eb2e3b8861313123346f679b6db05cd1a40c973

Observation f0cc8be1-9168-4401-b590-c2240d6e1ad1 · inbound

EgoCoT-Bench: Benchmarking Grounded and Verifiable Operation-Centric Chain of Thought Reasoning for MLLMs cites this paper.

EgoCoT-Bench: Benchmarking Grounded and Verifiable Operation-Centric Chain of Thought Reasoning for MLLMs EgoSchema: A Diagnostic Benchmark for Very Long-form Video Language Understanding

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-20T05:53:22.271439Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-20T05:53:05.450946Z digest=sha256:9d87691261c583a61b3e5aa003e897d3b745c5acf4c587d17870978e52ba6d5f

Observation ccad6fc8-7167-48d0-a6bc-5f44830d7c13 · inbound

EgoProx: Evaluating MLLMs on Egocentric 3D Proximity Reasoning Across a Cognitive Hierarchy cites this paper.

EgoProx: Evaluating MLLMs on Egocentric 3D Proximity Reasoning Across a Cognitive Hierarchy EgoSchema: A Diagnostic Benchmark for Very Long-form Video Language Understanding

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-06-30T13:34:40.524149Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-30T13:28:43.538541Z digest=sha256:00fd14088878311bf12d35de68ca326cd2e2c0baa73d1ac3a00c9c9267577b44

Observation dc68e372-e42d-4d31-9ffa-227accf43bb7 · inbound

TeachObs: A Human-Validated Benchmark for Multimodal Teaching Observation and Model Evaluation cites this paper.

TeachObs: A Human-Validated Benchmark for Multimodal Teaching Observation and Model Evaluation EgoSchema: A Diagnostic Benchmark for Very Long-form Video Language Understanding

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-06-28T23:12:46.641053Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-28T23:11:43.712391Z digest=sha256:8e4cce4d4c14beabbd5702adfd39784c0cacd919423167d139338848f930a25f

Observation 09494012-2fb3-4efb-8910-3abc33393d82 · inbound

Pause and Think: A Dataset and Benchmark for Video-Grounded Assistive Action Suggestion cites this paper.

Pause and Think: A Dataset and Benchmark for Video-Grounded Assistive Action Suggestion EgoSchema: A Diagnostic Benchmark for Very Long-form Video Language Understanding

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-06-28T19:12:34.434642Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-28T19:11:28.937135Z digest=sha256:4889f7f53cca3d0a9d49c6afa0759e92ad30d0201b135a32065c4988e90379ec

Observation b38608ab-01fb-4bc1-b61a-febc772f53a0 · inbound

Pause and Think: A Dataset and Benchmark for Video-Grounded Assistive Action Suggestion cites this paper.

Pause and Think: A Dataset and Benchmark for Video-Grounded Assistive Action Suggestion EgoSchema: A Diagnostic Benchmark for Very Long-form Video Language Understanding

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-06-30T11:44:38.506790Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-30T11:37:50.291543Z digest=sha256:5108b9e29c88f704ab93c1724ee833c1928cc5f7b838537163c12cd8b5e826bd

Observation a35d243a-9411-42c6-a6eb-2415516408d5 · inbound

SuperMemory-VQA: An Egocentric Visual Question-Answering Benchmark for Long-Horizon Memory cites this paper.

SuperMemory-VQA: An Egocentric Visual Question-Answering Benchmark for Long-Horizon Memory EgoSchema: A Diagnostic Benchmark for Very Long-form Video Language Understanding

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-06-28T19:02:34.145726Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-28T19:00:20.891691Z digest=sha256:800e02053bbbbb7553a31a7b764f3b5e99f07da8b1010f6e81196725aa389d20

Observation b33479a2-f95b-47b1-bb50-245f20f7b627 · inbound

NextMotionQA: Benchmarking and Judging Human Motion Understanding with Vision-Language Models cites this paper.

NextMotionQA: Benchmarking and Judging Human Motion Understanding with Vision-Language Models EgoSchema: A Diagnostic Benchmark for Very Long-form Video Language Understanding

Reference 44

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T07:26:46.156168Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-06-28T06:55:22.332372Z digest=sha256:e615e75cccd5a452bc8eb1a94b91942ee189c43950d8f2dd3f41ac7baeea7f8b

Observation 2ceec0df-85e7-4ce4-8fab-9895482f6318 · inbound

MOSS-Video-Preview: Toward Real-Time Video Understanding via Cross-Attention cites this paper.

MOSS-Video-Preview: Toward Real-Time Video Understanding via Cross-Attention EgoSchema: A Diagnostic Benchmark for Very Long-form Video Language Understanding

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-07-01T22:26:17.890023Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-28T15:22:31.310003Z digest=sha256:9dcc1a4ea65e4bbe112cca20ca08dd1b5072f116a22de86b7b3d2d592a7fce43

Observation 4f2851b0-9617-47dc-92f0-6fb59e21afad · inbound

EgoSAT: A Comprehensive Benchmark of Egocentric Streaming Interaction Understanding cites this paper.

EgoSAT: A Comprehensive Benchmark of Egocentric Streaming Interaction Understanding EgoSchema: A Diagnostic Benchmark for Very Long-form Video Language Understanding

Reference 24

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T16:39:57.307946Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-26T00:26:33.306128Z digest=sha256:56221afb036ad8d15586304f1737ab6a1f66272143321f052e6970630c042f6d

Observation 365b3d41-34c0-4c51-bf09-e04b99c0b6c3 · inbound

From Accuracy to Visual Dependence: Auditing and Filtering Modality Collapse in Traffic VideoQA cites this paper.

From Accuracy to Visual Dependence: Auditing and Filtering Modality Collapse in Traffic VideoQA EgoSchema: A Diagnostic Benchmark for Very Long-form Video Language Understanding

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-06-30T06:14:18.581710Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-30T06:14:11.109840Z digest=sha256:a02e1b5da48bfde2919c6837ed02a7367781eb7d0cab77d353a16a19e4b61649

Observation dbf7d4f3-fae6-4381-8f1b-f91477ad19b1 · inbound

MEMORA: Embodied Action Memory from Egocentric Videos for Reasoning and Planning cites this paper.

MEMORA: Embodied Action Memory from Egocentric Videos for Reasoning and Planning EgoSchema: A Diagnostic Benchmark for Very Long-form Video Language Understanding

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-02T02:42:47.589671Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T02:42:47.589671Z digest=sha256:9c10ed62173450b4fe7ebf196496365d23d62d8b9086da351f5bd26de91ec42f

Observation b7e99be9-fe47-41e4-a095-7abe1408cc75 · inbound

ObjectStream: Latent Objects as Memory Anchors for Streaming Video Understanding cites this paper.

ObjectStream: Latent Objects as Memory Anchors for Streaming Video Understanding EgoSchema: A Diagnostic Benchmark for Very Long-form Video Language Understanding

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-04T03:21:46.535881Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T03:21:46.535881Z digest=sha256:4cef0d6620a8600359e235fc12cf6951dd7a4426cd8af56a791112c0eda3cac2

Observation 6e0885d0-099b-498c-9184-bfdcca6e8f75 · inbound

CRAFT: Compression via Recursive Adaptive Fusion of Video Tokens for Vision-Language Models cites this paper.

CRAFT: Compression via Recursive Adaptive Fusion of Video Tokens for Vision-Language Models EgoSchema: A Diagnostic Benchmark for Very Long-form Video Language Understanding

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-04T23:46:51.588231Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T23:46:51.588231Z digest=sha256:ef3de4825d8530bc4c9f5df83a48597fc33e9160fe13498c8a7c36b085235963