Pith. sign in

Paper Citation Record · LEDGER

E.T. Bench: Towards Open-Ended Event-Level Video-Language Understanding

As of 15 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 8 inbound Pith citation observations for arXiv:2409.18111.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2409.18111 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 8 of 8 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 8 of 8 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-11T20:54:16.286499Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-28T23:12:46.642138Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 86425931-5824-470a-b991-7e2b0337dc5a · inbound

LinVT: Empower Your Image-level Large Language Model to Understand Videos cites this paper.

LinVT: Empower Your Image-level Large Language Model to Understand Videos E.T. Bench: Towards Open-Ended Event-Level Video-Language Understanding

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-11T20:54:16.286499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:54:16.286499Z digest=sha256:6c2f7499181301bbb13e74af012c8609dba0ce2a43664f264e0736681bb5cf8f

Observation 039de858-bae8-427b-a797-8f0160f23146 · inbound

DisTime: Distribution-based Time Representation for Video Large Language Models cites this paper.

DisTime: Distribution-based Time Representation for Video Large Language Models E.T. Bench: Towards Open-Ended Event-Level Video-Language Understanding

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T12:32:52.655077Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:32:52.655077Z digest=sha256:24b1cbf6ec30df8762a09aaef10d84403d9e07ddc782729f54a5f98925fde001

Observation ca08579a-f769-47a9-87cc-a7bb7a4e650c · inbound

VRBench: A Benchmark for Multi-Step Reasoning in Long Narrative Videos cites this paper.

VRBench: A Benchmark for Multi-Step Reasoning in Long Narrative Videos E.T. Bench: Towards Open-Ended Event-Level Video-Language Understanding

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T04:22:55.922688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:22:55.922688Z digest=sha256:bd22f28753212d08a3baf443b414810615c79372fd9a38da38bbb37d37c50127

Observation ec5d62f4-428d-4497-bbd9-08a8a5b08701 · inbound

AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding cites this paper.

AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding E.T. Bench: Towards Open-Ended Event-Level Video-Language Understanding

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:52.796096Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:29:52.796096Z digest=sha256:8ba9cce5dc3a710ea3481bc80a63a09540ad7f797e66edfa37ed07402659fb83

Observation 420a5d6b-82a0-4e9c-97f4-e00fa50e180c · inbound

ProactiveVideoQA: A Comprehensive Benchmark Evaluating Proactive Interactions in Video Large Language Models cites this paper.

ProactiveVideoQA: A Comprehensive Benchmark Evaluating Proactive Interactions in Video Large Language Models E.T. Bench: Towards Open-Ended Event-Level Video-Language Understanding

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T18:03:14.017689Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:03:14.017689Z digest=sha256:53c8c49159994f51eda3d782602ebf98d11e08fa6b687a68838a8c7b2b50bf83

Observation dbb30744-8dba-4a3a-814f-8920adb9b8f4 · inbound

Minerva-Ego: Spatiotemporal Hints for Egocentric Video Understanding cites this paper.

Minerva-Ego: Spatiotemporal Hints for Egocentric Video Understanding E.T. Bench: Towards Open-Ended Event-Level Video-Language Understanding

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-19T16:03:08.056726Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-19T16:02:53.887605Z digest=sha256:a09a4d976149cc670c298b316da8dc086c76f656fd64540c971ebcd2e4b774aa

Observation 08226c30-4fac-4025-8537-5582c9283b83 · inbound

An Efficient Streaming Video Understanding Framework with Agentic Control cites this paper.

An Efficient Streaming Video Understanding Framework with Agentic Control E.T. Bench: Towards Open-Ended Event-Level Video-Language Understanding

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-05-20T11:33:14.447270Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-05-20T11:30:22.151045Z digest=sha256:02dad14b01528b1d00122570e0091aefb314bba5c0968be71eda5c7160d2863b

Observation ca3797d4-6e56-45df-9a2e-bf35d2f14a0e · inbound

TeachObs: A Human-Validated Benchmark for Multimodal Teaching Observation and Model Evaluation cites this paper.

TeachObs: A Human-Validated Benchmark for Multimodal Teaching Observation and Model Evaluation E.T. Bench: Towards Open-Ended Event-Level Video-Language Understanding

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-06-28T23:12:46.643590Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-28T23:11:43.712391Z digest=sha256:362d3c75e36875e6f218034d2743592f9c295d7d5572e1faa62f03b9347a1596