Pith. sign in

Paper Citation Record · LEDGER

CATER: A diagnostic dataset for Compositional Actions and TEmporal Reasoning

As of 14 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 15 inbound Pith citation observations for arXiv:1910.04744.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
1910.04744 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 15 of 15 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 15 of 15 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-12T21:07:09.467237Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T12:49:53.191967Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 392126da-5990-48c3-a699-73335e1101e7 · inbound

TempCompass: Do Video LLMs Really Understand Videos? cites this paper.

TempCompass: Do Video LLMs Really Understand Videos? CATER: A diagnostic dataset for Compositional Actions and TEmporal Reasoning

Reference 85

Resolution
verified exact
arxiv_id, observed 2026-05-17T02:46:16.784145Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-17T02:46:16.632743Z digest=sha256:861feb1ff57178e70f1e48e41822541a8283e71eca736e294f74e3b1047bef1f

Observation 85907818-1608-44f8-b939-c07d911e3541 · inbound

VideoCogQA: A Controllable Benchmark for Evaluating Cognitive Abilities in Video-Language Models cites this paper.

VideoCogQA: A Controllable Benchmark for Evaluating Cognitive Abilities in Video-Language Models CATER: A diagnostic dataset for Compositional Actions and TEmporal Reasoning

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-12T21:07:09.467237Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T21:07:09.467237Z digest=sha256:fc6ca8dc594f3150754ba96b34375e0220a1d94ef7dd1277d8b93044c06c3377

Observation 13911945-1afd-4b63-80e6-a897e2f06dfb · inbound

Video Representation Learning with Joint-Embedding Predictive Architectures cites this paper.

Video Representation Learning with Joint-Embedding Predictive Architectures CATER: A diagnostic dataset for Compositional Actions and TEmporal Reasoning

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-11T15:33:34.688128Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:33:34.688128Z digest=sha256:b083cc6e7c8ac61376d6b216666fd66046c298fce74d5f600e12e385187896ff

Observation 172adfa6-e429-4450-90f1-f9b4c35f8bb4 · inbound

Do Language Models Understand Time? cites this paper.

Do Language Models Understand Time? CATER: A diagnostic dataset for Compositional Actions and TEmporal Reasoning

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-11T12:47:17.268614Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:47:17.268614Z digest=sha256:c3e8130548910b8e33e3bdf70d028c79e17e3c1d0fc0f617b0db9e3ffdcae6d8

Observation ad3ab6d0-712e-41b8-933e-0cba572da75a · inbound

An Empirical Study of Autoregressive Pre-training from Videos cites this paper.

An Empirical Study of Autoregressive Pre-training from Videos CATER: A diagnostic dataset for Compositional Actions and TEmporal Reasoning

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:08.433520Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:08.433520Z digest=sha256:447576323c3392506e76501aacbe3448b40e593a4cdd2be7b9bfed39ae71c81e

Observation 243d1345-a77f-4e04-8745-6af9140535df · inbound

Finding the Trigger: Causal Abductive Reasoning on Video Events cites this paper.

Finding the Trigger: Causal Abductive Reasoning on Video Events CATER: A diagnostic dataset for Compositional Actions and TEmporal Reasoning

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-10T20:11:06.725782Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:11:06.725782Z digest=sha256:1a72ae9fa0fd34649648965aa1fda7d0b2bdd09d6b7c3835273eded2529c2bf3

Observation 6c279471-d033-4ad8-8d46-16a9da93dcc7 · inbound

Do Vision-Language Models Have Internal World Models? Towards an Atomic Evaluation cites this paper.

Do Vision-Language Models Have Internal World Models? Towards an Atomic Evaluation CATER: A diagnostic dataset for Compositional Actions and TEmporal Reasoning

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-06T22:22:36.996473Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:22:36.996473Z digest=sha256:7e88c0a4fe497b8732e34334fc02083492ea62b61a862f49739c01572dc3cc60

Observation 3e3f57ba-3de1-4988-b913-7542a7a44a9b · inbound

Video Understanding by Design: How Datasets Shape Video Models cites this paper.

Video Understanding by Design: How Datasets Shape Video Models CATER: A diagnostic dataset for Compositional Actions and TEmporal Reasoning

Reference 146

Resolution
unresolved
no resolver link, observed 2026-08-04T19:37:41.978225Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:37:41.978225Z digest=sha256:dd9ea0851441cf3729cde1a08ab2902298ed8fd936a1bc3781709a0b2382d524

Observation 667d2b12-ff81-4155-952d-6a72bce92ebf · inbound

LychSim: A Controllable and Interactive Simulation Framework for Vision Research cites this paper.

LychSim: A Controllable and Interactive Simulation Framework for Vision Research CATER: A diagnostic dataset for Compositional Actions and TEmporal Reasoning

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-13T06:27:24.551372Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-13T06:26:14.825783Z digest=sha256:7acf6ee5e832c2cfcccc5f0add650211ddf7525bc62905040f8ae58085e2abff

Observation faf0e921-e144-4e6b-8fdb-c8eb70274db6 · inbound

VGenST-Bench: A Benchmark for Spatio-Temporal Reasoning via Active Video Synthesis cites this paper.

VGenST-Bench: A Benchmark for Spatio-Temporal Reasoning via Active Video Synthesis CATER: A diagnostic dataset for Compositional Actions and TEmporal Reasoning

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-22T07:14:42.896419Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-22T07:12:02.612292Z digest=sha256:4815d8b7aed66a74783fab75b47339e101041527993d55ff9995308c2b899dae

Observation 9758491e-52e3-419f-92c0-c140056475ef · inbound

What-If World: A Causal Benchmark for General World Models in Embodied Scenarios cites this paper.

What-If World: A Causal Benchmark for General World Models in Embodied Scenarios CATER: A diagnostic dataset for Compositional Actions and TEmporal Reasoning

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-06-29T18:23:50.381067Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-29T18:23:22.987086Z digest=sha256:3f6a95563005082d8e855959bbb76dfd7a24e2bf5d788c02b1a05182cd34af7f

Observation 50f1a3af-5596-4128-ad9d-c0b632926066 · inbound

ChronoPhyBench: Do MLLMs Truly Understand the World or Merely Exploit Language Priors? cites this paper.

ChronoPhyBench: Do MLLMs Truly Understand the World or Merely Exploit Language Priors? CATER: A diagnostic dataset for Compositional Actions and TEmporal Reasoning

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-07-02T20:27:22.646770Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-27T20:23:18.667677Z digest=sha256:a385680397761530942a8ed91acd120d416a17ab0b58ce035e25fb5f1fbe6942

Observation 3bbcbeb4-7029-4eac-abca-5c6929ff1cb9 · inbound

PhysEditWorld: A Large-Scale Dataset Toward Physics-Editable World Models cites this paper.

PhysEditWorld: A Large-Scale Dataset Toward Physics-Editable World Models CATER: A diagnostic dataset for Compositional Actions and TEmporal Reasoning

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-07-04T12:49:53.193845Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-26T05:46:21.198781Z digest=sha256:1af018899eb6c4a9de89dad0d0b6d1577b6d46ed86b806985e8d2bdc51a85b61

Observation fd5361db-c85a-4b70-8427-4e2c485d1979 · inbound

PhysEditWorld: A Large-Scale Dataset Toward Physics-Editable World Models cites this paper.

PhysEditWorld: A Large-Scale Dataset Toward Physics-Editable World Models CATER: A diagnostic dataset for Compositional Actions and TEmporal Reasoning

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-06-30T12:04:39.236002Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-30T10:19:06.268547Z digest=sha256:ac95a00a9104025c8eac4cbf772620100e3f90ee55e0a70b66b8ddeb03bc2966

Observation 92d5a29b-286d-4fa8-8f7b-61ca4ec6b850 · inbound

IMBench: A Benchmark for Intuitive Robotic Manipulation cites this paper.

IMBench: A Benchmark for Intuitive Robotic Manipulation CATER: A diagnostic dataset for Compositional Actions and TEmporal Reasoning

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-01T22:45:11.729883Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T22:45:11.729883Z digest=sha256:e66aa4cea08095edeffdff6f35adadb6ee4865ba1b74149c24ad360157bb7c6a