Pith. sign in

Paper Citation Record · LEDGER

A Shortcut-aware Video-QA Benchmark for Physical Understanding via Minimal Video Pairs

As of 14 August 2026, this Paper Citation Record lists 19 of 19 outbound references and 9 inbound Pith citation observations for arXiv:2506.09987.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.09987 v1

Coverage vector

measured 19 of 19 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T04:42:37.838863Z

measured 28 of 28 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 9 of 9 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-12T17:56:48.420938Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-30T10:34:36.459072Z

Reference resolution

19 of 19 outbound references displayed

  • verified exact0
  • verified fuzzy9
  • unresolved10
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation e1d51c16-d587-425b-8a09-5d466e971a4a · outbound

This paper cites an unresolved cited work.

A Shortcut-aware Video-QA Benchmark for Physical Understanding via Minimal Video Pairs Unresolved cited work

Reference 1

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:42:38.021779Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T04:42:37.791497Z digest=sha256:178963d9f0a45f8d52c4aafba30bbb8c16baf1ea213ece2c5c43ef36d39520ac

Observation b41858ef-cdc5-44d6-99b9-31afd1f77cdb · outbound

This paper cites an unresolved cited work.

A Shortcut-aware Video-QA Benchmark for Physical Understanding via Minimal Video Pairs Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:42:38.013541Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T04:42:37.794495Z digest=sha256:0901f23be4970d5d601c8c10fceca4a56a177ac10991e05eb022e0517e51bbf5

Observation 66c2b1e8-7371-4b74-8cc5-df5c8aeeaa2a · outbound

This paper cites Both the other options.

A Shortcut-aware Video-QA Benchmark for Physical Understanding via Minimal Video Pairs Both the other options

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:42:38.004229Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T04:42:37.797183Z digest=sha256:32aafbf2e75e3423a246c219233b51cbd20069ced33d0e50f8d134c9aab0a913

Observation bd009d75-00f7-4b7a-b8c9-29b293f3d716 · outbound

This paper cites an unresolved cited work.

A Shortcut-aware Video-QA Benchmark for Physical Understanding via Minimal Video Pairs Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:42:37.895789Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T04:42:37.833710Z digest=sha256:7bb846920c94e6a7299ac3fff41a34aa98e16f5ce85b3b96849746f03f722861

Observation c6261c81-d004-4762-a5f9-8001f49cfa31 · outbound

This paper cites red triangle.

A Shortcut-aware Video-QA Benchmark for Physical Understanding via Minimal Video Pairs red triangle

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:42:37.995692Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T04:42:37.800448Z digest=sha256:70967df2ba35d366a6539e0e85ccaff854c03cd516aa63357fb1f51405acce66

Observation f99e1750-b010-420e-b9d3-d8970378befc · outbound

This paper cites an unresolved cited work.

A Shortcut-aware Video-QA Benchmark for Physical Understanding via Minimal Video Pairs Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:42:37.987394Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T04:42:37.803395Z digest=sha256:703c7a22f130fbd5bcec26dda0f2191c020f6591dd22c232228d11623baf9f66

Observation 69ea94b9-b263-472e-9366-353ef21944d8 · outbound

This paper cites Move yel- low triangle to blue heart.

A Shortcut-aware Video-QA Benchmark for Physical Understanding via Minimal Video Pairs Move yel- low triangle to blue heart

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:42:37.979713Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T04:42:37.806541Z digest=sha256:3822356d4181321060fdd45e9150ab00c196f8fade6ae580380a1be4e32738b4

Observation 005c38f2-f694-44fe-9342-861277fdb40b · outbound

This paper cites spinning something so it continues spinning.

A Shortcut-aware Video-QA Benchmark for Physical Understanding via Minimal Video Pairs spinning something so it continues spinning

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:42:37.971538Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T04:42:37.809225Z digest=sha256:826ddcfaea010f90c303cb1d47d2b980587e56bd131396e0f497bcd995d731af

Observation b72bd6c6-65e7-4901-9a8f-65bf87f22f33 · outbound

This paper cites If no pairs fulfill this strict criterion, we relax it such that only one object must overlap: P ′ ={(x i, xj)|obj(v i)∩obj(v j)̸=∅}.

A Shortcut-aware Video-QA Benchmark for Physical Understanding via Minimal Video Pairs If no pairs fulfill this strict criterion, we relax it such that only one object must overlap: P ′ ={(x i, xj)|obj(v i)∩obj(v j)̸=∅}

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:42:37.963363Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T04:42:37.812436Z digest=sha256:fadc0e37f8e1da22d137f4f7a5254ca2686ad1358460ab46c416a9503bfb20ce

Observation 3f9c877e-4d02-463e-a496-bc3020564916 · outbound

This paper cites ,(xik , xjk ) sim(vim, vjm)≥sim(v im+1 , vjm+1 ).

A Shortcut-aware Video-QA Benchmark for Physical Understanding via Minimal Video Pairs ,(xik , xjk ) sim(vim, vjm)≥sim(v im+1 , vjm+1 )

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:42:37.955040Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T04:42:37.815170Z digest=sha256:fa3fb495e20db641f30d611e25d7445ae0316dfb82b3092c6081aecc124fe13e

Observation 6e3998be-cef0-4861-8f75-4d4d4d98171a · outbound

This paper cites an unresolved cited work.

A Shortcut-aware Video-QA Benchmark for Physical Understanding via Minimal Video Pairs Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:42:37.947008Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T04:42:37.817732Z digest=sha256:1c3345cc352a2b2f738c9e5949a52847ad84f0cc3c0ba831c561ed37b2e3f7bd

Observation 8e81b4fe-efa6-42bf-8555-74848e2652dd · outbound

This paper cites How many objects are moving when the video ends? A) 2 B) 3.

A Shortcut-aware Video-QA Benchmark for Physical Understanding via Minimal Video Pairs How many objects are moving when the video ends? A) 2 B) 3

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:42:37.939002Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T04:42:37.820670Z digest=sha256:b3208122ae5faa948a6f3b009e7d7a6b0c99d73ffa0ed6db8d508bada0571a29

Observation e75a8ed1-44d9-420d-bf2c-b7bf1387efa0 · outbound

This paper cites fuzzy subset.

A Shortcut-aware Video-QA Benchmark for Physical Understanding via Minimal Video Pairs fuzzy subset

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:42:37.930499Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T04:42:37.823555Z digest=sha256:831b14b1ae8daef9f4fdf0de1927a5152b57706691d67565a2a684f1829c840f

Observation 702d5902-e00f-43fa-9fe8-30291a4df7e4 · outbound

This paper cites an unresolved cited work.

A Shortcut-aware Video-QA Benchmark for Physical Understanding via Minimal Video Pairs Unresolved cited work

Reference 14

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:42:37.921742Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T04:42:37.825928Z digest=sha256:d650620328aa5df31345701de5931dd213851356bed78b3d2a387e53e192aa85

Observation babc19e3-831e-4baf-bcdc-c15f13d6c74a · outbound

This paper cites an unresolved cited work.

A Shortcut-aware Video-QA Benchmark for Physical Understanding via Minimal Video Pairs Unresolved cited work

Reference 15

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:42:37.912490Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T04:42:37.828749Z digest=sha256:2ef2ab369aa6a9cf51cb91948bd86328c5f3754d477782753aa4e0e4b9f03e48

Observation 6818c665-9438-4ae7-a293-f15165473c8f · outbound

This paper cites an unresolved cited work.

A Shortcut-aware Video-QA Benchmark for Physical Understanding via Minimal Video Pairs Unresolved cited work

Reference 16

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:42:37.903908Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T04:42:37.831292Z digest=sha256:84b7653196b90559da6cf78ecafa7159d37a3fca90a582808aa96a4b798e5316

Observation 0134955b-604f-4347-8bc9-54b88e603cb6 · outbound

This paper cites In order to push the field further we are now asking the models more and more nuanced questions, and the answer may lie only in a short span of a less than second.

A Shortcut-aware Video-QA Benchmark for Physical Understanding via Minimal Video Pairs In order to push the field further we are now asking the models more and more nuanced questions, and the answer may lie only in a short span of a less than second

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:42:37.886812Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T04:42:37.836507Z digest=sha256:ca7e18e4883dfc041d00c19996548146e47eb929a8d34f45549b2b4643cec10a

Observation 3dc7ba2b-2020-4a9b-9dca-c64f364d549c · outbound

This paper cites an unresolved cited work.

A Shortcut-aware Video-QA Benchmark for Physical Understanding via Minimal Video Pairs Unresolved cited work

Reference 19

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:42:37.876885Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T04:42:37.838863Z digest=sha256:a96ad07f2e98d1ff034b677e1b9b100a12b1c9e89dfbdb7aaca2f58581dc0d85

Observation 19cd38e0-24ab-4abf-a6f4-3b2267239606 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

A Shortcut-aware Video-QA Benchmark for Physical Understanding via Minimal Video Pairs LLaVA-OneVision: Easy Visual Task Transfer

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-07T04:42:37.786681Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:42:37.786681Z digest=sha256:0f25f70c24c81ac16a816ca9e6e4c258de28abe729993717b8bb6247ecc5444b

Pith citing papers

Observation e6b2e4ab-88ac-49fb-8b66-3b16879d33fa · inbound

Embodied AI Agents: Modeling the World cites this paper.

Embodied AI Agents: Modeling the World A Shortcut-aware Video-QA Benchmark for Physical Understanding via Minimal Video Pairs

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T22:10:28.130255Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:10:28.130255Z digest=sha256:f694d771e27a0232c3fbf848e1ffa963bb823268f88ea9172c3c4b4894890f01

Observation c04fef8c-e3c7-438a-8f02-d83742c4efab · inbound

VisPhyWorld: Probing Physical Reasoning via Code-Driven Video Reconstruction cites this paper.

VisPhyWorld: Probing Physical Reasoning via Code-Driven Video Reconstruction A Shortcut-aware Video-QA Benchmark for Physical Understanding via Minimal Video Pairs

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-22T11:11:27.528157Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-22T11:06:38.102348Z digest=sha256:e00980ff1311a5f247782b4330e08a2ee5b1654064e34c33e238552f4380573b

Observation 1d155d78-ff93-41bd-a425-c7d5f8c5fd6f · inbound

TOC-Bench: A Temporal Object Consistency Benchmark for Video Large Language Models cites this paper.

TOC-Bench: A Temporal Object Consistency Benchmark for Video Large Language Models A Shortcut-aware Video-QA Benchmark for Physical Understanding via Minimal Video Pairs

Reference 18

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T06:31:26.671992Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-12T04:13:21.487431Z digest=sha256:9a5734da9d1778d4e5491d078d08228fc2d9b1c4fe36f397b2a52b20dc0e2c6e

Observation 96a3f3a5-4152-47f6-a1d6-735db32b4ae3 · inbound

TOC-Bench: A Temporal Object Consistency Benchmark for Video Large Language Models cites this paper.

TOC-Bench: A Temporal Object Consistency Benchmark for Video Large Language Models A Shortcut-aware Video-QA Benchmark for Physical Understanding via Minimal Video Pairs

Reference 18

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T06:57:28.342496Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-13T06:53:42.726350Z digest=sha256:f0f08d7790b47233c18231e8f50db6d3a791ea95a4b9b4fbfd03016bc41bf5c2

Observation 7a01867c-a937-408a-8b7d-9184f6135dba · inbound

Learning Spatiotemporal Sensitivity in Video LLMs via Counterfactual Reinforcement Learning cites this paper.

Learning Spatiotemporal Sensitivity in Video LLMs via Counterfactual Reinforcement Learning A Shortcut-aware Video-QA Benchmark for Physical Understanding via Minimal Video Pairs

Reference 18

Resolution
metadata mismatch
arxiv_id, observed 2026-05-22T07:46:14.800731Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-22T07:45:56.473188Z digest=sha256:851bdd2f9a6ce733646dd3e827696e8065beff5084731ebb7552a8039568a41e

Observation 3f6b82da-50f9-4164-84f1-77cc48a726ec · inbound

VGenST-Bench: A Benchmark for Spatio-Temporal Reasoning via Active Video Synthesis cites this paper.

VGenST-Bench: A Benchmark for Spatio-Temporal Reasoning via Active Video Synthesis A Shortcut-aware Video-QA Benchmark for Physical Understanding via Minimal Video Pairs

Reference 36

Resolution
metadata mismatch
arxiv_id, observed 2026-05-22T07:14:42.783296Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-22T07:12:02.612292Z digest=sha256:4cc5ae4ea608413000107fb7fe691f6411b06b4d418cd73ce3c495e3664f1bfc

Observation 2a6531c2-e3a3-43da-8c2f-ded6c2f038b7 · inbound

How Should World Models Be Evaluated for Embodied Decision-Making? A Decision-Making-Centric Position cites this paper.

How Should World Models Be Evaluated for Embodied Decision-Making? A Decision-Making-Centric Position A Shortcut-aware Video-QA Benchmark for Physical Understanding via Minimal Video Pairs

Reference 34

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T10:34:36.460645Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-30T10:31:37.108792Z digest=sha256:001eaa95e80f2744aed4e543ff7aba51303d1881b03299da9805db12adab6649

Observation 4e1e5bf9-0685-4249-9df0-f4c04de7cc11 · inbound

AdaThinkV: Adaptive Thinking for Token-Efficient Video Reasoning cites this paper.

AdaThinkV: Adaptive Thinking for Token-Efficient Video Reasoning A Shortcut-aware Video-QA Benchmark for Physical Understanding via Minimal Video Pairs

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-04T17:24:16.066057Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T17:24:16.066057Z digest=sha256:bd864fa29e55dd8dfeb0f57e8729ec144a3232e4b842776ee7a30874cefcd2ad

Observation 74fed687-0104-4443-a078-658533090feb · inbound

FADE: From Passive Verification to Active Discovery in Counterfactual Video Understanding cites this paper.

FADE: From Passive Verification to Active Discovery in Counterfactual Video Understanding A Shortcut-aware Video-QA Benchmark for Physical Understanding via Minimal Video Pairs

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-12T17:56:48.420938Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T17:56:48.420938Z digest=sha256:ab49adaa70b776a4b20e9ca4cd533649c77834972d329635d9a2fc7f375caab4