Pith. sign in

Paper Citation Record · LEDGER

PHYBench: Holistic Evaluation of Physical Perception and Reasoning in Large Language Models

As of 16 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 33 inbound Pith citation observations for arXiv:2504.16074.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2504.16074 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 33 of 33 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 33 of 33 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:36:26.588538Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

1
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 34ceef5a-16a8-4f97-8686-d72cca1ab5ab · inbound

KORGym: A Dynamic Game Platform for LLM Reasoning Evaluation cites this paper.

KORGym: A Dynamic Game Platform for LLM Reasoning Evaluation PHYBench: Holistic Evaluation of Physical Perception and Reasoning in Large Language Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T15:36:26.588538Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:36:26.588538Z digest=sha256:c92310c67d2a3b84a45cf184e992bbeb4064aad8f12c314a02e057b30b5d02ef

Observation c0b9f034-2e97-42c1-bb1a-d5525b541192 · inbound

lmgame-Bench: How Good are LLMs at Playing Games? cites this paper.

lmgame-Bench: How Good are LLMs at Playing Games? PHYBench: Holistic Evaluation of Physical Perception and Reasoning in Large Language Models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T15:26:59.524310Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:26:59.524310Z digest=sha256:774e7d8c7d90a6397e6e69b38c00d1d47f69bcc240651695e317936ad1946b9c

Observation f59b5a75-24ae-458c-8033-15a5d622bfe8 · inbound

PhysicsArena: The First Multimodal Physics Reasoning Benchmark Exploring Variable, Process, and Solution Dimensions cites this paper.

PhysicsArena: The First Multimodal Physics Reasoning Benchmark Exploring Variable, Process, and Solution Dimensions PHYBench: Holistic Evaluation of Physical Perception and Reasoning in Large Language Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T15:20:26.948825Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:20:26.948825Z digest=sha256:d9062b6c50a89855de6cd3b27074ffc43d92e539bd379f97a7fa68e45849602b

Observation 3e72e41a-3861-42b4-b98c-1fab90fbf033 · inbound

PhyX: Does Your Model Have the "Wits" for Physical Reasoning? cites this paper.

PhyX: Does Your Model Have the "Wits" for Physical Reasoning? PHYBench: Holistic Evaluation of Physical Perception and Reasoning in Large Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T15:14:54.491719Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:14:54.491719Z digest=sha256:ca9997b9a81420f06fba402d3a16499cebe7c93bd8d2b0e89c57dd3a045192ab

Observation 137ad425-18a0-4f6a-bb26-72eb091899c9 · inbound

On Path to Multimodal Historical Reasoning: HistBench and HistAgent cites this paper.

On Path to Multimodal Historical Reasoning: HistBench and HistAgent PHYBench: Holistic Evaluation of Physical Perception and Reasoning in Large Language Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T14:01:14.955222Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:01:14.955222Z digest=sha256:e50e2032f994bd19a548cb3dc9a630fd6dc8f7c28b61a229a9123e5642f3e851

Observation b1401a96-4fdf-4fce-8ad7-6f18ae4f77ba · inbound

PhySense: Principle-Based Physics Reasoning Benchmarking for Large Language Models cites this paper.

PhySense: Principle-Based Physics Reasoning Benchmarking for Large Language Models PHYBench: Holistic Evaluation of Physical Perception and Reasoning in Large Language Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T12:22:25.013730Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:22:25.013730Z digest=sha256:b3c5c9ee29d0db8771e0775c1f87305c4f3686e599c57604380d3d2e002254fd

Observation cb23a208-8f20-4c44-afa0-73abc1de32fb · inbound

Generalizable LLM Learning of Graph Synthetic Data with Post-training Alignment cites this paper.

Generalizable LLM Learning of Graph Synthetic Data with Post-training Alignment PHYBench: Holistic Evaluation of Physical Perception and Reasoning in Large Language Models

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-07T12:02:44.713393Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:02:44.713393Z digest=sha256:94536e26074771bf89da6e50b35592c933d0b26bf5aa08d6039cd391c603abdd

Observation ef9be859-608e-4677-a58a-2d2f2b39bf22 · inbound

OneIG-Bench: Omni-dimensional Nuanced Evaluation for Image Generation cites this paper.

OneIG-Bench: Omni-dimensional Nuanced Evaluation for Image Generation PHYBench: Holistic Evaluation of Physical Perception and Reasoning in Large Language Models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T05:25:44.353897Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:25:44.353897Z digest=sha256:5153bd5d04945f391fe7f8204476029a947e3e057c4336939c588af5eba1c78f

Observation a7ec1267-8c38-4f5b-8ca6-c35061de6a75 · inbound

SciDA: Scientific Dynamic Assessor of LLMs cites this paper.

SciDA: Scientific Dynamic Assessor of LLMs PHYBench: Holistic Evaluation of Physical Perception and Reasoning in Large Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T00:40:18.603262Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:40:18.603262Z digest=sha256:dfc692c20253e95ad4aa7106ec6b4a6d2b7dcc031773e5b8af041c178f7c3d36

Observation bc12a3de-5e3b-4c8c-9447-ab0a6ee1263d · inbound

ABench-Physics: Benchmarking Physical Reasoning in LLMs via High-Difficulty and Dynamic Physics Problems cites this paper.

ABench-Physics: Benchmarking Physical Reasoning in LLMs via High-Difficulty and Dynamic Physics Problems PHYBench: Holistic Evaluation of Physical Perception and Reasoning in Large Language Models

Reference 1949

Resolution
unresolved
no resolver link, observed 2026-08-06T19:46:46.994892Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:46:46.994892Z digest=sha256:f2f327887b0efed3a2348bdba81bd15ad6c3d97ac75e353d6dbc17472cec2d7c

Observation f23f2682-f21a-46f2-be0b-74d2cd113e36 · inbound

Agentic Exploration of Physics Models cites this paper.

Agentic Exploration of Physics Models PHYBench: Holistic Evaluation of Physical Perception and Reasoning in Large Language Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-04T13:51:31.512492Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T13:51:31.512492Z digest=sha256:c353d62f2d24b865daab1d6919e69c75d110d8463c5e93212e2255c2bbe65403

Observation e0b744c6-cc25-48c7-86ca-448723632ff8 · inbound

Probing the Critical Point (CritPt) of AI Reasoning: a Frontier Physics Research Benchmark cites this paper.

Probing the Critical Point (CritPt) of AI Reasoning: a Frontier Physics Research Benchmark PHYBench: Holistic Evaluation of Physical Perception and Reasoning in Large Language Models

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-18T11:52:35.306254Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-18T11:52:10.205796Z digest=sha256:912a833ab91d8c68787bf6a9994a0ea3cac3d1115ce970f261695c559f4bfb9e

Observation cbac3c2e-1c69-46f2-8286-5d56524584fb · inbound

LLaDA2.0: Scaling Up Diffusion Language Models to 100B cites this paper.

LLaDA2.0: Scaling Up Diffusion Language Models to 100B PHYBench: Holistic Evaluation of Physical Perception and Reasoning in Large Language Models

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-14T18:53:21.116805Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-14T18:53:20.911374Z digest=sha256:c4d00647d0eb50700e526c2c4e39bf1be9b5c31646c33927e080e9ba3fe00731

Observation d603fdcf-2b82-42da-b44d-d8cc4cf1ac57 · inbound

Ministral 3 cites this paper.

Ministral 3 PHYBench: Holistic Evaluation of Physical Perception and Reasoning in Large Language Models

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-14T19:12:24.750113Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-14T19:12:24.627033Z digest=sha256:83bdf59fc5ecdc44f0eda9c9ccd86d6574f998c9e07cd9b196d6038c8e4cb356

Observation b7c46987-ace9-4a1a-9360-c9abda75a829 · inbound

Attention Sink Forges Native MoE in Attention Layers: Sink-Aware Training to Address Head Collapse cites this paper.

Attention Sink Forges Native MoE in Attention Layers: Sink-Aware Training to Address Head Collapse PHYBench: Holistic Evaluation of Physical Perception and Reasoning in Large Language Models

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-16T08:47:37.220461Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-16T08:47:29.236561Z digest=sha256:53a239c4b69b1b79a2621ad6414123488def223b39d817a4410b2ba46c58f916

Observation 5682c563-bb91-4c9e-bcff-234fb65fec30 · inbound

Attention Sink Forges Native MoE in Attention Layers: Sink-Aware Training to Address Head Collapse cites this paper.

Attention Sink Forges Native MoE in Attention Layers: Sink-Aware Training to Address Head Collapse PHYBench: Holistic Evaluation of Physical Perception and Reasoning in Large Language Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-03T05:50:24.771643Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T05:50:24.771643Z digest=sha256:44423d256cf7509c3c8e3d519c8c29fd76637c0a4befcbccaa05e8580d1291db

Observation d9d0573e-e1ac-475f-a076-319faa5b08f4 · inbound

OmniFysics: Towards Physical Intelligence Evolution via Omni-Modal Signal Processing and Network Optimization cites this paper.

OmniFysics: Towards Physical Intelligence Evolution via Omni-Modal Signal Processing and Network Optimization PHYBench: Holistic Evaluation of Physical Perception and Reasoning in Large Language Models

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-16T07:10:43.156734Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-16T07:09:46.254851Z digest=sha256:42eaaf5ac7d0cada91a1b3157dd69137c526b1129ae45c28436a72113f39b295

Observation af40aa66-2d6c-4ea5-aac5-f4df7fab6785 · inbound

Vision Language Models Cannot Reason About Physical Transformation cites this paper.

Vision Language Models Cannot Reason About Physical Transformation PHYBench: Holistic Evaluation of Physical Perception and Reasoning in Large Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-07-15T13:27:51.848177Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T13:27:51.848177Z digest=sha256:dc08feb74f9b844814d2a7e1df618b8d27cce696006bd5455a9d68a98aa28407

Observation 969d8e6d-5cbb-49a2-a9a8-34101c6babf9 · inbound

Seed1.8 Model Card: Towards Generalized Real-World Agency cites this paper.

Seed1.8 Model Card: Towards Generalized Real-World Agency PHYBench: Holistic Evaluation of Physical Perception and Reasoning in Large Language Models

Reference 55

Resolution
verified exact
arxiv_id, observed 2026-05-15T07:45:14.405531Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-15T07:44:02.827006Z digest=sha256:d6e1b7a8ee6066499e764982c557df855e4dc8fc0d83231a3ff920225bcd63af

Observation c20b9fcf-9eaf-4c23-ac30-c378aaf4a9db · inbound

PolyReal: A Benchmark for Real-World Polymer Science Workflows cites this paper.

PolyReal: A Benchmark for Real-World Polymer Science Workflows PHYBench: Holistic Evaluation of Physical Perception and Reasoning in Large Language Models

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-13T19:58:12.284099Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-13T19:55:17.366868Z digest=sha256:3bc5816f677f0c19967117c7e980ce5a88db15d53925389b11f56fb643aa63fe

Observation aaaff15b-3695-4c8d-a275-dc55b59c8288 · inbound

PRL-Bench: A Comprehensive Benchmark Evaluating LLMs' Capabilities in Frontier Physics Research cites this paper.

PRL-Bench: A Comprehensive Benchmark Evaluating LLMs' Capabilities in Frontier Physics Research PHYBench: Holistic Evaluation of Physical Perception and Reasoning in Large Language Models

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-10T11:20:11.381486Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-10T11:10:21.639856Z digest=sha256:38354baf66b5e2d945d5475b2a2a168cdb112359825fe4de0d126f81dec3e045

Observation f727223d-5410-41c5-ab9f-d9748d93aea9 · inbound

PhysCodeBench: Benchmarking Physics-Aware Symbolic Simulation of 3D Scenes via Self-Corrective Multi-Agent Refinement cites this paper.

PhysCodeBench: Benchmarking Physics-Aware Symbolic Simulation of 3D Scenes via Self-Corrective Multi-Agent Refinement PHYBench: Holistic Evaluation of Physical Perception and Reasoning in Large Language Models

Reference 22

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T21:16:13.110357Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-08T06:13:18.967603Z digest=sha256:1d5a13e25745e6f80a53d4657312d239f4b4c4d324d5ef7348edcf73e045b1c2

Observation 79a6c43f-2664-4ec9-b725-153a6c4c7309 · inbound

PhysNote: Self-Knowledge Notes for Evolvable Physical Reasoning in Vision-Language Model cites this paper.

PhysNote: Self-Knowledge Notes for Evolvable Physical Reasoning in Vision-Language Model PHYBench: Holistic Evaluation of Physical Perception and Reasoning in Large Language Models

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-11T21:56:26.628602Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-08T03:43:18.640857Z digest=sha256:69e615ca808a779811fa7bf10a3b0b86f9656806f1b7b3e386f3a1a870e25123

Observation 5d636051-6d1a-4be7-abd3-9f50fa66a361 · inbound

Heterogeneous Scientific Foundation Model Collaboration cites this paper.

Heterogeneous Scientific Foundation Model Collaboration PHYBench: Holistic Evaluation of Physical Perception and Reasoning in Large Language Models

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-09T04:30:10.962510Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-07T08:50:05.980191Z digest=sha256:1946d379b4620655b3bd87d4aefe0d17ffe576511a4ee4db62161f1c280d336c

Observation c55c8f39-8847-494f-ac22-95dec3f160dd · inbound

PDEAgent-Bench: A Multi-Metric, Multi-Library Benchmark for PDE Solver Generation cites this paper.

PDEAgent-Bench: A Multi-Metric, Multi-Library Benchmark for PDE Solver Generation PHYBench: Holistic Evaluation of Physical Perception and Reasoning in Large Language Models

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-12T07:31:26.675661Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-12T02:36:40.696567Z digest=sha256:a0acad5a8dfef4828b9fea1199bbc40f81ed16490b9cab53893372f9c2e29800

Observation 668e0963-bcde-40db-8433-fc7ccdcea260 · inbound

MPMWorlds: Material-Point-Method Simulations for Inferring and Extrapolating Physical Dynamics cites this paper.

MPMWorlds: Material-Point-Method Simulations for Inferring and Extrapolating Physical Dynamics PHYBench: Holistic Evaluation of Physical Perception and Reasoning in Large Language Models

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-07-02T01:16:24.687383Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-06-28T12:19:27.221596Z digest=sha256:c6251990d26137b0fd343fd05565d2ab04048f189cb89b3f156aae1efb244864

Observation 22bbed5b-1b7d-4d4c-87bf-c95da090f955 · inbound

ChronoPhyBench: Do MLLMs Truly Understand the World or Merely Exploit Language Priors? cites this paper.

ChronoPhyBench: Do MLLMs Truly Understand the World or Merely Exploit Language Priors? PHYBench: Holistic Evaluation of Physical Perception and Reasoning in Large Language Models

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-07-02T20:27:22.619660Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-06-27T20:23:18.667677Z digest=sha256:f3300db09ad8cdf26d57f5a6c1f187640ac55cc37e9dccd51b43f2fa82a49ba3

Observation 188ce006-ec3f-4cf6-8db8-cef93e7c21fe · inbound

Sci-Rho: A Multilingual Visually-Grounded Symbolic Benchmark for STEM Problems cites this paper.

Sci-Rho: A Multilingual Visually-Grounded Symbolic Benchmark for STEM Problems PHYBench: Holistic Evaluation of Physical Perception and Reasoning in Large Language Models

Reference 64

Resolution
verified exact
arxiv_id, observed 2026-07-02T20:47:22.878722Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-06-27T20:11:02.445626Z digest=sha256:2e3a1d8fdc94a3eb02f58659b92bb6c4c552c2d1bb8774167761272b4adb75bb

Observation 323ff441-fde6-4d5e-95e1-41ece48248e6 · inbound

Seed2.0 Model Card: Towards Intelligence Frontier for Real-World Complexity cites this paper.

Seed2.0 Model Card: Towards Intelligence Frontier for Real-World Complexity PHYBench: Holistic Evaluation of Physical Perception and Reasoning in Large Language Models

Reference 88

Resolution
verified exact
arxiv_id, observed 2026-07-02T19:07:17.387781Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-07-02T18:57:46.841456Z digest=sha256:3ab709587ee27fac2fb3918f9d21d7c945b7704f9476a0f79fefa8fb5554ee87

Observation 2bb6da02-4f4e-4cd2-bfe9-f1a20a654948 · inbound

Testing Frontier Large Language Models' Physics Literacy in Parallel Physical Worlds cites this paper.

Testing Frontier Large Language Models' Physics Literacy in Parallel Physical Worlds PHYBench: Holistic Evaluation of Physical Perception and Reasoning in Large Language Models

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-07-02T19:27:18.703824Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-07-02T19:18:43.558804Z digest=sha256:f96ec9396dbeaf7ab2c83e9ab833f4edcb2ddfb3aa2b2163d4040a201dc1c6c7

Observation 1d9d6a85-fc24-471f-84df-0d55338e06e6 · inbound

Imaging-101: Benchmarking LLM Coding Agents on Scientific Computational Imaging cites this paper.

Imaging-101: Benchmarking LLM Coding Agents on Scientific Computational Imaging PHYBench: Holistic Evaluation of Physical Perception and Reasoning in Large Language Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-07-14T09:17:00.467000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T09:17:00.467000Z digest=sha256:9703db1acdf4e4f43fcbdcd896cc9303af1503985a81a7e27d80c90c82e69edf

Observation a9f75970-1fd3-4dcf-9cb4-c4e41ddf764a · inbound

CLVisc Agent for autonomous relativistic hydrodynamics studies cites this paper.

CLVisc Agent for autonomous relativistic hydrodynamics studies PHYBench: Holistic Evaluation of Physical Perception and Reasoning in Large Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-01T00:52:32.240795Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:52:32.240795Z digest=sha256:44d993e3dbf7a208a2f5b5b0e703afa19f05ffdc8ad62bfc457f71c573c54336

Observation 6db93d44-64d1-43b9-b14c-0093100b1754 · inbound

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping cites this paper.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping PHYBench: Holistic Evaluation of Physical Perception and Reasoning in Large Language Models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T04:24:54.863248Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:24:54.863248Z digest=sha256:28bfaf12fd0a80536ae4faa809a9de6f514a320981aa13566482467504471051